Search PDF Indexer
Audience:
Low-code Engineers,System/Security AdministratorsSkill Prerequisites:
Search
The Search PDF Indexer add-on lets the site's search read the text inside PDF files, so users can find a PDF by words in its content and not only by its file name. It uses the PDFBox library, so you don't need to install an Adobe PDF IFilter on the server.
The add-on has no actions, fields or settings of its own. It replaces the parser behind the PDF Files file type that you turn on for a search behavior.
This add-on is the DnnSharp.PdfIndexerBox package, with product code SBSTPDF. It's installed separately and needs the Search module. It doesn't check the license when it runs. You can't see from the search settings whether it's installed, because PDF Files is listed either way. Use the Text Extraction tool to check. See Checking that it works.
What it does
Search reads each file type with a content parser. Without this add-on, the PDF Files type uses the Windows IFilter for PDF, which only works if an Adobe PDF IFilter is installed on the server. See Configure File Types.
When the add-on is installed, PDF Files uses the add-on's parser instead:
- What it indexes. The text of every page of a PDF, for files with the
.pdfextension or theapplication/pdftype. - Where from. PDFs in the folders that a behavior's content sources include, when PDF Files is checked in the behavior's File Types. See Content Sources.
- New behaviors. PDF Files is checked by default.
Setting up
- Install the add-on package on the site.
- In each search behavior that should search PDFs, open the Content Sources tab. Check that the folders with your PDFs are included and that PDF Files is checked under File Types.
- Reindex the behavior. PDFs that were indexed before the add-on was installed keep their old content until they're indexed again. You can reindex from the search admin, or with the Index Behavior action.
You don't need the Adobe PDF IFilter. If you installed it only for search, you can leave it or remove it. The add-on is used either way.
Checking that it works
Use the Text Extraction tool in the Content Sources tab to see the text that search reads from a PDF. See Text Extraction. If the text is there, the PDF can be indexed. Then search for a word that's only inside the PDF.
Considerations
- Text only. The parser reads text that is stored in the PDF. Scanned documents that are only images have no text to read, so their content can't be found. Run them through OCR first if you need their content.
- Errors are logged. If a PDF can't be read, for example because it's damaged, the error is written to the site log with the file name and folder, and its text isn't indexed. Indexing of other files goes on.
- Large files. Each PDF is read into memory. Large PDFs also use temporary files in
Portals/_default/temp, so the app pool needs write access to that folder. - One PDF parser. There's an older, separate package,
DnnSharp.SearchBoost.PdfIndexer. It's disabled in this version and isn't shipped. Don't install both.
To search from your own forms, see the Advanced Search add-on and the Search actions.
Revised 10/02/2026