Skip to main content
Version: 1.28 (Current)

Plain Text from PDF

Audience: Low-code Engineers

Skill Prerequisites: Tokens, Actions

Reads the text of a PDF file. The text of the whole document can be saved in a token, and a list of actions can run once for each page with that page's text. Optionally, the images in the PDF are saved to a portal folder.

The action reads the text that's stored in the PDF. It doesn't use OCR, so scanned documents and pages that are images return no text.

note

This action requires the PDF feature to be licensed.

Typical Use Cases​

  • Save the text of an uploaded PDF in a database, so it can be searched
  • Find a value in a document, such as an invoice number or a total, and store it in a token
  • Process a multi-page document page by page, for example saving one database row per page
  • Send the text of a PDF to an AI or text-processing service

Don't use it to​

  • Read scanned documents or photos of documents. They contain images, not text, and no OCR is done.
  • Copy pages into a new PDF. Use Extract PDF instead.
  • Keep the layout or formatting of the document. Only plain text is returned.
Action NameDescription
Extract PDFCopies selected pages of a PDF into a new file.
Check PDF SignatureChecks whether a PDF contains an electronic signature.
Run SQL QuerySaves the extracted text, or each page's text, in a database.
Log Debug MessageWrites the extracted text to the log while you build the flow.

Input Parameter Reference​

ParameterDescriptionSupports TokensDefaultRequired
Source fileThe PDF to read. Use one of the supported file identifiers.Yesempty stringYes
FolderThe portal folder where the images found in the PDF are saved. Select a folder from the list, or switch to expression mode and enter a folder path relative to the portal root. The folder must already exist. When the value is empty, images aren't extracted. See Image extraction.Yesnone selectedYes
On Process PageActions that run once for each page, in page order. The page's text is in the [PDF:PageText] token. See Tokens in On Process Page.NoemptyNo

Supported file identifiers​

The file must be in the portal's file system.

FormatExample
File ID77
Path relative to the portal rootInvoices/Invoice-1024.pdf
Path that includes the portal folder/Portals/0/Invoices/Invoice-1024.pdf
Absolute URL of a file on the sitehttps://example.com/Portals/0/Invoices/Invoice-1024.pdf
LinkClick URL/LinkClick.aspx?fileticket=...
Physical path inside the portal folderC:\inetpub\site\Portals\0\Invoices\Invoice-1024.pdf

For files uploaded through a Single File Upload field, the field's [FieldName:FileId] token is a convenient identifier.

Output Parameters Reference​

ParameterDescription
Store Entire Plain TextThe name of the token that stores the text of the whole document. The text of each page is followed by a line break. If images are extracted, a line for each image is added after the text of its page.

Tokens in On Process Page​

These tokens are available to the actions in On Process Page:

TokenDescription
[PDF:PageText]The text of the current page.
[PDF:PageNumber]The number of the current page. The first page is 1.
[PDF:PagesCount]The total number of pages in the document.

The actions run in a copy of the current context. They can read the tokens that existed before the action ran, but tokens they create aren't available after Plain Text from PDF finishes. To keep a result, save it somewhere, for example with Run SQL Query.

Image extraction​

When Folder has a value, each image found on a page is saved in that folder, named <PDF file name> - Page <n> - Image <i>.jpeg. For example, the first image on page 2 of Invoice-1024.pdf is saved as Invoice-1024.pdf - Page 2 - Image 1.jpeg.

A line like this is added to the Store Entire Plain Text token for each image, after the text of its page:

Image 1 : https://example.com/LinkClick.aspx?fileticket=...

Keep in mind:

  • The image data is saved as it's stored in the PDF, always with a .jpeg extension. Images stored in another format may not open correctly.
  • Images with the same name are overwritten, so reading the same PDF twice replaces the images from the first run.
  • The image URLs are built from the current web request. Image extraction may fail when the action runs without one, for example in a scheduled job.

Considerations​

  • No OCR. Text is read with the PdfPig library, in reading order. Pages that contain only images return empty text.
  • The action fails with File '...' not found. if the file can't be found, and with Folder '...' not found. if Folder is set to a folder that doesn't exist.
  • Password-protected PDFs aren't supported. There's no password parameter, so a PDF that needs a password to open can't be read.
  • The text of a large document can be long. Check the size of the database column before saving the whole text.
  • Conditions compare token values as quoted strings, for example [PDF:PageNumber] == "1".

Examples​

tip

To understand how to use the below examples, please see Running Examples.

1. Save the text of an uploaded PDF​

This action reads the PDF uploaded in the Document field, saves its images in Documents/Images, and saves the text in the DocumentText token. A later action can store [DocumentText] in the database.

{
"Title": "Plain Text from PDF",
"ActionType": "PlainTextFromPdf",
"Description": "Read the text of the uploaded document",
"Parameters": {
"FileIdentifier": "[Document:FileId]",
"Folder": {
"Expression": "",
"Value": "/Documents/Images",
"IsExpression": false,
"Parameters": {}
},
"StorePlainText": "DocumentText",
"OnProcessPage": []
}
}

2. Save each page in a database table​

This action runs a SQL query for every page of the PDF, saving the page number and text in a DocumentPages table. The values are passed as bound parameters, so quotes in the page text don't break the query. Images aren't needed here, but Folder still has a value because the editor requires one.

{
"Title": "Plain Text from PDF",
"ActionType": "PlainTextFromPdf",
"Description": "Save one row per page",
"Parameters": {
"FileIdentifier": "[Document:FileId]",
"Folder": {
"Expression": "",
"Value": "/Documents/Images",
"IsExpression": false,
"Parameters": {}
},
"StorePlainText": "",
"OnProcessPage": [
{
"Title": "Run SQL Query",
"ActionType": "RunSql",
"Description": "Insert the page text",
"Parameters": {
"SqlQuery": "INSERT INTO DocumentPages (DocumentId, PageNumber, PageText) VALUES (@DocumentId, @PageNumber, @PageText)",
"BindTokens": [
{
"name": "DocumentId",
"value": "[DocumentId]"
},
{
"name": "PageNumber",
"value": "[PDF:PageNumber]"
},
{
"name": "PageText",
"value": "[PDF:PageText]"
}
]
}
}
]
}
}

Revised 09/26/2026