Plain Text from PDF
Audience:
Low-code EngineersSkill Prerequisites:
Tokens,Actions
Reads the text of a PDF file. The text of the whole document can be saved in a token, and a list of actions can run once for each page with that page's text. Optionally, the images in the PDF are saved to a portal folder.
The action reads the text that's stored in the PDF. It doesn't use OCR, so scanned documents and pages that are images return no text.
This action requires the PDF feature to be licensed.
Typical Use Cases
- Save the text of an uploaded PDF in a database, so it can be searched
- Find a value in a document, such as an invoice number or a total, and store it in a token
- Process a multi-page document page by page, for example saving one database row per page
- Send the text of a PDF to an AI or text-processing service
Don't use it to
- Read scanned documents or photos of documents. They contain images, not text, and no OCR is done.
- Copy pages into a new PDF. Use Extract PDF instead.
- Keep the layout or formatting of the document. Only plain text is returned.
Related Actions
| Action Name | Description |
|---|---|
| Extract PDF | Copies selected pages of a PDF into a new file. |
| Check PDF Signature | Checks whether a PDF contains an electronic signature. |
| Run SQL Query | Saves the extracted text, or each page's text, in a database. |
| Log Debug Message | Writes the extracted text to the log while you build the flow. |
Input Parameter Reference
| Parameter | Description | Supports Tokens | Default | Required |
|---|---|---|---|---|
| Source file | The PDF to read. Use one of the supported file identifiers. | Yes | empty string | Yes |
| Folder | The portal folder where the images found in the PDF are saved. Select a folder from the list, or switch to expression mode and enter a folder path relative to the portal root. The folder must already exist. When the value is empty, images aren't extracted. See Image extraction. | Yes | none selected | Yes |
| On Process Page | Actions that run once for each page, in page order. The page's text is in the [PDF:PageText] token. See Tokens in On Process Page. | No | empty | No |
Supported file identifiers
The file must be in the portal's file system.
| Format | Example |
|---|---|
| File ID | 77 |
| Path relative to the portal root | Invoices/Invoice-1024.pdf |
| Path that includes the portal folder | /Portals/0/Invoices/Invoice-1024.pdf |
| Absolute URL of a file on the site | https://example.com/Portals/0/Invoices/Invoice-1024.pdf |
| LinkClick URL | /LinkClick.aspx?fileticket=... |
| Physical path inside the portal folder | C:\inetpub\site\Portals\0\Invoices\Invoice-1024.pdf |
For files uploaded through a Single File Upload field, the field's [FieldName:FileId] token is a convenient identifier.
Output Parameters Reference
| Parameter | Description |
|---|---|
| Store Entire Plain Text | The name of the token that stores the text of the whole document. The text of each page is followed by a line break. If images are extracted, a line for each image is added after the text of its page. |
Tokens in On Process Page
These tokens are available to the actions in On Process Page:
| Token | Description |
|---|---|
[PDF:PageText] | The text of the current page. |
[PDF:PageNumber] | The number of the current page. The first page is 1. |
[PDF:PagesCount] | The total number of pages in the document. |
The actions run in a copy of the current context. They can read the tokens that existed before the action ran, but tokens they create aren't available after Plain Text from PDF finishes. To keep a result, save it somewhere, for example with Run SQL Query.
Image extraction
When Folder has a value, each image found on a page is saved in that folder, named <PDF file name> - Page <n> - Image <i>.jpeg. For example, the first image on page 2 of Invoice-1024.pdf is saved as Invoice-1024.pdf - Page 2 - Image 1.jpeg.
A line like this is added to the Store Entire Plain Text token for each image, after the text of its page:
Image 1 : https://example.com/LinkClick.aspx?fileticket=...
Keep in mind:
- The image data is saved as it's stored in the PDF, always with a
.jpegextension. Images stored in another format may not open correctly. - Images with the same name are overwritten, so reading the same PDF twice replaces the images from the first run.
- The image URLs are built from the current web request. Image extraction may fail when the action runs without one, for example in a scheduled job.
Considerations
- No OCR. Text is read with the PdfPig library, in reading order. Pages that contain only images return empty text.
- The action fails with
File '...' not found.if the file can't be found, and withFolder '...' not found.ifFolderis set to a folder that doesn't exist. - Password-protected PDFs aren't supported. There's no password parameter, so a PDF that needs a password to open can't be read.
- The text of a large document can be long. Check the size of the database column before saving the whole text.
- Conditions compare token values as quoted strings, for example
[PDF:PageNumber] == "1".
Examples
To understand how to use the below examples, please see Running Examples.
1. Save the text of an uploaded PDF
This action reads the PDF uploaded in the Document field, saves its images in Documents/Images, and saves the text in the DocumentText token. A later action can store [DocumentText] in the database.
{
"Title": "Plain Text from PDF",
"ActionType": "PlainTextFromPdf",
"Description": "Read the text of the uploaded document",
"Parameters": {
"FileIdentifier": "[Document:FileId]",
"Folder": {
"Expression": "",
"Value": "/Documents/Images",
"IsExpression": false,
"Parameters": {}
},
"StorePlainText": "DocumentText",
"OnProcessPage": []
}
}
2. Save each page in a database table
This action runs a SQL query for every page of the PDF, saving the page number and text in a DocumentPages table. The values are passed as bound parameters, so quotes in the page text don't break the query. Images aren't needed here, but Folder still has a value because the editor requires one.
{
"Title": "Plain Text from PDF",
"ActionType": "PlainTextFromPdf",
"Description": "Save one row per page",
"Parameters": {
"FileIdentifier": "[Document:FileId]",
"Folder": {
"Expression": "",
"Value": "/Documents/Images",
"IsExpression": false,
"Parameters": {}
},
"StorePlainText": "",
"OnProcessPage": [
{
"Title": "Run SQL Query",
"ActionType": "RunSql",
"Description": "Insert the page text",
"Parameters": {
"SqlQuery": "INSERT INTO DocumentPages (DocumentId, PageNumber, PageText) VALUES (@DocumentId, @PageNumber, @PageText)",
"BindTokens": [
{
"name": "DocumentId",
"value": "[DocumentId]"
},
{
"name": "PageNumber",
"value": "[PDF:PageNumber]"
},
{
"name": "PageText",
"value": "[PDF:PageText]"
}
]
}
}
]
}
}
Revised 09/26/2026