Crawl Next Batch
Audience:
Low-code EngineersSkill Prerequisites:
Actions,Tokens,Automation Jobs
Fetches the next pages from the web crawler's queue and runs actions for each one. Pages get into the queue with Add Site to Crawl, and from the links found on pages already crawled.
Each run handles a small batch, so you normally run this action from a scheduled automation job. For the whole crawl flow and the crawl rules, see How crawling works.
This action is part of the Web Crawler add-on (PlantAnApp.WebCrawler). The add-on is installed separately and needs the WEBCRWL feature in your license. If it isn't licensed, the action fails with a "not licensed" error. If you don't see the Web Crawler actions, the add-on isn't installed.
Typical Use Cases
- Save each crawled page in a table, for example to search it with Index Rule
- Clean each page with Clean HTML and summarize it with Summarize Content
- Log broken links with On Page NotFound
- Collect the external links of a site with On External Link
- List the PDF files published on a site with On Process Document
Don't use it to
- Crawl a site that wasn't added. Use Add Site to Crawl first.
- Get one page right now. Use Fetch Content.
- Crawl a whole site in one run. Each run crawls at most one page per site. See Which pages are crawled.
Related Actions
| Action Name | Description |
|---|---|
| Add Site to Crawl | Adds a site and its crawl rules to the queue. |
| Remove Site to Crawl | Removes a site and its queued pages. |
| Clean HTML | Keeps only the main content of [Crawler:RawHtml]. |
| Fetch Content | Fetches a single page right away. |
| Summarize Content | Summarizes the text of a page with AI. |
| Execute Actions for each List Entry | Loops over a list, for example the links you collected in the page actions. |
Input Parameter Reference
| Parameter | Description | Supports Tokens | Default | Required |
|---|---|---|---|---|
| Batch Size | The most pages to crawl in one run, for example 20. Since each run takes at most one page per site, this is also the most sites handled in one run. A value of 0 or less uses the default. | Yes | 10 | No |
| Parallel Page Crawls Size | How many pages are fetched at the same time. A value of 0 or less uses the default. | Yes | 5 | No |
| Min Page Reindex Interval | How many hours to wait before a page is fetched again, for example 24. A site's own Min Page Reindex Interval, set in Add Site to Crawl, wins over this value. A value of 0 or less uses the default. | Yes | 24 | No |
| On Process Page | Actions to run for each HTML page that was fetched successfully. See Page tokens. | No | empty | No |
| On Process Document | Actions to run for each file of a type listed in the site's File Types, such as a PDF. See Document tokens. | No | empty | No |
| On Page NotFound | Actions to run for each page that returned 404 Not Found. See Error tokens. | No | empty | No |
| On Page Error | Actions to run for each page that returned another error status, for example 403 or 500. See Error tokens. | No | empty | No |
| On External Link | Actions to run for each link to another host found on a page. See External link tokens. | No | empty | No |
Output Parameters Reference
This action doesn't create tokens after it runs. The tokens below are only available inside the actions of each event.
Page tokens
Available in On Process Page.
| Token | Value |
|---|---|
[Crawler:RawHtml] | The HTML of the page, exactly as it was downloaded. It isn't cleaned. Use Clean HTML to keep only the main content. |
[Crawler:PageTitle] | The text of the page's <title>. If there's no title, the path of the address, for example /docs/guides/setup. |
[Crawler:PageUrl] | The full address of the page, for example https://www.example.com/docs/guides/setup. |
[Crawler:RelativePath] | The part of the address after the site's Base Url, for example /guides/setup. |
[Crawler:LastModified] | The date from the page's Last-Modified header, or else its Date header, in ISO 8601 format, for example 2026-09-28T10:15:00.0000000+00:00. Empty if the server sends neither. |
[Crawler:SiteId] | The internal number of the site. It's the value of Output Site Id Token Name in Add Site to Crawl. |
[Crawler:UserAssignedSiteId] | The Site Id you gave the site, for example ExampleDocs. Empty if none was set. |
[Crawler:BaseUrl] | The site's Base Url, without a trailing /. |
Document tokens
Available in On Process Document. The file itself isn't downloaded, only its address is known. Use [Crawler:DocumentUrl] to get the file with another action.
| Token | Value |
|---|---|
[Crawler:DocumentType] | The MIME type, for example application/pdf. |
[Crawler:DocumentTitle] | The file name from the address, for example price-list.pdf. |
[Crawler:DocumentUrl] | The full address of the file. |
[Crawler:RelativePath] | The part of the address after the site's Base Url. |
[Crawler:LastModified] | Always empty for documents. |
[Crawler:SiteId], [Crawler:UserAssignedSiteId], [Crawler:BaseUrl] | The same as for pages. |
Error tokens
Available in On Page NotFound and On Page Error: [Crawler:PageUrl], [Crawler:RelativePath], [Crawler:SiteId], [Crawler:UserAssignedSiteId] and [Crawler:BaseUrl]. There's no token with the status code.
External link tokens
Available in On External Link.
| Token | Value |
|---|---|
[Crawler:ExternalUrl] | The full address of the link, for example https://partner.example.org/pricing. |
[Crawler:ExternalDomain] | The host of the link, for example partner.example.org. |
[Crawler:FoundAtUrl] | The address of the page where the link was found. |
[Crawler:SiteId], [Crawler:UserAssignedSiteId], [Crawler:BaseUrl] | The same as for pages. |
Which pages are crawled
Each run takes, for each site, the one queued page that's been due the longest. Then it keeps up to Batch Size of these pages. Sites that were crawled least recently go first.
So each run crawls at most one page per site, whatever the Batch Size. With a single site, each run crawls one page. To crawl a site of 1,000 pages in a day, the job must run at least every minute or so. A larger Batch Size helps only when you crawl several sites.
A page is due when its next crawl time has passed:
- New pages, from Add Site to Crawl or from links, are due right away.
- After a page is fetched, it's due again after the reindex interval. This also happens when the fetch failed.
- A page that no longer matches the site's Restrict to Paths or Exclude Paths isn't fetched again.
The batch runs across all sites, whichever portal or module added them.
What happens for each page
- The crawler sends a
HEADrequest to get the content type. - For an HTML page, it downloads the page with a
GETrequest. For a file type listed in the site's File Types, it doesn't download anything. Other content types are skipped. - It saves the status of the page and when to fetch it again. When the site's Follow Links is
True, it adds the links found on the page to the queue. See What gets crawled. - When the whole batch is fetched, the events run one page at a time:
| Result | Event |
|---|---|
Status 200–299, HTML page | On Process Page |
Status 200–299, listed file type | On Process Document |
Status 404 | On Page NotFound |
Any other status, for example 401, 403 or 500 | On Page Error |
| No response, for example a timeout, a DNS error or a refused connection | No event. The error is saved with the page, and it's tried again after the reindex interval. |
| Content type that isn't HTML or a listed file type, or no content type | No event. The page is skipped. |
mailto: address | No event. |
After the page's own event, On External Link runs once for each link on the page to another host. It only runs when the site's Follow Links is True, because links are only read then. Links to the same host that are outside the site's paths don't count as external.
Redirects are followed automatically. The tokens show the address that was queued, not where it redirected to.
Considerations
- Run it from a scheduled job. Create an automation job with a schedule trigger, for example every minute, and put this action in it. A job runs without a web request, so tokens such as the current page or query string aren't available. The
[Crawler:...]tokens don't need them. - Tokens for each page are separate. Each page runs its actions on its own copy of the tokens. Tokens created for one page aren't there for the next page, or after the action. Lists are shared, so you can collect values with Add List Entry.
- Save the content yourself. The crawler doesn't store the pages. Save them in the On Process Page actions, for example with Run SQL Query. Use an insert-or-update query, since each page comes back after the reindex interval.
- Failing actions. If an action in an event fails, this action fails, and the rest of the batch doesn't run its events. Those pages are already marked as fetched, so they're only fetched again after the reindex interval. Wrap risky actions in Execute Actions with
On Erroractions. - External links repeat. On External Link runs for every page that has the link, and again each time the page is fetched. Store links with an insert-or-update query to avoid duplicates.
- Errors aren't raised. Pages that fail to download don't make the action fail. Use On Page Error and On Page NotFound to see failing pages. Network errors and timeouts don't run any event. They're saved in the
LastCrawledResultandLastCrawlStatusMessagecolumns of thepaa.WebCrawler_Pagetable, and some are also written to the log. - Servers without
HEAD. A server that doesn't answerHEADrequests with a content type is skipped for that page. - Raw HTML can be large.
[Crawler:RawHtml]holds the whole page, including scripts and styles. Clean it with Clean HTML before saving or summarizing it. - Security. The crawler fetches the addresses it was given, including internal ones, and follows redirects to any host. See Security.
Examples
To understand how to use the below examples, please see Running Examples.
1. Save each page in a table
This action crawls up to 10 pages. For each HTML page, it cleans the HTML and saves it in a CrawledPages table. It updates the row if the page was saved before. Broken links are written to the log.
Run it from a scheduled automation job. It expects a table with SiteId, Url, Title, Html and CrawledOn columns.
{
"Title": "Crawl Next Batch",
"ActionType": "WebCrawler.CrawlNextBatch",
"Description": "Crawl and save the next pages",
"Parameters": {
"BatchSize": 10,
"ParallelPageCrawlsSize": 5,
"MinPageReindexInterval": 24,
"OnProcessPage": [
{
"Title": "Clean HTML",
"ActionType": "WebCrawler.CleanHtml",
"Description": "Keep only the main content",
"Parameters": {
"RawHtml": "[Crawler:RawHtml]",
"StoreCleanHtml": "CleanHtml"
}
},
{
"Title": "Run SQL Query",
"ActionType": "RunSql",
"Description": "Insert or update the page",
"Parameters": {
"SqlQuery": "UPDATE CrawledPages SET Title = @Title, Html = @Html, CrawledOn = GETUTCDATE() WHERE Url = @Url; IF @@ROWCOUNT = 0 INSERT INTO CrawledPages (SiteId, Url, Title, Html, CrawledOn) VALUES (@SiteId, @Url, @Title, @Html, GETUTCDATE());",
"BindTokens": [
{
"name": "SiteId",
"value": "[Crawler:UserAssignedSiteId]"
},
{
"name": "Url",
"value": "[Crawler:PageUrl]"
},
{
"name": "Title",
"value": "[Crawler:PageTitle]"
},
{
"name": "Html",
"value": "[CleanHtml]"
}
]
}
}
],
"OnPageNotFound": [
{
"Title": "Log Error",
"ActionType": "LogError",
"Description": "Log the broken link",
"Parameters": {
"Message": "Crawler: page not found [Crawler:PageUrl] on site [Crawler:UserAssignedSiteId]"
}
}
]
}
}
2. Collect external links and PDF files
This action crawls up to 20 pages. It saves each external link once in an ExternalLinks table, and logs the address of each PDF file. The site must be added with Follow Links set to True and application/pdf in File Types.
{
"Title": "Crawl Next Batch",
"ActionType": "WebCrawler.CrawlNextBatch",
"Description": "Collect external links and PDF files",
"Parameters": {
"BatchSize": 20,
"OnExternalLink": [
{
"Title": "Run SQL Query",
"ActionType": "RunSql",
"Description": "Save the link once",
"Parameters": {
"SqlQuery": "IF NOT EXISTS (SELECT 1 FROM ExternalLinks WHERE Url = @Url) INSERT INTO ExternalLinks (Url, Domain, FoundAt) VALUES (@Url, @Domain, @FoundAt);",
"BindTokens": [
{
"name": "Url",
"value": "[Crawler:ExternalUrl]"
},
{
"name": "Domain",
"value": "[Crawler:ExternalDomain]"
},
{
"name": "FoundAt",
"value": "[Crawler:FoundAtUrl]"
}
]
}
}
],
"OnProcessDocument": [
{
"Title": "Log Error",
"ActionType": "LogError",
"Description": "Log the PDF file",
"Parameters": {
"Message": "Crawler found [Crawler:DocumentTitle] ([Crawler:DocumentType]) at [Crawler:DocumentUrl]"
}
}
]
}
}
Revised 09/28/2026