Add Site to Crawl
Audience:
Low-code EngineersSkill Prerequisites:
Actions,Tokens,Automation Jobs
Adds a website to the web crawler's queue. The action saves the site and its crawl rules, and queues the start page. It doesn't download anything itself.
The pages are fetched later by Crawl Next Batch, usually from a scheduled job. It follows the links it finds and runs your actions for each page. To stop crawling a site, use Remove Site to Crawl. See How crawling works.
This action is part of the Web Crawler add-on (PlantAnApp.WebCrawler). The add-on is installed separately and needs the WEBCRWL feature in your license. If it isn't licensed, the action fails with a "not licensed" error. If you don't see the Web Crawler actions, the add-on isn't installed.
Typical Use Cases
- Crawl your documentation or knowledge base, so its pages can be stored, searched or summarized
- Keep a copy of a site's pages up to date, since pages are fetched again after an interval
- Find the PDF files published on a site
- Collect the external links of a site, for example to check them
Don't use it to
- Get one page right now. Use Fetch Content.
- Clean HTML you already have. Use Clean HTML.
- Crawl internal or private addresses from user input. See Security.
Related Actions
| Action Name | Description |
|---|---|
| Crawl Next Batch | Fetches the next queued pages and runs actions for each one. |
| Remove Site to Crawl | Removes a site and its queued pages. |
| Fetch Content | Fetches a single page right away. |
| Clean HTML | Keeps only the main content of a page's HTML. |
| Summarize Content | Summarizes the text of a crawled page with AI. |
| Index Rule | Indexes rows of a table in a search behavior, for example the pages you saved while crawling. |
Input Parameter Reference
| Parameter | Description | Supports Tokens | Default | Required |
|---|---|---|---|---|
| Base Url | The address where crawling starts, for example https://www.example.com/docs/. It must be a full address that starts with http or https. Only pages on the same host and under this path are crawled. See What gets crawled. | Yes | empty string | Yes |
| Site Id | Your own ID for the site, for example ExampleDocs. It must be unique. It's needed to remove the site later, and adding a site with an existing Site Id updates it. See Site Id. | Yes | empty string | No |
| Min Page Reindex Interval | How many hours to wait before a page of this site is fetched again, for example 12. When set, it replaces the Min Page Reindex Interval of Crawl Next Batch for this site. | Yes | empty | No |
| Follow Links | When True, links found on each page are added to the queue. When False, only the start pages are crawled. Supports expressions. | No | True | No |
| Ignore NoFollow | Only shown when Follow Links is True. When True, links with rel="nofollow" are followed too. When False, they're skipped. Supports expressions. | No | False | No |
| Restrict to Paths | Only shown when Follow Links is True. Relative paths to limit the crawl to, for example guides or api/v2. When at least one is set, only pages under these paths are crawled, and these paths become the start pages instead of the Base Url. Each path must be a real page. A token can hold several paths, separated by commas or new lines. | Yes | empty | No |
| Exclude Paths | Only shown when Follow Links is True. Relative paths to skip, for example archive. Pages under these paths aren't crawled. A token can hold several paths, separated by commas or new lines. | Yes | empty | No |
| File Types | The file types to report besides HTML pages. The only choice is application/pdf. Files of these types run the On Process Document actions of Crawl Next Batch. | No | empty | No |
| Dynamic File Types | More file types, as MIME types separated by commas or new lines, for example [FileTypes]. Only application/pdf is recognized, so other types are skipped. | Yes | empty string | No |
| User Agent Header | The User-Agent header sent with each request, for example ExampleCrawler/1.0. It must be a single name/version value. If it's empty, no User-Agent header is sent. | Yes | empty string | No |
Output Parameters Reference
| Parameter | Description |
|---|---|
| Output Site Id Token Name | The name of a token that receives the site's internal number, for example CrawlSiteId. It's the same value as [Crawler:SiteId] in Crawl Next Batch, not your Site Id. |
How crawling works
Crawling a site takes three actions. Each one runs on its own, usually at different times.
- Add Site to Crawl saves the site and its rules in the
paa.WebCrawler_Sitetable. It queues the start page in thepaa.WebCrawler_Pagetable. The start page is the Base Url, or each path of Restrict to Paths when it's set. - Crawl Next Batch takes the pages that are due, fetches them, and runs your actions for each one. When Follow Links is
True, the links it finds are added to the queue. Each fetched page is scheduled to be fetched again after the reindex interval. You normally run it from a scheduled automation job, for example every minute. - Remove Site to Crawl deletes the site and all its queued pages.
A few things to know:
- Each page is stored once per site. When a link is found again, it isn't queued a second time.
- Pages are never removed from the queue while the site exists. A page that disappears keeps being fetched, and runs the On Page NotFound actions each time.
- The crawler only stores the queue and the crawl status. The page content isn't saved. Save what you need in the actions of Crawl Next Batch, for example with Run SQL Query.
- Crawling isn't tied to a portal or module. Crawl Next Batch crawls the queued pages of every site that was added, wherever it was added from.
What gets crawled
| Rule | How it works |
|---|---|
| Host | Only links to the same host as the Base Url are queued. www. is ignored, so example.com and www.example.com are the same host. The scheme and port aren't compared. |
| Path | Only links whose path starts with the Base Url path are queued. With a Base Url of https://www.example.com/docs, the page https://www.example.com/blog isn't crawled. |
| Restrict to Paths and Exclude Paths | The path after the Base Url must start with one of the Restrict to Paths, and mustn't start with any of the Exclude Paths. The check is a case-sensitive "starts with", so guide also matches guides and guide-old. |
| Links | Links come from <a href> elements, and respect the page's <base href>. Empty links, # links, javascript: and mailto: links are skipped. |
| NoFollow | Links with rel="nofollow" are skipped, unless Ignore NoFollow is True. |
| File types | Before fetching a page, the crawler asks the server for the content type with a HEAD request. HTML pages are fetched. PDF files are reported only when application/pdf is in File Types or Dynamic File Types. Other types are skipped. |
| Addresses | A trailing / is removed and query parameters are sorted by name. So https://www.example.com/docs/ and https://www.example.com/docs are one page, and so are page?b=2&a=1 and page?a=1&b=2. |
The crawler has no depth limit, page limit or delay between requests. It doesn't read robots.txt. Use Restrict to Paths, Exclude Paths and the job schedule to control how much is crawled.
Site Id
- When you set a Site Id that already exists, the site is updated with the new settings. All its queued pages and their crawl history are deleted, and the start pages are queued again. The whole site is crawled again from the start.
- When Site Id is empty, a new site is added each time the action runs, even for the same Base Url. Such sites can't be removed with Remove Site to Crawl, which needs the Site Id. Always set a Site Id.
Security
The crawler fetches any http or https address, including localhost, internal servers and private IP addresses. Redirects are followed automatically, also to other hosts. The raw HTML of each page is then available to your actions.
Don't build the Base Url from user input. If you must, check it first, for example with a condition that only allows your own domains.
Considerations
- Errors. The action fails when Base Url is empty (
There is no token defined for the BaseUrl), isn't a valid full address, or doesn't start withhttp(BaseUrl must be a well formed url starting with http or https.). It also fails when User Agent Header isn't a valid value, for example a browser string with spaces. - Run it once per site. Adding the site doesn't crawl it. Running the action again with the same Site Id restarts the crawl from the start, so don't run it on every crawl.
- Restrict to Paths replaces the start page. When it's set, the Base Url itself isn't queued. A restricted path that's also in Exclude Paths isn't queued.
- Reindex interval. The site's Min Page Reindex Interval always wins over the one of Crawl Next Batch. Leave it empty to use the Crawl Next Batch value.
- Automation jobs. This action can run anywhere, including a workflow or a scheduled job, since it doesn't need a web request.
Examples
To understand how to use the below examples, please see Running Examples.
1. Crawl the guides of a documentation site
This action adds the guides section of a site, skips its archive, and refetches pages every 12 hours. The internal site number is saved in the CrawlSiteId token.
{
"Title": "Add Site to Crawl",
"ActionType": "WebCrawler.AddSiteToCrawl",
"Description": "Queue the guides of the example site",
"Parameters": {
"BaseUrl": "https://www.example.com/docs/",
"UserAssignedSiteId": "ExampleDocs",
"MinPageReindexInterval": 12,
"FollowLinks": "1",
"IgnoreNoFollow": "0",
"RestrictToPaths": [
"guides"
],
"ExcludePaths": [
"guides/archive"
],
"UserAgentHeader": "ExampleCrawler/1.0",
"OutputSiteIdTokenName": "CrawlSiteId"
}
}
2. Crawl a whole site and report its PDF files
This action adds the site from the SiteUrl token, only when it's on the example.com domain. PDF files found on the site run the On Process Document actions of Crawl Next Batch.
{
"Title": "Add Site to Crawl",
"ActionType": "WebCrawler.AddSiteToCrawl",
"Description": "Queue the example site with its PDF files",
"Condition": "[SiteUrl].StartsWith(\"https://www.example.com/\")",
"Parameters": {
"BaseUrl": "[SiteUrl]",
"UserAssignedSiteId": "ExampleSite",
"FollowLinks": "1",
"IgnoreNoFollow": "0",
"FileTypes": [
{
"FileTypeId": "application/pdf"
}
]
}
}
Revised 09/28/2026