Skip to main content
Version: 1.28 (Current)

Add Site to Crawl

Audience: Low-code Engineers

Skill Prerequisites: Actions, Tokens, Automation Jobs

Adds a website to the web crawler's queue. The action saves the site and its crawl rules, and queues the start page. It doesn't download anything itself.

The pages are fetched later by Crawl Next Batch, usually from a scheduled job. It follows the links it finds and runs your actions for each page. To stop crawling a site, use Remove Site to Crawl. See How crawling works.

note

This action is part of the Web Crawler add-on (PlantAnApp.WebCrawler). The add-on is installed separately and needs the WEBCRWL feature in your license. If it isn't licensed, the action fails with a "not licensed" error. If you don't see the Web Crawler actions, the add-on isn't installed.

Typical Use Cases​

  • Crawl your documentation or knowledge base, so its pages can be stored, searched or summarized
  • Keep a copy of a site's pages up to date, since pages are fetched again after an interval
  • Find the PDF files published on a site
  • Collect the external links of a site, for example to check them

Don't use it to​

  • Get one page right now. Use Fetch Content.
  • Clean HTML you already have. Use Clean HTML.
  • Crawl internal or private addresses from user input. See Security.
Action NameDescription
Crawl Next BatchFetches the next queued pages and runs actions for each one.
Remove Site to CrawlRemoves a site and its queued pages.
Fetch ContentFetches a single page right away.
Clean HTMLKeeps only the main content of a page's HTML.
Summarize ContentSummarizes the text of a crawled page with AI.
Index RuleIndexes rows of a table in a search behavior, for example the pages you saved while crawling.

Input Parameter Reference​

ParameterDescriptionSupports TokensDefaultRequired
Base UrlThe address where crawling starts, for example https://www.example.com/docs/. It must be a full address that starts with http or https. Only pages on the same host and under this path are crawled. See What gets crawled.Yesempty stringYes
Site IdYour own ID for the site, for example ExampleDocs. It must be unique. It's needed to remove the site later, and adding a site with an existing Site Id updates it. See Site Id.Yesempty stringNo
Min Page Reindex IntervalHow many hours to wait before a page of this site is fetched again, for example 12. When set, it replaces the Min Page Reindex Interval of Crawl Next Batch for this site.YesemptyNo
Follow LinksWhen True, links found on each page are added to the queue. When False, only the start pages are crawled. Supports expressions.NoTrueNo
Ignore NoFollowOnly shown when Follow Links is True. When True, links with rel="nofollow" are followed too. When False, they're skipped. Supports expressions.NoFalseNo
Restrict to PathsOnly shown when Follow Links is True. Relative paths to limit the crawl to, for example guides or api/v2. When at least one is set, only pages under these paths are crawled, and these paths become the start pages instead of the Base Url. Each path must be a real page. A token can hold several paths, separated by commas or new lines.YesemptyNo
Exclude PathsOnly shown when Follow Links is True. Relative paths to skip, for example archive. Pages under these paths aren't crawled. A token can hold several paths, separated by commas or new lines.YesemptyNo
File TypesThe file types to report besides HTML pages. The only choice is application/pdf. Files of these types run the On Process Document actions of Crawl Next Batch.NoemptyNo
Dynamic File TypesMore file types, as MIME types separated by commas or new lines, for example [FileTypes]. Only application/pdf is recognized, so other types are skipped.Yesempty stringNo
User Agent HeaderThe User-Agent header sent with each request, for example ExampleCrawler/1.0. It must be a single name/version value. If it's empty, no User-Agent header is sent.Yesempty stringNo

Output Parameters Reference​

ParameterDescription
Output Site Id Token NameThe name of a token that receives the site's internal number, for example CrawlSiteId. It's the same value as [Crawler:SiteId] in Crawl Next Batch, not your Site Id.

How crawling works​

Crawling a site takes three actions. Each one runs on its own, usually at different times.

  1. Add Site to Crawl saves the site and its rules in the paa.WebCrawler_Site table. It queues the start page in the paa.WebCrawler_Page table. The start page is the Base Url, or each path of Restrict to Paths when it's set.
  2. Crawl Next Batch takes the pages that are due, fetches them, and runs your actions for each one. When Follow Links is True, the links it finds are added to the queue. Each fetched page is scheduled to be fetched again after the reindex interval. You normally run it from a scheduled automation job, for example every minute.
  3. Remove Site to Crawl deletes the site and all its queued pages.

A few things to know:

  • Each page is stored once per site. When a link is found again, it isn't queued a second time.
  • Pages are never removed from the queue while the site exists. A page that disappears keeps being fetched, and runs the On Page NotFound actions each time.
  • The crawler only stores the queue and the crawl status. The page content isn't saved. Save what you need in the actions of Crawl Next Batch, for example with Run SQL Query.
  • Crawling isn't tied to a portal or module. Crawl Next Batch crawls the queued pages of every site that was added, wherever it was added from.

What gets crawled​

RuleHow it works
HostOnly links to the same host as the Base Url are queued. www. is ignored, so example.com and www.example.com are the same host. The scheme and port aren't compared.
PathOnly links whose path starts with the Base Url path are queued. With a Base Url of https://www.example.com/docs, the page https://www.example.com/blog isn't crawled.
Restrict to Paths and Exclude PathsThe path after the Base Url must start with one of the Restrict to Paths, and mustn't start with any of the Exclude Paths. The check is a case-sensitive "starts with", so guide also matches guides and guide-old.
LinksLinks come from <a href> elements, and respect the page's <base href>. Empty links, # links, javascript: and mailto: links are skipped.
NoFollowLinks with rel="nofollow" are skipped, unless Ignore NoFollow is True.
File typesBefore fetching a page, the crawler asks the server for the content type with a HEAD request. HTML pages are fetched. PDF files are reported only when application/pdf is in File Types or Dynamic File Types. Other types are skipped.
AddressesA trailing / is removed and query parameters are sorted by name. So https://www.example.com/docs/ and https://www.example.com/docs are one page, and so are page?b=2&a=1 and page?a=1&b=2.

The crawler has no depth limit, page limit or delay between requests. It doesn't read robots.txt. Use Restrict to Paths, Exclude Paths and the job schedule to control how much is crawled.

Site Id​

  • When you set a Site Id that already exists, the site is updated with the new settings. All its queued pages and their crawl history are deleted, and the start pages are queued again. The whole site is crawled again from the start.
  • When Site Id is empty, a new site is added each time the action runs, even for the same Base Url. Such sites can't be removed with Remove Site to Crawl, which needs the Site Id. Always set a Site Id.

Security​

caution

The crawler fetches any http or https address, including localhost, internal servers and private IP addresses. Redirects are followed automatically, also to other hosts. The raw HTML of each page is then available to your actions.

Don't build the Base Url from user input. If you must, check it first, for example with a condition that only allows your own domains.

Considerations​

  • Errors. The action fails when Base Url is empty (There is no token defined for the BaseUrl), isn't a valid full address, or doesn't start with http (BaseUrl must be a well formed url starting with http or https.). It also fails when User Agent Header isn't a valid value, for example a browser string with spaces.
  • Run it once per site. Adding the site doesn't crawl it. Running the action again with the same Site Id restarts the crawl from the start, so don't run it on every crawl.
  • Restrict to Paths replaces the start page. When it's set, the Base Url itself isn't queued. A restricted path that's also in Exclude Paths isn't queued.
  • Reindex interval. The site's Min Page Reindex Interval always wins over the one of Crawl Next Batch. Leave it empty to use the Crawl Next Batch value.
  • Automation jobs. This action can run anywhere, including a workflow or a scheduled job, since it doesn't need a web request.

Examples​

tip

To understand how to use the below examples, please see Running Examples.

1. Crawl the guides of a documentation site​

This action adds the guides section of a site, skips its archive, and refetches pages every 12 hours. The internal site number is saved in the CrawlSiteId token.

{
"Title": "Add Site to Crawl",
"ActionType": "WebCrawler.AddSiteToCrawl",
"Description": "Queue the guides of the example site",
"Parameters": {
"BaseUrl": "https://www.example.com/docs/",
"UserAssignedSiteId": "ExampleDocs",
"MinPageReindexInterval": 12,
"FollowLinks": "1",
"IgnoreNoFollow": "0",
"RestrictToPaths": [
"guides"
],
"ExcludePaths": [
"guides/archive"
],
"UserAgentHeader": "ExampleCrawler/1.0",
"OutputSiteIdTokenName": "CrawlSiteId"
}
}

2. Crawl a whole site and report its PDF files​

This action adds the site from the SiteUrl token, only when it's on the example.com domain. PDF files found on the site run the On Process Document actions of Crawl Next Batch.

{
"Title": "Add Site to Crawl",
"ActionType": "WebCrawler.AddSiteToCrawl",
"Description": "Queue the example site with its PDF files",
"Condition": "[SiteUrl].StartsWith(\"https://www.example.com/\")",
"Parameters": {
"BaseUrl": "[SiteUrl]",
"UserAssignedSiteId": "ExampleSite",
"FollowLinks": "1",
"IgnoreNoFollow": "0",
"FileTypes": [
{
"FileTypeId": "application/pdf"
}
]
}
}

Revised 09/28/2026