Skip to main content
Version: 1.28 (Current)

Web Crawler

Audience: Low-code Engineers

Skill Prerequisites: Actions, Tokens, Automation Jobs

The Web Crawler add-on reads web pages from your server. Its actions can fetch a single page now, or crawl a whole site over time by following its links and running your actions for each page. They also extract the main content of a page, without the menus, headers, footers and scripts, so you can store, search or summarize it.

note

This add-on is the PlantAnApp.WebCrawler package. It's installed separately and needs the WEBCRWL feature package in your license. If it isn't licensed, the actions fail with a "not licensed" error. If you don't see these actions, the add-on isn't installed.

Choosing an action​

ActionWhat it doesUse it to
Fetch ContentDownloads one page right away and saves its title, cleaned HTML, plain text and last modified date in tokens.Get the text of an article to summarize or save.
Clean HTMLKeeps only the main content of HTML you already have, and removes navigation, scripts, styles, images and most attributes.Shrink a crawled page or a Server Request response before you save it or send it to AI.
Add Site to CrawlSaves a site and its crawl rules, and queues its start page. It doesn't download anything itself.Start crawling a documentation site or knowledge base.
Crawl Next BatchFetches the next pages that are due, queues the links it finds, and runs your actions for each page, file, error and external link.Process crawled pages from a scheduled automation job.
Remove Site to CrawlDeletes a site and all its queued pages and crawl history.Stop crawling a site that's no longer needed.

One page or a whole site?​

  • One page, now. Use Fetch Content. It sends a single GET request and returns cleaned content. Scripts on the page don't run, so you only get the HTML the server sends.
  • A whole site, over time. Use the three crawl actions together. The crawler follows links on the same host and under the start path, and fetches each page again after a reindex interval, so your copy stays up to date.
  • HTML you already have. Use Clean HTML. It's the same cleanup Fetch Content does, and works on [Crawler:RawHtml] in Crawl Next Batch too.

How the crawl actions work together​

  1. Add Site to Crawl runs once per site. Always give the site a Site Id. You need it to remove the site, and running the action again with the same Site Id restarts the crawl from the start.
  2. Crawl Next Batch runs from a scheduled automation job, for example every minute. Each run fetches at most one page per site, so the job schedule decides how fast a site is crawled. It crawls every site that was added, whichever portal or module added it.
  3. Remove Site to Crawl stops crawling a site, by its Site Id. To pause a crawl instead, disable the job.

The crawler only keeps the queue and the crawl status, not the pages. Save what you need in the On Process Page actions of Crawl Next Batch, for example with Run SQL Query, and use an insert-or-update query, since each page comes back after the reindex interval. You can then index the saved pages with Index Rule.

Security​

  • The requests come from your web server. Fetch Content and the crawler fetch any http or https address the server can reach, including localhost and internal addresses, and they follow redirects to other hosts. Don't build a URL from user input without checking it, for example with a condition that only allows your own domains.
  • The result isn't safe HTML. Fetch Content and Clean HTML remove scripts and event handlers, but keep javascript: and data: links. Run Sanitize Html before you show the content on a page.
  • The crawler has no limits of its own. There's no depth limit, page limit or delay between requests, and it doesn't read robots.txt. Control it with Restrict to Paths, Exclude Paths and the job schedule.

To search the pages you save, see the Search actions. To summarize them with AI, see the OpenAI add-on.

Revised 10/02/2026