Web Crawler
Audience:
Low-code EngineersSkill Prerequisites:
Actions,Tokens,Automation Jobs
The Web Crawler add-on reads web pages from your server. Its actions can fetch a single page now, or crawl a whole site over time by following its links and running your actions for each page. They also extract the main content of a page, without the menus, headers, footers and scripts, so you can store, search or summarize it.
This add-on is the PlantAnApp.WebCrawler package. It's installed separately and needs the WEBCRWL feature package in your license. If it isn't licensed, the actions fail with a "not licensed" error. If you don't see these actions, the add-on isn't installed.
Choosing an action
| Action | What it does | Use it to |
|---|---|---|
| Fetch Content | Downloads one page right away and saves its title, cleaned HTML, plain text and last modified date in tokens. | Get the text of an article to summarize or save. |
| Clean HTML | Keeps only the main content of HTML you already have, and removes navigation, scripts, styles, images and most attributes. | Shrink a crawled page or a Server Request response before you save it or send it to AI. |
| Add Site to Crawl | Saves a site and its crawl rules, and queues its start page. It doesn't download anything itself. | Start crawling a documentation site or knowledge base. |
| Crawl Next Batch | Fetches the next pages that are due, queues the links it finds, and runs your actions for each page, file, error and external link. | Process crawled pages from a scheduled automation job. |
| Remove Site to Crawl | Deletes a site and all its queued pages and crawl history. | Stop crawling a site that's no longer needed. |
One page or a whole site?
- One page, now. Use Fetch Content. It sends a single
GETrequest and returns cleaned content. Scripts on the page don't run, so you only get the HTML the server sends. - A whole site, over time. Use the three crawl actions together. The crawler follows links on the same host and under the start path, and fetches each page again after a reindex interval, so your copy stays up to date.
- HTML you already have. Use Clean HTML. It's the same cleanup Fetch Content does, and works on
[Crawler:RawHtml]in Crawl Next Batch too.
How the crawl actions work together
- Add Site to Crawl runs once per site. Always give the site a Site Id. You need it to remove the site, and running the action again with the same Site Id restarts the crawl from the start.
- Crawl Next Batch runs from a scheduled automation job, for example every minute. Each run fetches at most one page per site, so the job schedule decides how fast a site is crawled. It crawls every site that was added, whichever portal or module added it.
- Remove Site to Crawl stops crawling a site, by its Site Id. To pause a crawl instead, disable the job.
The crawler only keeps the queue and the crawl status, not the pages. Save what you need in the On Process Page actions of Crawl Next Batch, for example with Run SQL Query, and use an insert-or-update query, since each page comes back after the reindex interval. You can then index the saved pages with Index Rule.
Security
- The requests come from your web server. Fetch Content and the crawler fetch any
httporhttpsaddress the server can reach, includinglocalhostand internal addresses, and they follow redirects to other hosts. Don't build a URL from user input without checking it, for example with a condition that only allows your own domains. - The result isn't safe HTML. Fetch Content and Clean HTML remove scripts and event handlers, but keep
javascript:anddata:links. Run Sanitize Html before you show the content on a page. - The crawler has no limits of its own. There's no depth limit, page limit or delay between requests, and it doesn't read
robots.txt. Control it with Restrict to Paths, Exclude Paths and the job schedule.
To search the pages you save, see the Search actions. To summarize them with AI, see the OpenAI add-on.
Revised 10/02/2026