Fetch Content
Audience:
Low-code EngineersSkill Prerequisites:
Actions,Tokens,HTML
Downloads a web page and saves its main content in tokens. The page title, a simplified HTML version, a plain text version and the last modified date each go into their own token.
The action sends a GET request from the web server and removes things like headers, footers, navigation, scripts and images. It works the same way as Clean HTML. See How the content is extracted.
This action is part of the Web Crawler add-on (PlantAnApp.WebCrawler). The add-on is installed separately and needs the WEBCRWL feature in your license. If it isn't licensed, the action fails with a not-licensed error. If you don't see the Web Crawler actions, the add-on isn't installed.
Typical Use Cases
- Get the text of an article so you can summarize it with Summarize Content or send it to Chat
- Save the title and content of a page someone shared, for example in a bookmarks or research app
- Check when a page was last changed before you process it again
Don't use it to
- Call an API or get the raw response. Use Server Request instead.
- Crawl a whole site. Use Add Site to Crawl and Crawl Next Batch instead.
- Download files such as PDFs or images. The response is always read as text and parsed as HTML.
- Get pages that are built with JavaScript. Scripts don't run, so you only get the HTML the server sends.
- Make HTML safe to display. Use Sanitize Html on the result instead.
Related Actions
| Action Name | Description |
|---|---|
| Clean HTML | Extracts the main content from HTML you already have. |
| Sanitize Html | Removes unsafe markup. Use it before you display the fetched HTML. |
| Server Request | Sends any HTTP request and returns the raw response. |
| Summarize Content | Summarizes text with OpenAI. |
| Chat | Sends a message to an OpenAI chat model. |
| Add Site to Crawl | Adds a site to the list of sites to crawl. |
| Crawl Next Batch | Crawls the next batch of pages. |
| Remove Site to Crawl | Removes a site from the list of sites to crawl. |
Input Parameter Reference
| Parameter | Description | Supports Tokens | Default | Required |
|---|---|---|---|---|
| Url | The full address of the page, for example https://www.example.com/blog/post. Only http and https URLs work. | Yes | empty string | Yes |
| User Agent | The browser the request pretends to be: Win10Chrome116, Win10Chrome117, MacOsChrome116, MacOsChrome117 or GoogleBot. In expression mode, you can type your own user agent string. See User agents. | Yes, in expression mode | Win10Chrome116 | No |
| Store Title | The token name that gets the page title, for example PageTitle. See Page title. | No | empty string | No |
| Store HTML | The token name that gets the simplified HTML of the main content. | No | empty string | No |
| Store Plain Text | The token name that gets the text of the main content, without tags. | No | empty string | No |
| Store Last Modified | The token name that gets the date the page was last changed. It comes from the Last-Modified response header. If the server doesn't send it, the Date header is used, which is usually the current time. | No | empty string | No |
| On Error | Actions to run when the page can't be fetched. See Error handling. | No | Empty | No |
| Ignore Errors | Continues with the next actions when the page can't be fetched. On Error still runs. | No | false | No |
All the Store parameters are optional. Only the tokens you name are saved.
Output Parameters Reference
| Parameter | Description |
|---|---|
| Store Title | Holds the page title. |
| Store HTML | Holds the simplified HTML of the main content. |
| Store Plain Text | Holds the plain text of the main content. |
| Store Last Modified | Holds the last modified date. |
[StatusCode] | Set only when the action fails. It's the HTTP status code, for example 404, or 0 if no response was received. |
User agents
| Value | User agent sent |
|---|---|
Win10Chrome116 | Chrome 116 on Windows 10 |
Win10Chrome117 | Chrome 117 on Windows 10 |
MacOsChrome116 | Chrome 116 on macOS |
MacOsChrome117 | Chrome 117 on macOS |
GoogleBot | Google's crawler (Googlebot 2.1) |
Some sites return different content, or block the request, depending on the user agent. Try another value if a page comes back empty or blocked.
Page title
The title is taken from the first of these that exists:
- The
<title>tag. - The first
<h1>, then the first<h2>, and so on down to<h6>. - The path of the URL, for example
/blog/post.
HTML entities in the title aren't decoded. A title like Fish & Chips is saved as Fish & Chips.
How the content is extracted
The HTML is simplified the same way as Clean HTML:
- The main part of the page is used: the
<article>if there's exactly one, otherwise the<main>if there's exactly one, otherwise the<body>. nav,header,footer,script,style,linkandimgtags are removed with their content.- All attributes are removed except
href,titleandalt. Links stay as they are in the page. Relative links aren't turned into full URLs. - Empty tags are removed.
- A
divorspanwith only one child is unwrapped. Its child is moved to the end of the parent, so the order of the content can change. The plain text is affected too.
The plain text is built from the same content:
- A line break is added after
p,div,liandh1toh6tags. A space is added aftera,tdandbuttontags. - Line breaks and tabs inside the text are removed, except in
<pre>. Words that are only separated by a line break in the HTML end up joined. is removed and repeated spaces become one. Other HTML entities are decoded.
The response is read as text using the character set in the Content-Type header, or UTF-8 if there isn't one. A <meta charset> tag in the page isn't used.
Error handling
The action fails when:
- The URL isn't a full
httporhttpsURL. The message isURL ... is not valid. - The server can't be reached, the SSL certificate isn't valid, or there's no response within 100 seconds.
- The status code isn't
200. Other success codes, such as204, fail too. - The page contains Google reCAPTCHA. The message ends with
Captcha required.and[StatusCode]is200.
Then:
[StatusCode]is set, together with[Exception],[ExceptionType],[ExceptionMessage]and[ExceptionStack].[ExceptionMessage]ends withFetch Content URL:and the URL.- The
On Erroractions run. If one of them ends execution, for example Display Error Message, that's the result. - Otherwise, if
Ignore Errorsis checked, the next actions run. If not, the action fails. Administrators see the error. Other users seeError fetching content fromfollowed by the URL.
Considerations
- The request comes from your web server. Any URL the server can reach is fetched, including
localhost, internal network addresses and cloud metadata addresses. Redirects are followed automatically. Don't pass a URL typed by a user without checking it first, for example with a condition that only allows the domains you expect. - The HTML isn't sanitized.
javascript:anddata:links are kept, and so are tags likeiframeorformwhen they still have content. Run Sanitize Html before you show the result on a page. - Error details can include the page. When the status code isn't
200, the error message includes the response body, and it's in[ExceptionMessage]. Don't show[ExceptionMessage]to end users. - Error tokens stay.
[StatusCode]and the exception tokens remain after the action. If you already use a token namedStatusCode, it's overwritten when the action fails. - Compared to Server Request. Server Request supports other methods, headers, authentication, cookies and a timeout setting, and returns the raw response. Fetch Content only sends a
GETwith a user agent, and returns cleaned content. - Large pages are downloaded in full before they're processed. There's no size limit you can set.
Examples
To understand how to use the below examples, please see Running Examples.
1. Get the text of an article
This action fetches a page and saves its title, simplified HTML and plain text. You can then pass [ArticleText] to Summarize Content or Chat.
{
"Title": "Fetch Content",
"ActionType": "WebCrawler.FetchContent",
"Description": "Get the article content",
"Parameters": {
"Url": "https://www.example.com/blog/post",
"UserAgent": {
"Expression": "",
"Value": "Win10Chrome116",
"IsExpression": false,
"Parameters": {}
},
"StoreTitle": "ArticleTitle",
"StoreHtml": "ArticleHtml",
"StorePlainText": "ArticleText",
"StoreLastModified": "ArticleModified",
"IgnoreErrors": false
}
}
2. Log failed pages and continue
This action fetches the page in [PageUrl], but only when it's filled in. If the page can't be fetched, On Error logs the status code and the error, and Ignore Errors lets the next actions run.
{
"Title": "Fetch Content",
"ActionType": "WebCrawler.FetchContent",
"Description": "Fetch the page and log failures",
"Condition": "[PageUrl] != \"\"",
"Parameters": {
"Url": "[PageUrl]",
"StoreTitle": "PageTitle",
"StorePlainText": "PageText",
"IgnoreErrors": true,
"OnError": [
{
"Title": "Log Error",
"ActionType": "LogError",
"Parameters": {
"Message": "Fetch Content failed with status [StatusCode]: [ExceptionMessage]"
}
}
]
}
}
Revised 09/28/2026