Skip to main content
Version: 1.28 (Current)

Fetch Content

Audience: Low-code Engineers

Skill Prerequisites: Actions, Tokens, HTML

Downloads a web page and saves its main content in tokens. The page title, a simplified HTML version, a plain text version and the last modified date each go into their own token.

The action sends a GET request from the web server and removes things like headers, footers, navigation, scripts and images. It works the same way as Clean HTML. See How the content is extracted.

note

This action is part of the Web Crawler add-on (PlantAnApp.WebCrawler). The add-on is installed separately and needs the WEBCRWL feature in your license. If it isn't licensed, the action fails with a not-licensed error. If you don't see the Web Crawler actions, the add-on isn't installed.

Typical Use Cases​

  • Get the text of an article so you can summarize it with Summarize Content or send it to Chat
  • Save the title and content of a page someone shared, for example in a bookmarks or research app
  • Check when a page was last changed before you process it again

Don't use it to​

  • Call an API or get the raw response. Use Server Request instead.
  • Crawl a whole site. Use Add Site to Crawl and Crawl Next Batch instead.
  • Download files such as PDFs or images. The response is always read as text and parsed as HTML.
  • Get pages that are built with JavaScript. Scripts don't run, so you only get the HTML the server sends.
  • Make HTML safe to display. Use Sanitize Html on the result instead.
Action NameDescription
Clean HTMLExtracts the main content from HTML you already have.
Sanitize HtmlRemoves unsafe markup. Use it before you display the fetched HTML.
Server RequestSends any HTTP request and returns the raw response.
Summarize ContentSummarizes text with OpenAI.
ChatSends a message to an OpenAI chat model.
Add Site to CrawlAdds a site to the list of sites to crawl.
Crawl Next BatchCrawls the next batch of pages.
Remove Site to CrawlRemoves a site from the list of sites to crawl.

Input Parameter Reference​

ParameterDescriptionSupports TokensDefaultRequired
UrlThe full address of the page, for example https://www.example.com/blog/post. Only http and https URLs work.Yesempty stringYes
User AgentThe browser the request pretends to be: Win10Chrome116, Win10Chrome117, MacOsChrome116, MacOsChrome117 or GoogleBot. In expression mode, you can type your own user agent string. See User agents.Yes, in expression modeWin10Chrome116No
Store TitleThe token name that gets the page title, for example PageTitle. See Page title.Noempty stringNo
Store HTMLThe token name that gets the simplified HTML of the main content.Noempty stringNo
Store Plain TextThe token name that gets the text of the main content, without tags.Noempty stringNo
Store Last ModifiedThe token name that gets the date the page was last changed. It comes from the Last-Modified response header. If the server doesn't send it, the Date header is used, which is usually the current time.Noempty stringNo
On ErrorActions to run when the page can't be fetched. See Error handling.NoEmptyNo
Ignore ErrorsContinues with the next actions when the page can't be fetched. On Error still runs.NofalseNo

All the Store parameters are optional. Only the tokens you name are saved.

Output Parameters Reference​

ParameterDescription
Store TitleHolds the page title.
Store HTMLHolds the simplified HTML of the main content.
Store Plain TextHolds the plain text of the main content.
Store Last ModifiedHolds the last modified date.
[StatusCode]Set only when the action fails. It's the HTTP status code, for example 404, or 0 if no response was received.

User agents​

ValueUser agent sent
Win10Chrome116Chrome 116 on Windows 10
Win10Chrome117Chrome 117 on Windows 10
MacOsChrome116Chrome 116 on macOS
MacOsChrome117Chrome 117 on macOS
GoogleBotGoogle's crawler (Googlebot 2.1)

Some sites return different content, or block the request, depending on the user agent. Try another value if a page comes back empty or blocked.

Page title​

The title is taken from the first of these that exists:

  1. The <title> tag.
  2. The first <h1>, then the first <h2>, and so on down to <h6>.
  3. The path of the URL, for example /blog/post.

HTML entities in the title aren't decoded. A title like Fish & Chips is saved as Fish &amp; Chips.

How the content is extracted​

The HTML is simplified the same way as Clean HTML:

  • The main part of the page is used: the <article> if there's exactly one, otherwise the <main> if there's exactly one, otherwise the <body>.
  • nav, header, footer, script, style, link and img tags are removed with their content.
  • All attributes are removed except href, title and alt. Links stay as they are in the page. Relative links aren't turned into full URLs.
  • Empty tags are removed.
  • A div or span with only one child is unwrapped. Its child is moved to the end of the parent, so the order of the content can change. The plain text is affected too.

The plain text is built from the same content:

  • A line break is added after p, div, li and h1 to h6 tags. A space is added after a, td and button tags.
  • Line breaks and tabs inside the text are removed, except in <pre>. Words that are only separated by a line break in the HTML end up joined.
  • &nbsp; is removed and repeated spaces become one. Other HTML entities are decoded.

The response is read as text using the character set in the Content-Type header, or UTF-8 if there isn't one. A <meta charset> tag in the page isn't used.

Error handling​

The action fails when:

  • The URL isn't a full http or https URL. The message is URL ... is not valid.
  • The server can't be reached, the SSL certificate isn't valid, or there's no response within 100 seconds.
  • The status code isn't 200. Other success codes, such as 204, fail too.
  • The page contains Google reCAPTCHA. The message ends with Captcha required. and [StatusCode] is 200.

Then:

  1. [StatusCode] is set, together with [Exception], [ExceptionType], [ExceptionMessage] and [ExceptionStack]. [ExceptionMessage] ends with Fetch Content URL: and the URL.
  2. The On Error actions run. If one of them ends execution, for example Display Error Message, that's the result.
  3. Otherwise, if Ignore Errors is checked, the next actions run. If not, the action fails. Administrators see the error. Other users see Error fetching content from followed by the URL.

Considerations​

  • The request comes from your web server. Any URL the server can reach is fetched, including localhost, internal network addresses and cloud metadata addresses. Redirects are followed automatically. Don't pass a URL typed by a user without checking it first, for example with a condition that only allows the domains you expect.
  • The HTML isn't sanitized. javascript: and data: links are kept, and so are tags like iframe or form when they still have content. Run Sanitize Html before you show the result on a page.
  • Error details can include the page. When the status code isn't 200, the error message includes the response body, and it's in [ExceptionMessage]. Don't show [ExceptionMessage] to end users.
  • Error tokens stay. [StatusCode] and the exception tokens remain after the action. If you already use a token named StatusCode, it's overwritten when the action fails.
  • Compared to Server Request. Server Request supports other methods, headers, authentication, cookies and a timeout setting, and returns the raw response. Fetch Content only sends a GET with a user agent, and returns cleaned content.
  • Large pages are downloaded in full before they're processed. There's no size limit you can set.

Examples​

tip

To understand how to use the below examples, please see Running Examples.

1. Get the text of an article​

This action fetches a page and saves its title, simplified HTML and plain text. You can then pass [ArticleText] to Summarize Content or Chat.

{
"Title": "Fetch Content",
"ActionType": "WebCrawler.FetchContent",
"Description": "Get the article content",
"Parameters": {
"Url": "https://www.example.com/blog/post",
"UserAgent": {
"Expression": "",
"Value": "Win10Chrome116",
"IsExpression": false,
"Parameters": {}
},
"StoreTitle": "ArticleTitle",
"StoreHtml": "ArticleHtml",
"StorePlainText": "ArticleText",
"StoreLastModified": "ArticleModified",
"IgnoreErrors": false
}
}

2. Log failed pages and continue​

This action fetches the page in [PageUrl], but only when it's filled in. If the page can't be fetched, On Error logs the status code and the error, and Ignore Errors lets the next actions run.

{
"Title": "Fetch Content",
"ActionType": "WebCrawler.FetchContent",
"Description": "Fetch the page and log failures",
"Condition": "[PageUrl] != \"\"",
"Parameters": {
"Url": "[PageUrl]",
"StoreTitle": "PageTitle",
"StorePlainText": "PageText",
"IgnoreErrors": true,
"OnError": [
{
"Title": "Log Error",
"ActionType": "LogError",
"Parameters": {
"Message": "Fetch Content failed with status [StatusCode]: [ExceptionMessage]"
}
}
]
}
}

Revised 09/28/2026