Our crawler
We read published guidance automatically so we can tell you when something changes. This page explains exactly what our crawler does, so site owners can recognise it and control it.
How to identify us
Our crawler identifies itself in the user agent string as
SupportPortalBot/1.0 and links back to this page.
We do not disguise ourselves as a browser.
What we read, and why
We read official guidance pages about benefits, grants and support schemes, so that when a rate or a rule changes we can flag the pages on our site that depend on it. We keep a copy of the pages we rely on as evidence for what we published and when.
We do not crawl to republish your content. We quote briefly and always link back.
How we behave
- We respect
robots.txt, including any rules written specifically for our user agent. - We request pages slowly, one at a time per site, with a pause between requests.
- We honour caching: if a page has not changed since our last visit, we take the server's word for it and download nothing.
- We back off when a site is slow or returns errors, and stop entirely if asked to.
- We only request pages over the public internet. We do not attempt to reach private addresses or internal services.
- We do not submit forms, log in, or follow links that change anything.
How often
It depends on the page. Guidance that changes with the tax year is checked infrequently; a page announcing a scheme that opens and closes is checked more often. In practice most pages are visited no more than once a day, and usually much less.
If you want us to stop
Add a rule for our user agent to your robots.txt and we will stop on
our next visit. If you would like us to stop immediately, or you think we are
causing a problem, contact us and we
will remove your site from our list.
We would rather hear from you than be blocked at the firewall — if our crawler is behaving badly, that is a bug we want to fix.