44up.io · Crawler
StorePulse in your logs?
StorePulse is the measuring instrument of the 44up Observatory. It observes the publicly reachable websites of organic food delivery services in Germany, Austria and Switzerland: whether they are reachable, how accessible they are, how well search engines and AI assistants can read them, and how their range and prices develop.
How the crawler identifies itself
Mozilla/5.0 (compatible; StorePulse-<tool>/1.0; +https://44up.io)python-requests/<version> StorePulse-<tool>/1.0 (+https://44up.io) <tool> names the measurement, for example ClosureWatch, A11y, AIReadiness or Fetch. We do not present ourselves as a browser or as another vendor’s crawler.
Every request we send to a provider’s website or to a public directory carries this token, with two exceptions: if a page only builds its content with JavaScript, we render it in an automated Chromium browser, which identifies as a browser (HeadlessChrome). Page speed is measured for us by Google PageSpeed Insights; those requests come from Google. If a shop offers a prerendered version of its pages, we read that instead of rendering (parameter ?_escaped_fragment_=), under our own token.
What we fetch
| When (UTC) | What |
|---|---|
| daily 05:40 | Your robots.txt and the homepage; if the shop lives at a separate address, that address next. If the homepage does not answer, one attempt at the www variant or over http. No subpages. |
| daily 11:40 and 17:40 | Only websites whose morning check was inconclusive (timeout, error page): one re-check each. |
| Mondays 02:30 | robots.txt, sitemap.xml and the child sitemaps it links to, to detect new and removed pages. If we find no sitemap, the homepage once. |
| 1st of the month from 03:00 | Individual public pages: homepage, imprint or contact page, terms and conditions, FAQ, jobs or careers page, delivery area, accessibility statement, shop and product pages for prices (using the shop’s public product search where available), up to eight further subpages linked from the homepage, robots.txt, sitemap.xml, llms.txt. We look for some of these pages at common addresses (e.g. /karriere), so the odd 404 is expected. For a few providers, the public newsletter archive. Plus a Google PageSpeed Insights measurement of the homepage (mobile and desktop). |
| 1 January, April, July, October 04:00 | Homepage and up to six payment and shipping pages (e.g. /zahlungsarten, /agb). Also DNS and TLS data for your domain; those are not page requests. |
Requests come from a server in a data centre in Frankfurt am Main, Germany.
What we do not do
- No login, no shopping cart, no orders, no newsletter sign-ups.
- If an address answers HTTP 429, we ask it at most once more, never in a retry loop.
- We do not present ourselves as Googlebot, GPTBot or any other third-party crawler.
- We respect your robots.txt. A group
User-agent: StorePulseapplies to all our measurements; without one,User-agent: *applies. This includes the page-speed measurement: if your robots.txt disallows the fetch, we do not ask Google PageSpeed Insights.
Opt your website out
Two ways, neither needs an account:
User-agent: StorePulsewithDisallow: /in your robots.txt. Takes effect from the next run. From then on we only fetch your robots.txt.- An email to crawler@44up.io with your website’s address, sent from an address under the same domain. We confirm by reply and remove the website from all measurements, including fetching its robots.txt and PageSpeed Insights. Values already collected are no longer shown for your website; they remain only in anonymous market figures.