The Slow Shelf serves public domain literature as plain, readable web pages. No trackers, no analytics scripts, no advertising, no JavaScript. The texts come from Project Gutenberg and remain in the public domain.
It is also a measurement site. It exists to study how automated clients — search crawlers, AI training crawlers, and retrieval agents — behave when they encounter an ordinary text archive.
In 2026, automated traffic overtook human traffic on the web for the first time. Most published figures come from a small number of infrastructure providers reporting on their own networks, in aggregate. There is very little public data describing what automated clients actually do at the level of individual requests: how deep they crawl, how many connections they open in parallel, whether they reuse those connections, how often they return, and how their arrival patterns differ from human ones.
Those parameters matter, because web infrastructure was designed around human behaviour. Understanding how machine clients diverge from it is a prerequisite to building systems that serve both well.
This site records standard web server access logs. For each request:
No cookies are set. No client-side scripts run. There is nothing here that tracks a visitor across sites, because there is no mechanism on this site capable of doing so.
IP addresses are used for one purpose: verifying that clients claiming to be a particular crawler actually originate from that operator's published address ranges. Any dataset published from this work will have addresses removed or aggregated to network level before release.
Analysis is concerned with aggregate behaviour — distributions, rates, and patterns across client classes. Individual visitors are not of interest and are not profiled.
This site permits all crawlers, including AI training and retrieval crawlers. See robots.txt. This is deliberate: a site that blocks crawlers cannot observe them.
There is no rate limiting. Crawl at whatever rate your systems normally use — the natural rate is the measurement.
If you operate a crawler and would prefer your traffic excluded from the analysis, or if you have questions about the methodology, get in touch and it will be honoured.
Correspondence about this project is welcome, particularly from other people working on web traffic measurement.
Findings will be published openly, including the methodology and, where it can be done responsibly, the underlying data.