Crawler directory
Every crawler, and what its operator says it is for
One page per crawler, fetcher and agent, built only from the operator's own published documentation. Each entry carries the exact robots.txt token, the user-agent strings, the directive to copy, and how to verify a visit really came from who it claims.
Browse by category
Each category is its own page, so no single list has to carry every agent.
No bot categories are published yet. Everything in this family is held back until it clears the depth gate — sourced, dated, and reviewed by a person.
What the fetch is for
Almost every agent here is doing one of four things, and which one decides what allowing or blocking it costs you.
No bots are published yet. Everything in this family is held back until it clears the depth gate — sourced, dated, and reviewed by a person.
How this directory is built
An agent gets a page here only when its operator publishes documentation naming it and saying what it does. Every user-agent string, robots.txt token and verification method on these pages is transcribed from that documentation, with the URL and the date it was read printed at the bottom of the page. Tokens that circulate widely on third-party bot lists with no traceable operator page are deliberately absent: naming one in your robots.txt costs nothing, but nobody can honestly tell you what the rule achieves, so we do not.
The reasoning behind the four-way classification, and the case for and against each configuration, is in the AI crawler guide. These pages are its reference section.
