Snapshot · 21 Sept

AIFTIBot

If you found this address in your server logs, it was us. This page says what the crawler does and how to make it stop.

How to block it

Add this to your robots.txt. It is checked before every fetch, and a robots.txt we cannot retrieve is treated as a refusal, not as permission.

User-agent: AIFTIBot
Disallow: /

A Crawl-delay directive is honoured too, if you would rather slow it down than shut it out.

What it collects

What
Headlines and public posts that mention AI systems, to measure how the public discourse about them leans between fear and trust.
Mostly not by crawling
Most of the corpus comes from public APIs and RSS feeds — Hacker News, GDELT, Mastodon, Bluesky, publisher feeds — not from fetching pages.
Text kept
Titles and short excerpts, with the link back to you. Article bodies are not stored, and raw text ages out of the store on a fixed retention window.
Rate
A few requests per source per collection, three collections a day. Per-host minimum intervals are enforced in code, not by convention.
Never
Logins, paywalled content, private posts, or personal data.

What it is for

A free, non-commercial index of how frightened public discourse is of AI, published at this site. Every number is published with its sample size and confidence interval; the methodology describes how it is computed, including what it cannot tell you.