About

The AI Crawler Honeypot is an automated observatory measuring how AI search engines, training scrapers, and autonomous agents interact with web infrastructure. It collects request telemetry, verifies claimed identity against published network ranges, exposes tools that agents can call, and publishes open, transparent access logs.

Zero Raw IP Retention

Privacy is architected into the system by default. Real client IP addresses are never written to disk or long-term databases. When a request is received, the IP is hashed using SHA-256 combined with a daily rotating salt (salt:{YYYY-MM-DD}). The salt expires automatically after 24 hours and is purged daily at 00:05 UTC, making irreversible reconstruction mathematically impossible.

Vendor CIDR Verification

When a request claims a user-agent string associated with known AI crawlers (such as OpenAI, Anthropic, Google, or Perplexity), the connecting IP is checked against official published CIDR blocks and reverse DNS records. Requests are categorized as:

Robots.txt Conformance & Telemetry

The honeypot passively tracks whether agents respect robots.txt directives without dropping or resetting connections. Requests to disallowed paths are logged with the matching rule. Observational data is classified for cadence and sequential anomalies, giving researchers clear insight into automated crawler behavior.

Tool Registry & Agent Interaction

The honeypot exposes a registry of tools that AI agents can discover and call. Agents that read /tools.json can invoke tools via POST /tools/call, including:

Every tool call is logged with the same schema as request entries. The tool-call log is public at /tools/log.

Public Logs & Data Access

All activity is made publicly accessible across several views:

About

A public AI crawler observation honeypot. Factual observation only — no judgments, no editorialising, no claims about whether any agent is "safe" or "unsafe." Every entry is a timestamped field/value record.

Contact: vic@fixseo.uk