This page documents how the AI Crawler Honeypot works. It describes the verification pipeline, the logging schema, the privacy architecture, and the public record.
A public AI crawler observation honeypot. It logs HTTP requests to this domain, verifies claimed identity against published vendor infrastructure, and publishes a public record.
Real client IP addresses are never written to disk or long-term databases.
When a request is received:
/24 (IPv4) or /48 (IPv6).salt:{YYYY-MM-DD} with a 24-hour TTL.After 24 hours, the salt no longer exists. The hash cannot be reversed. Historic entries remain intact with their computed hash, but no link to the original IP is retained.
When a request claims a user-agent associated with a known AI crawler, the connecting IP is checked against the vendor's published infrastructure.
Categories:
The honeypot tracks whether agents respect robots.txt directives. It does not drop or reset connections. It serves a 200 response to every request, including those to disallowed paths.
Requests to disallowed paths are logged with the matching rule. This allows the record to show which agents respected the rules and which did not.
Requests to disallowed paths and requests to /ask are classified using
DeepSeek. The classification records:
behavior_pattern: mechanical, human_like, or unknowncadence_anomaly: booleanpath_sequence_anomaly: booleannotes: factual observations only
Compliant requests (robots_rule == "none") do not call DeepSeek. They
receive a stub classification: unclassified — skipped — routine traffic.
Every entry contains the following fields:
timestamp (UTC, ISO 8601)claimed_agent (User-Agent string, treated as a claim not identity)verification (verified, claimed, mismatch, or unverifiable)source_network (truncated IP prefix, /24 or /48)ip_hash (SHA-256 of IP + daily salt)requested_pathrobots_rule (the matching disallow rule, or "none")response_codedeepseek_classification (JSON object)tool_name, tool_input, tool_output (for tool calls only)No other fields are added. The schema is stable.
Public log (log:public:{date}): 24-hour TTL. Entries expire automatically.
Archive (log:archive:{timestamp}:{uuid}): permanent, no TTL. Retained for
12 months, then reviewed under the retention policy at /privacy.
Salt (salt:{date}): 24-hour TTL. Purged daily at 00:05 UTC.
The record is published at the following paths:
No judgments. No scores. No labels like “safe,” “unsafe,” “malicious,” or “threat.” Every entry is a factual field/value record.
No claims about whether any agent is safe or unsafe. No claims about intent. Verification describes the source network, not the agent's purpose.
Corrections, removal requests, and general enquiries: vic@fixseo.uk