# Everyone not named below. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # Data for our own pages. Private and one-time pages (quiz results, email links) are not blocked here: a blocked # page can never show its noindex, so they stay crawlable and say noindex themselves. Disallow: /api/ Disallow: /admin/ # Bulk training crawlers: refused. Each group repeats its own Disallow: a crawler that finds its own name # ignores the * group completely. User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: TikTokSpider Disallow: / User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: PetalBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Google-CloudVertexBot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Timpibot Disallow: / # Answer engines (GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, DuckAssistBot) # are absent on purpose: they fall through to * and may read the public site. # For agents that prefer plain text: /llms.txt is a map, /md/... are Markdown copies of pages. # They are not search results and are not in the sitemap. Sitemap: https://homerobotcompared.com/sitemap.xml