# lagrange.dev robots.txt — updated 2026-08-03 # Policy: allow search (traditional + AI answer engines), block AI training. # Goal: cut non-human crawl bandwidth while keeping search rankings # and AI answer-engine visibility fully intact. # ---------- Search engines: full access (do not touch) ---------- User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / # ---------- AI answer/search crawlers: KEEP for visibility ---------- # These power live citations and answers that send people to lagrange.dev. # Blocking the training crawlers below does NOT affect these. OpenAI and # Anthropic both confirm search visibility is governed only by the # SearchBot agents, never by GPTBot/ClaudeBot. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / # Mistral: both agents are search/user-initiated, not training. Keep. User-agent: MistralAI-Index Allow: / User-agent: MistralAI-User Allow: / # ---------- AI TRAINING crawlers: BLOCK ---------- # 2026-08-03: GPTBot and ClaudeBot are now blocked, reversing the 2026-07-20 # decision to allow them. Rationale — ClaudeBot alone is ~20% of identified # bot requests (up 66% MoM) and GPTBot ~9.6%, versus ~3.3% and ~2.8% for # their search counterparts. Crawl-to-refer ratios are roughly 2,237:1 # (ClaudeBot) and 217:1 (GPTBot) — effectively no referral traffic in return. # Both honor robots.txt, so this is enforceable today without Cloudflare. # Blocking Google-Extended does NOT affect Google Search rankings. # Major AI labs User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / User-agent: DeepSeekBot Disallow: / User-agent: PanguBot Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: AI2Bot Disallow: / User-agent: Ai2Bot-Dolma Disallow: / User-agent: Bytespider Disallow: / User-agent: TikTokSpider Disallow: / User-agent: CCBot Disallow: / User-agent: Amazonbot Disallow: / # Dataset / scraper crawlers feeding model training User-agent: Timpibot Disallow: / User-agent: FirecrawlAgent Disallow: / User-agent: img2dataset Disallow: / User-agent: LAIONDownloader Disallow: / User-agent: ICC-Crawler Disallow: / User-agent: Cotoyogi Disallow: / User-agent: Brightbot Disallow: / User-agent: FriendlyCrawler Disallow: / User-agent: ISSCyberRiskCrawler Disallow: / User-agent: Factset_spyderbot Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: PetalBot Disallow: / User-agent: Scrapy Disallow: / # ---------- SEO tool crawlers ---------- # Slowed, not blocked — your team may use Ahrefs/Semrush for its own audits. # If you do not use them, change these to Disallow: / User-agent: AhrefsBot Crawl-delay: 10 User-agent: SemrushBot Crawl-delay: 10 # Aggressive index crawlers with no marketing value to Lagrange: User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / # ---------- Everyone else ---------- User-agent: * Allow: / Sitemap: https://lagrange.dev/sitemap.xml Sitemap: https://lagrange.dev/sitemap.xml