SEO tools
Robots.txt Generator (with AI crawlers)
Free robots.txt generator with one-click rules for AI crawlers such as GPTBot, ClaudeBot, Google-Extended and PerplexityBot, plus sitemap and crawl rules.
Default rules
Private or low-value areas crawlers should skip.
- /admin
- /checkout
- /*?*
Paths to keep crawlable inside a disallowed folder.
One absolute URL per line.
Googlebot ignores Crawl-delay; Bing and Yandex respect it.
AI crawlers
Decide, bot by bot, who may use your content, to train models, or to cite you in live AI answers.
- GPTBotOpenAITraining
Collects content to train OpenAI's models.
- OAI-SearchBotOpenAIAI search
Indexes pages so they can appear in ChatGPT search results.
- ChatGPT-UserOpenAIUser-triggered
Fetches a page live when a ChatGPT user asks about it.
- ClaudeBotAnthropicTraining
Collects content to train Anthropic's Claude models.
- Claude-SearchBotAnthropicAI search
Indexes pages to improve Claude's search answers.
- Claude-UserAnthropicUser-triggered
Fetches a page live when a Claude user requests it.
- Google-ExtendedGoogleTraining
Controls use of your content for Gemini training; does not affect Google Search.
- PerplexityBotPerplexityAI search
Indexes pages so Perplexity can cite and link to them.
- Perplexity-UserPerplexityUser-triggered
Fetches a page live to answer a Perplexity user's question.
- Applebot-ExtendedAppleTraining
Controls use of your content to train Apple Intelligence models.
- CCBotCommon CrawlTraining
Open web archive widely used as AI training data.
- BytespiderByteDanceTraining
TikTok's parent company crawler, used for model training.
- AmazonbotAmazonTraining
Powers Alexa answers and may be used to train Amazon's models.
- Meta-ExternalAgentMetaTraining
Collects content to train Meta AI models.
- cohere-aiCohereTraining
Crawler associated with Cohere's language models.
- DuckAssistBotDuckDuckGoAI search
Fetches pages for DuckDuckGo's AI-assisted answers, with citations.
Search engine bots
These follow your default rules above. Never block them on a live site.
- Googlebot
- Bingbot
- Applebot
- YandexBot
- DuckDuckBot
- Baiduspider
Your robots.txt
# robots.txt, generated with the 67 Digital robots.txt generator
# https://67.digital
# Default rules for every crawler
User-agent: *
Disallow: /admin
Disallow: /checkout
Disallow: /*?*
# AI crawlers that may NOT use this site
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
User-agent: Amazonbot
User-agent: Meta-ExternalAgent
User-agent: cohere-ai
Disallow: /
# Allowed AI crawlers follow the default rules above:
# OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, DuckAssistBot
Sitemap: https://example.ae/sitemap.xmlUpload it to the root of your domain so it opens at /robots.txt. It is a request, not a lock, well-behaved bots obey it, but it does not protect private data.
Want this done for you?
Our team runs SEO, content and AI automation for brands across Abu Dhabi and Dubai.
Talk to 67 Digital →
