What is a Robots.txt File?
The robots.txt file is a plain-text document located in the root directory of your website (e.g., https://drillseo.com/robots.txt) that instructs search engine web crawlers which URLs and directories they are permitted or forbidden to access.
Standard Robots.txt Syntax Rules
- User-agent: Specifies which crawler the rule applies to (e.g.,
User-agent: GooglebotorUser-agent: *). - Disallow: Specifies URL paths that bots must not crawl.
- Allow: Explicitly permits crawling of specific sub-paths within a disallowed directory.
- Sitemap: Declares the absolute URL location of your XML sitemap index.
Recommended Production Robots.txt Configuration
User-agent: *
Disallow: /api/
Disallow: /admin/
Disallow: /search?
Allow: /
# Sitemap Index Declaration
Sitemap: https://drillseo.com/sitemap.xml
Managing AI Scrapers & LLM Crawlers
To control whether AI model training bots scrape your content, declare specific user-agent directives:
# Block AI Scrapers from Training on Content
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: PerplexityBot
Disallow: /


