ToolStop
Definition

What is robots.txt?

robots.txt is a plain-text file at the root of a website that tells web crawlers which paths they may or may not fetch.

Also called: robots file · crawler directives

The file lives at /robots.txt and uses a simple grammar: 'User-agent:' followed by 'Allow:' and 'Disallow:' rules. A single 'Disallow: /' locks every crawler out of the whole site — usually not what you want.

robots.txt is a request, not a wall. Well-behaved crawlers (Googlebot, Bingbot) obey it; malicious scrapers ignore it. Use HTTP auth or firewall rules for actual access control.

A typical setup allows everything, blocks internal search or admin URLs, and points to the XML sitemap with a 'Sitemap:' directive.

Try a tool that uses robots.txt

Related definitions