Home › Glossary › What Is robots.txt?

What Is robots.txt?

robots.txt is a standard text file placed at a website's root that tells automated crawlers which parts of the site they are permitted or not permitted to access.

How it works

A robots.txt file uses simple directives (Allow/Disallow) to specify rules for crawlers, either for all bots or for specific named user agents. It’s a voluntary standard — compliant crawlers respect it, but it is not a technical access control.

Why it matters for responsible scraping

Reputable scraping providers check and honor a site’s robots.txt rules as part of operating ethically and within a site’s stated preferences, alongside reviewing terms of service.