SEO

Robots.txt Explained Simply, With Common Mistakes to Avoid

What robots.txt does, how search engines use it, safe settings for WordPress sites, the difference between blocking crawling and preventing indexing, and mistakes that hide sites from Google.

Robots.txt Explained Simply, With Common Mistakes to Avoid
On this page
  1. What it looks like
  2. Crawling vs indexing
  3. Common mistakes
  4. How to check yours
  5. AI crawlers
Key takeaways
  • robots.txt controls crawling; noindex controls whether pages appear in search.
  • Never leave a staging rule that blocks the whole site, or block CSS and JavaScript.
  • Check it in Search Console after launches and migrations.

Robots.txt is a small text file at the root of your website that tells search engine crawlers which areas they may or may not crawl. Used correctly it's harmless and useful; used wrongly it can hide your whole site from Google.

What it looks like

A typical robots.txt contains rules such as allowing all crawlers, disallowing an admin or private folder, and pointing to your sitemap. On WordPress, a common safe setup allows everything except the admin area (while allowing the file WordPress uses for front-end features), plus a sitemap line.

Crawling vs indexing

This is the most misunderstood point:

  • robots.txt controls crawling: whether bots visit a URL
  • noindex controls indexing: whether a page appears in search results

If you block a page in robots.txt, Google can't see its noindex tag, and the URL might still appear in results if other sites link to it. To keep a page out of search results, use noindex and let it be crawled.

Common mistakes

  • Blocking the whole site with a leftover staging rule after launch
  • Blocking CSS and JavaScript files, which stops Google rendering pages properly
  • Using robots.txt to hide private content. It's public, and it doesn't secure anything. Use passwords for private areas.
  • Blocking pages you want removed from search instead of using noindex

How to check yours

  • Visit yourdomain.com/robots.txt
  • Use the robots.txt report and URL Inspection tool in Google Search Console
  • After launches and migrations, confirm nothing important is disallowed

AI crawlers

Some businesses choose whether to allow AI crawlers in robots.txt. Decide deliberately based on whether you want your content used and cited by AI tools; see AI search and your website.

See also XML sitemaps explained and the technical SEO audit.

Need help with your website?

I'm Sameer, a freelance WordPress developer building fast, SEO-friendly websites since 2020. Tell me what you need and I'll reply with a plan and a fixed quote within 24 hours.

Found this useful? Share it:
Contact

Let's build your next website

Available for freelance projects, agency white-label work and long-term maintenance. Feel free to pass this along to your team or company.

Your details are emailed to me, then WhatsApp opens so we can chat right away.

Chat now