Practical guide · Website Promotion

Robots.txt and Sitemaps

Robots.txt and XML sitemaps perform different jobs. Robots.txt manages crawler access to paths; a sitemap lists canonical URLs the site wants search engines to discover. Neither tool guarantees removal, indexing or…

Information checked: 21 July 2026.

Robots.txt and XML sitemaps perform different jobs. Robots.txt manages crawler access to paths; a sitemap lists canonical URLs the site wants search engines to discover. Neither tool guarantees removal, indexing or ranking.

In brief: Use robots.txt for crawling and page-level controls for indexing; use sitemaps for discovery.

Do not confuse the controls

Crawl and index controls

robots.txt Disallow
What it does: Requests that compliant crawlers avoid a path; What it does not do: Reliably remove a known URL from search
noindex
What it does: Requests that a crawled page not be indexed; What it does not do: Work when the page cannot be crawled
canonical
What it does: Signals the preferred version among duplicates; What it does not do: Guarantee Google selects that URL
XML sitemap
What it does: Lists preferred discoverable URLs; What it does not do: Force crawling or indexing
HTTP 404/410
What it does: States content is unavailable; What it does not do: Redirect users to a replacement

A safe robots.txt file

  • Keep it simple and host it at the root of the relevant hostname.
  • Do not block CSS or JavaScript needed to understand public pages.
  • Do not use it to protect confidential information.
  • Add the sitemap location when useful.
  • Test changes before deployment and monitor accidental broad rules.

A clean sitemap

  • Include canonical, indexable URLs returning 200.
  • Exclude redirects, 404 pages, noindex pages and duplicate parameters.
  • Use accurate last modification values only when content materially changes.
  • Split large or operationally distinct sitemap groups if useful.
  • Submit in Search Console and investigate processing errors.

What good looks like

The controls are coherent when a public page is crawlable, indexable and listed as intended, while private or retired content uses authentication, noindex or proper status responses rather than a misleading robots rule.

What often breaks

A staging rule or copied robots file can accidentally block an entire production site. Conversely, confidential URLs may remain publicly accessible because a team mistakes Disallow for access control.

Safe-release checks

  • Store the previous robots.txt version.
  • Test important allowed and blocked paths.
  • Compare sitemap URLs with status and canonical decisions.
  • Verify the live production hostname after deployment.

A practical next move

Test the production robots file and generated sitemap against a representative URL list. Confirm that private content is protected independently and that canonical public pages are neither blocked nor omitted accidentally.

Sources and date checked

Search, product and regulatory information was checked against the following primary sources. Interfaces and policies can change, so recheck operational details before acting. Date checked: 21 July 2026.

Keep the decision under your control

Retain the relevant accounts, source material, supplier terms and recovery information. Recheck changing prices, interfaces and rules before acting.