SmartEcomSEO
Technical SEO

Shopify robots.txt.liquid: What to Block (and What Never to Touch)

What Shopify's 2026 default robots.txt blocks, how robots.txt.liquid works, the few rules worth adding on big catalogs, and the edits that quietly cost traffic.

Illustration for “Shopify robots.txt.liquid: What to Block (and What Never to Touch)”
Key takeaways
  • Most Shopify stores should not edit robots.txt at all. The default already blocks sort, tag-combination and multi-filter URLs, and Shopify updates it.
  • If you customize robots.txt.liquid, loop over robots.default_groups and add rules. A pasted plain-text file freezes an old copy of Shopify's defaults.
  • Never block /cdn/, ?variant= or ?page= URLs: they carry product images, Shopping landing pages and the paths to deep products.
  • robots.txt stops crawling, not indexing. Blocking a URL hides its noindex and canonical from Google.
  • Google-Extended only controls Gemini training use. Blocking OAI-SearchBot removes you from ChatGPT search.

The best robots.txt edit for most Shopify stores is no edit. Shopify generates a sensible file, updates it, and already blocks the URL types that multiply on big catalogs. The traffic losses we see come from custom robots.txt.liquid templates: rules pasted from a forum in 2021, wildcards that catch real products, and "block the AI bots" edits that removed a store from ChatGPT search.

This guide covers what the 2026 default blocks, how the template works, the handful of rules worth adding on a large catalog, and the edits to avoid. We read live robots.txt files, including Shopify's Dawn demo store and several large brands, and Shopify's and Google's documentation in September 2026.

What Shopify's default robots.txt blocks in 2026#

Open yourstore.com/robots.txt first. The file changed this year. On the Dawn demo store, the new version starts with comments that point agents to /agents.md and Shopify's UCP endpoints, then a single group for all crawlers plus a group for Google's AdsBot. Many established stores still serve the older format, usually because a customized template froze it.

Here are the rule groups in the new default, abbreviated:

texttext
User-agent: *
Allow: /
Allow: /products/checkout
Allow: /*/products/checkout
# ...similar Allow lines for account and orders

# Private / transactional
Disallow: /admin
Disallow: /cart/
Disallow: /checkout
Disallow: /*/checkout
Disallow: /account
Disallow: /*/account

# Filters, sort, previews, language-picker crawl traps
Disallow: /collections/*sort_by*
Disallow: /*/collections/*sort_by*
Disallow: /collections/*+*
Disallow: /*/collections/*+*
Disallow: /collections/*filter*&*filter*
Disallow: /*/collections/*filter*&*filter*
Disallow: /blogs/*+*
Disallow: /*?*oseid=*
Disallow: /*?*preview_theme_id=*

Sitemap: https://yourstore.com/sitemap.xml

What each group does:

Rule groupExamplesWhy it's there
Transactional/cart/, /checkout, /orders, /accountPrivate, personalized pages with no search value
Internal endpoints/services, /sf_*, /cart.js, /recommendations/productsPlatform and AJAX endpoints
Crawl trapssort_by, tag combinations (+), two or more filter parametersURL types that multiply without limit
Previews and trackingpreview_theme_id, oseid, repeated ls parametersTheme previews and language-picker loops
Market subfolders/*/collections/*sort_by* and similarThe same rules under /en-ca/, /de/ and so on

Two changes from the older file matter for SEO. The new default no longer disallows /search or /policies/. Internal search pages are now crawlable unless your theme adds noindex (Shopify's help center gives the snippet, and we cover it in Shopify duplicate content). Policy pages being crawlable is useful, because Google reads return and shipping policies.

If your file looks different: older and customized versions#

Many live stores don't serve the 2026 format. When we compared the Dawn demo with large brands' files, the differences fell into a few groups:

What you might seeWhat it meansWhat to do
Disallow: /search and Disallow: /policies/Older default rules, usually frozen in a custom templateFine to keep /search blocked if search pages aren't indexed; consider unblocking /policies/
Disallow: */collections/*filter*&*filter* with a leading *The older way of writing the multi-filter rule; it also matches market subfoldersNothing, it still works
Extra groups for AhrefsBot, Nutch or other named botsSomeone added custom groupsCheck each group; a named group replaces the * rules for that bot
Disallow: /collections/*/products*A custom block on collection-path product URLsOnly safe if nothing in the theme links there
No Sitemap: lineA plain-text template that dropped itRebuild the template on the Liquid loop

The third row catches people out. A crawler follows the most specific group that names it and ignores the * group. If you add User-agent: AhrefsBot with only a crawl delay, that bot no longer sees any of your * disallows. Shopify's own file repeats the key rules in its AdsBot group for the same reason: its comment notes that AdsBot "ignores robots.txt unless specifically named".

Why the new file has Allow lines for /products/checkout#

This is a good lesson in how wildcards bite. Disallow: /*/checkout exists to block checkout pages in market subfolders like /en-ca/checkout. But /*/checkout also matches /products/checkout-counter-display and /collections/checkout-accessories, because Google treats a rule as a prefix.

Shopify fixes it with Allow: /products/checkout. Google applies the most specific matching rule, measured by path length, and uses the least restrictive rule when two are equally specific. /products/checkout is longer than /*/checkout, so the product stays crawlable. Keep that in mind before you write any rule with a leading /*/ or a word in the middle of a wildcard.

How robots.txt.liquid works#

To customize the file, open Online Store > Themes > Edit code, add a new template and choose robots.txt. Shopify creates templates/robots.txt.liquid with its default Liquid. It has access to the robots, group, rule, user_agent and sitemap objects, plus request.

Three facts from Shopify's documentation to know before you start:

  • Shopify Support won't help. The help center says Support "can't help with edits to the robots.txt.liquid file".
  • Use the Liquid objects. Shopify's developer docs say it's "strongly recommended to use the provided Liquid objects whenever possible" because "the default rules are updated regularly".
  • Reverting is deleting. Delete the template and Shopify serves its default file again.

The safe pattern: keep the defaults, add your rules#

This is the structure Shopify documents. It prints every default group and rule, then appends custom rules to the group for all crawlers:

liquidliquid
{% for group in robots.default_groups %}
  {{- group.user_agent }}

  {%- for rule in group.rules -%}
    {{ rule }}
  {%- endfor -%}

  {%- if group.user_agent.value == '*' -%}
    {{ 'Disallow: /collections/*filter.v.availability*' }}
    {{ 'Disallow: /*/collections/*filter.v.availability*' }}
  {%- endif -%}

  {%- if group.sitemap != blank -%}
    {{ group.sitemap }}
  {%- endif -%}
{% endfor %}

After saving, open /robots.txt and check that every custom rule sits on its own line. On one well-known brand's live file in September 2026 we found two rules fused into one line (...-remoteDisallow: /collections/*-loop). A line like that is one broken rule, not two working ones.

Removing a default rule#

Shopify's docs show wrapping the rule output in a condition. For example, if your live file still disallows /policies/ and you want policy pages crawlable:

liquidliquid
{%- for rule in group.rules -%}
  {%- unless rule.directive == 'Disallow' and rule.value == '/policies/' -%}
    {{ rule }}
  {%- endunless -%}
{%- endfor -%}

Remove default rules rarely. Each one exists because a URL type caused problems across millions of stores.

What to block on a large catalog#

Short list. Add a rule only when Search Console's Crawl stats show Googlebot spending real time on the URL type, and the URLs have no search value.

CandidateSuggested ruleBefore you add it
Single low-value filters (availability, sometimes price)Disallow: /collections/*filter.v.availability*Confirm no filtered URL is indexed and earning clicks
Internal search, if crawled heavilyDisallow: /search and Disallow: /*/searchAdd noindex first, wait for pages to drop out, then block
App pages with no content (wishlists, compare tools under /apps/ or /a/)A specific path, never /apps/ wholesaleSome apps serve real landing pages or reviews from these paths
Extra sitemapsSitemap: https://yourstore.com/...Only for sitemaps an app or headless page actually generates

A note on search: Shopify's search results link to products with tracking parameters like ?_pos=1&_sid=...&_ss=r. Those URLs canonicalize to the product, so they're a crawl-budget question, not an indexing one. Blocking them with a pattern such as /*?*_pos=* is optional. We've rarely needed it.

The filter decision is covered in depth in Shopify faceted navigation SEO, including when to build a real collection instead of a filter page.

What never to touch#

Don'tWhy
Replace the template with plain textYou lose Shopify's future updates to the default rules
Block /cdn/ or *.js / *.cssProduct images are served from /cdn/shop/ on your domain; Google also needs CSS and JS to render
Block ?variant=Shopping landing pages often use variant URLs, and Storebot-Google obeys robots.txt
Block ?page= or /collections/*?*Deep products are only linked from paginated collection pages
Block URLs you want noindexedGoogle can't see a noindex on a page it can't crawl
Remove the Sitemap: lineIt's the easiest way for every crawler to find your sitemap
Use crawl-delay for GooglebotGoogle ignores it
Block /collections/*/products/* before fixing linksIt hides the canonical; fix the theme's within links first

On that last row: some large brands do block collection-path product URLs, and it can work when nothing links to them. If your theme still links there, blocking stops Google from reading the canonical, and the internal links go nowhere. Fix the links first. The URL side of this is in Shopify URL structure.

AI crawlers: what each block actually does#

The default file has no AI-specific rules. If your team wants to control AI access, block the right token for the right reason:

TokenWhat it controlsEffect of blocking
GooglebotGoogle Search, including AI Overviews and AI ModeYou disappear from Google
Google-ExtendedUse of crawled content to train Gemini modelsGoogle says it doesn't affect Search inclusion or ranking
OAI-SearchBotChatGPT search resultsYou drop out of ChatGPT search answers
GPTBotOpenAI model trainingNo effect on ChatGPT search
Storebot-GoogleGoogle Shopping surfacesProducts can drop from Shopping

For most ecommerce brands we leave all of them allowed. Blocking training crawlers is a business decision; blocking search crawlers is almost always a mistake. More on AI visibility in how to rank in ChatGPT.

Multiple domains and Markets#

If you run Markets on separate domains and need different rules per domain, the template can branch on request.host. Shopify's docs reserve this for stores with distinct domains that need different crawling behavior. For subfolders, you don't need it: write each custom rule twice, once for /collections/... and once for /*/collections/..., the same way Shopify's defaults do.

How to test a change#

  1. Read the live file. curl -s https://yourstore.com/robots.txt shows exactly what crawlers see.
  2. Check the robots.txt report in Search Console (Settings > robots.txt). It shows the version Google fetched and any parse problems.
  3. Test real URLs. Use URL Inspection on one product, one collection, one paginated collection, one blog post and one policy page. None should say "Blocked by robots.txt".
  4. Wait before judging. Google generally caches robots.txt for up to 24 hours, and changes to crawling take longer to show up in Crawl stats.
  5. Stay under 500 KiB. Google ignores content past that size. A file that large on Shopify means something has gone wrong.

The short version#

Leave the defaults alone unless Crawl stats give you a reason. Put a monthly reminder on the calendar to read the live file after theme updates, app installs and Markets changes, because each of those is a moment a custom template can drift.

If you customize, loop robots.default_groups, add narrow rules for market subfolders too, and never block images, variants, pagination or pages you want noindexed. Then read the live file with your own eyes. This is one small part of our Shopify technical SEO work, one of the first things we check in a Shopify SEO audit, and item one in the Shopify SEO checklist. For how robots.txt fits with sitemaps, canonicals and the rest, see the Shopify technical SEO guide.

People also ask

Questions people ask about this

Can you edit robots.txt on Shopify?

Yes. Create a robots.txt.liquid template in your theme (Online Store > Themes > Edit code > Add a new template > robots.txt). Shopify recommends keeping its Liquid loop over robots.default_groups and adding or removing rules inside it, so you still receive Shopify's updates. Deleting the template restores the default file. Shopify Support doesn't help with edits to it.

What does Shopify's default robots.txt block?

Checkout, cart, account and order pages, internal endpoints, theme preview URLs, sort_by URLs, tag combinations with a plus sign and collection URLs with two or more filter parameters, with matching rules for market subfolders. The 2026 version we read on Shopify's Dawn demo store no longer blocks /search or /policies/. Check your own live file, because customized templates differ.

Should I block collection filter URLs in Shopify robots.txt?

Multi-filter URLs are already blocked. Single-filter URLs canonicalize to the parent collection on standard themes, which is usually enough. Block a single filter type only if Crawl stats show heavy crawling and it has no search demand, such as availability. If a filter value has real demand, build a proper collection for it instead of a filter page.

Does blocking a URL in robots.txt remove it from Google?

No. robots.txt controls crawling, not indexing. Google can still index a blocked URL from links, usually without a description. It also can't see a noindex tag or canonical on a page it isn't allowed to crawl. To remove pages, keep them crawlable and serve noindex, or delete them so they return a 404.

Should a Shopify store block AI crawlers like GPTBot?

Separate search crawlers from training crawlers. Blocking OAI-SearchBot removes you from ChatGPT search results, and blocking Googlebot removes you from Google, including AI Overviews. GPTBot and Google-Extended relate to model training; Google says Google-Extended doesn't affect Search. Most ecommerce brands we work with leave all of them allowed.

Next step

Want this done on your store?

Fixes for the crawl, index and schema problems Shopify creates on its own: duplicate product paths, tag pages, filter URLs, app bloat and Markets hreflang.

Talk to Zack directlyUsually replies within 24 hours
WeChatSearch this ID in WeChat
WhatsApp +86 186 8214 2136