Shopify robots.txt.liquid: What to Block (and What Never to Touch)
What Shopify's 2026 default robots.txt blocks, how robots.txt.liquid works, the few rules worth adding on big catalogs, and the edits that quietly cost traffic.

- Most Shopify stores should not edit robots.txt at all. The default already blocks sort, tag-combination and multi-filter URLs, and Shopify updates it.
- If you customize robots.txt.liquid, loop over robots.default_groups and add rules. A pasted plain-text file freezes an old copy of Shopify's defaults.
- Never block /cdn/, ?variant= or ?page= URLs: they carry product images, Shopping landing pages and the paths to deep products.
- robots.txt stops crawling, not indexing. Blocking a URL hides its noindex and canonical from Google.
- Google-Extended only controls Gemini training use. Blocking OAI-SearchBot removes you from ChatGPT search.
The best robots.txt edit for most Shopify stores is no edit. Shopify generates a sensible file, updates it, and already blocks the URL types that multiply on big catalogs. The traffic losses we see come from custom robots.txt.liquid templates: rules pasted from a forum in 2021, wildcards that catch real products, and "block the AI bots" edits that removed a store from ChatGPT search.
This guide covers what the 2026 default blocks, how the template works, the handful of rules worth adding on a large catalog, and the edits to avoid. We read live robots.txt files, including Shopify's Dawn demo store and several large brands, and Shopify's and Google's documentation in September 2026.
What Shopify's default robots.txt blocks in 2026#
Open yourstore.com/robots.txt first. The file changed this year. On the Dawn demo store, the new version starts with comments that point agents to /agents.md and Shopify's UCP endpoints, then a single group for all crawlers plus a group for Google's AdsBot. Many established stores still serve the older format, usually because a customized template froze it.
Here are the rule groups in the new default, abbreviated:
User-agent: *
Allow: /
Allow: /products/checkout
Allow: /*/products/checkout
# ...similar Allow lines for account and orders
# Private / transactional
Disallow: /admin
Disallow: /cart/
Disallow: /checkout
Disallow: /*/checkout
Disallow: /account
Disallow: /*/account
# Filters, sort, previews, language-picker crawl traps
Disallow: /collections/*sort_by*
Disallow: /*/collections/*sort_by*
Disallow: /collections/*+*
Disallow: /*/collections/*+*
Disallow: /collections/*filter*&*filter*
Disallow: /*/collections/*filter*&*filter*
Disallow: /blogs/*+*
Disallow: /*?*oseid=*
Disallow: /*?*preview_theme_id=*
Sitemap: https://yourstore.com/sitemap.xmlWhat each group does:
| Rule group | Examples | Why it's there |
|---|---|---|
| Transactional | /cart/, /checkout, /orders, /account | Private, personalized pages with no search value |
| Internal endpoints | /services, /sf_*, /cart.js, /recommendations/products | Platform and AJAX endpoints |
| Crawl traps | sort_by, tag combinations (+), two or more filter parameters | URL types that multiply without limit |
| Previews and tracking | preview_theme_id, oseid, repeated ls parameters | Theme previews and language-picker loops |
| Market subfolders | /*/collections/*sort_by* and similar | The same rules under /en-ca/, /de/ and so on |
Two changes from the older file matter for SEO. The new default no longer disallows /search or /policies/. Internal search pages are now crawlable unless your theme adds noindex (Shopify's help center gives the snippet, and we cover it in Shopify duplicate content). Policy pages being crawlable is useful, because Google reads return and shipping policies.
If your file looks different: older and customized versions#
Many live stores don't serve the 2026 format. When we compared the Dawn demo with large brands' files, the differences fell into a few groups:
| What you might see | What it means | What to do |
|---|---|---|
Disallow: /search and Disallow: /policies/ | Older default rules, usually frozen in a custom template | Fine to keep /search blocked if search pages aren't indexed; consider unblocking /policies/ |
Disallow: */collections/*filter*&*filter* with a leading * | The older way of writing the multi-filter rule; it also matches market subfolders | Nothing, it still works |
| Extra groups for AhrefsBot, Nutch or other named bots | Someone added custom groups | Check each group; a named group replaces the * rules for that bot |
Disallow: /collections/*/products* | A custom block on collection-path product URLs | Only safe if nothing in the theme links there |
No Sitemap: line | A plain-text template that dropped it | Rebuild the template on the Liquid loop |
The third row catches people out. A crawler follows the most specific group that names it and ignores the * group. If you add User-agent: AhrefsBot with only a crawl delay, that bot no longer sees any of your * disallows. Shopify's own file repeats the key rules in its AdsBot group for the same reason: its comment notes that AdsBot "ignores robots.txt unless specifically named".
Why the new file has Allow lines for /products/checkout#
This is a good lesson in how wildcards bite. Disallow: /*/checkout exists to block checkout pages in market subfolders like /en-ca/checkout. But /*/checkout also matches /products/checkout-counter-display and /collections/checkout-accessories, because Google treats a rule as a prefix.
Shopify fixes it with Allow: /products/checkout. Google applies the most specific matching rule, measured by path length, and uses the least restrictive rule when two are equally specific. /products/checkout is longer than /*/checkout, so the product stays crawlable. Keep that in mind before you write any rule with a leading /*/ or a word in the middle of a wildcard.
How robots.txt.liquid works#
To customize the file, open Online Store > Themes > Edit code, add a new template and choose robots.txt. Shopify creates templates/robots.txt.liquid with its default Liquid. It has access to the robots, group, rule, user_agent and sitemap objects, plus request.
Three facts from Shopify's documentation to know before you start:
- Shopify Support won't help. The help center says Support "can't help with edits to the robots.txt.liquid file".
- Use the Liquid objects. Shopify's developer docs say it's "strongly recommended to use the provided Liquid objects whenever possible" because "the default rules are updated regularly".
- Reverting is deleting. Delete the template and Shopify serves its default file again.
The safe pattern: keep the defaults, add your rules#
This is the structure Shopify documents. It prints every default group and rule, then appends custom rules to the group for all crawlers:
{% for group in robots.default_groups %}
{{- group.user_agent }}
{%- for rule in group.rules -%}
{{ rule }}
{%- endfor -%}
{%- if group.user_agent.value == '*' -%}
{{ 'Disallow: /collections/*filter.v.availability*' }}
{{ 'Disallow: /*/collections/*filter.v.availability*' }}
{%- endif -%}
{%- if group.sitemap != blank -%}
{{ group.sitemap }}
{%- endif -%}
{% endfor %}After saving, open /robots.txt and check that every custom rule sits on its own line. On one well-known brand's live file in September 2026 we found two rules fused into one line (...-remoteDisallow: /collections/*-loop). A line like that is one broken rule, not two working ones.
Removing a default rule#
Shopify's docs show wrapping the rule output in a condition. For example, if your live file still disallows /policies/ and you want policy pages crawlable:
{%- for rule in group.rules -%}
{%- unless rule.directive == 'Disallow' and rule.value == '/policies/' -%}
{{ rule }}
{%- endunless -%}
{%- endfor -%}Remove default rules rarely. Each one exists because a URL type caused problems across millions of stores.
What to block on a large catalog#
Short list. Add a rule only when Search Console's Crawl stats show Googlebot spending real time on the URL type, and the URLs have no search value.
| Candidate | Suggested rule | Before you add it |
|---|---|---|
| Single low-value filters (availability, sometimes price) | Disallow: /collections/*filter.v.availability* | Confirm no filtered URL is indexed and earning clicks |
| Internal search, if crawled heavily | Disallow: /search and Disallow: /*/search | Add noindex first, wait for pages to drop out, then block |
App pages with no content (wishlists, compare tools under /apps/ or /a/) | A specific path, never /apps/ wholesale | Some apps serve real landing pages or reviews from these paths |
| Extra sitemaps | Sitemap: https://yourstore.com/... | Only for sitemaps an app or headless page actually generates |
A note on search: Shopify's search results link to products with tracking parameters like ?_pos=1&_sid=...&_ss=r. Those URLs canonicalize to the product, so they're a crawl-budget question, not an indexing one. Blocking them with a pattern such as /*?*_pos=* is optional. We've rarely needed it.
The filter decision is covered in depth in Shopify faceted navigation SEO, including when to build a real collection instead of a filter page.
What never to touch#
| Don't | Why |
|---|---|
| Replace the template with plain text | You lose Shopify's future updates to the default rules |
Block /cdn/ or *.js / *.css | Product images are served from /cdn/shop/ on your domain; Google also needs CSS and JS to render |
Block ?variant= | Shopping landing pages often use variant URLs, and Storebot-Google obeys robots.txt |
Block ?page= or /collections/*?* | Deep products are only linked from paginated collection pages |
| Block URLs you want noindexed | Google can't see a noindex on a page it can't crawl |
Remove the Sitemap: line | It's the easiest way for every crawler to find your sitemap |
Use crawl-delay for Googlebot | Google ignores it |
Block /collections/*/products/* before fixing links | It hides the canonical; fix the theme's within links first |
On that last row: some large brands do block collection-path product URLs, and it can work when nothing links to them. If your theme still links there, blocking stops Google from reading the canonical, and the internal links go nowhere. Fix the links first. The URL side of this is in Shopify URL structure.
AI crawlers: what each block actually does#
The default file has no AI-specific rules. If your team wants to control AI access, block the right token for the right reason:
| Token | What it controls | Effect of blocking |
|---|---|---|
Googlebot | Google Search, including AI Overviews and AI Mode | You disappear from Google |
Google-Extended | Use of crawled content to train Gemini models | Google says it doesn't affect Search inclusion or ranking |
OAI-SearchBot | ChatGPT search results | You drop out of ChatGPT search answers |
GPTBot | OpenAI model training | No effect on ChatGPT search |
Storebot-Google | Google Shopping surfaces | Products can drop from Shopping |
For most ecommerce brands we leave all of them allowed. Blocking training crawlers is a business decision; blocking search crawlers is almost always a mistake. More on AI visibility in how to rank in ChatGPT.
Multiple domains and Markets#
If you run Markets on separate domains and need different rules per domain, the template can branch on request.host. Shopify's docs reserve this for stores with distinct domains that need different crawling behavior. For subfolders, you don't need it: write each custom rule twice, once for /collections/... and once for /*/collections/..., the same way Shopify's defaults do.
How to test a change#
- Read the live file.
curl -s https://yourstore.com/robots.txtshows exactly what crawlers see. - Check the robots.txt report in Search Console (Settings > robots.txt). It shows the version Google fetched and any parse problems.
- Test real URLs. Use URL Inspection on one product, one collection, one paginated collection, one blog post and one policy page. None should say "Blocked by robots.txt".
- Wait before judging. Google generally caches robots.txt for up to 24 hours, and changes to crawling take longer to show up in Crawl stats.
- Stay under 500 KiB. Google ignores content past that size. A file that large on Shopify means something has gone wrong.
The short version#
Leave the defaults alone unless Crawl stats give you a reason. Put a monthly reminder on the calendar to read the live file after theme updates, app installs and Markets changes, because each of those is a moment a custom template can drift.
If you customize, loop robots.default_groups, add narrow rules for market subfolders too, and never block images, variants, pagination or pages you want noindexed. Then read the live file with your own eyes. This is one small part of our Shopify technical SEO work, one of the first things we check in a Shopify SEO audit, and item one in the Shopify SEO checklist. For how robots.txt fits with sitemaps, canonicals and the rest, see the Shopify technical SEO guide.
Questions people ask about this
Can you edit robots.txt on Shopify?
Yes. Create a robots.txt.liquid template in your theme (Online Store > Themes > Edit code > Add a new template > robots.txt). Shopify recommends keeping its Liquid loop over robots.default_groups and adding or removing rules inside it, so you still receive Shopify's updates. Deleting the template restores the default file. Shopify Support doesn't help with edits to it.
What does Shopify's default robots.txt block?
Checkout, cart, account and order pages, internal endpoints, theme preview URLs, sort_by URLs, tag combinations with a plus sign and collection URLs with two or more filter parameters, with matching rules for market subfolders. The 2026 version we read on Shopify's Dawn demo store no longer blocks /search or /policies/. Check your own live file, because customized templates differ.
Should I block collection filter URLs in Shopify robots.txt?
Multi-filter URLs are already blocked. Single-filter URLs canonicalize to the parent collection on standard themes, which is usually enough. Block a single filter type only if Crawl stats show heavy crawling and it has no search demand, such as availability. If a filter value has real demand, build a proper collection for it instead of a filter page.
Does blocking a URL in robots.txt remove it from Google?
No. robots.txt controls crawling, not indexing. Google can still index a blocked URL from links, usually without a description. It also can't see a noindex tag or canonical on a page it isn't allowed to crawl. To remove pages, keep them crawlable and serve noindex, or delete them so they return a 404.
Should a Shopify store block AI crawlers like GPTBot?
Separate search crawlers from training crawlers. Blocking OAI-SearchBot removes you from ChatGPT search results, and blocking Googlebot removes you from Google, including AI Overviews. GPTBot and Google-Extended relate to model training; Google says Google-Extended doesn't affect Search. Most ecommerce brands we work with leave all of them allowed.
Want this done on your store?
Fixes for the crawl, index and schema problems Shopify creates on its own: duplicate product paths, tag pages, filter URLs, app bloat and Markets hreflang.
More from Technical SEO

The Shopify Apps Slowing Down Your Store and Quietly Hurting SEO (Plus a 20-Minute Audit)
Shopify apps slow down your store, leave code behind after you uninstall them, and sometimes change what Google sees. This 20-minute audit shows where to look and what to remove.

Crawled – Currently Not Indexed on Shopify: Why Google Skips Your Pages
Crawled – currently not indexed means Google fetched a page and chose to leave it out. On Shopify most of those URLs are harmless; here is how to find the few that cost you sales and fix them.

Shopify Speed Optimization: Fixing LCP on Collection Pages Without Killing Apps
Why Shopify collection pages fail LCP, how to find the real LCP element, and the fixes in order: image loading, animations, app scripts and Liquid render time.