Industry Report #004

The Locked Storefront: Retail Blocks the Agents It Will Soon Depend On

Ten of sixteen Fortune 500 retail consumer domains failed a basic crawlability probe, and every consumer storefront crawl was blocked outright. The measurable retail web is a set of small corporate sites averaging 309 pages, and across all 5,073 pieces of structured data they publish, not one annotation is answer-shaped.

Industry · Retail & E-commerce Dataset · Q3 2026 Sites analyzed · 11 Published · August 25, 2026
TL;DR The fast read. Each line is measured below.
  • 10/16 Fortune 500 retail consumer domains failed the crawlability probe outright: anti-bot walls, commerce sprawl, or no usable sitemap. Every consumer storefront crawl attempt that followed was blocked.
  • 0 of the cohort's 5,073 structured-data annotations are answer-shaped (FAQ / Q&A). No other cohort we have measured recorded a literal zero.
  • 31.7% of published content is link-invisible on average across the six measurable sites (live and in the sitemap, but nothing links to it), reaching 91% on one.
  • 34.3% of pages are dead-ends with no outbound editorial link, the highest rate of any cohort, three times the technology sector's.
  • 10.4% average agent-readiness score. Not one of the eleven sites exposes a content signal, an API catalog, or machine-readable content negotiation.
  • 309 pages on the average measurable retail site, the smallest cohort we have crawled. Where the section exists, investor relations absorbs nearly half of all pages: the readable retail web is written for shareholders, not shoppers.

The Finding

Every industry report in this series begins with what we measured. This one has to begin with what we could not.

We set out to crawl the Fortune 500 retail sector the way a link-following AI agent would: politely, honoring robots directives, reading public pages. Sixteen consumer retail domains went through the pre-flight crawlability probe. Ten failed it outright: anti-bot challenge pages, hard 403 walls, or no usable content inventory. Of the six that passed, several blocked the actual crawl anyway: a green probe on Tuesday became a wall of refused connections on Wednesday. Across three successive waves and twenty-three crawl attempts (consumer storefronts first, then the corporate subdomains of the same brands, then standalone corporate domains), not a single consumer storefront produced a measurable site. The corporate subdomains of consumer brands were walled just as uniformly as the storefronts they belong to. One consumer site did technically crawl, and yielded almost nothing but faceted search and SEO landing pages: commerce sprawl with no editorial structure to measure, so it was excluded.

What survived is the other retail web: eleven standalone corporate and investor-relations sites, averaging 309 pages each, totaling 3,403 pages and 4.4 million tokens. This is the entire Fortune 500 retail surface an unaided agent can actually read. And on that surface, the sector records a number no other cohort has produced: of 5,073 structured-data annotations across all eleven sites, exactly zero are answer-shaped.

Retail’s public face is closed to machines. Its open face has nothing a machine can quote. Both findings are the report.

The Wall Is the First Finding

For healthcare, finance, and technology, crawl failures were a footnote. For retail they are the phenomenon. The pattern in the probe data is clean: the more consumer-facing the domain, the harder the wall. Grocery, big-box, home improvement, apparel: the storefronts that transact with the public sit behind commercial anti-bot layers that classify any non-browser reader as a threat. This is a rational posture against scrapers and inventory bots. It also means that as AI assistants become a real referral channel, these companies’ primary web presence is a black box to the agents doing the referring: not poorly structured, not thin. Absent.

The escape hatch that worked is telling. The sites that crawled cleanly are the standalone corporate domains, the -inc.com and companies.com style properties that exist apart from the storefront. These are governed by a different team, serve a different audience, and face none of the storefront’s bot pressure. They are also small: the largest is 864 nodes, the smallest 11, and the median site in this cohort is just 52 pages of graph. Everything measured below describes this corporate web, with the standing caveat that it is the minority of retail’s content, measurable precisely because it is the part nobody thought to wall.

Three Tiers of Reachability

Every page falls into one of three reachability tiers, judged against the full link graph (navigation included) for reach and the content link graph for editorial authority:

  1. Reachable via editorial links. An in-content link path leads to the page. These are the pages a link-following agent finds on its own.
  2. Navigation-only. Reachable, but only through chrome: header, footer, or menu links that repeat site-wide. An agent that filters boilerplate, as most do, loses these.
  3. Link-invisible. Zero inbound links of any kind. The page exists as a URL and sits in the sitemap, but nothing on the site advertises it to a link-following crawler.

Link-invisibility requires the site’s full inventory (the XML sitemap) to measure. Only six of the eleven sites exposed one usable enough to count against (see Methodology). Averaged across those six:

Reachable via editorial links: 62.2% Reachable via editorial links 62.2% Navigation-only: 6.1% 6.1% Link-invisible: 31.7% Link-invisible 31.7%
  • Navigation-only (6.1%)
Mean across the 6 measurable retail F500 sites, Q3 2026.
View data
Segment Value
Reachable via editorial links 62.2%
Navigation-only 6.1%
Link-invisible 31.7%

Just under a third of the published corporate-retail web is invisible to any link-following agent, almost exactly the technology sector’s 32%, on a cohort a tenth the size. Here is how it distributes:

0% 20% 40% 60% 80% 100% sample mean 31.7% R9: 5.9% R9 R5: 24.8% R5 R7: 5% R7 R2: 22% R2 R1: 41.5% R1 R8: 91.2% R8 Link-invisible page rate (% of site's published inventory)
Each dot is one of the 6 measurable sites. Mean 31.7%.
View data
Site Link-invisible page rate (% of site's published inventory)
R7 5%
R9 5.9%
R2 22%
R5 24.8%
R1 41.5%
R8 91.2%
Sample mean 31.7%

The extreme case, R8 at 91%, is an apparel holding company whose sitemap advertises hundreds of pages while its editorial graph reaches barely thirty of them: the inventory is dominated by locale variants and archival pages that canonical deduplication collapses and that nothing on the live site links to. It is the most extreme gap between published and linked we have recorded. But the mid-table is the more representative story: two of the cohort’s largest, most actively maintained sites (a grocery IR hub and an apparel group) quietly orphan 22% and 25% of their inventory, the same publish-then-forget archive pattern finance and technology showed.

The Dead-End Signature

Retail adds a failure mode the other cohorts did not lead with. Across the cohort, 34.3% of pages are dead-ends: they have inbound links but no outbound editorial link at all. The page is a terminus. Technology’s rate was 10%; retail’s is more than three times that, the highest we have measured.

The cause is visible in what these sites are. Investor-relations pages are generated artifacts (a press release, an SEC filing, a quarterly-results page) stamped out by a publishing system that links to them from an index and never links out of them to anything. A link-following agent that arrives at one of these pages stops there. Combined with the orphan rate, the picture is a web of cul-de-sacs: on the average measurable retail site, an agent can neither discover a third of the content nor continue onward from a third of what it does reach.

Machine-Readability: Seven Sites Publish No Structured Data at All

Structured-data coverage in this cohort is not a spectrum. It is a switch, and for most of the cohort the switch is off:

0% 20% 40% 60% 80% 100% sample mean 35.4% R11: 0% R11 R10: 0% R10 R6: 0% R6 R5: 0% R5 R4: 0% R4 R9: 100% R9 R3: 0% R3 R7: 100% R7 R1: 0% R1 R8: 89.3% R8 R2: 100% R2 Pages carrying JSON-LD / schema.org structured data (%)
Each dot is one site. Mean 35.4% (n=11). Seven of eleven publish literally none.
View data
Site Pages carrying JSON-LD / schema.org structured data (%)
R1 0%
R3 0%
R4 0%
R5 0%
R6 0%
R10 0%
R11 0%
R8 89.3%
R2 100%
R7 100%
R9 100%
Sample mean 35.4%

Seven of eleven sites publish zero schema markup: not sparse coverage, none. Technology’s worst site had 20%. And on the four sites that do publish it, composition tells the rest:

Navigation
44.8%
WebPage, BreadcrumbList, ListItem, SearchAction
Identity
32.4%
Organization, PostalAddress, ContactPoint
Product / other
15.2%
Product, PropertyValueSpecification
Article / media
7.6%
Article, ImageObject
Answer (FAQ / Q&A)
0%
FAQPage, Question, Answer

Share of all schema.org annotations across the 11 sites, by type.

5,073 annotations cohort-wide. Answer-shaped markup: zero.

Over three quarters is navigation and identity scaffolding. The answer-shaped share, the FAQ and Q&A markup an AI assistant lifts directly into a reply, is 0.0%. Healthcare had a little, finance had a little, technology averaged 6%. Retail is the first cohort to record none anywhere. The sector most exposed to AI-mediated purchase decisions has, on its measurable web, written nothing in the one format built to be quoted.

Render dependency is the one signal that mostly clears: the cohort averages 11.4%, and eight of eleven sites sit under 10%. The exception is severe: one grocery corporate site (R4) serves 80% of its content only after JavaScript executes, the highest single-site rate we have recorded in any cohort. For that site the storefront wall and the render wall amount to the same thing: an agent that does not run a browser reads almost nothing.

0% 20% 40% 60% 80% 100% sample mean 11.4% R6: 7.7% R6 R7: 5% R7 R1: 2.5% R1 R5: 1.9% R5 R2: 0.8% R2 R3: 0.6% R3 R10: 0% R10 R9: 17.6% R9 R8: 0% R8 R11: 9.1% R11 R4: 79.8% R4 Pages whose content requires browser rendering (%)
Each dot is one site. Mean 11.4% (n=11). One outlier at 79.8%.
View data
Site Pages whose content requires browser rendering (%)
R8 0%
R10 0%
R3 0.6%
R2 0.8%
R5 1.9%
R1 2.5%
R7 5%
R6 7.7%
R11 9.1%
R9 17.6%
R4 79.8%
Sample mean 11.4%

Agent-Readiness: 10.4%, and the Interesting Failures Are Unanimous

The agent-readiness score checks seven concrete affordances a site can offer an AI agent, from llms.txt-style content signals to API catalogs to markdown content negotiation. The cohort averages 10.4%, and the shape of the failure matters more than the mean. Three checks fail on all eleven sites: no site publishes a content signal, an API catalog, or machine-readable content negotiation. The scattered passes (two sites each on four checks) are mostly incidental platform behavior rather than intent. The two best scores in the cohort, at 43%, belong to sites that also rank among the worst on other lenses. Readiness here is accidental, not designed.

Fragility: Small Webs, Heavy Hubs

Betweenness centrality counts how often a page sits on the shortest path between two others; when it concentrates in a few pages, the structure is fragile. On the average measurable site, the top 1% of pages carry 43.9% of all betweenness, in line with technology’s 47%, but on webs this small the concentration is more literal. On the smallest sites the “top 1%” is a single page: one blog index, one SEC-filings page, through which most of the site’s internal structure routes. The extreme is R10 at 90.9%: a thirteen-page site that is functionally a star around one hub.

0% 20% 40% 60% 80% 100% sample mean 43.9% R8: 22.2% R8 R2: 60.3% R2 R3: 20.4% R3 R6: 32.2% R6 R4: 60.1% R4 R7: 18.7% R7 R1: 30% R1 R5: 47.4% R5 R9: 56.5% R9 R10: 90.9% R10 Share of total betweenness carried by the top 1% of pages (%)
Each dot is one site (n=10; one site too small to compute). Higher means more fragile.
View data
Site Share of total betweenness carried by the top 1% of pages (%)
R7 18.7%
R3 20.4%
R8 22.2%
R1 30%
R6 32.2%
R5 47.4%
R9 56.5%
R4 60.1%
R2 60.3%
R10 90.9%
Sample mean 43.9%

The Industry Scorecard

The five-lens analysis assigns each site a green, amber, or red score per lens. Across all eleven:

Site Skeleton Size, density, and average path length. How big and how connected the site is at the body level. Circulation PageRank distribution and structural bottlenecks. How importance flows between pages, and which hubs hold it all together. Organs Community detection. Whether the site's topical clusters cleanly separate, or whether one mega-cluster dominates everything. Health Islands, orphans, and dead-ends. Where content is structurally dying: unreachable, unlinked, or terminating. Nervous Sys. Click depth, bridges, and cross-community linking. Whether the site is a well-designed building or a pile of disconnected rooms.
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
  • Green: healthy
  • Amber: moderate concern
  • Red: critical
Five-lens scorecard for each of the 11 retail sites.

Where technology failed on Health, retail fails on Neighborhoods: five reds and four ambers out of eleven on the lens that measures topical community structure. These sites are too small and too flat to form real content neighborhoods: a handful of sections around an IR hub, with little for a community-detection algorithm to find. Health continues the cross-cohort pattern (three reds, seven ambers, one green). Skeletons remain mostly fine, which by now is the least surprising finding in the series: every sector can build a tree; what happens on top of the tree is where they diverge.

Content Quality at Scale

Across all eleven sites combined: 3,403 pages and 4.4 million tokens of body text, the smallest cohort footprint to date by a wide margin.

Content metric Sample mean (n=11) Notes
Pages per site 309 Median 245, range 11–939
Avg word count per page 665 Below tech (1,099), above the thin-page bar
Avg token count per page ~1,103 Character-based estimate (length / 4)
Pages with thin content (under 200 words) 37.6% Nearly three times technology’s 14%
Internal-link share of all links 71.7% internal Comparable to prior cohorts
Title tag coverage 96.7% One site drags the average (63%)
Meta description coverage 48.1% Weakest of any cohort; three sites at 0%

Thin content at 37.6% is the highest of any cohort, consistent with what these sites are: press-release stubs, filing indexes, board-member bios. Meta descriptions are missing from half the cohort’s pages, and three sites have none at all. The corporate retail web is not just small; page for page, it gives an agent less to work with.

What Every Retail Site Shares

Healthcare organized around the newsroom, finance around jurisdiction, technology around the product line. The measurable retail web organizes around the investor. On the sites where an investor-relations section exists, it absorbs on average 46% of all pages, by far the largest single section, with careers and contact sections as the recurring supporting cast. Several sites’ top-level structure is effectively SEC filings, press releases, careers, contact.

This is the structural consequence of the wall. Retail companies do publish for machines: the storefront APIs, the feeds, the merchandising systems are all machine-first. But that machinery lives behind the anti-bot layer with the storefront itself. What is left on the open web is the site built for analysts and regulators, and its structure says so. An AI assistant asked about a retail company’s products, services, sustainability record, or store policies is left to read quarterly-results pages, because that is what the readable web contains.

What This Means for AI Search Readiness

For an AI agent that discovers and quotes content by following links, three implications follow.

1. The wall is a choice with a new cost. Anti-bot layers were priced against scrapers. They are now also pricing out assistant crawlers and agent traffic, the channel projected to mediate a growing share of purchase research. A sector that blocks every non-browser reader has opted out of that channel by default, storefront-first. The competitor that whitelists legitimate agent traffic, or publishes a crawlable product and policy surface outside the wall, gets quoted; the one that does not, does not exist in the answer.

2. The open surface answers the wrong question. What retail leaves readable is written for shareholders. Zero answer-shaped markup, 38% thin pages, half the inventory missing meta descriptions, a third of pages invisible and a third of the rest dead-ends: an assistant asked a consumer question about these companies finds a web that structurally cannot answer it.

3. Small is fixable. This is the cheapest re-wire in the series. The flip side of a 309-page site is that the remediation is bounded. Re-linking the invisible tier, adding outbound links to the press-release template, writing answer-shaped markup for the pages meant to be quoted: on sites this size, each is a project measured in weeks, not quarters. The same fixes on a technology site meant re-wiring thousands of pages. Retail’s measurable web could be made agent-legible faster than any cohort we have studied. The harder, strategic question is what to publish outside the wall in the first place.

The fix is structural, and every piece of it is measurable: expose a crawlable product and policy surface, re-link the orphaned archive, break the dead-end template, and add answer-shaped markup to the pages meant to be quoted. This is the entire premise of the Digital MRI service.

Methodology

Eleven anonymized Fortune 500 retail and e-commerce corporate sites, each run through the same pipeline: an HTTP-first adaptive crawl, main-content extraction (navigation, header, and footer links filtered out), dual-graph construction, and the five-lens topology analysis (Skeleton, Circulation, Organs, Health, Nervous System). The cohort was assembled in three waves: consumer storefront domains (blocked or unusable), corporate subdomains of consumer brands (uniformly blocked), and standalone corporate domains (usable). Sixteen consumer domains were probed pre-flight; ten failed on anti-bot or inventory grounds, and probe-passing sites were still blocked at crawl time in several cases. This selection process is itself a finding: the cohort is the measurable subset of Fortune 500 retail, which systematically excludes consumer storefronts. Aggregate figures are exploratory benchmarks for that corporate-retail web, not for retail web presence at large.

Reachability and invisibility use only the six sites whose crawl captured a usable sitemap; topology, content, and the scorecard use all eleven; fragility uses the ten sites large enough to compute. Several sites in this cohort are very small (11–52 pages); cohort means are reported alongside per-site distributions throughout because on samples this small the mean alone would mislead. The 91% invisibility outlier reflects a sitemap dominated by locale variants and archival URLs that the site’s live editorial graph no longer references; it is reported as measured, with this composition noted.

The dual-graph model. Each page is classified against two graphs from the same crawl: the full graph (navigation included) for reach, and the content graph (in-content editorial links only) for authority. A page is reachable via editorial links if it has an inbound content link, navigation-only if reachable in the full graph but not the content graph, and link-invisible if nothing links to it in either, surviving only in the sitemap. This matches the model used in IR-002 and IR-003; per-site rates are not directly comparable across cohorts, but the shared finding now holds across four sectors: a measurable share of F500 content sits outside any link path an unaided agent can walk.

Disclaimers:

  • Methods. PageRank, Louvain community detection, and betweenness centrality over crawled page structure.
  • Robots & ethics. Site-level robots directives respected; disallowed pages never fetched; anti-bot walls honored, never circumvented. Publicly accessible page structure only. No content, metadata, or user data stored.
  • Anonymization. Codenames R1-R11; the codename-to-domain mapping is intentionally not published. A site-level data appendix is available on request.
  • Navigation exclusion. Any link target appearing on more than 80% of a site’s pages is treated as global navigation and dropped from the content graph; the full graph keeps it for reachability only.
  • Token counts. Character-based estimates (page length divided by four), not tokenizer output.
  • Betweenness. Reported as the share of total betweenness centrality carried by each site’s top 1% of pages; not computed for one site too small to support it.
  • Scope. Statistical patterns for educational purposes only; not advice about any specific site or company.