Back to blog
11 min read

Beyond Schema: Why AI Doesn't Recommend Your Products (and the Semantic Shelf Fix)

Semantic Shelf Positioning Matrix diagram: classification accuracy vs evidence strength
  • ai-search-optimization
  • semantic-shelf
  • ai-visibility
  • agentic-ecommerce
  • generative-engine-optimization

Schema helps AI read your product, but it does not decide if AI picks your product. Structured data qualifies your page for enhanced results and clean fact extraction. The recommendation decision is separate: it depends on whether the model recognizes the product, files it on the correct Semantic Shelf, and finds enough credible outside evidence to prefer it over 40 alternatives.

I have spent the last quarter reading every AI visibility report merchants have sent me. Almost all of them lead with the same fix: "add more schema." That advice is not wrong, but it is not the reason your product is not being recommended. Two recent LinkedIn essays, The Schema Myth and Beyond the Schema by Nitin Kumar, unpack what actually moves the needle. Both draw on a pre-registered study by Lian Pham of TeleSuite covering 1,679 products across 46 buyer queries and 99,436 scored evidence snippets. This is my working translation of that research into a merchant-side playbook for AI search optimization.

The trench coat that disappeared

Imagine two waterproof trench coats. Both stores ship perfect Product schema: name, price, material, availability, ratings, offer details. Both pages pass every Google Rich Results test. Now ask ChatGPT or Gemini for "the best waterproof trench coat for commuting in London under $300." One coat gets recommended. The other never appears, not even in the candidate set of 40 that the model narrows to 3.

The losing merchant reads that outcome and blames schema completeness. That instinct is the mistake. The winning coat did not win because its markup was better. It won because the model already knew the coat existed, filed it under "waterproof commuter trench coat for professionals," and found supporting text that answered the buyer's real questions: fit over business clothing, packability, waterproof durability after washing. Schema handled facts. The Semantic Shelf and its supporting evidence handled the decision.

What schema actually does (and where it stops)

Schema versus Recommendation Evidence: 2x2 matrix showing product schema completeness against product-specific recommendation evidence
The Schema Myth: complete markup with insufficient product-specific evidence is technically ready and commercially unproven.

Structured data does three specific jobs, and stops there:

  • Reduces ambiguity when a machine parses your page.
  • Improves attribute accuracy for price, availability, ratings, shipping.
  • Qualifies you for enhanced results like product snippets and shopping cards.

Google is explicit that valid markup does not guarantee a rich result, and OpenAI is equally clear that a well-formed product feed helps ChatGPT index and present your catalog but does not guarantee selection. Feeds and schema make you legible. They do not make you preferred. A merchant can declare that a trench coat is waterproof, durable, and highly rated; schema makes that easier to extract. It does not establish how the coat performs in heavy rain, how sizing compares with three other brands, or why it deserves one of three places in a recommendation. Those answers come from evidence that lives outside your own page.

The four-layer AI comprehension model

Recommendation happens after comprehension: 4 decision layers (Access, Interpretation, Candidacy, Selection) and their typical failures
The four decision layers a recommendation system runs before it names your product. A system can pass Access and Interpretation and still fail at Candidacy or Selection.

Nitin's core insight in The Schema Myth is that "AI comprehension" is not a single event. It is four sequential layers, and most merchants only optimize for the first two.

LayerQuestion the model answersWhere merchants usually stop
1. OrganizationIs this a real, credible brand?Yes, most brands get here.
2. ProductWhat specifically is this product, for whom, and when should it be considered?Half of brands stop here, and it is the biggest leak.
3. AttributesWhat are the concrete facts (price, size, material, warranty)?Schema handles this well.
4. RecommendationDoes this product deserve one of the 3 slots in the answer?Almost no one measures this.

A system can answer layer 1 correctly and fail every layer after it. That happens the moment your product names are written for conversion inside a familiar website. "Launch Faster," "The Essential Edit," or "Transit 02" are strong on-site names. Outside the site, in a fragment surfaced by a model, they carry no category information. The company is understood. The product is unresolved.

Brand-to-Product Translation Loss, explained with a worked example

Nitin coins a metric for this leak that I have started using with every merchant I audit. He calls it Brand-to-Product Translation Loss, the share of organization-level recognition that fails to convert into correct product-level understanding.

The formula is simple:

Brand-to-Product Translation Loss = 1 − (correctly understood product instances / organization recognition instances)

Here is the worked example. A retailer runs 80 buyer-intent prompts across ChatGPT, Gemini, and Perplexity. The organization is recognized correctly in 68 of those instances. The priority product, however, is correctly understood in only 34.

Translation Loss = 1 − (34 / 68) = 50%. Half of the brand equity that arrives at the product page never makes it back out as product understanding.

Now the recommendation math. From those 34 correctly understood instances, the product earns 10 recommendations, a 29.4% conversion rate at the shortlist stage. If better product naming, category language, and buyer context raise correct understanding from 34 to 51 instances, and the shortlist conversion holds, expected recommendations rise from 10 to roughly 15. That is a 50% recommendation lift with zero new publishing volume. The fix is not more content. The fix is repairing the layer 2 identity leak.

The Semantic Shelf: the position AI assigns your product

Semantic Shelf Positioning Matrix: classification accuracy vs evidence strength 2x2
Semantic Shelf Positioning Matrix. Classification and authority are separate problems, identify which one limits AI visibility before producing more content.

Once the product is correctly recognized, the next question is where the model files it. Nitin calls this position the Semantic Shelf. A Semantic Shelf is the category, use case, audience, and comparison group the model associates with your product. It decides which alternatives enter the same candidate set and which buyer questions your product is even eligible to answer.

Two variables set the quality of the shelf:

  1. Classification accuracy. Does AI understand what the product is, who it serves, and when it should be considered?
  2. Evidence strength. Does credible public information support the claims required for recommendation?

Combine them and you get a 2x2 that explains almost every AI visibility outcome I see:

ClassificationWeak evidenceStrong evidence
CorrectUnderstood but unproven. In the candidate set, rarely shortlisted.Recommended. This is the target quadrant.
WrongInvisible. Not even considered.Loud in the wrong room. High mention volume, wrong buyer context.

Most "we are not being recommended" complaints turn out to be a wrong-classification problem, not a schema problem. And more content published against the wrong shelf makes the situation worse, not better.

See where AI files your products on the Semantic Shelf.

TeleScope maps how ChatGPT, Gemini, and Perplexity currently classify your catalog, and where the gap sits between the shelf you occupy and the shelf you want to own.

Run a free TeleScope scan →

What the TeleSuite study of 1,679 products actually found

Lian Pham's pre-registered study is worth reading in full. It covered 46 shopping queries, 1,679 products, and 99,436 scored text snippets. Recommendation outcomes came from GPT-5 and Gemini 2.5 under fixed, no-browsing conditions, with the instrument and pass-fail rules frozen before results were captured. Four findings changed how I think about AI search optimization:

  • Familiarity beats evidence at the gate. Model familiarity with the product correlated with recommendation at r=0.176. Supporting-text density correlated at only r=0.077. Recognition is a gate; evidence helps after you clear it.
  • Evidence still matters, in matched pairs. Across 507 pairs of products with similar semantic proximity, the recommended product had 1.6 more supporting snippets on average. The difference was statistically significant at p=0.002. Recommended products had more supporting text in 53.8% of pairs, versus 42.8% for the unpicked product.
  • Raw mention volume delivered no advantage. Generic publicity was not a lever, and blended coverage across editorial, retail, Reddit, and YouTube did not distinguish winners reliably. The signal came from text located near the meaning of the query, not sheer quantity.
  • Negative surrounding text quietly punishes. Recommended products sat in a text environment that was 24.8% negative. Unpicked products sat in 29.0% negative material. A brand can accumulate mentions and still fail if the sentiment around it erodes recommendation confidence.

Read the raw research at Zenodo record 21417361 if you want the full instrument, the pre-registration, and the query set. It is the most useful piece of AI shopping research I have seen this year, and it is free.

How to fix classification before you publish another word

Publishing more content against a wrong shelf multiplies the wrong signal. Fix classification first. Five moves, in order:

  1. Rewrite the product name for portable identity. The first visible sentence of every priority product page must state product type, target audience, primary use case, and one meaningful constraint. "Transit 02" becomes "Transit 02: a waterproof commuter trench coat for professional women, sized for layering, under $300." That opening sentence is what a model quotes when a fragment gets surfaced outside your site.
  2. Align retailer and marketplace copy to the same sentence. If your Shopify page says "waterproof commuter trench coat" and your Nordstrom description says "everyday classic outerwear," the model sees two products, not one. Alignment across retailer copy is the single highest-return classification fix I have seen.
  3. Corroborate through third-party evidence near the buyer's meaning. Reviews, editorial coverage, creator posts, and Reddit threads that repeat the same category language are worth more than 10 broad brand mentions. Encourage reviewers to answer the specific buyer questions (fit, durability, climate suitability), not to praise the brand generically.
  4. Measure by prompt, not by page. Track model recognition, classification accuracy, candidate-set inclusion, evidence sentiment, and recommendation frequency per priority product, per query, per model, per season. A single AI visibility score hides every real failure point.
  5. Sweep the sentiment environment. Given the 24.8% versus 29.0% negative sentiment gap, a quarterly review sweep is table stakes. Unanswered negative reviews, out-of-date press, and Reddit complaints act as a silent recommendation tax.

Keep the schema. Keep the feeds. Protect crawlability. Those assets do the layer 1 and layer 3 work well. Reinvest the rest of your budget in layer 2 (product identity) and layer 4 (evidence density near buyer meaning). That is where recommendation lives.

Where I place this in the wider AI search optimization stack

SEO tools optimize pages for retrieval. AEO tools tune answers for AI engines. Neither measures the layer under both: whether the model actually files your product on the shelf where buyers are shopping. That layer is what TeleSuite built TeleScope to measure. And it is the argument Rosmon Sidhik and Akanksha Lokam develop, in collaboration with Nitin, in The Machine-Readable Brand: The Future of Fashion in the Age of AI Agents. Read the book if you own a catalog. It is the clearest treatment I have seen of the shift from ranking to recommendation.

FAQ

Schema helps ChatGPT read your product accurately. It does not decide whether ChatGPT picks your product. Recognition, correct classification on the Semantic Shelf, and supporting evidence near the buyer's query all sit above schema in the decision path.

What is a Semantic Shelf?

A Semantic Shelf is the category, audience, use case, and comparison group that an AI system assigns to your product. It sets which alternatives enter your candidate set and which buyer prompts your product can answer. Two variables control it: classification accuracy and evidence strength.

What is Brand-to-Product Translation Loss?

Brand-to-Product Translation Loss is the share of organization-level recognition that fails to become correct product-level understanding. The formula is 1 minus (correctly understood product instances divided by organization recognition instances). A 50% loss is common and typically caused by unclear product names, thin category language, and inconsistent retailer copy.

Does raw brand mention volume help AI recommendations?

No. Lian Pham's study of 1,679 products found that raw mention volume showed a slightly negative correlation with recommendation. The useful signal is text located near the meaning of the buyer's query, not total quantity of generic coverage.

How do I measure AI product recommendation performance?

Run 40 to 80 buyer-intent prompts across ChatGPT, Gemini, and Perplexity. For each prompt, capture whether the model recognized the product, classified it correctly, included it in the shortlist, and finally recommended it. Track sentiment of the surrounding text. Do this per product, per query, per model, per quarter.

Do product feeds fix the recommendation problem?

Feeds fix legibility and freshness. They do not fix recognition, classification, or evidence. If your feed is perfect and your product still does not appear in ChatGPT's answers for its buyer prompt, you have a Semantic Shelf problem, not a feed problem.

The Machine-Readable Brand book cover

The Machine-Readable Brand

The full playbook for winning AI product recommendations, by Rosmon Sidhik and Akanksha Lokam with Nitin Kumar.

Get the book on Amazon →

See how AI classifies your brand

TeleScope maps your product's Semantic Shelf across ChatGPT, Gemini, and Perplexity in minutes. You get the two-gaps map: where AI files your catalog today, where you want it filed, and the exact evidence to close the gap.

Run a free TeleScope scan →

Read the research: Lian Pham's pre-registered study on Zenodo.
Go deeper: The Machine-Readable Brand, by Rosmon Sidhik and Akanksha Lokam with Nitin Kumar.

More articles