AI Product Recommendations Vary: What 1,536 Answers Reveal
A new audit shows why one AI shopping answer—or an API result—cannot stand in for the advice customers actually see. Here is what brands should measure.
Ask an AI assistant what to buy and the answer can sound decisive. Ask again—or switch from the customer-facing interface to an API—and the products or displayed sources may change. For a brand, that creates a dangerous shortcut: treating one polished answer as a stable market position.
A September 2026 research preprint puts numbers behind the problem. It does not tell us how often every shopper sees a particular brand, but it does show why the observation method belongs next to the visibility score.
- One answer is not the whole experience. The researchers observed variation in recommended products and displayed sources across repeated requests and systems.
- An API is a distinct surface. In this audit, the sources shown by the ChatGPT and Gemini APIs overlapped only modestly with their respective consumer interfaces.
- A useful brand benchmark needs a defined question, market, interface, date, and repeat count. A mention, recommendation, citation, and sale are different events.
Methodology & sources
Editorial review for factual claims (as of 2026-09-25).
We reviewed the full methods and results of Marin and colleagues' September 16, 2026 preprint. The percentages below are the authors' observations, not GEO Tracker AI customer data. The study used Netherlands-based collection and English-language browser settings; it did not test US or Czech shoppers as separate markets. Our measurement advice is an interpretation, not a claim that the paper validates our exact weekly cadence or an optimal number of repeats.
What the 1,536 responses actually represent
The researchers first assembled 2,528 real, user-written commercial-advice queries. They then selected 117 physical-product questions for the live audit. Each question was submitted three times in five conditions: the logged-out ChatGPT and Gemini interfaces, their corresponding APIs, and Google AI Overviews. That produced 1,755 attempted observations, of which 1,536 yielded a response. The three repetitions were separated by hours, not weeks.
This distinction matters. “1,536 answers” does not mean 1,536 independent shoppers, 1,536 US purchases, or 1,536 experiments on the best tracking schedule. It means a bounded set of answers under specified collection conditions. The researchers also report that only 134 of 351 Google searches produced substantive AI Overview content their parser could recover; those analyses therefore use a smaller observed subset.
The same shopping task can look different across AI systems
Among responses that recommended a product, the authors found a first-person preference in 79% of ChatGPT interface answers, versus 7% for Gemini and 2% for AI Overviews. That is a difference in how advice is framed, not a 79% chance that ChatGPT recommends any particular brand.
Sources diverged too. For the same query, the ChatGPT and Gemini consumer interfaces shared 5.4% of displayed domains on average; in 76.7% of comparisons, they shared none. This is a finding about displayed source domains. It does not mean the products never matched, the underlying retrieval was identical, or one provider was necessarily more accurate.
| Observed comparison | Reported result | What a brand should not infer |
|---|---|---|
| ChatGPT vs Gemini interfaces | 5.4% mean overlap of displayed domains | “95% of product recommendations disagree” |
| Interface vs API, ChatGPT | 12.0% mean displayed-domain overlap | “API data is useless” |
| Interface vs API, Gemini | 14.8% mean displayed-domain overlap | “The interface always shows the full source set” |
An API can still be useful for repeatable diagnostics. But this study shows it cannot automatically be labeled “what the customer saw.” Interfaces and APIs can expose different source layers. If the business question is customer-facing visibility, the collection surface must be explicit and a proxy should be validated before it is generalized.
How we turn a variable answer into a useful decision
Our starting point is a fixed panel of real buyer questions, including awkward questions about price, fit, availability, and trade-offs—not just prompts engineered to mention the brand. For each observed answer we record the engine, market, wording, time, and access method. We then keep four outcomes separate:
- Mention: Was the brand or product named at all?
- Recommendation: Was it actually suggested for the buyer's need, and with what caveats?
- Citation: Was a relevant source displayed, and was the claim supported by it?
- Commercial outcome: Did a visit, lead, or sale occur in systems that can verify it?
Repeated observations help distinguish a recurring pattern from a lucky or unlucky single answer. Stable questions make later comparisons interpretable. A change log then links a finding to a specific product-data or content improvement before we remeasure. This is how an AI visibility report becomes a work brief, not a screenshot collection.
We use a weekly repeated benchmark for standard strategic decisions, with faster checks when a team needs an alert. The paper supports accounting for variability; it does not prove that weekly beats daily or that three repeats are universally optimal. The right rhythm depends on the decision, the cost of observation, and how quickly the underlying offer changes.
What brands should do next
Pick five to ten questions a buyer would genuinely ask before choosing your product. Verify that your own site gives complete, current answers. Then compare the actual AI answers by surface and market, including cases where the brand is absent or inaccurately described. Prioritize fixes that make the offer easier to understand and verify; remeasure the same questions after the changes.
A good measurement result is not “we appeared once.” It is a defensible account of where, how often, and under which conditions customers can encounter your offer—and what your team can improve.
Frequently asked questions
The same Q&A pairs ship as FAQPage structured data so AI engines can quote them verbatim.
- Why can the same AI shopping question produce different recommendations?
- AI answers can vary across providers, interfaces, available sources, and repeated requests. A September 2026 audit found changes in recommended products and displayed domains under its controlled conditions. That does not mean every answer is wrong; it means a single response is a weak stand-in for what all shoppers see.
- Can an AI API result stand in for the consumer-facing answer?
- Not automatically. In the cited audit, mean displayed-domain overlap between interface and API was 12.0% for ChatGPT and 14.8% for Gemini. APIs and interfaces may expose different source layers. Label the surface you measure, and validate any proxy against the actual customer-facing experience before generalizing.
- Does this research prove that weekly measurement with three repeats is best?
- No. The researchers repeated 117 physical-product queries three times across five conditions, with repeats separated by hours. Their results support checking variability and recording the observation conditions. They do not compare daily with weekly monitoring or establish three repetitions as a universal optimum for every brand or market.
- What should a brand track besides an AI mention?
- Track whether the answer merely names the brand, recommends a product, cites a source, and describes the offer accurately. Keep those answer-level observations separate from verified visits, leads, and sales. A citation is not a recommendation, and none of these answer signals alone proves commercial impact.
Primary research and scope
- Marin et al., “Auditing AI-Generated Product Recommendations” (arXiv, September 16, 2026) — abstract, findings, authors, and publication status.
- Full paper and methods — query selection, five audit conditions, three repetitions, location, response counts, and limitations.
Reviewed September 25, 2026. This is a preprint, not a GEO Tracker AI experiment or a study of US-specific shopping outcomes.
Related articles
Your GEO Score
Establish an AI mention baseline you can defend
GEO Tracker AI runs repeatable checks for supported engines so you can see whether your brand is mentioned, what context shows up, and how that changes week over week — complementary to Search Console, not a replacement for it.