Study · 19 July 2026
Three assistants, the same 40 questions, less than a third agreement
In two earlier articles we said we would split our source study by assistant and publish it. This is that, with one honest amendment: half of what we promised turns out to be unanswerable, and the half that is answerable produced a cleaner result than the question we originally asked.
- 30.2%
- named by all three
- 46.7%
- named by exactly one
- 64.1%
- agreement on identical sources
- ~60%
- of the picture from one assistant
The half of the promise we cannot keep
What we said we would find out was which of the three assistants leans hardest on Reddit. Having opened up how our own measurement works to do it, that question has no answer for two of the three, and the reason is worth explaining because it is a property of these systems and not a flaw in ours.
Our check runs one live web search and hands the identical results to ChatGPT and Gemini. Perplexity does not accept results handed to it, it always runs its own retrieval, so it gets the bare question. That means asking which of ChatGPT and Gemini prefers Reddit is like asking which of two people reading the same newspaper prefers a different newspaper. They read what they are given.
What this does make possible is better. Because ChatGPT and Gemini read identical sources, any difference between what they recommend is the model’s own judgement with retrieval held constant. That is a controlled experiment, and it is not one you can run by using the consumer apps. Then Perplexity, doing its own retrieval, tells you what changes when the sources change too. So instead of one weak answer we get two clean ones: how much disagreement comes from the model, and how much comes from the sources.
How it ran
The same 40 software buying questions as our first study, spread across twenty-odd categories. Each got one live search. The results went to ChatGPT and Gemini, the bare question went to Perplexity, and this time we extracted the recommended products from each assistant’s own answer separately rather than from the merged text. All 40questions returned answers. On one, “stripe alternatives for saas”, Perplexity named nothing at all, which leaves 39 questions where all three said something and agreement can be computed.
Across those, the assistants named 437 distinct products once name variants are collapsed. That collapsing is not a detail, and the next section explains why we are telling you about it.
A correction to our own first pass
Our initial analysis of this data said the assistants agreed 25.9% of the time. That number was wrong in two directions that both exaggerated disagreement, and we caught it by reading the raw rows instead of trusting the aggregate. Both errors are easy to make and worth naming, because any study of this kind that does not mention them has probably made them.
Name variants counted as different products. One assistant wrote “Profound”, another “Profound AI”, a third “Profound AI / Profound”. That is one company and it was scored as three, so three assistants naming the same product registered as perfect disagreement. The same happened with Otterly, Scrunch and every product whose name ends in AI. The fix was to normalise to a brand key before comparing, while keeping genuinely distinct products apart: “Bigin by Zoho CRM” stays separate from “Zoho CRM”, because Bigin is a different product.
An empty answer counted as disagreement. On the Stripe question, ChatGPT and Gemini named the identical eight products and Perplexity returned nothing. The naive maths recorded that as 0% agreement on a question where the two assistants that answered agreed completely. Excluding questions where any assistant said nothing, and reporting them separately, fixes it.
Corrected, agreement rose from 25.9% to 30.2%. The conclusion did not change, but a number we would have published was off by four points, and it was off in the direction that made our own product look more necessary. That is exactly the direction a measurement company should check itself hardest in.
Finding 1: they agree on under a third
Of the 437 distinct products named across 39 questions, all three assistants named the same product 132 times, 30.2%. Two of the three agreed on another 101 (23.1%). And 204 products, 46.7% of everything named, came from exactly one assistant.
Say that last one slowly. Nearly half of all the products recommended across this study were recommended by a single assistant, and the other two never mentioned them at all. Not ranked them lower. Did not mention them.
The mean per-question overlap tells the same story from the other side: on a typical question, about 31.7% of the products named were named by all three. The most agreed question in the set, “best hosting for nextjs apps”, reached 63%. The least, “best design tools for non designers”, managed 7%: fourteen products named, one of them by everybody.
Finding 2: identical sources, and still only 64% agreement
This is the controlled result and the most interesting number in the study. ChatGPT and Gemini were handed the same search results, word for word, for every question. On that identical evidence they agreed on 64.1% of the products they named.
So more than a third of the disagreement between two assistants has nothing to do with what they read. Given the same pages, two models make meaningfully different decisions about which products in those pages are worth putting in front of a buyer. Whatever that judgement is made of, training, tuning, how each model weighs a hedged sentence against a confident one, it is doing about a third of the work of deciding whether your product gets recommended.
That reframes a common piece of advice. “Get into the sources and the assistants will follow” is right, and it is not sufficient. Being in the retrieved pages is necessary. It buys you a coin flip weighted somewhere around two thirds, not a guarantee.
Finding 3: different sources roughly halve it again
Perplexity, running its own retrieval on the same question, agreed with the other two on 37.3% of names. Compared with the 64% between two models sharing sources, that is close to half.
It also produced the most singular recommendations by a distance. Of the 204 products named by exactly one assistant, 122 came from Perplexity, against 42 from ChatGPT and 40 from Gemini. That is a threefold gap, and it is the expected consequence of being the only one of the three fetching its own pages: it finds things the shared search did not surface.
Whether that is a feature depends on what you want. If you are a buyer, Perplexity is showing you a wider field, including smaller products the others miss. If you are a small product, Perplexity is the assistant most likely to be the one that has heard of you, and the one whose result you can least explain, because its retrieval is not visible from outside.
Finding 4: one assistant shows you about 60% of the picture
The practical consequence, and the number to remember from this article. If you check exactly one assistant to see how your category is being described, here is the share of all named products you would actually see:
| Perplexity | 65.2% |
| Gemini | 60.2% |
| ChatGPT | 58.1% |
Whichever one you pick, roughly 40% of what is being said about your category is invisible to you. And it is not a random 40%: it is disproportionately the products that only one assistant surfaces, which for a small company is frequently the category in which you yourself appear.
The failure mode this produces is specific. A founder checks ChatGPT, sees their name, and stops worrying. What they have established is that they are visible on one assistant, on one phrasing, on one day. Gemini, which their buyer may well have met inside a Google search without choosing it, might not name them at all, and nothing in that founder’s process would ever reveal it.
Two questions, taken apart
Percentages hide what disagreement actually looks like, so here are two questions in full, exactly as the three assistants answered them.
“simple crm for small business”
ChatGPT
HubSpot CRM, Zoho CRM, OnePageCRM
Gemini
Bigin by Zoho CRM, Salesforce Starter Suite, OnePageCRM, HubSpot's free CRM, Zoho CRM
Perplexity
Less Annoying CRM, Bigin by Zoho CRM, HubSpot CRM, Pipedrive, Freshsales, monday CRM
Nine distinct products between them and one, OnePageCRM, that reaches two of the three. Perplexity is the only one to mention Pipedrive, Freshsales, monday or Less Annoying CRM, which between them are a large fraction of the small-business CRM market. A founder at any of those four could check ChatGPT, find themselves absent, and conclude they have an AI visibility problem, when what they have is an assistant problem: they are being recommended, just not there.
“best ai visibility tracking tools”, our own category
ChatGPT
Otterly AI, Scrunch AI, Knowatoa, Peec AI, Profound, AthenaHQ, Rankscale AI
Gemini
Semrush, Hall, Nightwatch, Writesonic, Peec AI, Profound, Scrunch AI, Otterly AI
Perplexity
Profound, Frase, Otterly AI, Peec AI, ZipTie, Semrush, Ahrefs, Scrunch AI
This one is qualitatively different rather than just numerically. ChatGPT answered entirely with purpose-built AI visibility products. Gemini and Perplexity both reached for large general SEO suites, Semrush and Ahrefs, alongside the specialists. Those are not slightly different shortlists, they are different interpretations of what the buyer is asking for: one read it as “a new category of tool”, the others as “a feature my existing SEO vendor might have”.
For anyone selling in that category, the two readings imply completely different competitive positions and completely different pages to write. And you would only ever discover that there are two readings by asking more than one assistant. Glotier, incidentally, appears in none of the three, which we would rather state here than have a reader notice.
They also disagree about how many to name
A smaller pattern, consistent across the set. ChatGPT named 265 products across the questions, an average of 6.63 per answer. Gemini named 273, averaging 6.83. Perplexity named 286, averaging 7.15, and was also the only one to return an empty answer.
The spread is not dramatic and it points the same way as everything else: Perplexity casts the widest net, ChatGPT the narrowest. A narrower shortlist is harsher if you are outside it and better if you are in it, because being one of six is a stronger position than being one of eight. Neither is right. They are different editorial temperaments applied to the same job.
Why identical evidence produces different answers
This section is reasoning rather than measurement, and it is labelled as such because we cannot see inside these models and neither can anyone else selling you an opinion about them. But the 64% figure demands some explanation, and a few mechanisms are plausible enough to be worth naming.
What counts as a recommendation is a judgement call. A retrieved page might mention nine products, describe four in detail and endorse one. Deciding how far down that list to go before the answer stops being useful is genuinely subjective, and two models can draw it in different places while both being reasonable.
Prior familiarity leaks in. Even a grounded answer is written by a model that has opinions from training. A product it already recognises may get promoted out of a passing mention, and one it has never encountered may get skipped despite being on the page. That would explain some of the pattern where the better-known name survives into both answers and the obscure one appears in only one.
Length and structure force choices. An assistant writing a short answer must cut. Which three of nine candidates survive a cut is exactly the kind of decision where small differences in preference produce large differences in output.
None of these are levers you can pull. That is the point of separating the 64% from the 37%: roughly a third of your fate with a given assistant is decided by something you cannot influence, and the remaining two thirds is retrieval, which you can. Spending your effort on the two thirds is not a compromise, it is the only rational allocation.
What to do with this
Stop treating one assistant as the assistant
The single most common measurement error in this subject is generalising from one engine, and this study puts a number on the cost: you are seeing about 60% of the field. Ask at least ChatGPT, Gemini and Perplexity, and treat a name that appears on one as a partial result rather than a verdict.
Expect the answers to differ, and stop trying to reconcile them
People see two assistants disagree and assume one is broken. Neither is. At 64% agreement on identical evidence, disagreement is the normal operating state of these systems. Measure the spread instead of hunting for the correct answer.
Work on sources first anyway
A third of the disagreement is model judgement you cannot influence. Two thirds traces to what was retrieved, and that you can influence. Sources remain the highest-leverage move even though they are not a guarantee.
Watch Perplexity if you are small
It produced three times as many single-assistant recommendations as either of the others. If any assistant is going to surface a product nobody else has heard of, on this evidence it is the one running its own search.
Count across questions and across engines
One question on one assistant is two sampling errors stacked. The unit that means something is the fraction of your buyers' questions where you are named, measured on more than one engine, tracked over time.
Narrow questions get agreement, fuzzy ones do not
Sorting the questions by three-way overlap produces a pattern clear enough to plan around. The questions where the assistants agreed most were the technically narrow ones: “best hosting for nextjs apps” at 63%, “best project management tools for small teams” at 60%, “best merchant of record for indie saas” at 56%.
The questions where they agreed least were the broad and subjective ones: “best design tools for non designers” at 7%, fourteen products named and one shared. “How to track brand mentions in chatgpt” at 8%. “Simple crm for small business” at 11%.
The mechanism is not mysterious. “Merchant of record for indie SaaS” has a small, well-defined set of correct answers, and any competent process converges on roughly the same names. “Design tools for non designers” has fifty defensible answers, and which four you surface depends entirely on judgement calls about who counts as a non designer and what counts as design.
If your category is the narrow kind, the good news is that the shortlist is stable and the bad news is that it is short and hard to enter. If your category is the broad kind, the shortlist is a lottery with many tickets: you will be named somewhere, inconsistently, and coverage rather than any single answer is the only thing worth measuring. Knowing which of those you are in changes what a disappointing result means.
Run a version of this yourself
You can get most of the value by hand in an afternoon, and doing it once tells you more about your position than reading about ours.
Take five questions your buyers ask, without your product name in them. Put each to ChatGPT, Gemini and Perplexity with web search enabled. Write down every product each one names, in three columns. Then count three things: how many products appear in all three columns, how many appear in only one, and whether yours appears anywhere.
Two pieces of advice from having done it forty times. Normalise the names before you compare, or you will record the same disagreement artefact we did and conclude the assistants agree far less than they do. And treat an empty or evasive answer as missing data rather than as a zero, for the same reason.
What you will end up with is your own version of the number that matters: not whether you are visible, but on how many of your buyers’ questions, on how many of the engines they might use.
What this study cannot tell you
One run, one day, 40 questions, English, software categories. Enough to see a shape, not enough to put a confidence interval on any single percentage, and we would expect the exact figures to move on a rerun.
The comparison is also between our panel’s configuration of these assistants and not between the consumer apps you would open in a browser. Those apps run their own retrieval, their own system prompts and sometimes different models, so the answers you get by hand will not match these exactly. What transfers is the structural finding, that agreement between assistants is low and that a meaningful part of it is model rather than retrieval, which is a property of the systems rather than of our setup.
And brand normalisation is a judgement call. We collapsed “Profound” and “Profound AI” and kept “Bigin by Zoho CRM” separate from “Zoho CRM”. Reasonable people would draw one or two of those lines differently, which would move the agreement figure by a point or two in either direction. It would not move it from a third to a half.
Related reading
The sources behind these answers, including why Reddit fed 38 of 40, are in what sources AI actually cites. The framework this measurement sits inside is in AI visibility, where this study is the evidence behind the part called consistency.
One assistant shows you 60% of the picture. See all three, on your own product. Free, no account.