Great Expectations was named in 5 of the 6 answers we recorded, more than any other data quality tool. We ran this on 2026-08-04: six buyer questions about data quality software, each put to ChatGPT and Gemini with one live web search behind it, with every product named and every source domain logged.
No other tool came close to that reach. Three products tied for second at 3 of 6 answers each (50%): Monte Carlo, OvalEdge, and Soda Core. Metaplane was the only other tool to appear more than once, named in 2 of 6 answers (33%). Everything else, seventeen separate tools, was named exactly once.
Across the six questions, 22 distinct products were named in total. Seventeen of those 22 (77%) showed up in a single answer and never again. So the "field" that ChatGPT and Gemini describe for data quality software is really five recurring names plus a long tail that changes with the exact question a buyer types.
How the shortlist changed across the six questions
The wording of the question rewrote the list. The bare "best data quality software" query returned 8 names (Monte Carlo, Atlan, DQLabs, OvalEdge, Great Expectations, Soda Core, Metaplane, Ataccama ONE), and "data quality software recommendations" also returned 8. But "most affordable data quality software" returned only 3, and "best data quality software for small teams" returned only 3. Same category, one live search each, and the answer size more than doubled depending on phrasing.
Three of the 8 names in that bare "best" answer appeared only there and nowhere else in the six: Atlan, DQLabs, and Ataccama ONE. The bare query pulls in one-off enterprise and catalog names that never resurface once the buyer adds a qualifier.
Adding the word "free" swapped the whole shortlist. "Best free data quality software" returned 7 names, and five of them appeared in no other answer: OpenMetadata, Amundsen, DQOps, Deequ, and DataCleaner, all open-source. The commercial observability tools that led the bare "best" query, Monte Carlo among them, vanished the moment the buyer said "free."
"Most affordable data quality software" returned just 3 names: Great Expectations, OvalEdge, and Collate. Collate appeared in this one answer and nowhere else across the six. A vendor's entire presence in the "affordable" conversation can rest on a single retrieved page.
"Best data quality software for small teams" was the only one of the six questions where Great Expectations did not appear at all. Its 3 names were Metaplane, Talend Data Quality, and OpenRefine. Of those three, only Metaplane appeared anywhere else in the six answers; Talend Data Quality and OpenRefine each showed up here and nowhere else.
"Data quality software recommendations" was the only answer to surface the large enterprise suites. Informatica Data Quality and Observability, SAS Viya, Talend Data Fabric, and Microsoft Data Quality were each named in this question and in none of the other five, alongside Acceldata, also a single-question name. Four of the biggest names in the category, one mention apiece, all riding on a single question's phrasing. The "what data quality software should I use" answer, meanwhile, returned 4 names (Monte Carlo, Great Expectations, Soda Core, dbt Core), and only dbt Core was unique to it.
Great Expectations is the exception that proves the pattern. It was named in 5 of 6 answers: "best," "what should I use," "best free," "most affordable," and "recommendations." It is the only tool that read as both a top pick and a free/affordable pick at once, which is why it survived nearly every rephrasing while everyone else got filtered in or out by a single word.
The source map: who fed the answers
26 distinct domains fed these six answers, and one appeared in every single one: reddit.com, cited in 6 of 6 answers. A community thread, not a vendor page and not a review directory, is the single most-cited source behind these AI recommendations.
The next two most-cited domains are vendor sites, and both belong to products that got named. montecarlo.ai was cited in 5 of 6 answers, and Monte Carlo the product was named in 3. ovaledge.com was also cited in 5 of 6 answers, and OvalEdge was named in 3. This is the key insight: when a vendor's own pages are the ones being retrieved, that vendor tends to land in the answer. The content and the citation move together.
atlan.com was cited in 4 of 6 answers, more than any vendor except Monte Carlo and OvalEdge, yet Atlan the product was named only once. Atlan publishes comparison and roundup content that these assistants retrieve and quote, but that content mostly recommends other tools. gable.ai and lakefs.io tell the same story from further down: each was cited in 2 of 6 answers, and neither Gable nor lakeFS was named in any of them. Getting cited is not the same as getting named.
Review directories barely registered. G2 appeared in 1 of 6 answers, and Capterra did not appear at all. softwarereviews.com and technologyadvice.com each appeared once. In a category most people assume is ruled by G2 grids, the answers leaned on Reddit (6 of 6), vendor blogs, Gartner (2 of 6), and Quora (1 of 6) instead.
What this means if you sell data quality software
Market share did not decide these answers. The source list did. Informatica, SAS, Microsoft, and Talend are among the largest names in the category, and each was named exactly once, only in the "recommendations" question. Great Expectations, an open-source project, was named 5 times because it is what the retrieved Reddit threads and free-tool roundups actually mention.
So the job is not "be bigger," it is "be in these exact pages." Reddit fed 6 of 6 answers, so a genuine, visible presence in the data-engineering threads matters more than another paid directory slot. Monte Carlo and OvalEdge show the other half of the play: their own domains were each cited in 5 of 6 answers and their products were named, so publishing content that AI retrieves works when that content is about your own product.
The Atlan result is the warning attached to that play: 4 citations, 1 naming. Publishing a comparison post that ranks your rivals gets you retrieved and gets them named. If you want to be the answer, the page that gets cited has to be about you, not a neutral roundup of the field.
The full list, counted
Below the top five, seventeen tools tied at a single mention each (17%), so the four single-mention rows here are examples from that tie, not a clean 6th-to-9th ranking. Note also that Talend appears under two separate labels the assistants used, Talend Data Fabric and Talend Data Quality, each named once. We keep them as separate rows because that is how they were recorded, and adding them together would invent a "Talend = 2" figure the data does not show.
| Product | Named in | Share |
|---|---|---|
| Great Expectations | 5 of 6 | 83% |
| Monte Carlo | 3 of 6 | 50% |
| OvalEdge | 3 of 6 | 50% |
| Soda Core | 3 of 6 | 50% |
| Metaplane | 2 of 6 | 33% |
| Collate | 1 of 6 | 17% |
| OpenMetadata | 1 of 6 | 17% |
| Talend Data Fabric | 1 of 6 | 17% |
| Talend Data Quality | 1 of 6 | 17% |
Where we fit, and where we don't
Glotier does not sell data quality software, and we appear in 0 of the 6 answers. That is correct: we are not in this category, so ChatGPT and Gemini should not name us here, and they don't.
But running this measurement is exactly what Glotier does, pointed at your category instead of this one. We asked six buyer questions, recorded all 22 products and all 26 domains, and showed which pages fed each answer. You can run the same check on "best [your category]" and see whether you are the Great Expectations of your space, named in 5 of 6, or one of the seventeen names mentioned once. The check is free, takes about a minute, and needs no account.