Measured · 2026-08-04

Who ChatGPT and Gemini recommend for data catalog software

We put 6 buyer questions about data catalog software to ChatGPT and Gemini. One tool tied at the top, each named in 5 of 6 answers, and none in all six. Here is the full list, the pages the answers were built from, and what it means whether or not your product is on it.

Answers read
6
Products named
21
Top source
quora.com

Atlan was named in 5 of the 6 answers we recorded, more than any other data catalog tool. We put six buyer-style questions about data catalog software to ChatGPT and Gemini on August 4, 2026, gave each question a single live web search, and wrote down every product the assistants named and every source page behind their answers. Atlan led at 5 of 6, Alation came next at 4 of 6, and three more tools tied behind them at 3 of 6.

The names that came up most

Atlan was the one product the assistants reached for almost every time, landing in 5 of the 6 answers (83%). The only question it missed was "best free data catalog software," which we come back to below.

Alation was named in 4 of the 6 answers (67%), the second most of any tool. It shows up under two labels in the raw data, "Alation" in 4 answers and "Alation Data Catalog" in 1, and we have kept those as separate rows rather than adding them into a number the measurement never produced.

Three tools tied at 3 of the 6 answers (50%) each: Amundsen, Collibra Data Catalog, and Talend. Collibra also appears under a second label, plain "Collibra," in 2 more answers, so the Collibra brand is present in 5 of the 6 answers across its two names while neither label on its own passes 3 of 6.

Below that, seven products were each named in 2 of the 6 answers (33%): Apache Atlas, Ataccama, Collibra, Data.world, DataHub, Informatica, and Informatica IDMC. Informatica is the third brand split across two labels, "Informatica" in 2 answers and "Informatica IDMC" in 2 more, covering 4 of the 6 questions between them without either label passing 2.

The answer changed a lot with the wording

This is the part worth staring at: the shortlist was not stable across the six questions. The bare "best data catalog software" question returned the shortest list, just 5 names (Collibra Data Catalog, Alation Data Catalog, Atlan, Coalesce Catalog, and DataHub), and Coalesce Catalog appeared in that one question and nowhere else in the six.

"Best free data catalog software" produced a completely different set of 4 names, and every one of them was open source: DataHub, OpenMetadata, Apache Atlas, and Amundsen. Atlan, Alation, and Collibra, the three commercial leaders everywhere else, were all absent from the free answer. OpenMetadata was named in this one question only.

"Most affordable data catalog software" did not return that open-source set at all. It brought back 8 names, and they were mostly the same paid vendors as the general questions: Alation, Collibra, Informatica, Atlan, Talend, AWS Glue, Data.world, and Ataccama. AWS Glue appeared in this question only. So "free" and "affordable" pulled two different worlds from the same assistants: free meant open-source projects, affordable meant the enterprise names with a price tag.

Amundsen was the open-source name with the widest reach, showing up in 3 of the 6 answers (small teams, what should I use, and free), while OpenMetadata reached only the free question. The narrower the wording, the more one-off names appeared: across the six questions, 9 products were named in a single answer each, and which question surfaced them tells the story. Coalesce Catalog came up under "best," Secoda and Microsoft Purview under "for small teams," Marquez under "what should I use," OpenMetadata under "free," AWS Glue under "most affordable," and Metaphor and Dataedo under "recommendations." A vendor that only shows up for one phrasing is one phrasing away from being invisible.

Where the answers came from

Two community sites fed every single answer: reddit.com appeared as a source in 6 of 6 answers, and quora.com in 6 of 6 as well. Whatever else changed, buyer threads on Reddit and Quora were behind every response ChatGPT and Gemini gave.

The most-cited non-community source was montecarlo.ai, a third-party roundup blog, present in 5 of the 6 answers. Other roundups and blogs filled most of the rest: integrate.io and notion.castordoc.com in 2 answers each, then barc.com, dataversity.net, thectoclub.com, acceldata.io, dqlabs.ai, and lakefs.io in 1 answer each.

Here is the insight we would flag to any vendor in this category: the two most-named products had their own website sitting in the sources. atlan.com was a source in 3 of the 6 answers and Atlan was named in 5; alation.com was a source in 4 of the 6 answers and Alation was named in 4. When your own domain is in the citation set, you tend to be in the answer. collibra.com, coalesce.io, dataedo.com, datahub.com, and open-metadata.org each fed 1 answer too, and each of those products was named.

Being cited is not the same as being named, though. Two sources, actian.com (1 of 6) and notion.castordoc.com (2 of 6), fed answers without any matching product landing in the results, so a domain in the source set helps but does not guarantee a mention.

Worth saying plainly: no review directory showed up. G2 and Capterra were not sources in any of the 6 answers. The pages doing the work were community threads (reddit.com, quora.com, and stackoverflow.com once), roundup blogs, a code host (github.com once), and vendors' own sites, not the directories a buyer might expect.

What we would do if we sold a data catalog

The source list decided who got named, not market share or who has the biggest sales team. So we would work the sources: get named and get accurate in the specific Reddit and Quora threads that fed all 6 answers, earn a line in the montecarlo.ai roundup that fed 5 of 6, and make our own domain strong enough to be cited the way atlan.com (3 of 6) and alation.com (4 of 6) were.

We would also test the wording, not just the category. Being the "affordable" answer (8 names) is a different fight from being the "free" answer (4 open-source names), and a tool can win one and be shut out of the other. If you sell a paid catalog, "best free" may simply not be your race, and that is worth knowing rather than guessing.

The full list, counted

ProductNamed inShare
Atlan5 of 683%
Alation4 of 667%
Amundsen3 of 650%
Collibra Data Catalog3 of 650%
Talend3 of 650%
Apache Atlas2 of 633%
Ataccama2 of 633%
Collibra2 of 633%
Data.world2 of 633%
DataHub2 of 633%
Informatica2 of 633%
Informatica IDMC2 of 633%

Nine more products were named once each (17%): Alation Data Catalog, AWS Glue, Coalesce Catalog, Dataedo, Marquez, Metaphor, Microsoft Purview, OpenMetadata, and Secoda.

Where Glotier fits

Glotier does not sell a data catalog, so we are correctly nowhere in these 6 answers, and we would not want to fake our way in. But we ran this exactly the way we run it for a customer's own category: six buyer questions, the real assistants, one live search each, every product and every source written down. The finding that Atlan sits in 5 of 6 answers while G2 and Capterra sit in none of them is the kind of thing you can only see by measuring, not by guessing. If you want the same read on your own category, the check is free, needs no account, and takes about a minute.

The questions we asked

One live web search per question, put to serper+or:chatgpt,gemini on 2026-08-04. 6 of 6 came back with an answer we could read. Whether a product was named is decided by looking for it in the answer text, not by asking a model for its opinion.

  1. best data catalog software
  2. best data catalog software for small teams
  3. what data catalog software should I use
  4. best free data catalog software
  5. most affordable data catalog software
  6. data catalog software recommendations

Questions people ask

What is the best data catalog software according to ChatGPT and Gemini?
In our August 4, 2026 measurement, Atlan was named in 5 of the 6 answers (83%), more than any other tool. Alation followed at 4 of 6 (67%), and Amundsen, Collibra Data Catalog, and Talend tied at 3 of 6 (50%). This reflects how often each was named across six buyer questions, not a quality ranking.
What is the best free data catalog software?
For the specific question "best free data catalog software," ChatGPT and Gemini named only open-source projects: DataHub, OpenMetadata, Apache Atlas, and Amundsen (4 names in total). None of the paid leaders (Atlan, Alation, or Collibra) appeared in the free answer, even though they dominated the other five questions.
What is the most affordable data catalog software?
The "most affordable" question returned 8 names, and they were mostly commercial vendors, not the open-source free set: Alation, Collibra, Informatica, Atlan, Talend, AWS Glue, Data.world, and Ataccama. AWS Glue showed up only for this question. Worth noting, "free" and "affordable" produced two different lists, so it is worth measuring each separately.
Do ChatGPT and Gemini use G2 or Capterra to recommend data catalog tools?
Not in this measurement. G2 and Capterra were not cited in any of the 6 answers. The sources were community threads (reddit.com and quora.com each fed all 6 answers, stackoverflow.com fed 1), roundup blogs (montecarlo.ai fed 5 of 6), a code host (github.com, 1), and vendors' own websites.
How do I get my data catalog software recommended by AI assistants?
Be present in the sources the answers actually pull from. The two most-named tools each had their own domain in the source set: atlan.com fed 3 of 6 answers and Atlan was named in 5, alation.com fed 4 of 6 and Alation was named in 4. Reddit and Quora fed all 6. In this data the source list, not market share, decided who got named.

Do you sell in data catalog software? Find out whether you are in that list.

Paste your domain and watch the same run happen for your own buyer questions: which of the three assistants names you, who gets named instead, and the exact pages those answers were built from. Free, no card, no account for the first check.

Show my visibility

For reference, Atlan was named in 5 of the 6 answers we read.

Start with Solo

Get Glotier Solo

Everything on this page is one measurement, taken by hand, on one day. Solo runs it for your product every day and writes the work it points to.

  • The Citation Agent: which of your pages an assistant can actually cite, and the fix for each
  • A paragraph-quality read of your pages: what an assistant can lift whole, and where a claim is missing its source
  • An X agent and a Reddit agent: the live threads worth answering in your category, with a reply drafted for each
  • A daily check, and an article written from what it measured that day, ready to publish
  • The same buyer questions re-asked every day, with every source page behind each answer

Track 3 products and up to 150 buyer questions across ChatGPT, Gemini and Perplexity. $39/month.

Cancel any time. Not for you? Email us within 7 days of a charge and we refund it in full. Refund policy

Not ready to pay? The check is free, with no card and no account. Run it on your own product.