Guide
How to check an AI visibility tool is telling the truth
Six tests, ten minutes, before you pay anybody. Run them on us as well. The last section is the two we would fail.
Why this is worth ten minutes
Every tool in this category is reporting on something you cannot see. You are not in the conversation the assistant had, you did not run the search, and you have no independent record of what came back. That is an unusual amount of trust to hand over for a monthly fee, and unusually easy to check.
Fabricating an assistant’s answer would be trivial and undetectable if you never look at one. We are not suggesting anyone does. We are pointing out that the difference between a tool that shows its working and one that does not is the difference between a claim and an audit.
The six tests
- 1
Ask to see the raw answer, verbatim
Not a summary, not a sentiment label, the assistant's actual reply. Any tool doing this work has it. A tool that will only show you a processed number has decided what you are allowed to verify, and that decision tells you more than the number does.
- 2
Ask for the source URLs behind each answer
This is the test most tools fail, and it is also the only output you can act on. A score tells you that you are losing. The source list tells you where, and it is the difference between a dashboard and a plan.
- 3
Ask whether the question set is frozen
If questions are regenerated each run, week-over-week movement is noise dressed as progress. Ask directly. A regenerated set is easy to ship and easy not to mention, and it quietly makes every trend chart meaningless.
- 4
Ask whether it ran a live search or answered from training data
A model answering from memory describes the web as it was months ago. That is a different measurement from a live retrieval, and both get sold as AI visibility. It changes what the number means entirely.
- 5
Reproduce one result by hand
Open the assistant, ask one of the questions yourself, and compare. Answers vary run to run so it will not match exactly, and it does not need to: you are checking whether you are in the same universe, not auditing to the decimal.
- 6
Ask what they publish about themselves
A vendor in this category can measure their own visibility as easily as yours. If they never mention their own number, that is usually because it is not flattering, which is fine, but it means they are asking you to be more transparent than they are.
Run them on us
The free check shows the verbatim answer from each assistant and every source URL behind it, with no account and no card, so tests one, two, four and five take one query. The question set on a paid plan is frozen, which covers test three. And test six is this: across 21 buying questions in our own category, AI names Glotier in 0 of them. That number is on our pricing page too, not only in a guide where a buyer would never look.
The two we would fail
We measure ChatGPT, Gemini and Perplexity. We do not measure Google AI Overviews, so if that is your main channel, our numbers are inference rather than measurement and you should weight them accordingly.
And our published category studies are 18 answers across 3 categories, measured on 26 July 2026. Reading all 131 sources by hand is what makes them useful and also what keeps them small. That is enough to disprove a confident claim and not enough to make one, which we say on the studies themselves rather than only here.
If a vendor cannot name two things their own data does not cover, they have not looked hard enough or they are not telling you.
Do this automatically
Get Glotier Solo
Every test on this page can be run against our free check before you pay anything: verbatim answers, source URLs, no account. Solo is what you buy once you want that same measurement repeated on a frozen question set every day.
- The same buyer questions re-asked every day, so you see the moment an answer changes
- Every source page behind each answer, which is where the work actually is
- An article written from what it measured that day, ready for you to publish
- The subreddits and threads worth answering, with a reply drafted for each
Track 3 products and up to 150 buyer questions across ChatGPT, Gemini and Perplexity. $39/month.
Cancel any time. Not for you? Email us within 7 days of a charge and we refund it in full. Refund policy
Not ready to pay? The check is free, with no card and no account. Run it on your own product.
Questions buyers ask
- How do I know an AI visibility tool is telling the truth?
- Ask to see the raw answer. Any tool measuring this has the assistant's reply in hand, so showing it costs them nothing, and a tool that will only show you a processed score has made a choice about what you are allowed to check. Then reproduce one result yourself: open the assistant, ask the same question, and compare.
- Why do two tools give me different numbers for the same brand?
- Different questions, different assistants, different countries, and different rules for what counts as a mention. All four are legitimate reasons to differ and none of them is visible in a headline figure. This is why comparing two vendors' scores is meaningless and comparing one vendor's number against its own history is the only thing that works.
- Can these tools fake the data?
- Fabricating an assistant's answer would be easy and hard to detect if you never see the answer, which is exactly why seeing it matters. A tool that shows you the verbatim reply and the source URLs has made itself checkable. Nothing else in the buying process tells you this much for ten minutes of work.
- What sample size should I expect?
- Smaller than you would like, everywhere, including here. Every question costs a real API call, so nobody is running thousands. What matters more is whether the set is frozen between runs. Ten questions asked identically every week is a measurement; a hundred regenerated each time is not.
- What does Glotier fail at?
- Two things worth knowing before you buy. We measure ChatGPT, Gemini and Perplexity, not Google AI Overviews, so any read-across is inference. And our published category studies are 18 answers across 3 categories, taken on one day: enough to disprove a sweeping claim, not enough to make one. We say so on the pages themselves.
Keep reading