Guide

What an AI visibility score actually measures

There is no industry definition, so two tools can hand you 34 and 71 for the same product on the same day and both be right by their own maths. Here is how to tell which kind you are looking at, including when it is ours.

Why the number on its own is not the point

A score is a compression. Somebody decided which questions to ask, which assistants to ask, in which country, and how to weight what came back, and then handed you one figure. Every one of those decisions changes the result, and none of them is visible on the dial.

The practical consequence is that comparing scores across two tools is meaningless. The only comparison that means anything is one tool’s number against its own number a month ago, and even that only holds if the question set was frozen in between. A tool that quietly changes its questions has changed its scale, and your improvement is an artefact.

The four questions to ask before trusting one

  1. 1

    What exact questions did you ask, and can I see them?

    If the question set is hidden you cannot tell whether it reflects how your buyers actually search or whether it was chosen to flatter somebody. Ten questions you can read beat a hundred you cannot.

  2. 2

    Which assistants, and did you run a live search?

    An answer from a model's training data describes the web as it was months ago. An answer from a live web search describes it now. These are not the same measurement and the difference is invisible in a score.

  3. 3

    Can I see the source pages behind each answer?

    This is the one that separates a dashboard from something you can act on. The score tells you that you are losing; the source list tells you where. Most tools stop at the score.

  4. 4

    Is the question set frozen between runs?

    If it is regenerated each time, week-over-week movement is noise. Ask directly, because a regenerated set is easy to ship and easy not to mention.

Those four work on us as well as on anyone else, and you should use them that way.

A count beats a composite

The alternative to a score is embarrassingly simple: a count, with the working attached. You asked twenty-one questions. You were named in three. Here are the twenty-one questions, here are the answers verbatim, and here is every source page each answer was built from.

A count cannot be tuned. Nobody can adjust a weighting to make three out of twenty-one look better. It is also immediately actionable in a way a dial is not, because the eighteen you lost come with the pages that beat you attached.

When we measured 3 categories end to end on 26 July 2026, that structure is what made the findings usable: 18 answers, 131 cited sources read one by one, and a different gatekeeper deciding each category. No score would have surfaced that. It would have compressed three completely different situations into three similar-looking numbers.

Our own number, since it is the fair test

Across 21 buying questions in our own category, AI names Glotier in 0 of them. We publish it on the product pages as well as here. If we ran a composite score we could almost certainly present that more kindly, which is the argument against composite scores in one sentence.

Do this automatically

Get Glotier Solo

A score you cannot audit is somebody's opinion with a decimal point. Solo gives you the count, the frozen question set, the verbatim answers and every source behind them, re-run every day so movement means something.

  • The same buyer questions re-asked every day, so you see the moment an answer changes
  • Every source page behind each answer, which is where the work actually is
  • An article written from what it measured that day, ready for you to publish
  • The subreddits and threads worth answering, with a reply drafted for each

Track 3 products and up to 150 buyer questions across ChatGPT, Gemini and Perplexity. $39/month.

Cancel any time. Not for you? Email us within 7 days of a charge and we refund it in full. Refund policy

Not ready to pay? The check is free, with no card and no account. Run it on your own product.

Questions people ask next

What is an AI visibility score?
A single number a tool puts on how often AI names your brand. There is no industry definition, so two tools can hand you 34 and 71 for the same product on the same day and both be internally consistent. It is a vendor's opinion expressed as a figure, which is why the formula matters more than the number.
Is a good AI visibility score worth anything?
Only if you know what went into it and it is measured the same way every time. A score built from a fixed question set, run on a schedule, is genuinely useful for direction. A score that blends mentions, sentiment, share of voice and an authority estimate into one figure tells you almost nothing you can act on, because when it moves you cannot tell which part moved.
Why do different tools give me completely different scores?
Because they ask different questions, of different assistants, in different countries, and weight the results differently. The assistants themselves disagree far more than they agree, so even two honest tools asking the same question of a different model will diverge. Comparing scores across tools is meaningless; comparing one tool's score against itself over time is the only thing that works.
What should I ask a vendor before trusting their score?
Four things. What exact questions did you ask, and can I see them. Which assistants, and did you run a live web search or answer from training data. Can I see the source pages behind each answer. And is the question set frozen, so this week's number is comparable to last week's. A vendor who will not answer all four is selling you a dial.
What does Glotier show instead?
A count and the working. You asked N questions, you were named in M of them, here are the answers verbatim and here is every source page each answer was built from. Our own category currently reads 0 of 21, which is not flattering and is published anyway. A count cannot be tuned; a composite can.

Keep reading