back to writing

SpotAQ / AI Visibility / Product

I built a tool to measure AI visibility. Turns out it had been lying to me.

I pointed SpotAQ at itself, got a zero, and audited the measurement code. 38% of the score was measuring my own infrastructure and the shape of the category, not how AI talks about the brand.

I'm building SpotAQ, a tool that checks whether AI assistants recommend your product. It puts real buyer questions to ChatGPT, Perplexity and Google AI, then reports whether you got named, whether you got recommended, and which sources the answer was built from.

Last week I finally pointed it at itself. Across 24 answers, AI recommended us zero times.

My first reaction wasn't "we're invisible." It was "can I even trust this number?" So I went through the measurement code. I found four problems, and I think every one of them is a trap you can fall into any time you build your own metric.

1. The more carefully a user configured it, the worse their data got

Users can set keywords they want tracked. But in the code those keywords replaced the base question set instead of adding to it. A project that tracked one keyword went from 5 questions down to 1. Three answers across three engines, so the mention rate could only ever read 0%, 33%, 67% or 100%.

A real improvement worth five points was mathematically invisible. And the users who configured the tool most carefully got the noisiest data. Existing tests never caught it because they only ever asserted prompts[0].

2. Category structure was leaking into the score

18% of the total came from "competitor pressure", computed as 100 - distinctCompetitors * 12.

How many tools an AI lists in one answer is a function of the category and the model's answer style. It is not a measure of your performance. Nine competitors zeroed that component outright. Worse: if you were doing well and got listed among the recommendations, the other tools in that list dragged your score down.

Competitor names come back as free text from an LLM, and nothing normalised them. One rival spelled "ChatGPT", "chatgpt ", "ChatGPT." and "chatgpt.com" counted as four competitors. 48 points, one company.

3. Failed requests were recorded as "the AI didn't mention you"

Three engines, and only one of them isolated per-request failures. The other two batched with Promise.all and no per-item catch, so a single flaky prompt rejected the whole batch. The rows that did get written carried mentioned: false.

So an API timeout and a genuine visibility drop produced the same number. On a product whose entire pitch is "we'll show you whether your work moved the needle."

4. And 8% of the score was API uptime

There was a "confidence" component measuring how many prompts returned an answer at all. Infrastructure health, sitting inside a customer-facing score.

Sentiment was another 12%, but it scored "unknown" as 45 and defaulted to 40 with no data, so mostly it measured whether extraction succeeded.

38% of the score was measuring my own infrastructure and the shape of the category, not how AI talks about the brand.

The fix that mattered most

The fix that mattered most wasn't reweighting. It was this: drop a component entirely when there's no evidence for it and re-normalise the rest, instead of substituting a default. That 40-point fallback for missing sentiment was the whole problem in miniature — an invented number that looked exactly like a real one.

Concretely:

  • competitor pressure became "share of answers where a rival wins and you don't"
  • confidence left the total and is reported separately as coverage
  • failed requests are excluded from every metric and surfaced as "this run covered 79%"
  • keywords append to the base set instead of replacing it

The composition changed, so old scores aren't comparable to new ones. I bumped a measurement version, history segments on it automatically, and the UI says plainly that the baseline reset is a methodology change rather than lost data.

That zero is now a zero I trust.

The transferable bit

Every so often, ask "if this number moves, what could have caused it?" and write out the full list. If anything on that list isn't something your user can act on, it doesn't belong in the score. Mine had three such things and they added up to more than a third of it.

Happy to go into any of it in more detail. Still figuring out how to climb off the zero, which is a different problem.

I am applying this while improving SpotAQ: making AI visibility easier to understand, improve, and act on.

Visit SpotAQFollow the build