THESYNTHETIC AGENT
The Synthetic Agent / Essay

No SEO Tool Can Give
You a Truthful AI
Visibility Score Yet.

AI visibility can be measured, but one opaque score hides model volatility, prompt selection, repeated runs and methodological choices.

The problem is not measurement. It is fake precision.
Here’s the point

No SEO tool can reduce AI visibility to one fully truthful score yet. The underlying systems vary by model, prompt, run, interface, geography and time. Tools can measure samples extremely well. The problem starts when the sample is presented as objective reality.

I want AI visibility tools to exist.

I want them to get really good.

I also want us to stop pretending the measurement problem is solved because somebody put a number from 0 to 100 in a circle.

The thing being measured is unstable.

AI answers are non-deterministic enough that repeated runs can cite different sources. Different engines can use completely different evidence. Consumer apps can behave differently from API calls. Search features can change by geography, device, account state and product version.

That does not make measurement impossible.

It makes methodology mandatory.

A score hides choices.

Every AI visibility score has choices buried underneath it.

Which models?

Which prompts?

How many runs?

What date?

What location?

Was web search forced?

Did a brand mention count the same as a citation?

Did a negative mention count as visibility?

Did an answer naming five companies score the same as one naming only you?

Was the first-party website required?

Those choices are the method.

If the dashboard hides them, the score looks more objective than it is.

The Prefer study shows why repeated measurement matters.

In 960 answers, two runs of the same question shared only 37% of cited sites on ChatGPT and 35% on Gemini. Even when answers are generated within seconds of each other, the source set can move.

So a tool that runs one prompt once and reports “visibility: 0” has learned almost nothing.

The honest product is a measurement system, not an oracle.

A strong tool can absolutely tell you:

How often you appeared across a defined prompt set.

How frequently you were cited.

Which domains supported the answer.

How results changed across models.

How volatile those results were across repeated runs.

How the trend moved over time under the same protocol.

That is useful as hell.

Just call it what it is.

Visibility is not authority. Authority is not conversion.

This is another place a single score collapses too much.

A model can know your company and never recommend it.

It can recommend you and cite somebody else.

It can cite you frequently and send no traffic.

It can send traffic that never converts.

Those are different stages.

Measure them separately.

The right dashboard should expose the raw evidence.

When I build the research for The Synthetic Agent, I want the summary to be clickable all the way down to the prompt, model, run, raw answer and cited URLs.

If a number surprises you, you should be able to inspect how it happened.

That is the standard.

Not fake precision.

Inspectable precision.

Sources checked Sep 24, 2026

Prefer: 960-answer AI citation study

Prefer: repeat-run answer volatility

Google: generative AI Search Console reports