“AI ranking position” is mostly a fake precision problem. Generative answers are not a stable ordered SERP. Different engines cite different sources, and repeated runs of the same question can change the evidence set. One screenshot is not a ranking.
Somebody is going to sell you “position tracking for ChatGPT” this week.
I am not saying the tools are useless.
I am saying the word position is doing a lot of dishonest work.
A generated answer is not ten blue links.
Traditional rank tracking works because the thing being measured is ordered. A page appears in a position. The position can move. You can sample it repeatedly and argue about personalization later.
AI answers are assembled.
The model may search. It may not. It may cite three sources or nineteen. It may name a company without citing its site. It may use one source to establish a fact and another to support the recommendation.
That is not a list.
It is a synthesis.
The engines do not even agree on the evidence.
Prefer’s September 2026 study ran 80 questions three times each across ChatGPT, Perplexity, Gemini and Claude. Across 960 answers, 1,329 different sites were cited.
72.7% of those sites were cited by only one engine.
Only 2.2% appeared across all four.
If four engines can answer the same question from substantially different source sets, what exactly is “position three” supposed to mean?
The same engine can change its mind seconds later.
The study repeated each question three times.
Two runs of the same question shared only 37% of cited sites on ChatGPT and 35% on Gemini. Claude was more consistent at 68%. Perplexity was dramatically more stable at 96% in that dataset.
That means one run is not a measurement.
It is a sample.
And one screenshot is barely even that.
There are things worth tracking.
This is where I do think the tooling gets interesting.
I would track whether the entity appears at all, how often it survives repeated runs, what the model says about it and which sources keep showing up underneath the answer. I would also keep an eye on how much the result changes when the prompt, model or run changes.
That tells me something real. A single “position” number usually does not.
I do not care what we call the metric.
Just do not smuggle SERP certainty into a system that does not behave like a SERP.
A useful AI visibility metric should expose how it was produced: models, prompts, runs, dates, geography, retrieval settings and the raw answers underneath the summary.
If it cannot show you that, the number is decoration.
Google’s own documentation points the same direction.
Google says AI Overviews and AI Mode can use different models and techniques, so the responses and links they show will vary. Google also says there is no special technical optimization required beyond the same foundational SEO work that makes pages eligible for Search.
That should make us less obsessed with inventing a new ranking primitive and more obsessed with building better evidence.
There is a better question than “What position am I?”
Ask this instead:
How reliably does this system understand and surface me when the question should reasonably include me?
That question admits uncertainty.
It also gives us something useful to measure.
Prefer: 960-answer AI citation study
