THESYNTHETIC AGENT
The Synthetic Agent / Essay

AI Search Doesn’t
Agree on What
Authority Is.

A September study of 960 AI answers found major differences in which sites ChatGPT, Perplexity, Gemini and Claude cite; and striking instability between repeated runs.

Authority is not one leaderboard. It is a consensus being assembled in real time.
Here’s the point

There is no single “AI ranking.” ChatGPT, Gemini, Claude and Perplexity can answer the same question with different sources, and even the same engine can change its sources from one run to the next. AI visibility is probabilistic, not a fixed leaderboard.

I keep seeing people talk about “ranking in ChatGPT” like somebody secretly rebuilt Google with a different logo.

That is not what is happening.

There is no universal AI results page.

There is no stable position three.

And there is definitely no magic schema field that makes four different models suddenly agree you are the authority.

One web. Four engines. Different evidence.

A September 2026 study from Prefer sent the same 80 questions to ChatGPT, Perplexity, Gemini and Claude three times each. That produced 960 answers and 1,329 different cited sites.

The number that jumped out at me was not how often the engines searched.

It was how little they agreed.

966 of the 1,329 cited sites; 72.7%; were cited by only one engine. Just 29 sites, or 2.2%, were cited by all four.

Same internet. Same general questions. Completely different source mixes.

That should kill the idea that GEO is just “SEO but for ChatGPT.”

The platforms have different habits.

Perplexity cited far more sources per answer than ChatGPT in the study. Gemini leaned heavily on YouTube. Reddit appeared frequently in Perplexity and Gemini, but Claude did not cite Reddit at all in the dataset. LinkedIn showed up heavily in Perplexity and almost nowhere else.

That is not a minor implementation detail.

It means a brand can look strong inside one model’s evidence ecosystem and almost invisible inside another.

If you optimize for one engine’s habits, you can accidentally overfit the wrong machine.

It gets weirder: the same engine can disagree with itself.

The researchers repeated every question three times.

Two runs of the same question shared only about 37% of cited sites on ChatGPT and 35% on Gemini. Perplexity was dramatically more consistent at 96% in this particular study.

So if somebody screenshots one AI answer and declares themselves “#1 in ChatGPT,” I would not take that very seriously.

One answer is an anecdote.

A useful measurement system needs repeated runs, multiple prompt phrasings, multiple engines, dates and enough observations to distinguish a pattern from a lucky retrieval.

Authority is becoming probabilistic.

This is the part I find more interesting than rankings.

Traditional search trained us to think of authority as a relatively stable score that expresses itself through position.

AI systems assemble authority at answer time.

They retrieve. They compare. They decide which sources belong in the response. The set can change by engine, prompt, model version, location and even repeated run.

That makes authority less like owning a position and more like building enough evidence that you keep getting invited into the room.

The biggest sources still have gravity.

The study was not pure chaos.

Prefer found that 23 of the 25 most-cited sites appeared across three or four engines.

That matters.

The long tail was fragmented, but the strongest sources traveled.

My takeaway is not “AI citations are random.” My takeaway is that broad authority creates cross-model gravity.

Big, trusted, frequently corroborated sources are easier for multiple systems to agree on.

For everybody else, the job is to build that kind of gravity deliberately.

What actually travels across models?

I would focus on things that remain useful no matter which engine wins the week.

Original information. Clear authorship. A coherent entity. Third-party corroboration. Firsthand experience. Strong topical depth. Real references from other sites. Consistent facts across the web. Pages that answer the damn question instead of circling it for 1,500 words.

Not because those are secret AI hacks.

Because they are the ingredients of evidence.

This is why “best” is a dangerous word.

When an AI model recommends a professional, company or product, the output can look authoritative enough that people treat it like a ranking.

It is not.

It is a generated answer produced from a changing retrieval process.

That does not make the answer useless. It means we should stop pretending one model response is an objective scoreboard.

The research I want to do through The Synthetic Agent is built around that distinction. I care less about who “wins” one prompt and more about what evidence repeatedly makes an entity understandable enough to appear across models, prompts and time.

That is a much harder question.

It is also the one that actually matters.

Sources checked Sep 24, 2026

Prefer: How AI engines search and cite: 960 answers measured