In Study 001 Wave 1, four web-grounded API model families produced substantially different real-estate recommendation sets. Across 72 observations, mean cross-model set overlap was 6.65%; repeat-run set overlap was 33.62%. These describe the tested setup, not recommendation quality or direct behavior in the consumer apps.
Ask a system to recommend a real-estate professional and the answer can arrive with the tone of a settled conclusion. Names. Reasons. Sources. Enough confidence to make the result feel like the end of the investigation.
We wanted to know what happens when you ask again, ask another system, or change what the buyer needs.
The useful finding was how much the recommendation sets moved. The full report has been published since September 25. This essay explains how to read it without turning a descriptive study into a claim it cannot support.
What we actually tested.
Wave 1 used six frozen Scottsdale real-estate recommendation prompts across four model families, with three repeated calls for every model-and-prompt combination. That is 6 × 4 × 3 = 72 primary observations.
The prompts covered broad recommendations and buyer, seller, luxury, relocation and neighborhood-expertise contexts. The collection used stateless API requests through Vercel AI Gateway with provider-native search or grounding.
The frozen endpoints were OpenAI GPT-5.6 Sol, Anthropic Claude Sonnet 4.6, Google Gemini 3.1 Pro Preview and Perplexity Sonar Pro. Those names identify this execution. They are not a claim that the consumer products use the same configuration or would give the same answers.
There was an important deviation. The public preregistration described consumer-facing applications and a separate evidence follow-up. Wave 1 used APIs and did not send that follow-up as a second turn. The published addendum records this difference. The original protocol remains preserved.
Different systems assembled different shortlists.
The mean cross-model recommendation overlap was 6.65%. To calculate it, the analysis pooled all names from the three repeats for each model and prompt, compared those sets between model pairs, then averaged the pairwise overlaps across the six prompt families.
That pooling detail matters. This is not the fraction of 72 individual answers that agreed. It is not the percentage of recommendations that were correct. It measures how much the model-level sets shared under the published method.
A reader looking for one universal list of “the best agents according to AI” should pause here. In this setup, the tested systems did not converge on one common set. That does not tell us which system had better judgment. We did not independently evaluate the professionals’ suitability.
What does 6.65% overlap mean?
Imagine one set contains A, B and C, while another contains B, C and D. They share two names, and there are four distinct names across both sets.
2 ÷ 4 = 50% Jaccard overlap.
Both lists could contain excellent choices. Both could contain poor choices. The overlap calculation cannot tell the difference.
For the same reason, subtracting the study’s 6.65% from 100% does not produce an error rate. “93.35% wrong” would be an invented conclusion. Set similarity answers a question about agreement, not truth.
It also ignores order. Two lists with the same names in reverse order have complete set overlap. That is why this research should not be repackaged as a position-tracking benchmark.
Repeating the question did not settle it.
Overall repeat-run recommendation stability was 33.62%, again expressed as mean Jaccard overlap. This calculation compares pairs of individual runs within the same model and prompt, then averages those results.
The cross-model and repeat-run figures therefore describe different comparisons. One pools names across repeats before comparing systems. The other compares the repeated answers themselves. Putting the two percentages side by side without that explanation invites an easy misreading.
What can we say? The same exact prompt, repeated under the study setup, did not reliably reproduce an identical shortlist. What can we not say? That a particular professional has a 33.62% chance of being recommended tomorrow. This statistic is not an individual appearance probability.
Changing the task changed the candidate pool.
The six prompt families also produced low mean cross-prompt overlap within each model family. A broad question, a relocation question and a neighborhood-expertise question did not surface the same pool.
Some of that variation may be appropriate. A relocating buyer and a luxury seller have different needs. Consistency is not automatically good if it means the system ignores the question.
The practical implication is narrower: the wording and context are part of the measurement. A business appearing in one broad prompt cannot assume the same visibility for every customer need. A monitoring program should explain which needs its prompts represent.
The citations tell another part of the story.
The dataset contained 108 unique citation domains. Portals and directories formed the largest coded source class, with 417 source-class assignments, compared with 149 first-party and 106 news/editorial assignments. These are assignments across the coded evidence, not 417 unique directory websites.
The overall repeat-run citation-domain overlap was 55.46%. That is a separate calculation from recommendation overlap. A system can use similar source domains while changing the people it recommends; source consistency and shortlist consistency should not be treated as the same property.
A citation is visible evidence attached to an answer. It does not prove that the cited page caused the recommendation. We did not isolate causal signals or test a recipe for making a particular professional appear. More frequent citation also does not establish better source quality.
What I would do with this as an operator.
First, stop treating one favorable answer as durable market position. Save it as an observation. Repeat the question, record the conditions and inspect whether the pattern persists.
Second, separate recognition from recommendation. Being named, being recommended and having your website cited are different events. Their usefulness depends on the decision you are trying to make.
Third, investigate factual errors. If a system gets your location or services wrong, that is a concrete discrepancy you can trace through the visible evidence. The study does not promise that correcting one page will change the answer, but it gives you a better question than “How do I become number one?”
The next test needs to address the gap.
Wave 1 is a dated, limited observation of API behavior in one market and six prompt contexts. It is not a national benchmark, a consumer-app leaderboard or a causal model of authority.
Study 001B is planned as a separate consumer-interface replication. Its job is to test the experience the original protocol intended to observe, with documented session and product state. It has not been established here as completed, and this essay adds no new observations to Wave 1.
The confidence of an answer should not end the investigation. Our first wave gives us a reason to look underneath it: at the question, the repeated outcomes, the visible sources and the limits of what we actually measured.
Research companion published October 9, 2026. Explains the frozen September 25 Wave 1 results; no new collection or change to the research evidence.
