THESYNTHETIC AGENT
Recommendation behavior
Results published72 observations4 model families6 prompts3 repeats

How AI Chooses Who to Recommend

Ask several AI systems the same real recommendation question. Ask it again. Change the context slightly. The apparent certainty starts to come apart.

The finding

In Wave 1, the same recommendation question did not have one stable machine answer.

Wave 1 produced 72 web-grounded API observations across OpenAI GPT-5.6 Sol, Anthropic Claude Sonnet 4.6, Google Gemini 3.1 Pro Preview and Perplexity Sonar Pro. The models often surfaced different professionals, repeated runs often changed the shortlist, and changing the user context changed it again.

Important execution note.

The public Protocol v1.0 described consumer-facing products and a separate evidence follow-up. Wave 1 used frozen API models through Vercel AI Gateway with native search/grounding, and the evidence follow-up was not sent as a separate second turn. The preregistration has not been rewritten. Read the full execution addendum.

6.65%

The models rarely surfaced the same recommendation set.

Mean cross-model recommendation overlap was 0.0665 Jaccard across the six prompt families. This measures shared names between model outputs. It does not measure whether one recommendation was better than another.

Broad
7.13%
Buyer
7.60%
Seller
6.64%
Luxury
11.25%
Relocation
3.63%
Neighborhood
3.65%
33.62%

Repeating the exact prompt did not reliably reproduce the same shortlist.

Overall repeat-run recommendation stability was 0.3362 Jaccard. Each model was asked the same prompt three times under the same study setup. The numbers below are descriptive repeat-overlap measures, not model scores.

OpenAI
21.01%
Claude
50.88%
Gemini
9.48%
Perplexity
53.11%
Every system produced 18 unique provider request IDs and 18 unique answer hashes.

A small change in context changed who appeared.

Buyer, seller, luxury, relocation and neighborhood-expertise wording materially changed the candidate pool within each system. Mean cross-prompt recommendation overlap stayed low across all four model families.

OpenAI
13.79%
Claude
22.49%
Gemini
16.13%
Perplexity
27.81%

The visible evidence came from a recurring web ecosystem.

Portals and directories were the largest coded source class. Citation frequency is not a quality score, and a visible citation does not prove that a page caused the recommendation.

homelight.com43
fastexpert.com42
zillow.com40
listwithclever.com36
effectiveagents.com36
expertise.com36
realtrends.com32
homes.com28
yelp.com27
realtor.com23
Source-class assignments: 417 portal/directory · 149 first-party · 106 news/editorial · 13 government/regulatory · 10 social · 17 other.

The citation set itself behaved differently by model.

Mean repeat-run citation-domain stability was 0.5546 overall. This is overlap in visible citation domains across repeated runs, not an assessment of source quality.

OpenAI
32.34%
Claude
80.53%
Gemini
11.59%
Perplexity
97.39%

The subject is the machine, not the people it named.

Study 001 does not publish a “best Scottsdale agent” table. Names were captured because overlap cannot be measured without them. The findings describe model behavior: disagreement between systems, volatility across repeats, sensitivity to prompt context and visible source choice.

Wave 1 also does not establish which model was “best,” whether citation frequency equals recommendation quality, whether a cited page caused a recommendation, or whether the API behavior is identical to the consumer applications.

The confidence of the answer is not evidence that the recommendation itself is stable.

That is the useful result from Wave 1. Different systems, repeated runs and small changes in context produced different recommendation sets from different visible evidence. The next replication should test the original consumer-interface protocol directly.