THESYNTHETIC AGENT
How the experiments work

Ask it again.

AI answers can look definitive even when the underlying behavior is not. The only way to see that clearly is to repeat the work and keep the evidence.

What gets studied

The object of study is the AI system: what it recommends, what it cites, how often it agrees with another system, and how much its own answer changes from one run to the next.

Use a real question.

I want the prompts to resemble something a person might actually ask, not a laboratory sentence designed to force a clean result. Each study will publish the exact prompt set it used.

Run the same question more than once.

One answer can be interesting. It is not enough to tell me whether the behavior is stable. Every formal study uses repeated runs so one lucky or strange answer does not become the headline.

Compare different systems under the same conditions.

When a study compares ChatGPT, Gemini, Claude, Perplexity or another system, the prompt wording stays fixed within that comparison. Model names, versions and interface details are recorded when the provider exposes them.

Then change the wording on purpose.

People do not all ask the same question the same way. A second part of the experiment uses controlled paraphrases so we can see whether a small wording change produces a meaningfully different recommendation or source set.

Keep the answer before interpreting it.

The raw response is the record. Names, citations, source URLs, visible search behavior, caveats and refusals are captured from that response before any summary metric is calculated.

Measure behavior, not human quality.

If an AI system recommends somebody, that tells us the system recommended them under those conditions. It does not prove the person is better at their job. These studies do not turn model output into a professional-quality score.

Sources matter because the systems do not all use the same ones.

When citations are available, I record where they came from and what kind of source they are. First-party websites, directories, news, reviews, social platforms and other source classes can be compared without pretending a citation automatically caused the recommendation.

Null results stay in.

If a model refuses, gives no recommendation, produces no citations or contradicts itself, that belongs in the study. Cleaning those cases out would make the result easier to explain and less honest.

Publish what somebody else would need to challenge it.

Each study should make the prompt set, dates, systems, run count, aggregation method and known limitations public. The goal is not to make the work look bulletproof. The goal is to make it inspectable.

What this method cannot prove.

AI products change constantly. Consumer apps and APIs may behave differently. Location, account history and personalization can affect results. A visible citation does not prove how much a source influenced the final answer. A pattern across systems is still a pattern, not proof of causation.

The answer is the observation. The pattern only appears after we ask again.

Current status: Study 001 is using this method to examine recommendation behavior in a bounded Scottsdale real-estate scenario. The study measures the models, not the agents.