The seductive answer is neat because nobody had to live it
Ask a model to role-play a first-time parent, a price-sensitive commuter or a loyal customer and it will often produce fluent reasons, objections and language. That fluency is productive in an early workshop. It can reveal a missing assumption, force a team to state a segment and generate counterarguments. It becomes risky when the output is summarized as what customers think.
A customer has a context: an actual budget, an actual phone, a distracted evening, a competing offer, a local reference, a private discomfort and the freedom to ignore your category completely. A synthetic persona has instructions and a learned pattern of text. Those are different evidence sources, even if their sentences look similarly plausible.
Research does not support a shortcut to representativeness
Studies on language models as survey respondents show why confidence needs a boundary. A 2023 evaluation found response-order and labeling biases in common survey-prompting methods. A later cross-domain preprint reports that the tested models exaggerated demographic differences and could direct segment decisions wrongly in its settings. These are bounded research findings, not proof that every simulation fails. They are enough reason not to call generated responses a representative customer sample.
The failure mode matters creatively. A model tends to explain a segment as a clean type. Real people are inconsistent; a new campaign can encounter a reaction the model has not seen, and no persona prompt can turn that absence into field evidence.
Source context: Questioning the Survey Responses of Large Language Models · When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Use simulation for friction, not validation
Give the tool roles that do not pretend to discover a market. It can turn a draft interview guide into possible misunderstandings, list the assumptions inside a positioning statement, suggest edge cases for an interactive experience, or help create plain-language alternatives. Ask it to show its uncertainty and prompt it with competing interpretations. Treat its output as a prompt for human research, not a vote.
Then take the most consequential unknown to people. Five recorded conversations with carefully recruited participants can overturn a beautiful simulated consensus. A short intercept around the actual asset can expose whether an image reads as intended. Behavioral data from a bounded pilot can show whether a supposed friction point actually changes completion. The right field method depends on the decision, not on the most impressive synthetic persona deck.
- Allowed early use: generate hypotheses, language alternatives, edge cases and interview probes.
- Not a substitute: incidence, preference share, willingness to pay, message recall, cultural fit or consent to a likeness use.
- Required label in internal work: ‘simulated prompts for research planning; not respondent evidence.’
Keep a two-column research board
Our recommended workflow names each item either ‘assumption to test’ or ‘observed evidence.’ A simulated audience may occupy the first column only. For the second, record who was contacted, the recruitment logic, the stimulus they saw, the question asked and the limits of the sample. This is deliberately ordinary research hygiene, but it prevents a board full of polished generative output from gaining the authority of interviews it never conducted.
The payoff is better creative judgment. Teams stop asking a fictional average person to approve work, and start asking living people the narrow questions only they can answer.
Sources & evidence limits
Source-backed facts are distinguished from the editorial workflow proposed here. Brand and agency accounts document their own work, not independent proof of performance. Read the linked source for its scope.
- Questioning the Survey Responses of Large Language Models
An empirical evaluation identifying ordering and labeling biases in LLM survey responses.
Checked 2026-09-19 · Source published 2023-06-13 - When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
A preprint’s bounded cross-domain findings on demographic over-determination and segment-targeting errors.
Checked 2026-09-19 · Source published 2026-07-28