Key takeaways
- Label generated responses as synthetic output and keep them out of the customer-evidence repository.
- Use models for hypothesis breadth, guide critique, edge cases, and rehearsal—not as proof of what people believe or do.
- Validate consequential conclusions with real participants, behavior, transactions, or other appropriate observed evidence.
- Record the model, prompt, context, sampling method, runs, selection, and human interpretation so the exercise can be audited.
01
Plausible speech is not lived experience
A generated persona can explain why “a procurement manager” rejected a message in fluent, specific language. The fluency makes the answer feel discovered. It was not observed from a procurement manager; it was produced from a model, a prompt, and the material supplied to it.
Research on synthetic HCI data shows possible utility but also explicit limitations and misuse risk. Critical work on synthetic users argues that user insight should remain traceable to real user data rather than simulation. Treat model output as an artifact from a tool, not a participant transcript.
Swipe to compare every column
| Use | What it can contribute | What it cannot establish |
|---|---|---|
| Interview-guide critique | Missing questions and assumptions | What customers will actually say |
| Scenario generation | Broader situations to investigate | Population frequency |
| Message rehearsal | Plausible objections and wording | Preference or conversion |
| Research synthesis | Structure over supplied evidence | New facts absent from evidence |
02
Keep synthetic work in a separate evidence lane
Label every output “synthetic” in the file, repository, presentation, and chart. Record the model and version, system instructions, prompt, supplied evidence, date, temperature or sampling settings where available, number of runs, selection process, analyst, and intended use.
Do not merge generated quotes with interview quotes, count agents as respondents, calculate synthetic percentages beside survey results, or describe the exercise as customer research. If the model was grounded in real studies, cite those studies and keep generated interpretation distinct from the source observations.
03
Use simulation to make human research sharper
Ask a model to identify assumptions in a discussion guide, produce counterexamples, vary situational constraints, rehearse interviewer follow-ups, or challenge an interpretation against supplied transcripts. These uses can make a researcher better prepared without pretending the model represents a population.
Turn synthetic output into a question: “Do buyers experience this constraint?” Then choose evidence capable of answering it—interviews, contextual observation, support records, product behavior, controlled experiments, or a representative survey. The appropriate method depends on the decision and stakes.
04
Report what changed in the research plan
A useful synthetic exercise should leave a visible trace: new recruitment criteria, a revised question, an edge case added to usability testing, or an assumption marked for validation. Report the tool’s contribution and the later human evidence separately.
Audit decks and generated summaries for synthetic material that has lost its label. The governance test is simple: could a decision-maker mistake this for something a real customer said, did, or preferred? If yes, the distinction is not clear enough.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- CHI 2023: Evaluating LLMs in generating synthetic HCI research data
- Jansen et al.: A critical analysis of synthetic users
- NIST: Generative AI Profile
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



