🚀

Synthetic Audience is now available on AI Creative Insights

Synthetic Users in UX Research: Possibilities, Risks, and Best Practices

Synthetic Users in UX Research: Possibilities, Risks, and Best Practices

Synthetic Users in UX Research: Possibilities, Risks, and Best Practices

Synthetic users are AI-generated profiles built on large language models that simulate how a user group might think, feel, or respond, without studying real people. They can support early hypothesis generation and desk research, but they cannot reliably replace real-user research due to sycophancy bias, shallow prioritization, and an inability to reflect lived behavioral experience.

Synthetic Users in UX Research

Tag

Technology

Date

Read Time

8 Min

Content

Senior Growth Marketer

Summary:


  • Synthetic users are AI-generated profiles built on large language models that simulate how a user group might think and respond, without studying real people.

  • They matter because teams under research pressure are already using them, often without guardrails.

  • Among the known failure modes are sycophancy, flattened priorities, the reinforcement of stereotypes, and claims relating to validation which no one can check independently.

  • Reserve their use for early hypothesis development and check this with actual participants before making any design decisions.


A product manager needs user input by Thursday. Recruiting takes three weeks. A synthetic user platform returns twelve interview transcripts in twenty minutes, and all of them sound plausible.

That is the appeal, and it is real. The problem is that plausible and accurate are different properties, and synthetic output is optimized for the first.

This guide covers what synthetic users are, where the evidence says they fail, and how to build a workflow that captures the speed benefit without letting simulated opinions quietly become design decisions.

What Are Synthetic Users in UX Research?

A synthetic user is an AI-generated profile, built on a large language model, that simulates a user group's thoughts, needs, and experiences without any real person being studied.

The distinction from a traditional persona is that synthetic users are conversational. A persona is a static document in a slide deck. A synthetic user answers back. You can interview it, probe it, and receive something that reads like a transcript from a real session.

That interactivity is exactly what makes them persuasive and what makes them risky. A document invites scrutiny. A conversation invites belief.

They are generated in two ways. Dedicated platforms package the workflow with persona management and study structure. General-purpose chatbots do the same thing with a well-written prompt, which is how unofficial synthetic research enters organizations without ever being labeled as such. The category has moved fast since synthetic users first entered UX testing, and the tooling is now easy enough that adoption usually precedes any policy about it.


How Synthetic Users Are Generated

The basic mechanism is simple. Specify a target user group and a research goal, and the model produces persona profiles plus interview-style responses in their voice.

What differs between tools is grounding.

Open-domain generation relies on the model's general training data. You describe "freelance graphic designers who use invoicing software" and the model assembles a plausible profile from patterns across the internet. Nothing about your product, your market, or your actual users informs the output.

Grounded generation conditions the model on primary data you upload: past interview transcripts, survey results, support tickets, analytics. The output still comes from a model, but it is anchored to evidence from your real users.

The difference matters more than most vendor comparisons suggest. Grounded output is closer to a search interface over research you already have. Open-domain output is closer to a well-informed guess wearing a persona's name. Only one is traceable back to a real person, and that should drive how much weight the output carries.

Teams already familiar with how generative AI fits into user research will recognize the pattern: the model is excellent at producing fluent, structured output and indifferent to whether that output is true.

Where Synthetic Users Show Real Promise

Three use cases hold up under scrutiny.

  • Desk research and unfamiliar domains: When entering a new category, synthetic users can help a team build vocabulary and a rough mental model quickly. This is the most defensible use, because you are treating the output as background reading rather than evidence.

  • Drafting interview guides and topic lists: Asking a synthetic user about a domain surfaces question angles a researcher might not have considered. The value is in improving the questions you take to real people, not in the answers themselves.

  • Proto-personas and hypothesis maps: Early artifacts that exist to be tested, not believed. A proto-persona built from synthetic input is a starting assumption with a label on it, which is a legitimate research object as long as the label stays attached. The guidance on user personas applies, with one addition: mark which parts came from a model.

Every one of these sits before real research begins. That is the pattern to hold onto.

The Validity Problem: Why Synthetic Users Can Mislead

Sycophancy

Synthetic users like things. They like your concept, your feature, and your redesign, and they explain their enthusiasm articulately.

This is a documented property of the underlying models rather than a quirk of any one tool. Anthropic researchers evaluating five state-of-the-art AI assistants across four free-form text generation tasks found consistent sycophancy: models shifted their assessments toward whatever view the user appeared to hold, gave more positive feedback when the user signaled they liked something, and revised correct answers when challenged.

A synthetic user is a language model wearing a persona. It inherits that tendency completely. If your prompt describes a feature you clearly believe in, you will get a supportive respondent.

Shallow prioritization

Ask a synthetic user what they need and you get a list. Ask which item matters most and the ranking tends to be flat, generic, or shaped by whatever order you mentioned things.

Real prioritization comes from constraint. People rank needs because they have limited time, money, and attention, and because they have lived the trade-off. A model has no such constraint, so it produces the appearance of prioritization without the pressure that makes it meaningful.

The imagined-experience problem

A synthetic user has never used your product, waited on hold, abandoned a checkout, or given up on a setup flow at 11pm. It can describe those experiences fluently because it has read millions of descriptions of similar experiences. It cannot report one.

This is why synthetic output cannot produce behavioral data. Everything it says is a reconstruction of what such a person would probably say, which is a different object from what a person did. The distinction is covered thoroughly in the wider comparison of synthetic data vs real data, and it applies with particular force in UX, where the gap between stated and observed behavior is the whole reason usability testing exists.

Bias Laundering and Stereotype Reinforcement

Synthetic personas reflect the dominant voices in their training data, which means they represent well-documented groups well and everyone else poorly.

The evidence here is direct. In an audit comparing 1,512 LLM-generated personas against 756 self-descriptions written by 126 real participants, researchers found that models disproportionately foregrounded racial markers, overproduced culturally coded language, and produced personas that were syntactically elaborate but narratively reductive. The output was fluent and superficially positive while flattening the people it claimed to represent.

The practical consequence is uncomfortable. Synthetic users are most tempting where recruitment is hardest: rare medical conditions, low-income users, users with disabilities, specialist professionals, non-English-speaking markets. Those are the same groups least represented in training data, so the output is most likely to be stereotype exactly where you are least able to notice.

Bias in research is not new, and the mechanics of sampling bias are well understood. What is new is that the bias arrives pre-packaged as a confident quote, which makes it far harder to challenge in a stakeholder meeting.

Circularity and Unverifiable Validation

Vendor validation claims in this category are almost entirely proprietary. A platform reports that its synthetic responses correlate strongly with real ones, without publishing the study, the sample, or the questions.

Academic practice is not much tidier. A review of 63 synthetic persona studies from leading AI venues found that 43 percent targeted undifferentiated general populations and only 35 percent explicitly discussed representativeness. If published research is inconsistent about who these personas are supposed to represent, vendor marketing claims deserve at least as much skepticism.

There is also a circularity problem. The output is generated from generalized internet data, then presented as specific insight about your users. If you validate it against other model output, you have confirmed consistency, not accuracy. Only real people can close that loop.

Synthetic Users vs Real User Research

The core distinction is short. Real research captures lived experience in context. Synthetic output reflects statistical averages of training data.

That difference compounds when segments matter. One cross-domain benchmark of LLM-simulated survey responses found that models inflate gaps between segments by two to fourfold compared with real respondents, treating demographic attributes as far more decisive than they actually are. A synthetic study can hand you a clean segment difference that does not exist, which is worse than no data because it is actionable.

Academic work on synthetic survey respondents points the same way. In a large test against the American National Election Study, synthetic responses showed markedly less variation than real ones and produced regression coefficients that differed significantly from the real estimates. Averages survived. Everything you would use to distinguish one group from another did not.

None of this makes synthetic users useless. It makes them unsuitable as evidence, which is a narrower claim and a more useful one.

Best Practices for Using Synthetic Users Responsibly

Treat every synthetic output as a hypothesis

Not a finding, not a data point, not "what users said." Language discipline does real work here, because the moment a synthetic quote appears in a deck without a qualifier, it becomes evidence in everyone's memory.

Label it clearly, every time

Any artifact containing synthetic output should say so on the same slide, not in an appendix. Stakeholders who see an interview quote assume a person said it. That assumption is the mechanism by which synthetic research does damage.

Avoid it entirely for niche, specialized, or vulnerable populations

Where training data coverage is thin, output quality drops and so does your ability to spot the drop. These are also the populations where getting it wrong causes the most harm.

Ground it in your own data when you use it

If your platform supports uploading real transcripts and study data, use that mode. It does not make the output evidence, but it keeps the hypotheses tethered to something real.

Write down the rule before you need it

Decide as a team which decisions may be informed by synthetic input and which always require real participants, and do it before a deadline makes the question urgent. The principles in this guide to responsible and ethical research practices translate directly, particularly around transparency about how evidence was produced.

When Not to Use Synthetic Users

1. High-stakes product or business decisions - Anything expensive, public, or hard to reverse needs real participants. Gartner predicts that by 2027, 60 percent of data and analytics leaders will face critical failures in managing synthetic data, with consequences for governance and accuracy. The failure mode is organizational: synthetic output enters the decision pipeline unlabeled, and later nobody can say which conclusions rested on real evidence.

2. Concept testing - This one deserves a specific warning. Synthetic users rate nearly any concept favorably, which makes them structurally unsuited to a method whose entire purpose is to separate good ideas from bad ones. A synthetic concept test will validate your weakest idea with the same enthusiasm as your strongest. Run concept testing for UX with people who have something to lose by being wrong.

3. Sensitive or vulnerable-population topics - Health, finance, grief, accessibility, safety. These require real, human-moderated research with consent and escalation paths. Simulation is not an acceptable substitute at any stage.

4. Anything requiring behavioral evidence - If the question is what people do rather than what they would say, no amount of prompt engineering closes that gap.

The honest summary of what AI can and cannot contribute here is covered well in this look at artificial intelligence in user research, which separates the tasks AI genuinely accelerates from the ones it only appears to.

Building a Workflow That Pairs Synthetic Screening With Real Validation

A workable pattern has three parts.

1. Restrict synthetic use to the earliest stage: Topic exploration, question generation, and domain familiarization. Nothing that produces a conclusion.

2. Install a mandatory validation gate: No synthetic-derived hypothesis informs a design or business decision until real participants have tested it. Make the gate explicit and named, so skipping it requires a decision rather than an oversight.

3. Define eligibility upfront: Write a short list of which decisions can take synthetic input and which always require real people. High-stakes, sensitive, accessibility-related, and anything shipping to a market you have not researched go on the second list automatically.

The gate is the part that fails in practice. Under deadline pressure, "we will validate this later" becomes "we shipped it." Naming a specific method for the validation step makes it likelier to happen, which is why teams that pair synthetic screening with a standing process for user interviews tend to hold the line better than teams relying on good intentions.

Frequently Asked Questions

1. What are synthetic users in UX research?

AI-generated profiles built on large language models that simulate a user group's thoughts and responses without any real person being studied. Unlike static personas, they are conversational and can be interviewed.

2. How accurate are synthetic users compared to real users?

They approximate broad averages reasonably but distort differences between groups, with benchmark work finding that models inflate segment gaps by two to fourfold. They also skew positive and cannot report actual behavior.

3. Can synthetic users replace real user research?

No. They cannot produce behavioral evidence, genuine prioritization, or unexpected findings, and they systematically favor whatever the prompt appears to want.

4. What are the biggest risks of using synthetic users?

Sycophancy, flattened prioritization, stereotype reinforcement for underrepresented groups, unverifiable validation claims, and synthetic output being mistaken for real participant quotes once it reaches a slide deck.

5. When is it appropriate to use synthetic users?

Desk research, unfamiliar domains, interview guide drafting, and proto-personas that are explicitly labeled as assumptions to be tested. Never as the basis for a design or business decision.

6. How do synthetic users introduce bias into research?

They reflect the dominant voices in training data, so they underrepresent marginalized and niche groups and tend toward stereotype where coverage is thin, presenting that bias as a confident first-person quote.

7. What's the difference between synthetic users and AI-moderated research?

Synthetic users are simulated and involve no real people. AI-moderated research uses AI to interview real, consented participants and adapt questions in real time based on their actual answers.

The Practical Takeaway

Synthetic users are a reasonable tool for the first hour of a project and a poor one for every hour after that. They accelerate the part of research where you are still forming questions, and they actively mislead in the part where you are answering them.

The teams handling this well are not the ones with the strictest ban or the most enthusiastic adoption. They are the ones who wrote down which decisions require real participants, labeled every synthetic artifact as synthetic, and built a validation step fast enough that skipping it never feels necessary. That last part is what makes the policy hold, and it is where user experience testing practice has to do the work that no guardrail can do on its own.


From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.

From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.

From Emotion to Action, With Insights That Speak Your Language.

Start turning customer signals into smarter decisions.