Why I’m reconsidering a principle I’ve believed for years.
I’ve always believed that imperfect research is better than no research at all.
Synthetic users have made me reconsider that. When synthetic responses are treated like findings from real users, they can create something far more dangerous than uncertainty: false confidence.
What I mean by “imperfect research”
Until synthetic users entered the picture, “imperfect research” still meant some contact with the thing you were trying to understand: actual users, their behavior, and their experiences. Maybe the sample is small, recruitment is messy, or the prototype is rough. Sometimes you only have time for five interviews. These limitations matter. They affect what you can conclude, but one thing remains true: you observed or talked to actual people.
No study is perfect. Every study has limitations that shape what we can responsibly conclude. As researchers, our job is to understand those limitations and avoid claiming more than the research supports.
When the imperfections are known and understood and we’re still learning directly from the people or behavior we’re trying to understand, imperfect research is better than no research.
Synthetic users change the equation
Using a synthetic participant isn’t the same as working with a small sample of users, conducting a flawed interview with a real user, or relying on a convenience sample.
It’s a simulation of what a user might say. The response depends on the model, how the synthetic user is constructed, and the data used to ground it.
Keep in mind, a synthetic user isn’t a database lookup of what similar people have said. Even when it’s grounded in real research, an LLM is still generating a predicted response based on patterns it has learned plus whatever context has been supplied.
And sometimes, the “data” used to create synthetic users is really just the team’s assumptions documented as proto personas. Rather than being grounded in actual research artifacts such as interview transcripts, survey results, analytics, or customer-support logs, the synthetic participant may simply be built from assumptions and unvalidated opinions.
Teams can build synthetic participants from their own assumptions, then treat the responses as if those assumptions had been tested. What comes back may look like new insight, but it can simply be the team’s original assumptions dressed up with more detail and confidence, under the guise of being “informed.”
When no research is conducted, at least the absence of research is visible. The team might make a decision based on assumptions, but those assumptions haven’t taken on the appearance of research findings. Synthetic-user “findings” can make those same assumptions look validated.
Imagine these two situations:
“We haven’t talked to customers yet, but we believe they’ll understand this pricing model.”
versus:
“We tested the pricing model with 50 synthetic customers, and 78% understood it.”
The second statement feels much more evidentiary. There are participants and a percentage. Maybe even quotes or themes.
But depending on how those synthetic users were constructed, the organization may not actually know much more about its customers than it did in the first scenario. It may simply feel more certain.
That’s the part that worries me. I’m not here to help teams build the wrong thing with increased confidence.
I’d rather have some contact with users than have a team make decisions entirely from assumptions.
I don’t believe synthetic users belong in the same category as imperfect research at all. It’s a fundamentally different kind of compromise, with different risks.
When close isn’t close enough
This is why I was especially interested in a recent MeasuringU study from Lucas Plabst, Jeff Sauro, and Jim Lewis. They created digital twins for 420 people who had already rated the usability of ChatGPT, Claude, Gemini, or Grok. A digital twin is a type of synthetic user built from data about a specific real person.
The researchers created two versions of each person’s twin. The “summary twin” received a 700-word persona generated from that participant’s survey responses. The “full twin” received the same summary plus nearly all of the participant’s original survey responses. In both cases, the participant’s System Usability Scale (SUS) responses were withheld, and the twins were asked to predict them.
At first glance, the synthetic results looked pretty good. The summary twins averaged 4.4 points higher than the humans on the 100-point SUS scale. The more detailed “full twins” were 6.7 points higher. As the researchers point out, you could describe those results as 95.6% and 93.3% accurate.
That sounds reassuring. Until you look at what those differences did to the interpretation.
The scores from the real participants resulted in grades ranging from B+ to A. Every full-twin mean translated to an A+. A team looking at the real participant results might see products that are performing well but still have room for improvement. Looking at the synthetic results, they could walk away thinking everything is superb.
That gets at something I think matters far more than whether synthetic results are numerically similar to human results: are they decision-equivalent?
Two sets of numbers can be relatively close and still tell a different story. If the story changes enough to affect what gets fixed, prioritized, funded, or ignored, then “93% accurate” isn’t nearly as reassuring as it sounds.
One other finding caught my attention:
Giving the digital twins more information about the actual participants didn’t make them more accurate.
The full twins had access to nearly all of each person’s survey responses, yet on average, their SUS scores were actually farther from the human scores than those from the simpler summary twins.
A much larger recent study published in Science Advances raises a related concern. It tested digital twins across 19 preregistered studies and 164 different outcomes. Each twin was built using more than 500 previous answers from the real person it was meant to represent.
Despite all that information, the twins’ responses correlated only weakly with the responses of their human counterparts. Even more surprising, on the researchers’ exploratory individual-level accuracy measure, twins built from those 500+ answers weren’t significantly more accurate than much simpler personas based on just 14 demographic characteristics.
The researchers describe today’s digital twins as “funhouse mirrors.” They may resemble the people they’re meant to represent, but they systematically distort them.
Neither study settles the question of what synthetic users may eventually be capable of. But together they illustrate the risk I’m worried about: synthetic results don’t have to be wildly wrong to send a team toward the wrong conclusion.
One good use case for synthetic users
Even though I don’t believe in using synthetic users as a substitute for talking to real people, they can be useful before the research even starts.
If you’re researching an unfamiliar industry or audience, rehearsing an interview with an AI-generated “participant” can help you get comfortable with the terminology and catch questions that aren’t as clear as you thought. You may also realize pretty quickly where you need to do a little more homework before sitting down with a real participant.
That’s the one genuinely good use case I’ve found for synthetic users: as a low-stakes practice partner before the real research begins.
Use AI to practice, then talk to actual humans.
Where I land
Imperfect research = incomplete insight into real people.
Synthetic research = generated responses in place of responses from real people.
We keep debating whether synthetic users are good or bad at research, but we’re lumping together two very different jobs. They can help researchers prepare to do research without being credible substitutes for the people we’re researching. The distinction becomes much more consequential once synthetic output starts being treated like research findings.
I will always believe that teams should do research with actual people whenever possible. But I don’t think every substitute for research deserves to be treated as simply another form of imperfect research.
Let’s call it what it is: simulated insights.
When considering using synthetic users in place of actual research participants, ask yourself, “If this synthetic user is wrong, what happens next?” If the answer is, “We change the product roadmap,” we should be asking whether those “findings” are reliable enough to act on.
Imperfect research involves real people. Synthetic research makes predictions about them. That difference matters most when those results shape what gets built.
Originally published on Making Research Matter, where I write about what it takes for research to actually lead to action.
Synthetic users are making me rethink “imperfect research” was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.