Synthetic User

Synthetic Patients: What They Are and When They Work

An explainer on synthetic patients.

Synthetic Patients blog image

If you build anything patient-facing, a clinical training tool, a diagnostic AI, a health app, you’ve probably run into the term “synthetic patients” and weren’t sure if it meant an AI chatbot standing in for a sick person, a spreadsheet of invented medical records, or both.

It’s both. A synthetic patient is an AI-generated stand-in for a real one, built either as a simulated conversational agent that talks and reacts like someone with a specific condition, or as a synthetic dataset of patient records: invented but statistically realistic charts, labs, and histories. Both exist for the same reason. Real patient data and real patient time are expensive, slow, and legally protected.

This piece covers both types, what the research says about how well they hold up, and exactly where the line sits between “good enough for training and testing” and “not good enough for a real clinical decision.”

What are synthetic patients?

A synthetic patient is either a simulated patient interaction or a synthetic patient dataset, and the two solve different problems. In the medical education literature, simulated patients are also called virtual patients, and Synthea explicitly describes its own output as “realistic but not real” data, which is a useful way to hold both categories in mind.

Simulated patient interactions are conversational AI agents built to act, talk, and respond like a patient with a defined condition, history, and personality. Medical schools use them to let students practice diagnosis and difficult conversations without a live person in the room. This category sits inside the wider synthetic user research field, the same AI-persona approach used for product and market research, just aimed at a clinical use case instead of a commercial one.

Synthetic patient datasets are invented but statistically realistic patient records: demographics, labs, diagnoses, and outcomes generated from real electronic health record (EHR) patterns without containing any real person’s data. Synthea, the open-source generator from MITRE, is the most widely used example, and it’s built the same way most synthetic data generation tools are: model the statistical shape of a real population, then generate new records that fit that shape without mapping back to anyone real.

One is a role-player. The other is a spreadsheet, closer to a digital twin of a patient population than an individual you could interview. People searching “synthetic patients” usually mean the first, so that’s the focus here, with the second covered where it’s relevant.

Can AI actually simulate a patient?

Yes, and the better systems don’t just generate plausible dialogue. They ground it in real clinical data, which is what separates a modern simulated patient from a generic chatbot playing a role.

AIPatient, a system described in Communications Medicine, pairs large language models with a knowledge graph built from de-identified discharge notes in the MIMIC-III database. Six task-specific agents split up the reasoning work, and unlike earlier virtual patient systems, which mostly leaned on a small, fixed set of hand-built cases, AIPatient is grounded in real, de-identified patient records. That grounding is what separates a modern synthetic patient from a scripted chatbot: the responses trace back to actual clinical presentations, not just generative AI text that merely sounds plausible.

Other systems take a narrower approach. SimPath, built for therapist training, retrieves from 5,810 synthetic patient profiles generated from 9,527 real, anonymized therapy transcripts across 20 DSM-5 diagnostic categories, so each session pulls a clinically grounded persona rather than improvising one from scratch.

How accurate are AI patients?

Accurate enough for training and low-stakes testing. Not accurate enough, yet, for anything a real clinical decision depends on.

On the high end, AIPatient’s own benchmarking found it held up consistently across three axes: how readable its output is, how robust it stays across different question types, and how stable that performance holds over repeated runs. In a separate study comparing four leading LLMs generating synthetic patient-physician conversations for plastic surgery scenarios, three clinically trained, blinded raters scored transcripts across seven criteria including medical accuracy, realism, and empathy. Every model averaged above 4.5 out of 5.

On synthetic datasets specifically, Synthea’s records have been validated against real clinical quality measures and found to closely resemble actual patients, and IBM researchers trained a risk-prediction neural network on 500,000 Synthea records that hit 94% accuracy.

The honest caveat: those numbers come from controlled, well-structured test cases. Real patients aren’t that tidy. A 2026 framework called VeriSim injected realistic “patient noise”, things like forgotten symptom onset, low health literacy, and stigma-driven non-disclosure, into simulated patients and re-ran the same models. Diagnostic accuracy dropped 15 to 25%, and conversations ran 34 to 55% longer. Smaller, 7-billion-parameter models degraded 40% more than larger, 70-billion-plus models. Clean benchmarks and messy reality are two different tests, and synthetic patients score much better on the first one.

Are synthetic patient personas reliable?

Synthetic patient reliability comes down to the stakes of the question being asked. For repetition and exposure, yes. For a single high-stakes judgment call, no.

Simulation earns its place specifically because live clinician studies are costly and hard to get past an ethics board, so synthetic patients fill the gap where the goal is practice reps, not a verdict. That’s a different job than being right once, under pressure, on a real person.

The risk shows up when that distinction gets blurred. Researchers building Roleplay-doh, a tool for creating LLM-simulated patients for counselor training, flagged that trainees can become overconfident practicing on an AI patient and undersupport a real one later. Their fix was structural: real counselors still need to pass traditional certification and background checks before they see real patients. Synthetic reps build skill. They don’t substitute for the credential.

Synthetic patients vs. real patients: what’s the difference?

DimensionSynthetic patientReal patient
AvailabilityOn demand, no scheduling or no-showsRequires recruitment, consent, and coordination
Cost per sessionNear-zero marginal costTime, incentives, and staff overhead
Privacy exposureNo protected health information involvedGoverned by HIPAA, GDPR, and IRB approval
ConsistencySame case replays identically every timeNatural variation session to session
What it capturesApproximated behavior and communication patternsActual physiology and lived experience
Best used forTraining, drills, early-stage testingDiagnosis, treatment decisions, safety validation

The two aren’t competing for the same job. One builds reps at scale. The other produces the ground truth those reps are aiming at. The same tradeoff shows up outside healthcare too. Synthetic users versus real users walks through the identical calculation for product and messaging research.

When should you use synthetic patient research?

Four situations, in order of how established the practice is:

Clinician training. Practicing breaking bad news, running a diagnostic interview, or holding a difficult conversation, all without risking an actual patient’s experience on a trainee’s first attempt. A 2024 study built exactly this kind of tool for medical schools, generating a range of patient demographics and disease states for trainees to practice on before a real encounter.

Testing diagnostic AI before deployment. Running a model against hundreds of synthetic cases, including noisy, realistic ones like the VeriSim benchmark above, before it ever touches a real patient interaction.

Generating training data without touching real records. Synthetic datasets like Synthea let teams build and test machine learning models without the privacy exposure of real patient data, which matters most when the eventual system needs broad, diverse training data that real records can’t legally or practically provide.

Testing patient-facing messaging and communication. This one sits outside clinical decision-making entirely: does a patient-facing message, an app screen, or a piece of care instruction actually land the way it’s intended, before real people see it. Articos’s synthetic user methodology, benchmarked against expert human research across 46 studies at 86% recall, works on this same principle: read a message, react to it, flag confusion, all in real time. For an early concept, like a new patient onboarding flow, Articos’s concept testing platform is built for that “does this land at all” stage. For headline- and copy-level work specifically, the messaging testing platform narrows further to just that question.

We ran exactly this test in Articos: 12 synthetic patients, split evenly across older adults, younger adults, low health literacy, and high health literacy, reacting to a plain-language versus a clinical-phrasing version of the same medication instruction.

12 synthetic patient personas across age and health-literacy roles used to test medication instruction comprehension in Articos

Plain language won on comprehension: 10 of 12 personas said they’d skim the dense clinical handout and default to the bottle label first, and vague terms like “as needed” and “persistent” got flagged as unclear by 9 of 12. But plain wording alone didn’t fix the riskier decision.

Articos research report verdict showing plain-language medication instructions outperformed clinical phrasing across synthetic patient personas

Five of the 12 personas still couldn’t tell whether a symptom meant “monitor it” or “call now” unless the instruction named the threshold directly, and 9 of 12 said their real fallback was calling the pharmacist rather than trusting their own read of the label. One persona, Sanjay, put the gap plainly: “I’d look for whether the instruction tells me two separate things: stop it, or keep taking it and call. Those are not the same to me.”

Goal score table showing plain language wins on dose comprehension but stays mixed on safe stop-or-call escalation decisions

That’s the same failure mode VeriSim measured externally: ambiguity, not vocabulary, is what breaks comprehension. Our own test lines up with the argument running through this whole piece: synthetic patients are reliable for testing whether wording lands, and that reliability has a ceiling. Comprehension and safe escalation are two different questions, and this one test alone couldn’t confirm the second without pairing plain language with explicit thresholds, which is a design fix, not a synthetic-versus-real-patient question.

Teams that want the same question answered by real people instead of synthetic ones have that option too. Wynter runs message tests against verified human B2B panels with a 12-to-48-hour turnaround, which is a fair, real alternative when the messaging decision is high enough stakes to want a human panel from the start.

The distinction that matters across all of this is the question being asked. “Will people understand this onboarding screen” is a reaction-and-messaging question. “Is this treatment plan sound” is a clinical one, and no synthetic persona, ours included, should be anywhere near answering it.

Where do synthetic patients fall short?

Anywhere the answer has to be right about one specific real person.

Diagnosis, treatment efficacy, drug and clinical trials, and any AI validation that feeds a safety decision all require real patient data. The AIPatient team named their own limits directly: their system draws on MIMIC-III, which skews toward a single hospital’s population, and it doesn’t yet model social determinants of health like income or living conditions, both of which change how real patients describe symptoms and follow care plans.

There’s also a human cost to substitution that goes beyond accuracy. The team behind an AI system for simulating difficult medical conversations put it plainly: standardized patients are often working actors whose livelihoods depend on collaborations with medical schools, and even a highly capable AI system cannot replicate the value of an actual human interaction. Synthetic patients are a tool for building skill faster. They’re not a replacement for the people and the real encounters that skill eventually has to serve.

What are the ethics of using synthetic patients?

Synthetic patients solve one ethical problem and create a smaller set of new ones.

The problem they solve: privacy. Training on synthetic records or synthetic conversations means no real protected health information changes hands, which sidesteps a real chunk of HIPAA and GDPR exposure that comes with recruiting actual patients for training purposes.

The problems they introduce are about diligence, not deception. A recent review of AI ethics in medical research names the core issues directly: privacy, bias, accountability, and informed consent all need active management, since a model trained on a narrow or non-diverse population can quietly reproduce that narrowness at scale, and a system’s “black box” reasoning can make it hard to say who’s accountable when it’s wrong. Managing that means validating the underlying data, disclosing when a training tool is synthetic, and keeping a human in the loop wherever the output touches an actual patient.

What’s the cheapest or free option for synthetic patient research?

On the dataset side, free. Synthea is open-source under an Apache 2.0 license, maintained by the nonprofit MITRE Corporation, and its output is explicitly free of cost, privacy, and security restrictions for research, industry, and education use. If you need synthetic patient records rather than a conversational patient, that’s the starting point, and it’s what most of the accuracy research cited above is built on.

On the conversational side, there isn’t a free, clinically validated equivalent yet. AIPatient and SimPath are research systems, not public products. The realistic low-cost path is prompting a general-purpose LLM directly to role-play a patient, but that’s exactly the setup VeriSim tested as the weak case: smaller, ungrounded models degraded the most under realistic patient noise. Free and fast is available. Free, fast, and clinically grounded is not, at least not yet.

Should you combine synthetic and human research instead of choosing one?

For most healthcare product and messaging work, yes. The teams that get the most out of synthetic patients don’t pick a side, they sequence the two: synthetic first, at scale, to cover breadth and iterate fast without scheduling delays, then real patients or clinicians to confirm anything that’s surprising, borderline, or headed toward an actual clinical or safety decision. That sequencing logic is also how Articos frames its own synthetic user product for teams doing primary research: run the synthetic pass to narrow down what’s worth testing, then spend real recruitment budget confirming the findings that actually matter, instead of using it to validate everything from scratch.

The one place this sequencing doesn’t apply: if the output touches a diagnosis, a treatment decision, or anything regulatory, synthetic patients are the practice round, not the record. Real patients, under proper ethical oversight, make the actual call.

How do you decide between synthetic and real patient research?

Ask what the output feeds. If it’s practice reps or a reaction to a piece of messaging, and the cost of being wrong is a revision, synthetic patients are the faster, cheaper option, and the research above shows they hold up well for exactly that. If the output feeds a diagnosis or a treatment decision, anything where a wrong answer could hurt a real person, use real patients under proper ethical oversight. Synthetic patients are reliable for the job they’re built for. Using one to make an actual clinical call is a scope mistake, not a reliability problem.