All articles

Can AI Agents Actually Mirror Real People?

The real payoff of an agent simulation is coverage: every objection, segment, and failure mode on the table before you build.

ai-agentsmarket-validationllm-simulationfoundersresearchproduct-validation

Most founders validate an idea the same way. You describe it to five friends, three of whom already like you, and you read approval into their politeness. Then you build.

The problem isn't the sample size of five. It's that the five are the wrong five, and they're being kind. What you actually want, before you write a line of code, is a room full of skeptical strangers. The customer who churned. The privacy hawk. The person who reads the terms of service. You want them arguing about whether they'd ever use the thing.

That's what LLM-agent simulations promise: a cast of personas that debate your question and surface the objections you didn't think to ask about. Do these agents mirror how actual people respond, or are you just watching a model flatter your prompt in a hundred different fonts?

The honest answer is partly, and it depends on what you ask of it.

The case that it works

The claim behind agent simulations is simple. A large model trained on a big slice of the internet has absorbed not just facts but the distribution of human opinion. It learns how different kinds of people tend to answer different kinds of questions. Prompt it to speak as a specific person and it can reconstruct a plausible slice of that distribution.

This isn't a vendor talking point. Independent researchers have tested it.

Argyle and colleagues, writing in Political Analysis in 2023, coined the term "silicon sampling." They conditioned GPT-3 on thousands of real demographic profiles drawn from US surveys and found that the model reproduced the actual response and vote distributions of those subgroups closely enough to earn a name: "algorithmic fidelity." Not one flattened average, but the shape of disagreement across groups.

Horton, in a 2023 NBER working paper, ran classic behavioral-economics experiments on LLM agents instead of human subjects. The agents reproduced qualitatively similar results to the original studies on fairness, framing, and status-quo bias. He called them "homo silicus," a cheap, fast stand-in for the human subject pool.

The idea of a full simulated society traces to Park's 2023 "generative agents": a small town of twenty-five characters who planned their days, remembered past events, and formed opinions. They were believable enough that human raters scored their behavior as more human than the behavior of real people asked to role-play the same scripts.

The most striking result came a year later. In 2024, Park and a team from Stanford and Google DeepMind built agents from two-hour interviews with 1,052 real people, then had each agent take the General Social Survey. The agents matched their humans' answers about 85% as well as the humans matched themselves two weeks later. The gap between an agent and its person was nearly as small as the gap between a person and their own future self.

That is the number to hold onto. Human opinion isn't stable, and measured against that noisy baseline, well-built agents get close.

What 100 agents actually buy you

So the agents are worth listening to. Here's where the framing usually goes wrong.

What you buy with a hundred agents is not a verdict. It's divergence. A hundred personas pushing on the same question from a hundred directions will drag objections, segments, trust conditions, and failure modes into the open, where you can see them before you write code instead of after you ship. The value is coverage. One run should leave you with a map of the argument, not a single percentage.

Here's a run of ours. We filled a room with 100 guests and staff of a busy neighborhood café and asked: if the owner puts a voice AI assistant at every table, so guests order and call a human just by talking to it, who embraces it, who walks, and what happens to tips and the six servers?

The guest split turned out to be the easy part. A large bloc adopts fast, a loyal minority defects. The headline finding the run surfaced sat somewhere else: the losses that decide whether this works are the ones nobody can see. In the simulation, regulars who feel pushed onto a machine quietly stop coming back, and the six servers don't get laid off, they leave over the following year as tips soften and a departure goes unfilled. That reframes the owner's job. What the run suggests decides survival isn't which guests like the robot, it's a written guarantee protecting the servers' headcount and tip pool, plus a way to catch the slow bleed early. Skip that, the simulation warns, and you risk bleeding your best people within a year.

And the guest split was only the thing we asked about. The room surfaced a pile of things we didn't. The one the owner still quotes is about language. Non-native speakers and tourists with limited English were among the AI's strongest advocates, an adoption driver nobody had weighted. It lets them order slowly and repeat themselves without embarrassment, turning a group that usually dreads ordering into likely promoters. The kitchen was a quiet winner too: the kitchen-staff personas predicted remakes cut roughly in half from clean AI tickets, and floated routing part of that saving into server pay to offset the tip dip. It also drew a hard line on allergies: humans must handle every one, because in the run the AI was caught mishearing and bluffing allergen questions, and one safety incident ends the experiment. And it flagged a threat we hadn't modeled, a competitive-response scenario: rival cafés two blocks away print "human service" signage, ready to harvest every alienated regular and every server whose tips drop.

None of that was in the question, and that's the whole point of running it.

Where it breaks

None of that makes the agents oracles. The same literature is candid about the failure modes, and any founder using this should know them.

The training data is skewed. Models learn from text that got written down and posted, which over-represents people who write online, in English, from wealthy countries. Social scientists call this the WEIRD problem: Western, Educated, Industrialized, Rich, Democratic. If your customer is a 58-year-old contractor who has never posted a review in his life, the model has thinner ground to stand on.

Alignment flattens the edges. The tuning that makes a model helpful and inoffensive also pulls it toward the agreeable middle. Real populations contain cranks, zealots, and committed contrarians. A tuned model tends to round them off, so a simulated audience can read as calmer and more reasonable than the real thing. Quieter audiences hide the objections that kill products.

Stated preferences aren't revealed preferences. People say they will pay for privacy and then click "accept all." They say they would switch and then never do. An agent can tell you what a person would say, not what they would do once money and friction enter the picture. Santurkar and colleagues showed a related gap in 2023: different models reflect different opinion distributions, leaning toward some demographic groups over others. Which model you ask changes the answer you get.

Genuinely novel situations are extrapolation, not measurement. When you ask about a product no one in the training data has lived through, the model isn't recalling how people reacted. It's guessing, plausibly, from adjacent cases. That guess can be useful, but it's a hypothesis, not a readout.

That model-dependence is one we take on ourselves, because you never touch it. In PredictAible you don't pick the model, you pick how big a room to fill: 25, 100, or 500 agents. The debate model is ours. Before publishing this piece we took one of our 100-agent runs and re-ran the identical seed from scratch, and the decisive conclusion held. That's how we separate a stable signal from a one-off artifact. A conclusion that survives a fresh replication is more likely reflecting how people actually talk about trust and money than the noise of a single run; one that doesn't is a soft spot that needs real human data before anyone bets on it.

The limits all point the same way. Agent simulations are good at surfacing the range of reactions and bad at pinning down exact percentages.

Why a run isn't cheap

PredictAible runs on Anthropic's Claude models. The analytical core, the part that casts the personas, builds the influence graph, runs the analysis, and verifies the result, sits on their frontier tier. A single 100-agent run is more than eight hundred model calls; the agents write close to a hundred thousand tokens of debate, and the analysis stages read all of it. You could wire something that looks similar on top of bargain models and charge less for it. We don't, and that's deliberate. What you're paying for is the quality of the reasoning underneath, not a thin wrapper around the cheapest API on the market.

Use it as a floodlight, not an oracle

Do not use an agent simulation as an oracle. It will not tell you that "42.3% of your market will buy." That number is false precision, the product of one model's skew, one prompt's framing, and a population the training data only partly covers. Bet on it and you're betting on the wrong thing.

Use it as a floodlight. Point it at your idea early and it lights up the objections, the segments, the attack vectors, and the trust conditions you would otherwise discover the hard way, from real customers, at real cost. Read the output for structure, not for numbers: the adopter groups tell you who to design for, the dealbreaker tells you what not to ship, the threat model tells you what to harden. The slow bleed of servers our café run flagged, no dramatic layoffs but quiet attrition a year out, is exactly the kind of thing a founder would rather learn in an afternoon than in a post-mortem.

Agents do not replace talking to customers. They tell you which customers to talk to, and which questions to bring.

Try it

Pick a decision you're currently guessing at. Frame it as a question, choose a society to put in the room, set it to 25, 100, or 500 agents, and watch them argue it out. Count how many objections land that you'd never have written down yourself.

Start a simulation →