Logo
FrontierNews.ai

Why AI Still Can't Predict What You'll Actually Do: The 85% Accuracy Breakthrough

Artificial intelligence excels at reasoning, but struggles to predict what real people will actually do in real situations. A new approach called "behavioral foundation models" trains AI systems on causal data from controlled experiments to replicate human decisions with 85% accuracy, significantly outperforming frontier large language models (LLMs) like ChatGPT and Claude on niche populations.

Why Do Frontier AI Models Fail at Predicting Human Behavior?

The gap between AI reasoning power and human behavior prediction is stark. Joon Sung Park, co-founder and CEO of Simile AI, explained that frontier models like ChatGPT and Claude score around 50 to 60 percent accuracy on general population samples when predicting human behavior. On niche populations that businesses actually care about, performance drops dramatically to 20 to 30 percent.

"If the model has to learn the underlying physics of the world that it's operating in, then I think you can just prompt your way into getting the actions out of it. I don't think the model has yet learned the complete mapping of social physics of humanity," said Joon Sung Park.

Joon Sung Park, Co-founder and CEO of Simile AI

The reason is philosophical as much as technical. Frontier models are trained to be super-rational, objective machines. They optimize for logical consistency and factual accuracy. But humans are irrational, inconsistent, and driven by emotions, trauma, and social context. Park noted that modeling human behavior requires capturing these imperfections: "If I make some mistakes, the model has to make the same kind of mistake. That's very hard".

Park

How Do Behavioral Foundation Models Actually Work?

Simile's approach relies on three types of data, each serving a distinct purpose in building accurate digital twins of real people:

  • Interview Data: Qualitative life-story information including childhood memories, trauma, and formative experiences that provide texture difficult to predict from structured data alone.
  • Observational Behavioral Data: Transaction records and web-scraped behavior that establish baseline statistics of what people actually do in their daily lives.
  • Causal Mechanism Data: Results from randomized controlled trials where variables are manipulated to reveal why behavior changes when conditions shift.

The third category is both the most important and the hardest to acquire. Decision-makers don't just want to know what will happen; they want to know what to do now to change the outcome. That requires understanding cause and effect, which only controlled experiments can reveal.

Simile validated this approach with a landmark study involving 1,000 people representatively sampled from the US population. Researchers spent two hours collecting wide-ranging data from each participant, including interviews and behavioral data. After two weeks, they built digital twins from the collected information and brought participants back to complete surveys, behavioral economics games, personality assessments, and published randomized controlled trials. The digital twins replicated their source individuals' attitudes and behaviors with 85 percent accuracy, matching how accurately people replicate themselves.

What Makes This Different From Personal AI Assistants?

Park's vision prioritizes simulation over the personal assistant approach that companies like OpenAI are pursuing. While assistants like OpenClaw can execute tasks, they lack a deep model of their users. His example illustrates the problem: an assistant that orders Hawaiian pizza for someone who hates pineapple on pizza has failed not because it can't execute the order, but because it lacks a model of the user's preferences and personality.

"Our bet was that this technology around simulation creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in," explained Park.

Joon Sung Park, Co-founder and CEO of Simile AI

Park traces this philosophy to a thought experiment he and his co-founders conducted during his Stanford PhD program. They imagined fast-forwarding ten years and asking what single application of foundation models would have mattered most. The runner-up was personal assistants that automate tasks on your behalf. The winner was simulation: recreating the world we live in with enough fidelity to answer counterfactual questions that surveys and focus groups cannot.

How Can Organizations Use Behavioral Foundation Models?

The practical applications extend far beyond academic interest. Simile's data collection strategy reflects real-world deployment at scale. The company runs its own randomized controlled trials in lab and virtual settings, with a critical design principle: stakes must be real. Park noted that "what makes the difference between what is attitudinal versus behavioral is if the stake in your decision is real".

Park

Beyond internal experiments, Simile partners with firms for behavioral data and operates a panel reaching tens of millions of people globally, with weekly data collection on tens of thousands of participants. A third research effort harvested tens of thousands of professionally designed randomized controlled trials from the Open Science Foundation's pre-registration platform, created in response to the replication crisis in social science. Post-training on this corpus of real experiments significantly improved the model's ability to predict human behavior.

Park's ambition scales from immediate applications like marketing concept testing to eventually simulating 8 billion people to tackle systemic challenges like climate change. The core insight remains: accurate simulation of human behavior requires not just better prompting or larger models, but new data collection methods, new training objectives, and a willingness to model human irrationality rather than super-rationality.