I.
When people working at frontier labs get asked questions like ‘why is developing artificial intelligence so important and worth the associated risks?”, the most convincing response will be that a sufficiently powerful model will revolutionise science and accelerate research, resulting in the solving of all of our problems.
This mission has not been lost on Psychology and its consequences (Behavioural Science, Economics, Anthropology, Sociology, Political Science, any of the other disciplines that are basically the same thing at their core). The past two years have seen a surge in interest in the ability for LLMs, and LLM-based agents, to simulate human participants at an accurate enough level for us to do real research with them. These will often be called ‘synthetic participants’.
Funnily enough, the papers advocating for this direction in social science research often claim the reason this would be such a big deal is that it cuts research costs down, or makes research way faster to do, or can be added as a bonus factor into decisions that previously wouldn’t have had any empirical evidence to go off.
This seems to me to ignore the obvious reason sufficiently accurate synthetic participants would accelerate Psychology 100x: they would allow you to study the stuff that you can’t get sign-off to study with real people, i.e. darker aspects of human behaviour, or anything that might meaningfully change someone’s life.
The pressing psychological questions of our lifetimes could be things like:
What drives someone to commit a terrorist attack?
Why do people kill themselves?
What predicts sexual assault/rape/perversion?
How do people get cripplingly addicted to AGI-based wire-heading?
How do dictators come to power?
Or from a more positive standpoint:
What’s the optimal amount of sex to be having?
What’s the best home environment for a child to grow up in?
What happens when you randomly assign children to poverty/wealth?
What happens if you don’t treat mental illness at all?
Can we induce traumatic experiences to better understand PTSD and then reduce it?
For all of these, Psychology can already do a decent job of talking around the things that probably matter in some sense, i.e. it’s easy to tell what is correlated with those things. But for causation you tend to want to run experiments. And it’s difficult to convince anyone to let you run an experiment on an intervention designed to stop terrorist attacks, as several participants successfully pulling off real-world terrorist attacks would reflect poorly on the university’s reputation, amongst other secondary issues.
So I am an unashamed supporter of attempts to use AI to completely redefine psychological and behavioural research. If it works, it won’t just make research cheaper and faster, it will fling the field forward so much that everything achieved by this point will seem utterly inconsequential. The models and theories you could actually draw out causally would be incomprehensibly better. In which case we are all out of a job but that’s a question for another day.
So anyway, that my hope for the long-term futu.. wait what?
How foolish of me to not realise the primary use of AI-replicated humans in this economy would be market research. Apologies.
In any case, who’s the guy in the photo?
The company that handled the data was Simile, a year-old start-up in Palo Alto, Calif. Simile says its technology can predict human behavior by tapping into its growing database of A.I. agents. That allows clients to test the reaction to new products or evaluate how their brand is being perceived by consumers. And, Simile promises, they can do so faster and cheaper than by running traditional market research studies.
“If you’re able to simulate the world, you can basically test out countless interventions,” said Joon Sung Park, Simile’s chief executive and co-founder. “This is sort of the window into humanity.”
Sounds like a big deal, surely this is making waves?
On Thursday, Simile is announcing $200 million in new funding at a $2 billion valuation. Less than six months ago, the company raised $100 million. Greenoaks led the latest round, which also included the existing investors Index Ventures, Hanabi, A*, Bain Capital Ventures, CVS Health Ventures and Definition.
Nice. And how far along are they?
Simile’s responses are now between 85 percent and 99 percent accurate, depending on the scenario and the population surveyed, according to the company’s head of product, Mihika Kapoor. The start-up says it already has enough data to tap into the A.I.-simulated opinions of very specific population groups.
“Simile’s responses are now between 85 percent and 99 percent accurate.”
Given no new theories of psychology have been published on the back of such insane technology, I’m telling you now that this cannot be what it seems.
Yet a $2B valuation and an NYT feature suggests I’m wrong.
One of us is lying here. Let’s find out if it’s me.
II.
Beginning with a number audit, so we can work out what we’re actually arguing over, before we get into whether their methodology suggests such a claim is possible.
“Simile’s responses are now between 85 percent and 99 percent accurate.”
Ok, what does that even mean? Accurate with regards to what?
I’m assuming we’re suggesting that Simile predicts human behaviour accurately 85-99% of the time. But again, this isn’t a complete sentence.
Let’s say we agree on Friday night that you and I are going to get coffee on Monday Morning, and I secretly write down what I think you’ll get up to on the weekend, and seal it in an envelope. On our Monday coffee date I allow to tell the story of your weekend, and at the end I excitedly open my envelope to reveal to you I completely predicted your weekend schedule. You went to the gym Saturday morning, then lunch with your sister, then park with the kids on Sunday, then watched the game after.
But you say I barely predicted anything, you do that every weekend anyway. And importantly, I didn’t predict what you’d do at the gym, and how quickly you gave up on your sets. I didn’t say what dish you ordered, or even what restaurant you’d pick. So how can I have perfectly predicted anything?
And even if I got those parts too, there’s always further minute detail to go into, who gets to decide when enough is enough? Who gets to decide when I’m fully right? Or half right? Or like 85% right? 85% of what?
‘Human behaviour’, whatever that is, is unquantifiable in a holistic way. So whenever you do see a number pertaining to it, and they’ve chosen not to tell you where it’s come from, you may want to be suspicious.
And wait a second, even if we could allow that number to exist, what do they mean their predictions are ‘accurate’ in the first place? Which humans are they talking about? They must be talking about some real people, who are actually doing the things the model is predicting them to do, right?
That sentence we started attacking was actually written by the NYT/Dealbook journalist covering Simile, so I will ease off on that and allow Park himself to explain it now:
(warning: the quotes throughout this will spell the word ‘behavior’, whereas I write ‘behaviour’. explicitly pointing this out may have made it worse. sorry.)
So that is when we decided that we actually want to validate that the simulations are accurate. So we went out and actually created simulations of thousand people of the US population.
We demonstrated that using our architecture and the models we can actually predict people’s behaviors 85% as accurately as people replicate their own.
When we saw that we thought okay this is something that we feel comfortable providing to our users as a platform for simulating their really important decisions.
This was enough to find the paper the number is based off, research which is allegedly worth $2B.
Simile’s website has a research section where they map out their progress. And this is the paper where they, in their own words, ‘proved’ that AI-based simulation can be accurate.
And here we are, right in the abstract:
The generative agents replicate participants’ responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later, and perform comparably in predicting personality traits and outcomes in experimental replications. Our architecture reduces accuracy biases across racial and ideological groups compared to agents given demographic descriptions. This work provides a foundation for new tools that can help investigate individual and collective behavior.
Here is my attempt to explain exactly what they did in this study, as clearly as I can:
(you can skip the painstaking detail here if you want, but that will require you to blindly trust some of my interpretations later)
They recruit a representative sample of ~1,000 Americans, paid for their time.
Each participant does a 2-hour interview about themselves with an AI interviewer.
The questions came from an abbreviated American Voices Project interview script. They covered life history, relationships, work, neighbourhood, health, finances, politics and personal values.
Each participant completed the General Social Survey questions.
These are ~180 multiple-choice questions about the person's background and opinions. Examples of the kinds of questions: what their job is, their marital status, how many children they have, how often they attend religious services, whether the government spends too much or too little on things like the military, education, or welfare, whether they think abortion should be legal in various circumstances, how much confidence they have in institutions like banks, Congress, or the press, whether they own a gun, how happy they are, their political party and how liberal or conservative they consider themselves.
Each participant does a big-five personality test.
Each participant does five Behavioural Economics games:
Dictator game: how much of $5 to keep and how much to give another person.
Trust game, first player: how much of $3 to send another person, who would receive three times the amount sent.
Trust game, second player: how much of the money received to return to the first player.
Public goods game: how much of $4 to contribute to a four-person fund. Contributions were doubled and divided equally among the group.
Prisoner’s dilemma: whether to cooperate or defect, with payments depending on both players’ choices.
Each participant does five more behavioural experiments from a replication project.
The descriptions are too long for this so if you want to find the detail, ctrl+f the paper for Ames & Fiske. The point of them being here is to see if the agents replicate the treatment effect, which you can’t do for the first 5 economic games.
After that’s all done, participants writes a short paragraph
about themselves. They were instructed to describe themselves as they might to a stranger, including information about their personal background, personality, and demographics.
Two weeks later, those participants come back, and repeat steps 3-6.
You are then provided with this visual to explain what is going to happen:
The 2-hour interviews alone are fed into models to generative individual digital twin agents for each participant. With the idea being that those agents are then handed the same battery of questionnaires, to see if they will answer the same things, and thus, in some sense, predict their behaviour.
I can see some hands going up already. Let’s finish the procedure first.
First, note the sentence at the bottom: compare actual to simulated responses, adjusting for participant self-consistency.
This is why they come back 2 weeks later, so we can adjust for self-consistency. You ask why would you do that, and the authors answer:
Since individuals often exhibit inconsistency in their responses over time in both survey and behavioral studies, we use participants’ own attitudinal and behavioral consistency as a normalization factor: the probability of accurately simulating an individual’s attitudes or behaviors depends on how consistent those attitudes and behaviors are over time.
To account for these varying levels in self-consistency, we asked each participant to complete our battery twice, two weeks apart. Our main dependent variable is normalized accuracy, calculated as the agent’s accuracy in predicting the individual’s responses divided by the individual’s own replication accuracy.
I’ve read this and I’ve read this and I’ve read this. On the positive side, it’s reminded me why I left academic writing in my past. On the negative side, I have no idea what it means. Within-human reliability is something you see measured often enough, and it does make sense to suggest you should account for this when trying to predict. Because it’s unfair to punish an agent for not predicting something that is inconsistently predictable.
But when you take this number and divide by it, rather than just offering at as context, you’re putting pressure on yourself to be clear over what your new number means. And you still need a good reason for focusing on just this normalised variable over the agent’s raw ability to predict answers.
People are inconsistent. No doubts about that. But why exactly should that inconsistency make an agent’s incorrect predictions count for less?
You could argue that it’s impressive that an agent can tell if a human will fluctuate on a specific question, and in a particular direction. But that’s not what happened here. Look, here comes that 85% figure, and an explanation of how they actually do the math:
For the GSS, the generative agents predicted participants’ responses with an average normalized accuracy of 0.85 (std = 0.11), calculated from a raw accuracy of 68.85% (std = 6.01) divided by participants’ replication accuracy of 81.25% (std = 8.11).
And this is where they place the raw number i.e. how well the model predicted the responses, straight up. The answer is 68.85% of the time. Here’s that compared to some others arms, such as just giving an LLM some basic demographic, and giving an LLM that short paragraph they wrote about themselves:
So we have a sub-result here that the interview technique, the thing that really changes this from a basic multilinear regression to this deeper type of Synthetic Twin, improved prediction by about 10 percentage points, which seems great. But it seems relevant to me that the interview questions…
…came from an abbreviated American Voices Project interview script. They covered life history, relationships, work, neighbourhood, health, finances, politics and personal values.
Which is … basically what the GSS covers. So what does this finding mean? Does it mean LLMs can predict human behaviour? Or does it mean for any given individual, their qualitative and quantitative responses to basic questions in similar areas are going to be pretty similar? Simile claim their models can predict how people will react to new stimuli, but this has nothing to do with that.
In any case, that 69% is thrown to the side as not a fair representation of how good the model is. A reminder again of the math:
For the GSS, the generative agents predicted participants’ responses with an average normalized accuracy of 0.85 (std = 0.11), calculated from a raw accuracy of 68.85% (std = 6.01) divided by participants’ replication accuracy of 81.25% (std = 8.11).
So the normalised score is that simple division on those two numbers, done at the very end. I’m not dying on the hill of this being inappropriate, but you need to be very careful around what this actually means, and the way they bleed ‘normalisation’ and ‘accuracy’ into the same sentence makes things really confusing here.
Let’s say there’s a GSS question with three options: A, B and C. Pretend it’s a political question where the answers represent Left-wing, Centrist, Right-wing viewpoints.
A participant answers Option B in the first round. The agent predicts their answer as being C on the first round, and gets it wrong. But two weeks later the participant comes back and changes to A.
By this method, the agent is compensated for that through the normalised accuracy process, i.e. because it’s harsh to expect a correct prediction when the participant themself is answering inconsistently.
Yet in this example, the prediction could be interpreted as even more wrong than before, while the agent’s ‘ability to predict participant responses’ will have increased when you do the math at the end.
But wait … it’s actually worse than that. Because this method doesn’t even measure response replication like that example. The simple division equation means this isn’t done on a question-by-question basis. Compensation is done across that participant’s entire answer set.
Meaning the real issue is that the denominator counts how many answers the person repeated, and thus also how many they changed, without caring whether the changes have anything to do with the ones the agent got wrong.
I know this is a mess, and trust me we are going to way further down the rabbit hole from here, but for now I think this example will help make it clear why this is a big deal:
Suppose there are just 10 questions. The agent gets 6 right and 4 wrong.
Two weeks later, the person changes 2 answers. They therefore repeat 8 of their original 10 answers.
The normalised score is 6 ÷ 8 = 75%, compared with raw accuracy of 60%.
But consider two possibilities (I can’t believe there’s no table support on substack):
Yes, the 75% will literally mean “generative agents replicate participants’ responses on the GSS 75% as accurately as participants replicate their own answers” and that’s technically true. But by reporting it, you’re suggesting that I’m supposed to think something on the back of that. So what does the 75% really mean in any practical sense?
We don’t know and the authors don’t know either, because the information contained in that first column does not exist in this methodology, because it is done with one big division at the end.
My real point is this: while it seems like a rigorous research practice, asking those people to come back and fill in the surveys again could have only possibly pushed up the normalised accuracy score of the agent. Literally. In fact, if they came back, answered even more stuff differently, further emphasising that no one really knows what they think about anything, then the agent would suddenly have a near perfect ability to predict responses (even though the raw prediction accuracy doesn’t change at all).
If they wanted that 85% to sound even more impressive? Invite them back for a third time, fourth time.
And you know what would actually happen eventually? The human response replication would dip below 69%, and now agents are predicting behaviour correctly more than 100% of the time. God the singularity is amazing, we are breaking the nature of probability wide open.
Now the riposte to that is that I’m being unfair and that a score above 100% is perfectly logical, it just means the model gets the first round answer right more than the participant repeats it. But my issue is how this number gets treated as if the top bound actually is 100%. When it escapes to the NYT but also within the research itself.
Back to that one sentence again:
For the GSS, the generative agents predicted participants’ responses with an average normalized accuracy of 0.85 (std = 0.11), calculated from a raw accuracy of 68.85% (std = 6.01) divided by participants’ replication accuracy of 81.25% (std = 8.11).
Which becomes this sentence in the abstract:
The generative agents replicate participants’ responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later, and perform comparably in predicting personality traits and outcomes in experimental replications.
Can you really use the words ‘85% as accurately as participants replicate their own answers’ without realising how misleading they are? Accuracy in this specific formulation can exceed 100% but ‘accuracy’ as it is actually used as a word in real life cannot. The people reading this abstract live in real life.
And when that reader knows this experiment has something to do with real people and AI agent guessing their answers, you don’t think you’re implying that the word ‘accurately’ is the same as the agent guessing the human’s answers ‘correctly’?
I predict _.
III.
Amongst all of that, you will have noticed that they’ve pivoted to only using the GSS to advertise their findings now, even though they used the entire barrage I showed you earlier. They say they ‘perform comparably’ on those other ones.
I think one fair reason the personality and economic game results feature less is because they’re captured by continuous measures, and this means correlations, which tend to get confusing when it comes to interpretation. Big five tests feature 40 total questions, combining into 5 constructs. If an agent generally gets where you place on each, while individually getting all 40 question predictions slightly wrong, I would still be impressed. But it does mean you can’t call a high correlation ‘accuracy’, so it would be very confusing if they started to claim they were predicting personality or something like that.
But something very interesting then appears when they get to the normalisation part here:
For the Big Five, the generative agents achieved a normalized correlation of 0.80 (std = 1.88), based on a raw correlation of r = 0.78 (std = 0.70) divided by participants’ replication correlation of 0.95 (std = 0.76).
…
The third component involved a series of five well-known economic games… On average, the generative agents achieved a normalized correlation of 0.66 (std = 2.83), derived from a raw correlation of r = 0.66 (std = 0.95) divided by participants’ replication correlation of 0.99 (std = 1.00).
Do you notice what the normalisation process does this time? Basically nothing.
Look at the participants’ replication correlations. More consistent than on the GSS. In fact, almost completely consistent.
What’s going on?
You may say that it’s possible people’s views on politics/religion/happiness/whatever will vary more than their personality and what they would do in lab-based economic games. And I say yeah, that is possible.
But consider another possibility. You have to keep in mind we aren’t actually measuring what these people do in their lives. What protests they go to/guns they own etc. We are measuring what button they will click to represent their view on that thing.
Now re-read exactly how they did this:
The first component of our evaluation, the GSS, is widely used across sociology, political science, social psychology, and other social sciences to assess respondents’ demographic backgrounds, behaviors, attitudes, and beliefs on a broad range of topics, including public policy, race relations, gender roles, and religion. Our evaluation focused on 177 core GSS questions, which we used to establish a benchmark for measuring the agents’ predictive accuracy.
If I paid you a flat fee to be a research participant online for me online, and you know that fee is at the end of all of these questionnaires and games, and one of the questionnaires I force you through is 177 multiple choice questions long, and you tell me that you will actually read all 177 questions you respond to, which mental institution would you prefer me to check you into?
I’ve never paid attention to a survey longer than 50 questions in my entire life. I simply respect myself too much. Yet this headline number is based off one this long? Imagine the mental state of the participant who’s 100 in and just loaded up another page of questions.
Imagine that same participant coming back two weeks later, and realising they have to fill out the entire thing again.
Would you be at all surprised if the average person doing that doesn’t end up dwelling on each question to long off enough to give a proper answer? Wouldn’t you expect their answers to vary out of pure fatigue across the two survey completions?
In fact, when you’ve just been hit with 200 questions on your social values, a quick personality test and one round of prisoner’s dilemma probably feel like respite.
Then add into the mix that the first time participants did the GSS, they did it directly after a 2-hour interview with an AI interviewer throwing questions on the same topics at them. So they probably aren’t keen to extend this research session any longer than they have to.
So what is likely to happen to the participant replication %? And what does that mean for what is going to get called agent prediction accuracy?
And what does that mean, all the way down the line, when it ends up in the New York Times, and the researchers allow their prediction accuracy number to be converted to ‘ability to predict human behaviour’?
IV.
This is now about to go several layers deeper. I do not judge you for backing out now, but I offer this to those who believe psychology has something to offer the world and is worth saving. By the end it will be clear that Simile are not the problem at all.
Ok, here’s the methodology again, bolding mine:
Since individuals often exhibit inconsistency in their responses over time in both survey and behavioral studies, we use participants’ own attitudinal and behavioral consistency as a normalization factor: the probability of accurately simulating an individual’s attitudes or behaviors depends on how consistent those attitudes and behaviors are over time.
It is time for us to now start treating words like they mean something. And the million-dollar two-billion-dollar word here is ‘behaviour’.
Simile CEO Joon Sung Park said this in his quote from earlier:
We demonstrated that using our architecture and the models we can actually predict people’s behaviors 85% as accurately as people replicate their own.
This is where we go from misleading to lying.
Hilariously, the only way he could possibly argue this statement to be true is if by behaviour he means the physical action of the participant clicking a button, rather than the behaviour the button is meant to represent via the topic of the survey. But even this would then have to be subjected to all of the issues with the 85% laid out above.
Let’s just say he means the survey measures, i.e. the GSS. What on earth does this have to do with behaviour? This is a survey of attitudes. Is having an attitude a behaviour?
Of course not. His own research recognises this. That’s why the bolded passages from the paper above refer to ‘attitudinal and behavioral consistency’. Where attitudinal refers to the GSS, and behavioural refers to the games. And the number for those was not 85%.
But this is the most important point that I hope will seem obvious: even if the behavioural parts of the study also had a 85% normalised accuracy, that sentence that Park says up there would still be a lie.
Games like the dictator game get referred to as behavioural because they involve the participants doing something. Which is fine. But an issue arises when you take that word and start applying it to other stuff, and this is where the issue stops being Simile’s and starts being Psychology’s.
This research is trying to get an agent to predict how someone will act in a specific lab game and that is it. Even if you do it perfectly, it tells you exactly nothing about a general ability to predict behaviour, and definitely nothing about any kind of real-life behaviour.
Psychology lab research is based on the notion that measuring what people do in hypotheticals might have something to do with how they would behave in a real version. This isn’t just a potentially faulty assumption, it’s actually a complete non-starter in ways that Psychology itself tends to not understand. Which has led to Park and Simile thinking they can use these tasks in their research when they really can’t.
Here is an example of what I mean.
What happens in the original story of the prisoner’s dilemma? Two prisoners are questioned separately. Each can either stay silent or betray the other. If both stay silent, they each get 1 year. If one betrays and the other stays silent, the betrayer goes free and the other gets 3 years. If both betray, they each get 2 years. They cannot communicate or know what the other will choose.
Game theory says they should always both betray and arrive at an inefficient equilibrium. Research findings show that people, when playing this in a lab study, co-operate more than economic rationality suggests.
Now, tell me what you think the limitation of that lab finding about what people would do is.
The thing I want to get across is that if you say ‘the situation is only a hypothetical so we can’t know for sure if that’s what people would do’ - then you are wrong. Psychology as Science will tell you that you’re exactly right, and you’ll get marks on your paper. But this is a lie.
Consider the story more. The police set up the entire incentive structure, and then inform the prisoners of this up front. They are very happy to let one man go if he betrays his partner. What does this tell you?
It tells you that this has to be a situation that fundamentally doesn’t matter that much to anyone. Why else would the police be designing a literal game? More importantly, it should just straight up confuse you when laid out, because this is a situation which does not happen. Because to be living in a world where lucky strategy gets you off scot-free, you must also be living in a world where prison time is basically not a big deal and who cares who ends up in prison.
I think it’s clear what I’m getting at, but one more example.
Another way the prisoner’s dilemma is often played out in the lab is through financial incentives. The Simile paper does this but they don’t supply enough data to break it down , so I’ll use this Nature paper as an example.
In this version, before participants choose whether to cooperate or defect, they are told that if both cooperate, they each get 5 points. If one defects while the other cooperates, the defector gets 7 points and the cooperator gets 1. If both defect, they each get 3 points. These points accumulate and determine how much real money they are paid when they are done.
In this study, the average participant played the game 3,720 times. And the average payout, in total, was $87.03. So in an average game, they are coming away with 2 cents.
(fyi they don’t say if average here means mean or median)
Now, tell me again what the limitation of observing how people co-operate in a given trial of this game is.
Step 1: Even though they are playing for money, and you should always want more money, this is a trivial amount. Although to be fair, the point is they should work out a consistent strategy, so over the ~4,000 games your decisions do matter more, probably the difference between earning ~$100 and ~$60 overall. How quickly do you see that return? Well there were max 20 sessions of 20 games, each lasting 35-minutes. So that’s almost 12 hours. Note that the participants don’t live in the lab, so add travel and admin of actually going. They get an extra $20 for attending all sessions, but this does not depend on their gameplay strategy. All in all, we are back to this being a really trivial amount of money for the effort you are putting in. Which means this study is limited because this is essentially a hypothetical to how they would behave if you were offering them an amount of money that really mattered to them in any given trial.
Right? Wrong.
Step 2: This is wrong because if you made this into a real-life dilemma, and multiplied the money 10000x, you cannot possibly be studying the same people. Because if you ever actually got offered a basic prisoner’s dilemma game, where the stakes were something like defect $1000, co-operate $600, the difference wouldn’t be that you now might now really dwell on the efficient strategy, and thus expose the limitation of the lab study when you defect way more often. The difference would be that you now must have found yourself in a world where money doesn’t matter. $1000 cannot be worth anything if some weird experimenter is just handing it out to you.
I know this sounds either obvious, or pointless, or both, but it’s very important in order to understand how to treat research findings.
A lab version of the prisoner’s dilemma is not a hypothetical version of real life. It is a hypothetical version of a hypothetical universe.
And the findings of the study relate to humans in that other universe, that neither you nor I exist in.
This is why calling these hypotheticals is wrong. We should borrow from Baudrillard, and call them Hyper-Hypotheticals.
And now what’s happening, is that Park and Simile are getting LLMs to predict what participants are going to do in a hyper-hypothetical, divide that by participant replication (for some reason), and turn that into a claim that they have devised a model that can accurately predict human behaviour.
Or phrase it this way: they are simulating what people might do in a simulation of a simulated universe, and those are all separate layers of simulation, not the same one. And they are getting it right 2/3rds of the time when their training data already contains everything known about how people tend to simulate themselves in simulated environments.
But wait, it doesn’t stop there. That’s only for the research the company is based on. Simile, the company, are for simulating new behaviours. So you can throw your new intervention/policy/product into their simulated society of agents, and find out what would really happen in real life.
And it’s going to be 85% accurate, right? Right?
Oh wait, that number was based on the weird interview input that covered the same topics as the outcomes they were measuring, and didn’t observe behaviour as a consequence of an intervention at all? And Simile aren’t going to interview a thousand people just for your new request are they? So what does this research even have to do with their actual product in any shape or form?
“If you’re able to simulate the world, you can basically test out countless interventions,” said Joon Sung Park, Simile’s chief executive and co-founder. “This is sort of the window into humanity.”
I’ve looked through the window, Dr. Park. I don’t think I’d recommend it.
V.
The thing that’s partly funny at this point is that I literally don’t care about Simile. Honestly, wish them well, hope that valuation goes to $10B. Let me know if you need help coming up with even more headline-grabbing numbers.
But in general any query you would consider using a simulated batch of participants like this for, you will generally get as much success from one LLM chat query. Same inputs and educated guessing either way. The only difference is the agents might get you a more logical distribution, but this is also largely done by just telling them to, so you can just add that to your own prompt too.
Anyway, reason I don’t care about them is because they’re just a market research company at the end of the day, and they’re just going to be selling their ‘simulations of the future’ to other knowledge workers who actually pay for market research.
And when it comes to this stuff, I think of a Tim Dillon observation on people who get scammed:
Look, if you think the Long Island medium is talking your dead daughter, you sort of deserve to get your money stolen.
If you're in the back of a Italian restaurant in Long Island and you've paid $300 to sit there with a bunch of other meatballs and this woman walks out and she pretends like she's talking to your daughter Nicole who died in a boating accident and you believe that … well you know, or maybe you don't know, but I know, that you should get robbed.
You should have your money taken from you because it'll just go to something else stupid, right? You're not gonna invest in the next thing that becomes Uber.
My mother took money and bought beanie babies, you know at McDonald's and tried to retire on them like collecting them. Would her money have been any worse giving it to Gary Vaynerchuck or the Long Island medium?
In this metaphor, Simile are the medium and everyday market research is a beanie baby.
(to be honest, I’ve been close to enough to market research that is so remarkably bad and hacked towards a certain result that I always wonder, why not just make it up? literally why not just make up stats about your brand and put them on billboards? it’d be cheaper and quicker. no one checks that academics analyse their data right, you think they’re checking you?)
The reason I’m 6,000 words into a post right now is this:
VI.
A few weeks back, Scott Alexander argued that the replication crisis gave people the impression that all psychology research is bad when it is, in fact, mostly good. He points out that it wasn’t really Psychology as Science that had a replication crisis, it was specific areas of social psychology like priming. As evidence for his thesis, he examines a modern Psychology textbook, where he observes:
Social psychology - which includes social priming, power poses, nudges, stereotype threat, the Stanford Prison experiment, the Wansink portion size study, and everything else you love to hate - is in chapter 12 of 16. The other fifteen chapters are comparatively innocuous.
But then an interesting note comes about the ‘good’ 15 chapters:
Not all of this is perfect. I don’t think anything that “Neo-Freudians: Adler, Erikson, Jung, and Horney” said has held up very well.
This is where we see the general acceptance of what it means to be ‘good’ or ‘acceptable’ psychology i.e. whether it is measurable and replicable within the confines of a research paper. And this is the lie that allows a company like Simile to genuinely believe they are on the forefront of insight about human nature.
Now what I’m saying about the limited nature of psychology experiments is already incorporated into this viewpoint, i.e. that experiments should be evaluated on whether they have good construct validity, are you measuring what you’re implying your measuring, and external validity, will this lab result translate to some real-world equivalent. But my issue is that debating those things assume that you are correct to even be operating within that paradigm. You’re debating the answers without debating the question.
Here’s an example.
Data Colada have done another expose on the Dan Ariely stuff. Study says people signing up-front makes them more honest in some self-reporting that happens later. Maybe something to do with exposing yourself makes you feel vulnerable so you’re less inclined to try get away with something, or maybe declaring honesty makes you identify as an honest person, so you behave accordingly. If you’re a government and a tax policy can only work properly if people are honest when they fill out return forms, this might be worth a try.
So it turns out someone involved fabricated the numbers to get the result.
What happens now is that researchers and practitioners who based real life programs off of this result are angry at the fraud, and tell their stories of taking this intervention out into the wild and finding no effect. The fraud makes this all make sense.
But this was already a crazy thing to do. The finding was based on a highly contrived experiment involving insurance, and policymakers saw that finding, switched it to a context of tax returns, upgraded it to a sample adults who care a lot about their tax returns, run it in the third-world, and feel ‘betrayed’ when it doesn’t work. (true story)
Pre-registration, bigger sample sizes, and better peer review don’t change that. It’s all still buying the same story, i.e. that psychology, economics, behavioural science if treated properly can be examined like chemistry or biology. They can’t. If I combine two chemical compounds together, and they produce some new compound, that’s a finding. That’s a rule. The same reaction would happen in Hawaii or Siberia, or in 10,000 BC. But you can’t similarly try to model human behaviour in dynamic, culturally-specific contexts and expect universal, replicable results.
When Alexander suggests that priming should be removed from what we consider good psychology, he is basing that off of it not replicating. But the crisis of things weirdly not replicating should have been a strong hint that the approach itself is off.
But when you take those failed replications and decide the fundamentally most important thing about whether something is a legitimate finding is replication, you create a paradigm where things like the GSS, and games like the Prisoner’s Dilemma count as leading examples of the science. But you have no idea if those constructs mean anything to anyone. You don’t even stop to ask, you just run the numbers.
And this is the exact mindset that leads to the field pushing towards synthetic data as legitimate research practice. Because if you want, you could make your own simulated sample, and replicate Simile’s results. You could hire a human sample online, have them fill in a survey, and probably predict their answers really well. But what is totally and completely absent is any question of whether anything the agents or the human are doing has anything to do with anything people do in their lives.
VII.
The reason I pick on Scott singling our the Neo-Freudians is not holding up well is because this fundamentally comes down to whether you, as a researcher, policymaker, just general enthusiast, have a coherent model of humans and human behaviour.
I would say at least Jung has that, and saying it ‘doesn’t hold up well’ is a misunderstanding of what we’re doing here. The reason there is no empirical, experimental, replicable evidence for what Jung is saying because Psychology is not a discipline that can be analysed like that. If we are pontificating on things like the unconscious, desire, motivation, why people actually do anything, you are operating outside of hard scientific research, and if you want to operate within research, then you can’t claim to ever be touching those things in any real way.
I’m not saying you have to accept Jung, but the process of advancing Psychology is to explain why don’t accept him with your own model.
And my issue is that the research process that ends up with Simile and synthetic participants, by which I mean the current discipline of Psychology, is that involves no coherent model of itself at all. Which is why synthetic participants trained by us can never provide anything useful, beyond what you could guess with your own intuition, or a basic ChatGPT prompt.
I do believe eventually there likely will be a form of synthetic participant and simulated world that teaches us about human nature and behaviour, but the only logical route there is an AI that is basically super-intelligence, that constructs its own model of humans from the ground up, and replaces all the research we already have.
Because in this current model, we don’t understand people enough, and we don’t understand agents enough, to make head nor tails of what we should be putting in the model, and what anything that happens within the model means.
Here’s a question you need to be able to answer if you want to try simulated environments for behavioural research:
If I simulate a town of 1,000 agents, based of the personas of real people I did interviews with, who are the behavioural findings about?
Simile suggest their simulated worlds can let you know what the effects of a policy, or a marketing campaign, would be. So let’s say I introduce gun control as a new policy to my simulated town. The agents are politically distributed, some happy, some angry. Then 10 of them decide they’re not going to have it, sneak out at night, and shoot up the entire mayor’s office.
The question is: is that a finding about:
Humans
LLM-based Agents that think they’re humans
LLM-based Agents themselves
LLM-based Agents that know they aren’t humans but are playing along
What LLMs think about Humans
What LLMs think the Human researchers want to see
I have no idea, and again, Simile can’t possibly know either. Because to even attempt to answer that question you need a deep model of human psychology and agent psychology that you don’t have, and the research paradigm you are doing this from would not allow you to have. Because it requires you to believe things that you can’t put in a paper, because they aren’t measurable and thus not replicable.
Another issue is that rare people do actually do extreme things sometimes. It’s extremely easy to model a population based on something like political preferences, and thus guess how a policy will go down. But if you are going to claim that you can simulate human behaviour, you don’t just get to take the easy bits and use that as evidence your model is good.
If you model an American town, and at no point in your 20 year simulation does a disgruntled kid shoot up a school, or a boss inappropriately use his power, or does anything weird or horrible happen, then you haven’t built a generalisable model of anything. What you’re going to do is just repeat back what the training data had in it, which is why the interview-method and GSS outcome is the only way they get that result.
But if you as a researcher think you can take that into your practice to learn something new, or work out how people might respond to something new, you have lost complete sight of what you are doing.
VIII.
When I suggest that Psychology isn’t science, and that any attempt to build it from within science is bound to fail (while declaring it’s succeeding), a justified riposte would be to say that the alternative, i.e. the old way of writing, theorising, and basically guessing, hasn’t exactly produced anything great either.
And maybe it hasn’t. But I think that’s because what we are trying to do here is infinitely hard.
Psychology is prone to envy of physics or biology, for being gold-standard hard sciences, where the real smart people go and do world-changing things.
But despite the name hard science, those are far easier to practice.
Because you can measure things, and validate them, and replicate them. Because they relate to the external world that science exists in. But it is as of now unknowable how the internal world is structured. There can only be guesses, but doesn’t mean you shouldn’t try theorise.
However, if you attempt to theorise from the ground-up, and allowing the numbers to tell you what’s going on, you’ll always end up at least one layer of simulation deeper than you think you are.
At least that’s true for now.
In that when I say Psychology isn’t a science, I just mean it isn’t this science. There could well be a way to measure desire, motivation, emotion, behaviour and predict the future versions of all those things. There could definitely be some way to create a digital version of your psyche that market researchers can use to their advantage. But it requires a total reinvention of the field, because if that science exists, we do not know how to do it yet.
Super-intelligence could well take care of that.
Which does unfortunately make one think: what it would feel like to be a perfectly-modelled agent, masquerading in a simulated world built to predict human behaviour?
It’d probably feel a bit like this.








