A new study led by Harvard psychologist Ashwini Ashokkumar and published in the journal Nature suggests that large language models (LLMs) such as GPT-4 can surprisingly well predict the outcomes of many experiments in social sciences.
However, the results come with a caveat: a system that predicts human reactions is not necessarily a system that understands human behavior. “Synthetic respondents” or “silicon samples” are not a direct replacement for real people.
Ashokkumar and her colleagues collected 70 real experiments conducted in the United States, involving nearly 120,000 participants. They provided GPT-4 with descriptions of hypothetical respondents along with experimental messages and survey questions, asking it to assess how these people would react under different conditions. Comparing GPT-4’s predictions with real results, the researchers found a strong correlation: the model could often distinguish more effective and less effective interventions. This is an impressive result, suggesting that LLMs can capture significant patterns in the social world, at least in the text-based surveys studied in the research. However, this is not proof that AI has found a reliable way to replace human-involved research.
Useful predictions do not equal understanding
American scientists Lisa Messeri and Molly J. Crockett warn that artificial intelligence (AI) systems can create “illusions of understanding” — results that seem convincing and useful but prompt users to overestimate what was actually understood. LLMs can generate plausible explanations or convincing predictions, but this may reflect complex pattern matching rather than a true understanding of the mechanisms behind observed behavior.
For example, the new study showed that GPT-4 often accurately assessed the likely effects of different interventions but systematically overestimated their magnitude by approximately twice compared to real results. This difference is critical: the tool can inform researchers that message X is likely to work better than message Y, but it may remain unreliable regarding whether the difference will be minor, moderate, or substantial.
A powerful tool for pilot studies
When used carefully, this information can still be very valuable. Researchers often conduct small pilot studies before launching expensive experiments. Such studies help refine interventions and assess whether the expected effect is large enough to justify conducting a larger study.
Predictions generated by LLMs can complement these pilot studies. For example, researchers can model how different demographic groups will respond to several versions of vaccination messages, workplace interventions, or different policy wording options. Meanwhile, the study showed that combining LLM predictions with human predictions was more accurate than each of these sources separately. The most useful future may lie not in AI replacing researchers or study participants but in helping determine where to direct limited human resources more effectively.
The temptation of “silicon sampling”
The idea of “synthetic respondents” or “silicon sampling” is interesting not only from a scientific point of view. It is increasingly being discussed in the field of sociological surveys, marketing research, and public consultations, where proponents see opportunities for faster and cheaper testing. At the same time, critics warn that this could undermine trust if simulations are passed off as real public opinion.
For a politician who wants to know the public’s reaction to a new tax policy or a company evaluating a new advertising campaign, LLMs can provide a quick and plausible answer. But this is not the same as measuring public opinion. A regular survey collects responses from people living in a particular society at a specific time. A synthetic sample, on the other hand, is based on patterns embedded in the model’s training data, query constructions, and its limitations. It can reproduce individual elements of human judgment but lacks life experience, local knowledge, and real interest in the issue being studied.
This gap can be particularly important for new issues, marginalized communities, rapidly developing events, and population groups underrepresented in online data. Ashokkumar and her colleagues found that the model generally worked well with different demographic groups but also recorded some differences in accuracy favoring samples of white respondents and supporters of the Republican Party in the U.S. Without careful calibration, synthetic respondents can reproduce dominant patterns in available data, smoothing out differences or minority positions.
Similar applications — and associated risks — also arise in AI-supported expert panels, forecasting, and Delphi method consensus tools. They can make expert forecasting and discussion processes faster and more accessible, but model-formed consensus can also hide real and useful differences rather than help resolve them. “Silicon sampling” can be useful for generating hypotheses and testing assumptions, but it cannot replace research involving real people. The risk lies not in using synthetic agents but in mistakenly taking the model-created substitute for the actual population.
The risk of optimizing harmful persuasion
The same predictive ability of AI can be misused. The study authors tested whether GPT-4 could identify social media content likely to reduce intentions to get vaccinated against COVID-19. Although the model may refuse to directly create anti-vaccine messages, it can still help identify which harmful messages from an existing set of options are most likely to be the most effective. This underscores the need for protective mechanisms that go beyond blocking obviously harmful queries. Systems may also need protection against using AI to rank, optimize, or target harmful persuasion.
Studies using commercial LLMs such as GPT-4 are also vulnerable to changes made by model developers. Such models can change or even stop working without notice, making it difficult or even impossible for other researchers to verify or replicate the results obtained.
The main conclusion is not that AI prediction is futile or, conversely, magical. Large language models can become a valuable tool for social sciences, helping researchers test new ideas, prioritize interventions, and model different scenarios with minimal costs. At the same time, researchers warn: the more accurate AI predictions become, the easier it is to mistakenly take them as evidence of a true understanding of human behavior.
Source: Phys.org



