ChatGPT and Gemini Don’t Take Risks Like Humans, and They Don’t Even Agree With Each Other

Large language models are increasingly being used to simulate people in research, test hypothetical decisions, and predict how humans might respond to different choices. But a new study suggests researchers should be careful about treating artificial intelligence as a stand-in for human behavior.

Researchers compared the risk preferences produced by three large language models with decisions made by real people in Sydney, Hong Kong, and Nanjing. The models were given demographic profiles and asked to make choices involving lotteries with different levels of risk.

Human and artificial intelligence comparing a guaranteed reward with a riskier choice, illustrating research on differences between AI and human risk preferences.

Researchers found that large language models did not consistently reproduce human risk preferences, with different AI models showing markedly different patterns of risk-taking.

The results showed a striking divide. The two GPT models tended to be more risk-averse than the human participants, while Google’s Gemini tended to be more willing to take risks.

Even the language used to prompt the models changed their apparent attitudes toward risk.

The findings, published in Communications in Transportation Research, suggest that an AI model’s apparent decision-making style may depend as much on the model and prompt as on the characteristics of the person it is supposed to represent.

Key Takeaways

  • Researchers compared three large language models with human decision data from Sydney, Hong Kong, and Nanjing.
  • The two GPT models generally produced more risk-averse choices than the human benchmarks.
  • Gemini showed the opposite tendency and was more risk-seeking.
  • Switching prompts from English to Chinese generally shifted simulated decisions in a more conservative direction.
  • The models struggled to reproduce the diversity of risk preferences found among actual people.
  • The findings do not establish that one AI model is safer, more rational, or better at making real-world decisions than another.

Researchers Tested Whether AI Could Behave Like Real People

The study by Liu and colleagues (2026) examined an increasingly important use of large language models: using them to simulate human behavior.

Instead of simply asking an AI system what it would do, the researchers constructed role-playing prompts based on demographic information from real survey participants in Sydney, Hong Kong, and Nanjing.

The models were then presented with abstract lottery choices.

These types of experiments are commonly used to measure risk preferences. A participant might have to choose between an option offering a relatively predictable outcome and another offering a potentially larger payoff accompanied by greater uncertainty.

Researchers analyzed those decisions using a standard economic framework known as constant relative risk aversion, or CRRA.

The purpose was not to determine which model made the “correct” choices. It was to see whether the distribution of AI-generated decisions resembled the risk preferences observed among real humans.

It often did not.

GPT Models Were More Cautious, While Gemini Took More Risks

One of the clearest findings was that there was no single “AI risk preference.”

The two GPT models tested by the researchers tended to produce choices that were more risk-averse than the human benchmarks. Gemini moved in the opposite direction, producing choices that were more risk-seeking.

That distinction matters because large language models are sometimes discussed as though they represent a common form of artificial intelligence whose responses can be generalized across systems.

The experiment suggests otherwise.

Different model families can generate systematically different behavioral patterns even when they are presented with similar decision problems.

This also means that a researcher using one model to simulate human participants could potentially reach different conclusions from a researcher using another model.

The Language of the Prompt Changed the Results

The model itself was not the only factor that mattered.

Researchers also examined what happened when prompts were presented in different languages.

Switching from English to Chinese generally shifted the simulated risk preferences in a more conservative direction.

That is particularly important for research attempting to use AI-generated participants across countries or cultures.

Ideally, changing the language of an otherwise comparable experiment would not fundamentally change the underlying risk characteristics the model is supposed to reproduce. But the findings suggest language can become part of the behavioral outcome.

The researchers therefore cautioned that model-generated behavior can be influenced by linguistic as well as model-specific factors.

AI Also Struggled to Reproduce Differences Between People

Matching the average human response is only part of the challenge.

Real populations contain people with very different attitudes toward uncertainty. Some readily accept risk for the possibility of a larger reward, while others strongly prefer more predictable outcomes.

The researchers found that the large language models did not reliably reproduce this variation.

Depending on the model and condition, simulated risk preferences could become too concentrated around similar responses or display patterns that differed from the human comparison groups.

This is an important limitation if researchers want to create thousands of simulated people and use their responses as substitutes for actual survey participants.

A model might produce a plausible average while still failing to recreate the underlying diversity of the population.

This Does Not Mean ChatGPT Is “Cautious” and Gemini Is “Reckless”

The results are tempting to interpret as personality traits.

They should not be.

The experiment measured outputs produced under specific prompts and lottery-choice tasks. It did not establish that GPT systems possess an inherently cautious personality or that Gemini is inherently reckless.

Large language models do not experience uncertainty, fear of loss, financial pressure, or anticipation of reward in the way humans do.

The study instead demonstrates that different systems can produce measurable patterns that resemble different risk attitudes when placed into experimental decision-making tasks.

Those patterns may also change with prompting, language, model updates, or experimental design.

The distinction becomes especially important when AI systems are used as simulated human subjects.

How Does This Compare With Previous Research?

The new findings fit a growing body of research showing that artificial intelligence can reproduce some features of human decision-making while diverging substantially on others.

Xiao and Wang (2025) compared large language models with human responses across a range of social decision scenarios. Their results showed that the models could display systematic behavioral tendencies while still differing from human participants in important ways, including responses involving risk and social context.

That finding is broadly consistent with the new study’s conclusion that apparently coherent AI decision patterns should not automatically be treated as accurate representations of human behavior.

Suh and Atlas (2026) also examined time and risk preferences across generative AI systems. Their results suggested that behavioral patterns can vary substantially across models and do not necessarily reproduce the distributions observed among humans.

Taken together, the studies point toward a similar conclusion: AI-generated decisions can look systematic without necessarily being human-like.

The Liu study adds an important cross-cultural dimension by showing that the language used in the prompt can also alter the apparent risk profile produced by a model.

Why Researchers Want to Simulate Humans With AI

There is an obvious attraction to using large language models as synthetic research participants.

Traditional behavioral studies can require recruiting hundreds or thousands of people, paying participants, administering surveys, and collecting data over weeks or months.

An AI system can generate thousands of responses almost immediately.

Synthetic participants could therefore be useful for testing questionnaires, exploring hypotheses, running preliminary experiments, or identifying questions worth investigating with real participants.

But the usefulness of that approach depends on whether simulated populations behave sufficiently like the populations they are intended to represent.

The new findings show why that assumption needs to be tested rather than taken for granted.

If one model is systematically more cautious than humans and another systematically more willing to accept risk, simply changing the AI system could alter the apparent outcome of an experiment.

The Study Has Important Limitations

The researchers examined risk preferences through controlled lottery-choice tasks. Real-world decisions are considerably more complicated.

Choosing between hypothetical lotteries is not the same as deciding whether to change careers, make an investment, trust another person, or accept a real-world risk.

The study also evaluated particular model configurations. Large language models are updated frequently, meaning results obtained from one version should not automatically be assumed to apply to later versions.

Prompt design presents another challenge.

The finding that changing language altered risk preferences illustrates how apparently minor experimental choices can affect AI-generated behavior.

And although the models were given demographic profiles derived from real participants, role-playing a person based on demographic characteristics does not mean the model reproduces that individual’s psychology.

What the Evidence Actually Shows

The study does not demonstrate that artificial intelligence makes better or worse decisions than humans.

It addresses a narrower question: can off-the-shelf large language models reliably reproduce human risk preferences when researchers ask them to simulate people?

For the models and tasks tested, the answer was not consistently.

The two GPT models leaned toward greater risk aversion than the human benchmarks, while Gemini leaned toward greater risk-taking. Prompt language shifted the results, and the models struggled to reproduce the full diversity of human risk preferences.

That does not make AI-generated participants useless. It does mean researchers may need to calibrate and validate them against real human data before treating synthetic responses as substitutes for actual people.

For now, an AI model can produce something that looks like a human decision.

Whether it represents how humans would actually decide is a separate question.

References

Liu, J., Song, B., Dixit, V., Wu, C., & Jian, S. (2026). Can large language models capture human risk preferences? A cross-cultural study. Communications in Transportation Research, 6(2), Article 9640025. https://doi.org/10.26599/COMMTR.2026.9640025

Suh, W., & Atlas, S. A. (2026). Time and risk preferences among generative AI systems. Journal of Business Research. https://doi.org/10.1016/j.jbusres.2026.116274

Xiao, F., & Wang, X. T. (2025). Evaluating the ability of large language models to predict human social decisions. Scientific Reports, 15, Article 32290. https://doi.org/10.1038/s41598-025-17188-7