The experiments compare people with an advanced LLM called o3 across situations that demand planning and back-and-forth interaction. When targets’ mental states were visible, o3 performed very well. When persuaders needed to ask questions or infer those states before persuading, o3 struggled, while humans managed the multistep planning more reliably. In later tests where targets were human or where real belief change was measured, o3 often matched or beat human persuaders, especially when it could rely on surface-level rhetorical moves rather than probing another’s mind.

These findings matter for anyone thinking about how AI will influence choices and norms. They suggest models can be persuasive without modeling internal beliefs the way people do, which raises questions about when an AI’s influence reflects true understanding versus effective tactics. Click through to explore how this work connects to human potential, growth, and inclusive decision-making, and what it means for designing systems that respect people’s goals and mental states.

Abstract
A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, noninteractive question-and-answer benchmarks. However, theoretical work in the field suggests that first-personal interaction is a crucial part of ToM and that such predictive, spectatorial tasks may fail to evaluate it. We address this gap with a novel ToM task that requires an agent to persuade a target to choose one of three policy proposals by strategically revealing information. Success depends on a persuader’s sensitivity to a given target’s knowledge states (what the target knows about the policies) and motivational states (how much the target values different outcomes). We varied whether these states were Revealed to persuaders or Hidden, in which case persuaders had to inquire about or infer them. In Experiment 1, participants persuaded a bot programmed to make only rational inferences. OpenAI’s reasoning model o3 excelled in the Revealed condition but performed no better than chance in the Hidden condition, suggesting difficulty with the multistep planning required to elicit and use mental state information. Humans performed moderately well in both conditions, indicating an ability to engage such planning. In Experiment 2, where a human target role-played the bot, and in Experiment 3, where we measured whether human targets’ real beliefs changed, o3 outperformed human persuaders, significantly so in the Revealed conditions. These results suggest that effective persuasion can occur without explicit ToM reasoning (e.g., through rhetorical strategies) and that o3 excels at this form of persuasion. Overall, our results caution against attributing human-like PToM to LLMs based on predictive, spectatorial benchmarks while highlighting their potential to influence people’s beliefs and behavior.

Read Full Article (External Site)