Logo
FrontierNews.ai

ChatGPT Boosts Student Work Quality, But Causal Reasoning Unlocks Rarer Ideas

A new classroom experiment reveals that ChatGPT and critical-thinking training deliver complementary benefits, not interchangeable ones. OpenAI and Bocconi University tested 1,053 first-year economics and business students on a marketing assignment, comparing the effects of GPT-4o access, structured causal-reasoning training, both interventions, or neither.

What Did ChatGPT Actually Do for Student Performance?

ChatGPT access produced measurable improvements in how students' written recommendations were evaluated. The model increased the awareness-and-usage score by an estimated 0.862 points from a baseline of 2.09 on a five-point scale, representing a 41% improvement in task performance. Students with GPT access also produced clearer logical structure and responses that more closely resembled expert recommendations, suggesting the model helped students articulate ideas more professionally.

The experiment used a rigorous 2x2 design, assigning 249 students to a control group, 256 to causal training alone, 197 to GPT access alone, and 351 to both interventions. Each student had 45 minutes to write a maximum of 180 words recommending ways to increase alumni awareness of the university's merchandise store. Master's students and domain experts then rated the submissions across multiple dimensions.

Why Did Causal Training Produce Different Results?

Causal-reasoning training took a different approach. Rather than providing a tool, the intervention taught students a structured method for thinking through cause-and-effect relationships. Students played a game about causal chains, identifying mechanisms, constructing coherent explanations, and specifying conditions under which claims might fail. The control group played a placebo version without reasoning instructions.

The results showed that causal training increased the semantic diversity of ideas both within individual answers and across students' responses. This means students trained in causal reasoning produced ideas that were less typical of their peers' work, measuring distance from the sample's common solutions rather than verified novelty in the wider world. However, causal training did not increase conventional evaluation scores on the rubric, which rewarded alignment with familiar marketing goals rather than originality.

How Did the Two Interventions Interact?

Students who received both ChatGPT access and causal training kept the diversity benefits of reasoning training while also gaining the output quality improvements associated with the model. The combined condition strengthened some measures of coherent logic and falsification reasoning, suggesting that GPT could help express the reasoning introduced by the causal intervention. However, the conventional evaluation score was largely attributable to GPT access alone, with no statistically meaningful additional scoring effect from the interaction.

This finding highlights a critical distinction: more ideas and more varied ideas remained separate outcomes. GPT alone did not increase between-student diversity, meaning the model's strength lay in quantity and polish rather than originality.

What Are the Key Limitations of This Study?

The researchers were careful to note what their findings do and do not prove. The experiment measured output from a single class exercise, not durable learning or long-term retention. The study did not test whether students recalled information weeks later, whether they could apply the skills to a different subject, or whether critical-thinking ability improved months after the intervention. The assignment itself favored a language model, since it involved widely documented marketing concepts, a tightly specified problem, and a short written recommendation.

Additionally, the rubric used to evaluate answers rewarded alignment with established marketing goals more readily than novelty. In follow-up analysis, mechanism identification, falsifiability, and between-response diversity were negatively associated with the rubric score, meaning that more original answers often scored lower under the conventional grading system.

Steps to Interpret AI and Reasoning Training Results Separately

  • Measure Output Quality Independently: Evaluate how well a response addresses the assigned task using a predefined rubric, separate from measures of originality or reasoning depth.
  • Track Reasoning Explicitly: Assess whether students identify causal mechanisms, construct falsifiable claims, and explain the conditions under which their recommendations might fail, even if these elements don't boost conventional scores.
  • Assess Originality Carefully: Measure semantic diversity and distance from typical solutions, but recognize that uncommon ideas can be distinctive without being feasible, useful, or correct.
  • Test for Lasting Learning: Conduct follow-up assessments weeks or months later to determine whether students retain knowledge or can apply skills to new problems, rather than relying on immediate performance.

The immediate implication for educators and workplace assessment is about what a score actually represents. A polished, expert-like response can demonstrate strong task performance while revealing little about retention or independent expertise. Conversely, a more unusual answer can score poorly when the rubric values conformity to established criteria.

For organizations considering AI tools in training or hiring, the findings suggest that ChatGPT and reasoning training serve different purposes. If the goal is to improve the quality and coherence of written work, AI access delivers measurable gains. If the goal is to develop independent, creative thinking, structured reasoning training may be more effective, though it requires assessment methods that reward originality rather than conformity.