What makes a good student paper, and which parts of it can AI improve? A randomised experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment.

But the study's most interesting finding is not the grade bump. The experiment exposed what the grading rubric actually rewards — and there is something uncomfortable there.

This guide unpacks the experiment, the finding, and what was not shown.

1. How the experiment was set up

In November 2025, 13 sections of an introductory management course were randomly split into four groups:

  • Control — no intervention
  • Causal reasoning lesson — a short teaching intervention
  • GPT-4o access
  • Both

The task: write marketing recommendations for the university's merchandise shop in up to 180 words. The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome.

2. Why this experiment stands apart

Most claims about AI's effect on education are observational: "grades went up", "assignments changed". Observation does not show cause — the grades could have risen for another reason.

This study sits in a different class: it is a randomised experiment. Sections of the same course were assigned to four groups by lot, so the difference between groups comes from the design rather than from students selecting themselves. That is the only method that supports a causal claim.

The second point of rigour is in the marking: the main performance score came from human graders who did not know which group each text belonged to. Blind assessment removes the "I marked it differently because I knew AI was used" effect.

The third is scale: 1,053 students is a large number for a classroom experiment of this kind.

Those three are reason enough to take the finding seriously. But the same rigour also clarifies the limits — as section six shows, even a well-designed experiment only demonstrates what it measured. What was measured here is the grade on this assignment.

3. What GPT-4o did to grades

Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their answers contained about two more ideas on average, were more logically coherent, and more closely matched the recommendations of three subject-matter experts.

Even after the researchers controlled for argumentation quality, number of ideas, idea diversity and text properties, a measurable GPT-4o advantage remained. The authors attribute it to higher content quality — not to greater student knowledge.

4. The lesson did not raise scores but changed the thinking

The causal reasoning lesson did not raise traditional scores. Work from those students actually scored slightly worse on average.

But the same students:

  • More often explained why a proposed action should work,
  • More often stated under what conditions it might fail,
  • Generated more diverse ideas that diverged from what their peers wrote.

Combining the lesson with GPT-4o added no further boost to traditional scores. On causal reasoning markers the two complemented each other, and the greater idea diversity from the lesson group held up.

5. The real finding: what the rubric punishes

This is the most important part of the study. Looking at which properties correlated with the score produces two lists.

Raised the score: more ideas, more coherent arguments, greater idea diversity within a single answer.

Lowered the score: stronger falsifiability, more detailed explanations of how proposed actions would work, greater divergence from other students' ideas.

In other words, the student who writes "this recommendation fails under these conditions" scores lower than the one who does not. Originality was not penalised across the board; but on this assignment the traditional score mainly rewarded well-structured answers that stayed inside the expected solution space.

The authors' conclusion is plain: diversity and originality need to be explicitly built into grading criteria if they are supposed to count.

And the sentence that follows from it: current grading systems measure polish, structure and completeness, but not learning and understanding. That is what makes AI so easy to use as a cheating tool. The model is already excellent at everything the rubric rewards.

6. What the study does not show

This section should not be skipped, because it seriously bounds the result.

No follow-up test. Students were never asked to demonstrate what they understood or retained without ChatGPT. The authors explicitly acknowledge it remains unclear whether the GPT advantage came from knowledge students actually gained or simply reflected better output with assistance. The study shows improved graded performance on this assignment, not improved learning.

Narrow scope. One university, freshmen, a narrow marketing task. Randomisation happened across 13 class sections rather than individually among all 1,053 students.

There is a conflict of interest. The main performance score came from human graders blind to group assignment — that part is sound. But additional measures such as causal reasoning and idea diversity were evaluated using models from OpenAI and Anthropic. More importantly: OpenAI provides the technology being studied and was involved in the research. Several authors work at OpenAI or were employed there during the study.

7. What remains when the AI is taken away

Other studies answer that question, and the picture is darker. Research so far clearly suggests that what matters is not whether students use AI but whether it supports their own thinking or replaces it. When it replaces it, students suffer.

  • A study covering more than 500,000 US college grades: after ChatGPT's launch, top grades increased most in writing- and programming-heavy courses with a large homework component.
  • Controlled experiments: after AI was taken away, participants performed worse than a control group. The drop was steepest among those who had mainly used GPT to get direct answers.
  • A 30-month study of more than 26,000 students in China: homework got better and faster with AI, but exam scores dropped. On later entrance exams, results were 18 to 24 percent lower over the long term.

The most instructive detail in that last study: students who spent roughly the same amount of time on homework as non-users despite having AI access did not show comparable declines. The harm comes not from the tool but from spending the time it saves on something other than thinking.

8. What a teacher can do

The practical takeaway is not a ban but the rubric. The authors' recommendation is direct: if diversity and originality are supposed to count, write them into the grading criteria. That can be done without changing the assignment itself.

Three additional criteria look useful in concrete terms. Falsifiability: scoring the answer to "under what conditions does your recommendation fail" rewards exactly where model output is weakest. Mechanism: asking "how does this action produce the outcome" measures the link between ideas rather than the count of them. Divergence: if going somewhere different from the rest of the class is rewarded instead of penalised, a tool that converges on the average loses its edge.

What the three share: none of them measures what the model is good at, and all of them require the student's own thinking. All three can be added to existing assignments.

9. For the student

The finding from the 30-month study in China is the most practical guide here: the harm comes not from the tool but from spending the time it saves on something other than thinking. Students who kept putting in the same hours showed no decline.

So the question is not "should I use it" but "what do I do with the time I gain". The difference between having a model draft something and then thinking over it, versus submitting its output as is, shows up not in the grade but in the next exam.

What to watch

The same argument arrives everywhere; the models are the same models and the assignments are the same assignments. Three things will decide it.

First, the rubric itself. The study's most solid finding is not about AI but about grading: a system that measures polish is defenceless against a tool that is good at producing polish. Banning the tool without changing the measure does not solve the problem, it hides it.

Second, the kind of evidence. The question to ask of every future claim: was anything measured after the AI was taken away? If not, that study measured output, not learning.

Third, who is doing the measuring. In a study where the company supplying the technology took part in the research, the framing of the result deserves attention. Indeed OpenAI presents these findings mainly as a problem for grading systems — which is the least troubling reading available for its own product.