The AI Tutor Nobody Spoke To: A Middle School Experiment

Oreopoulos and Low (2026) spent two years tracking students across 18 Tennessee middle schools (roughly ages 11 to 14). During remedial maths classes, a randomly selected group of students was given Khan Academy with its AI tutor, Khanmigo, configured to coach rather than spoon-feed answers. While 96% of students gave it an initial try, engagement rapidly evaporated.

The median student messaged the bot on just a third of their practice days, and in only 17% of the sessions in which they had made an error. When they did interact, their messages were mostly bare answers or clicks on suggested prompts. The students did improve in maths, but their gains were no greater than those earlier studies found for Khan Academy practice without an AI assistant. As the authors put it, “seeking help with one’s own confusion remained a choice, and most students declined it most of the time.”

Because these were middle-school students already behind in the subject, it would be tempting to dismiss this as a school problem with little relevance to a university lecture hall. University students are older, they chose their programme, and we expect them to manage their own learning. If any group should take up an optional AI tutor, it is them, right?

The University Paradox: Integration Without Engagement

In a randomised trial with nearly 2,400 undergraduates at the University of Maryland (Liu et al., 2026), the AI tutor sat right inside the learning management system and drew on the instructors’ own course materials. On paper, it was the seamless integration we have been asking for. Yet only 15% of students ever used it, mostly to ask for answers and explanations, and mostly around exams. Most instructors never even switched on its tutoring mode, so the study says less about Socratic design than about what happens when an institution simply makes an AI tutor available.

The outcomes were more worrying than low uptake alone. Final grades dropped, participation in the learning platform fell sharply, and first-generation students lost more than twice as much as their peers. Students also reported asking their instructors fewer questions. And the drop was not confined to the students who used the tool. The authors suspect that simply offering an AI tutor shifted the norms around AI in those courses. Offering a tool sends a message, whether we intend one or not.

It is a working paper, it relies on course grades rather than an independent test, and the mechanism remains unproven. But it is a useful warning against mistaking availability for a teaching and learning strategy.

Flipping the Script: The Power of Mandatory Structure

The positive results in higher education stem from a different setup. Kestin et al. (2025) found that an AI tutor outperformed in-class active learning in a Harvard physics course. The difference? The AI tutor was not an optional add-on at the edge of the course. It was the lesson: a carefully scaffolded sequence that students worked through. That is one course at one highly selective university, but the design point stands.

A separate, shorter experiment by Oreopoulos and colleagues, with over 6,000 middle-school students in Tennessee and a different maths platform (Oreopoulos, Liut, Sungu & Low, 2026), found the most encouraging results when AI support was built into a mastery workflow that required three correct answers in a row, although the gains a week later were modest. The authors describe this as structure “that turns mistakes into productive learning moments.”

Students seem to sense this need for structure themselves. In Hong Kong, secondary school students viewed ChatGPT as a “supportive but bounded” learning companion and preferred teacher-guided use to an outright ban (Ma et al., 2026). The study relies on self-reported data from a secondary school, so caution is warranted, but the sentiment is clear: students are not asking to be left alone with the tool. In effect, they are asking educators to show them where and how it fits into their learning.

But what is a well-designed AI tutor worth when an ordinary GenAI chatbot is just a tab away? These trials rarely ask what students do when other chatbots are within reach alongside the teacher or uni provided AI tutor interface. For our students, the alternative to a patient, Socratic AI tutor is not “no AI” but the rapidly improving Claude, ChatGPT, Gemini or Copilot in their default mode, which answer instantly. An AI tutor can be pedagogically excellent and still irrelevant if students simply open the other tab. The Maryland trial is one of the few that ran in that messier reality, and the authors are explicit about it: they measured the effect of offering their AI tutor “in a setting where other AI tools were already widely available.” Evaluating AI tutors as if students had no other options is tidy research, but it ignores the reality of access.

That other tab is also where much of our course material ends up. In my experience, few institutions offer AI access woven into the places where students and lecturers already work: the learning platform, lecture recordings, the helpdesk, the website. So people carry that context out themselves, uploading slides, transcripts and recordings into chatbots the institution has no oversight of, often against its own policy. Those who do this well tend to be the students who can afford a premium subscription and know how to feed it the right material, so the gap widens twice, in access and in ability. Integration is no cure, as the Maryland AI tutor shows. But the real choice is not between a well-designed institutional tool and nothing. It is between a tool we can shape and one we cannot.

The True Flaw: Redesigning the Course, Not the Bot

In Different Students, Same Chatbot I argued that the difference lies in students’ self-regulation. These studies add the other half. Under the pressure of grades and deadlines, even the most organised student makes a rational calculation, and a tool that withholds answers tends to lose out to one that readily provides them. A Socratic design inside an optional tool is a weak nudge competing with every other chatbot. To benefit, a student must first choose the tutor mode, and then have the patience to play along.

This is the tricky part. The direct answer wins because it offers relief, and relief arrives at once. The Socratic AI tutor offers something else: the satisfaction of working something out for yourself with a some support scaffolded to get there. That is arguably the better reward. It lasts longer, it is genuinely yours, and the ability to get there carries over to the next problem, with or without a chatbot. But it is delayed, and it comes only after a stretch of confusion that no one enjoys while it lasts. Students who have rarely reached that moment have little reason to believe it is worth waiting for.

That moment does not have to be a grand breakthrough, though. It can be small: one step solved, one connection made. A good AI tutor makes those steps small enough to succeed at, so the satisfaction comes often, rather than once at the end. It is also stronger when it builds on what students already bring. Recognising your own knowledge in a new problem, from an earlier course, a placement or simply life, feels different from receiving an answer from outside. An AI tutor that starts from what students know, rather than from what it wants to explain, lets them experience that more often. Still, an AI tutor cannot sell that pleasure by withholding answers alone. Nor can it offer what a teacher can: the embodied side of learning and the connections that come with it, caring enough about a student to show up for them, and nudging when needed. Teachers and courses can, by making sure students reach that moment often enough to know it exists.

For higher education, this points to a few things:

  • Measure effect, not just access or integration. An AI tutor that works in a trial but loses to the chatbot in the next tab has not solved much. What counts is whether students use it and learn from it, in the setting they actually work in.
  • Build the effort and the small wins into the course. Ask for a genuine attempt before AI help, break problems into steps students can succeed at, and start from what they already know.
  • Keep teachers close. In the Maryland trial, first-generation students lost the most, and students overall turned to their instructors less. An AI tutor should support human contact, not replace it.

As Plato and Xenophon describe him, Socrates did not wait in a corner for people to approach him; he went out to find them in the marketplace. Today’s AI tutors mostly wait. That passivity, not the Socratic questions they ask, is the real design flaw.

References

  • Barshay, J. (2026, September 21). Students didn’t get answers from Khanmigo. They didn’t want its questions, either. The Hechinger Report. https://hechingerreport.org/proof-points-khanmigo-math-ai-tutor/
  • Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  • Liu, J., Sweet, T., Chen, M. H., Engelberg, J., Masters, M. C., Clark, M., Persaud, A., Lancaster, A., Hollingsworth, J. K., & Rice, J. K. (2026). The effects of course-integrated AI tutoring on student performance and engagement: A randomized university trial (EdWorkingPaper 26-1598). Annenberg Institute at Brown University. https://doi.org/10.26300/y3f8-vh05
  • Ma, Q., Chen, S., Lo, W. M. S., Zou, D., & Pérez-Paredes, P. (2026). Generative AI as a learning companion for self-directed second language learning among secondary school students: A mixed-methods study. International Journal of Applied Linguistics. https://doi.org/10.1111/ijal.70375
  • Oreopoulos, P., & Low, N. (2026). One click away: AI tutoring with Khanmigo in a two-year school experiment (NBER Working Paper 35620). https://www.nber.org/papers/w35620
  • Oreopoulos, P., Liut, M., Sungu, A., & Low, N. (2026). Making AI tutoring productive: Evidence from a mastery-based math practice experiment (NBER Working Paper 35621). https://www.nber.org/papers/w35621