That’s the oral exam’s comeback story in miniature. Institutions reached for a 2,500-year-old assessment format even in lecture halls of six hundred students, because it felt like the obvious move. Obvious and sufficient turned out to be different things.

A very old idea

Oral exams go back to ancient Greece: a teacher asks, a student answers, on the spot, with nowhere to hide and no draft to revise (Brownlie, 2026). Written exams won out over the centuries for a mostly logistical reason: one proctor can watch two hundred students in a hall, while a viva needs a room, an examiner, and real time, one student at a time. Its victory had more to do with what massification could afford than with how well it revealed learning, so the oral format survived only in the corners where cost mattered less β€” languages, medicine, doctoral defences.

Generative AI just made writing cheap enough that the written exam’s old advantage, scale, stopped being an advantage at all. Institutions reached back past a century of written assessment to the format from before it, and for a moment, that felt like a solution.

Why it’s such an easy reach

A few things make the oral exam the reflexive answer, even in large classrooms, even for academics who know better.

It feels unfakeable because it draws on the oldest model we have of proving knowledge: a live person answering a live question. No new software, no new training, no curriculum redesign β€” five minutes of talking gets bolted onto whatever’s already there, and the assignment suddenly has teeth again.

Oral exams weren’t the only fix on the table. Universities also leaned harder on the room itself: Stanford wrapped up a three-year pilot by opening proctoring to every in-person test this September, and Princeton’s faculty approved a similar shift a month later, both citing AI as the reason (Whitford, 2026). That route asks for booked halls, trained invigilators, and a whole infrastructure of surveillance, in service of a written exam whose validity problems (Dawson et al., 2024) don’t go away just because someone is watching. Pulling assessment the other way, back into take-home essays and assignments, runs into a different problem: there is no reliable way to tell, after the fact, which parts of an unsupervised piece of work came from a student and which from a chatbot, whatever the assignment brief says (Liu & Lodge, 2026). The oral exam sat between those two costs: no new buildings, no need to trust what happened alone in a dorm room β€” just a room, a student, and a conversation.

It also relocates the anxiety somewhere that feels more human. Detection software accuses students algorithmically, at scale, with error rates universities are now walking back (Perkins et al., 2024). A conversation, by contrast, puts a person in the room making a judgement about another person. That reads as fairer, and often is.

There’s a subtler pull too: a fluent answer feels like knowing. That feeling is exactly what makes the format persuasive, and exactly what makes it fragile as evidence, because feeling convinced and being right rest on different kinds of proof (Dawson et al., 2024).

The problems it imports

Phillip Dawson and colleagues have spent several years arguing that higher education spends its cheating-prevention energy on the wrong target. Their case: what matters is whether the assessment evidence lets you validly infer a student has learned what the degree says they’ve learned, and an oral exam doesn’t automatically deliver that any more than a written one does β€” it just changes what gets hidden and what gets seen (Dawson et al., 2024).

Fawns, Boud, and Dawson’s newer framework sorts assessment evidence into four kinds of proxy β€” product, process, performance, and practice β€” and the oral exam sits squarely in the performance category (Fawns et al., 2026). Performance evidence is volatile by nature: it needs sampling across multiple moments to mean much, since a single five-to-twenty-minute conversation captures a snapshot rather than a pattern. It also has a habit of collapsing back into product whenever what’s really on display is a pre-rehearsed script delivered live, which is precisely what a student who used AI to draft and memorise convincing answers is doing. The format built to catch AI-assisted work can be beaten by the most boring possible move: rehearsing the AI’s answer.

The numbers make this almost comic. A Norwegian study of oral exams in secondary schools found student presentations ranging from five to sixteen minutes, and follow-up discussions from seven to twenty-three; with the same examiners, some students fielded fewer than ten questions while others fielded nearly fifty (Brownlie, 2026). Two students, same subject, same exam on paper, remarkably different tests in the room.

Whose rug, exactly

This format doesn’t put the same weight on everyone. The assessment literature tends to mention that in a subordinate clause, when it deserves the main sentence.

For a lot of teaching staff, this is not the first time the ground has moved. First the tools meant to verify learning stopped being trustworthy β€” detectors with real error rates, wrongful accusations, a whole industry walking back its own confidence. Now the fix for that turns out to need constant, personal, real-time judgement, exam after exam, student after student, with no shared training in how to ask a fair follow-up question or how much silence to allow before it counts against someone. The Norwegian numbers above β€” some students getting ten questions, others fifty, from the same examiners β€” aren’t evidence of bad teachers. They’re evidence of a format handed to people without the calibration it needs, on top of everything else already on their plate this year.

For students, the ground moves differently, and not evenly. Brownlie (2026) cites research showing university students report more anxiety in oral exams than written ones, and that students answering in a second language can know the material thoroughly and still need more time than the room allows to say so β€” their hesitation read as not-knowing when it’s actually translation. Fawns and colleagues push the point further: authentic, in-person, high-pressure formats can be genuinely harder to access for students juggling paid work or caring responsibilities, for disabled and neurodivergent students navigating conditions “authentic” settings weren’t built around, and for anyone whose confidence speaking to power was never nurtured in the first place β€” a very different inheritance than a classmate who’s been performing fluency in seminar rooms since childhood (Fawns et al., 2024). None of that shows up in a transcript. It shows up as a slightly longer pause, read by an unprepared examiner as evasiveness rather than as the entirely ordinary cost of thinking in a second register, under a clock, in front of someone with a grade in their hand.

Put together: the students with the least practiced confidence and the least institutional slack are asked to prove their integrity live, on the spot, at exactly the moment their instructors are least equipped to judge that performance consistently. That’s the rug pull.

What actual aid looks like

This argues for treating the oral exam with the same care any high-stakes format deserves, and Brownlie’s own piece already sketches most of it.

Tell students the purpose, the marking criteria, and the kinds of questions they’ll face before they walk in β€” an oral exam should not also be a test of guessing the rules (Brownlie, 2026). Give them somewhere to practise speaking under mild pressure well before the exam that counts, so the first time anyone speaks aloud about their work isn’t the day it’s graded. Build in real thinking time, and offer adjustments that cost almost nothing to give: choice of time slot, sitting or standing, a moment to gather thoughts before answering. Use the same, agreed follow-up prompts for every student instead of leaving each examiner to improvise in the moment, and have colleagues moderate a sample of the recordings together, so the burden of judgement is shared rather than carried alone behind a closed door.

The five minutes work best as a conversation about the work. Hattie’s synthesis of what actually moves student achievement, across thousands of studies, keeps landing on the same underlying condition: feedback that both teacher and student can see, close in time to the work (Hattie & Timperley, 2007; Hattie, 2012). Brownlie’s own recommendation does exactly that β€” students submit an annotated draft, then talk for five minutes about one revision they made, one source they chose, and how AI shaped their thinking along the way (Brownlie, 2026). That functions as a feedback loop wearing an exam’s clothes, and it’s the part of this whole comeback worth keeping, for the people on both sides of the table.

Sampled across multiple moments instead of staged once, paired with other kinds of evidence instead of asked to carry the whole validity argument alone, run with enough shared calibration that no single overstretched examiner is inventing the rules mid-conversation: that’s where the oral exam earns its return. Deployed as the newest security theatre, unsupported and alone, it costs its heaviest toll on exactly the people least able to pay it, and it still lasts only as long as it takes someone to build the glasses for it. This year, that was about twelve months.


References & Further Reading

Brownlie, N. (2026, August 20). Oral exams are making a comeback to stop AI cheating. But they have their own problems. The Conversation. https://theconversation.com/oral-exams-are-making-a-comeback-to-stop-ai-cheating-but-they-have-their-own-problems-289707

Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. Assessment & Evaluation in Higher Education. Advance online publication. https://doi.org/10.1080/02602938.2024.2386662

Fawns, T., Bearman, M., Dawson, P., Nieminen, J. H., Ashford-Rowe, K., Willey, K., Jensen, L. X., Damşa, C., & Press, N. (2024). Authentic assessment: From panacea to criticality. Assessment & Evaluation in Higher Education. Advance online publication. https://doi.org/10.1080/02602938.2024.2404634

Fawns, T., Boud, D., & Dawson, P. (2026). Identifying what our students have learned: A framework for practical assessment validation. Assessment & Evaluation in Higher Education. Advance online publication. https://doi.org/10.1080/02602938.2026.2620053

Hattie, J. (2012). Visible learning for teachers: Maximizing impact on learning. Routledge.

Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487

Liu, D., & Lodge, J. M. (2026, August 14). A wholesale ban of take-home assessments is not the answer [LinkedIn post]. LinkedIn. https://www.linkedin.com/pulse/wholesale-ban-take-home-assessments-answer-danny-liu-s7hie/

McIntosh, N. (2026, July 1). The exam room is no longer secure [LinkedIn post]. LinkedIn. https://www.linkedin.com/feed/update/urn:li:activity:7478021657945849858/

Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). GenAI detection tools, adversarial techniques and implications for inclusivity in higher education (arXiv:2403.19148). arXiv. https://doi.org/10.48550/arXiv.2403.19148

Whitford, E. (2026, June 22). Canβ€”and shouldβ€”honor codes survive in the AI age? Inside Higher Ed. https://www.insidehighered.com/news/faculty/learning-assessment/2026/06/22/can-and-should-honor-codes-survive-ai-age