Why an AI tutor needs more than a conversation
A clear explanation is not a diagnosis. What a tutoring product needs besides good answers, and what the recent trials actually showed.
By Lucas Hsu · · 4 min read
A student asks for an explanation of eigenvectors. The tutor gives a clear answer, checks a small example, and responds patiently to follow-up questions. This can be a valuable interaction.
Tomorrow, the student opens the app again. What should happen?
The answer depends on what survived the conversation, what the student can now do independently, and what the course expects next. Conversational skill helps with teaching. A complete learning product also needs ways to diagnose, choose activities, and evaluate the consequences.
The learner cannot supply every diagnosis
Asking students what they need preserves agency and reveals context. It also assumes they can accurately judge their knowledge. Sometimes they can. Sometimes the information available during study makes that judgment misleading.
Koriat and Bjork (2005) showed how judgments of learning can be inflated when information is visible during study but unavailable at the eventual test. A solution can feel understandable while its essential steps remain difficult to generate independently. Their experiments concerned memory judgments; applying the result to a particular tutoring interface is an inference that needs testing.
For a tutor, the practical question is what to do after an explanation. One option is to ask whether it made sense. Another is to ask the learner to solve a nearby problem and explain the decisive step. These responses provide different evidence.
Recent trials make design hard to ignore
The empirical picture supports careful distinctions between kinds of AI assistance. In a randomised study of undergraduate physics learning, Kestin and colleagues (2025) found that a custom AI tutor produced stronger immediate learning outcomes than the in-class active-learning comparison. The tutor incorporated pedagogical design, and the study examined particular lessons in a particular course. It does not establish that any chatbot will outperform any teacher.
A field experiment by Bastani and colleagues (2025) involved nearly a thousand high-school mathematics students. Access to AI improved performance during assisted practice. However, students using the less constrained interface subsequently performed worse when the assistance was removed; the tutor designed with safeguards largely mitigated that harm.
The crucial outcome was what students could do afterwards on their own.
These studies used different populations, tasks, and comparisons. They should not be combined into a single verdict on AI tutoring. Together, they motivate a more specific question: which interaction designs produce durable, independent capability?
A learning history should change the interaction
Consider a learner who repeatedly confuses linear independence with orthogonality. A conversational response can explain the difference. A system with relevant history could also schedule a later comparison, remove the familiar wording, and check whether the distinction holds in a new example.
This ambition has predecessors. Corbett and Anderson's (1994) knowledge-tracing work modelled students' changing knowledge of programming rules and used those estimates to individualise exercises. It demonstrates that persistent learner modelling predates language-model interfaces by decades. It also reminds us that useful state has to be tied to specific skills and observable responses.
A transcript is useful evidence, but it does not automatically provide such a model. “I understand” might indicate comprehension, politeness, or a wish to move on. A correct answer might be independent, copied, or reconstructed after a hint. A sound interpretation depends on how the response was produced.
What we want the architecture to do
For Memoza, the design goal is a loop that can select an appropriate activity, support the learner when needed, assess the response, and update future recommendations. Conversation belongs inside that loop. So do course information, grading, and delayed checks.
This is an architectural proposal, not a claim that every diagnostic capability described here is already implemented or experimentally validated. Each part can fail. A mistaken grade can produce a mistaken knowledge estimate; a mistaken estimate can redirect study away from what matters.
The learner should retain control. Recommendations should be explainable, uncertainty should be visible when relevant, and an override should be possible when the student knows something the system does not. A seminar tomorrow or a newly announced syllabus change can legitimately alter the plan.
The quality of an AI tutor ultimately extends beyond its most impressive answer. We should also ask what the learner attempts next, how much help that attempt requires, and whether the capability remains when the conversation is over.
References
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122 Field experiment.
Corbett, A. T., & Anderson, J. R. (1994). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253–278. https://doi.org/10.1007/BF01099821 Model and empirical research.
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6 Randomized controlled trial.
Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187–194. https://doi.org/10.1037/0278-7393.31.2.187 Experiments.
Put it into practice. Memoza marks your answers against a pre-validated solution and shows where the marks went.
Answer three questions, not a chat