The promise of artificial intelligence in the classroom was met with a mixture of utopian optimism and skepticism when generative AI entered the mainstream in late 2022. As students quickly mastered the art of using tools like ChatGPT to bypass critical thinking and generate instant answers for assignments, educators raised alarms about the potential erosion of foundational academic skills. In response, technology developers pivoted toward creating specialized educational AI—tools designed not to do the work for the student, but to coach them through it. Khan Academy, a leader in digital learning, introduced "Khanmigo," an AI-powered tutor built on the Socratic method, intended to guide students through complex problems without providing the answers outright. However, a significant new study conducted over two years in Tennessee suggests that the mere existence of sophisticated AI technology is insufficient to improve student learning outcomes if engagement remains low.
A Chronology of the AI Classroom Integration
The trajectory of AI in education has moved at a breakneck pace since the launch of OpenAI’s ChatGPT in November 2022. By early 2023, the educational sector was already witnessing the fallout of unchecked generative AI use, including plummeting test scores and a decline in homework integrity. Khan Academy, under the leadership of Sal Khan, sought to position its AI tool as a constructive alternative to the "answer-machine" models that were disrupting traditional classrooms.
In 2023, the organization released Khanmigo, marketed as a secure, pedagogical-focused chatbot that would mimic a human tutor’s approach by asking guiding questions rather than providing solutions. Researchers from the University of Toronto, led by Philip Oreopoulos and Nina Low, launched a randomized controlled trial (RCT) spanning the 2024 to 2026 academic years. Their objective was to measure the impact of this tool on middle school students who were performing at least one grade level below their peers. By tracking 18 schools in Tennessee, the study provided a rare, long-term look at how AI actually functions when removed from the controlled environment of a tech demonstration and placed into the hands of real students in a remedial setting.
The Findings: Engagement as the Primary Hurdle
The data emerging from the Tennessee trial presents a sobering reality for ed-tech developers. While nearly every student assigned to the Khanmigo intervention attempted to use the tool at least once, the novelty quickly wore off. When students realized that Khanmigo was programmed to withhold the answers they were seeking, their interest in the tool plummeted.
The researchers noted that for many students, the desire to find the path of least resistance superseded the desire for learning. "Seeking help with one’s own confusion remained a choice, and most students declined it most of the time," noted Oreopoulos and Low in their NBER paper. When faced with a choice between the frustrating process of working through a math problem with a Socratic AI and simply skipping the task or guessing, the majority of students opted to disengage from the AI tool entirely.
Crucially, the study compared the outcomes of students using the Khan Academy platform with the AI assistant against those using traditional digital remedial tools like IXL, Zearn, and DeltaMath. The students utilizing Khan Academy did show some improvement in math performance compared to the control group, but the researchers found no statistically significant difference between those who had access to the AI tutor and those who used standard Khan Academy practice sessions. In short, the AI tutor added no measurable value over existing, non-AI-driven digital learning platforms.

Official Responses and the Pivot
Sal Khan, the CEO of Khan Academy, has been transparent about the study’s results, choosing to embrace the data as a critical feedback loop for product development. In a public-facing blog post, Khan acknowledged that the findings reflected the internal usage metrics the company had observed throughout the rollout. He emphasized that the primary goal of the initial launch was safety and "doing no harm"—prioritizing student data privacy and preventing the "hallucinations" common in early large language models over immediate performance gains.
Khan maintains that the early, perhaps underwhelming, results are part of the iterative process of software engineering. "It’s allowed us to learn and hopefully make the new version even more helpful," Khan noted in an interview. The organization is now moving away from the "opt-in" model of AI assistance. In newer iterations of the platform, the AI is integrated directly into the workflow. Rather than a student having to navigate to a separate tab to seek help, the AI now functions as a reactive agent that appears automatically after a student incorrectly answers a problem.
Furthermore, Khan Academy is exploring gamification and incentive-based structures, such as offering students credit for using the AI tutor to correct their errors. By allowing a "bot-assisted redo" to count toward the mastery requirements—usually a streak of five correct answers—the organization hopes to align student incentives with the use of the tutor.
Broader Implications for Education Policy
The failure of the initial Khanmigo rollout to outperform traditional digital tools highlights a significant disconnect between the technological potential of AI and the psychological reality of student learning. Education policy experts point to three primary implications for school districts currently rushing to adopt AI tools:
- The "Active Learning" Requirement: Technology that functions as a passive repository of information is rarely effective for remediation. AI tutors must be designed to force active participation. If the tool allows a student to remain passive, it will inevitably be bypassed or ignored by students looking to minimize their workload.
- The Integration Problem: For an AI tool to be effective, it must be frictionless. As the Tennessee study demonstrated, requiring a student to move between different interfaces or make an active choice to seek help is a barrier that most struggling students will not overcome. Future iterations must ensure that AI support is context-aware and embedded directly into the primary learning task.
- Instructional Design vs. Tech Power: The study reinforces the long-standing belief that the efficacy of an educational tool is determined more by how it is integrated into the classroom pedagogy than by the sophistication of the underlying algorithm. Teachers, not just technology, remain the central drivers of student achievement. Without a teacher to frame the AI as a valuable, mandatory, or high-value resource, students are unlikely to adopt it as a legitimate study aid.
Looking Toward the Future
As the 2026-2027 school year progresses, the focus will shift to whether these updated, more intrusive, and incentivized versions of AI tutors can move the needle on standardized test scores. The fundamental challenge remains: how to balance the "human element" of teaching with the automated precision of AI.
The findings from the Tennessee trial serve as a cautionary tale for the education technology sector. The "AI revolution" in schools will not be won simply by providing students with more powerful tools; it will be won by designing systems that fundamentally change student behavior in ways that correlate with long-term cognitive development. The debate is no longer about whether AI can be a tutor—it is about whether AI can successfully navigate the human resistance to the difficult work of learning. As schools continue to integrate these systems, researchers, educators, and policymakers will be closely watching to see if the next generation of AI tools can finally turn engagement into mastery.









Leave a Reply