The rapid integration of generative artificial intelligence into educational settings has been heralded by many as a revolutionary shift, promising to alleviate the administrative burdens on educators and provide students with hyper-personalized learning experiences. Proponents of the technology argue that AI can serve as a powerful force multiplier, enabling teachers to generate comprehensive lesson plans, create diverse classroom materials, and deliver instantaneous feedback to students in a fraction of the time it would take a human. However, a groundbreaking randomized controlled trial—one of the first of its kind to rigorously test AI implementation in real-world classrooms—presents a far more sobering reality. The research suggests that rather than acting as a catalyst for excellence, artificial intelligence may, in some contexts, act as a "crutch" that undermines student motivation and actively erodes the quality of instruction.
The study, led by Alp Sungu, an assistant professor at the Wharton School of the University of Pennsylvania, reveals a troubling correlation between the use of AI teaching assistants and a decline in student engagement. According to the findings, students whose teachers were provided with AI-driven tools felt significantly less motivated to learn compared to those in traditional classrooms. This "motivation gap" was not a minor anomaly; it was particularly pronounced in classrooms led by instructors who were already struggling before the experiment began. For these "weaker" instructors, the reliance on AI did not merely fail to improve outcomes—it appeared to actively damage them, resulting in lower scores on standardized final exams and a measurable decrease in student confidence.
The Methodology: A Large-Scale Randomized Controlled Trial
The research, detailed in a draft titled "Generative AI Can Harm Teaching," was released online in June 2024. Conducted by Sungu in collaboration with a team of researchers from the University of Pennsylvania, including the renowned educational psychologist Angela Duckworth, the study involved a massive sample size of 193 teachers and over 2,800 middle and high school students. The experiment took place within a private school chain in Turkey, providing a controlled environment that followed the country’s national curriculum.
The timeline of the study was structured to capture the medium-term effects of AI integration. During the spring semester, teachers were randomly assigned to one of two groups. The experimental group was given access to a customized teaching assistant built on the ChatGPT framework, specifically tailored to align with Turkey’s educational standards. The control group, meanwhile, was instructed to continue their teaching practices as usual. Over a period of ten weeks, the teachers in the experimental group utilized the AI tool primarily for logistical and pedagogical preparation, such as generating lecture notes, designing homework assignments, and drafting examination questions.
To ensure the validity of the results, academic achievement was not measured by the teachers themselves, which could have introduced grading bias. Instead, researchers utilized externally administered standardized exams. This objective metric allowed the team to isolate the impact of AI on actual knowledge retention and application, independent of the teachers’ subjective assessments.
The Motivation Gap and the Loss of Personal Voice
One of the most striking findings of the study was the impact of AI on the "affective" side of learning. Students in the AI-integrated classrooms consistently rated their sessions as less enjoyable, less interesting, and less important than those in the control group. While the decline in intrinsic motivation was described as "modest" in the aggregate, it was notably higher among students whose teachers had been heavy users of AI even before the formal experiment began.
Professor Sungu suggests that this decline in interest stems from a loss of the "human element" in teaching. When an instructor relies heavily on AI to generate content, the resulting materials often lack a distinct personal voice or a unique teaching style. "When you start using AI-generated material, you’re losing your personal voice," Sungu observed. "It might be technically good enough, but it doesn’t really carry your own style. If everything is very uniform, it just becomes a bit more boring."
In pedagogical theory, the relationship between a teacher and a student is often built on the teacher’s ability to contextualize information through their own experiences and passion. When AI generates a syllabus or a lecture, it tends to produce "average" or "uniform" content that lacks the idiosyncratic nuances that keep students engaged. This uniformity can lead to a sterilized classroom environment where the material feels disconnected from the person standing at the front of the room.
The "Crutch" Effect: Why Weaker Instructors Suffer Most
Perhaps the most significant contribution of this research is the identification of how AI interacts with varying levels of teacher proficiency. The data indicated that average academic achievement across the entire sample did not see a drastic shift. However, when the researchers segmented the data by teacher performance, a clear pattern emerged: students of lower-performing teachers saw a decline in both achievement and self-confidence when AI was introduced.
The researchers hypothesize that stronger teachers use AI as a "co-pilot," treating the generated output as a rough first draft that requires rigorous editing, adaptation, and human oversight. In contrast, weaker instructors may be more inclined to use AI as an "answer machine" or a "delegation tool." By taking AI-generated assignments and lecture notes at face value and presenting them without modification, these instructors inadvertently lower the quality of their teaching.

"Teachers, just like students or coders, might be using AI as a crutch," Sungu explained. "Instead of doing the actual work, they’re using AI to delegate the task, and that lowers the quality of their teaching." This suggests a "Matthew Effect" in educational technology, where the "rich" (skilled teachers) get "richer" by using tools effectively, while the "poor" (less skilled teachers) see their performance further eroded by a reliance on automation.
Contextualizing the Findings: A Growing Body of Evidence
This study does not exist in a vacuum. It follows a previous, widely discussed 2024 study by Sungu which explored how students’ direct use of AI can harm learning. In that research, it was found that students who used ChatGPT to solve math problems performed better on practice sets but significantly worse on subsequent tests where the AI was unavailable. The conclusion was that AI served as a shortcut to the answer rather than a bridge to understanding.
The new research on teachers suggests a parallel phenomenon. Just as students use AI to bypass the cognitive struggle of learning, teachers may be using it to bypass the intellectual labor of lesson preparation. This labor—the act of thinking through how to explain a concept or how to structure a quiz—is precisely what prepares a teacher to handle the complexities of a live classroom. When that preparation is outsourced to an algorithm, the teacher enters the classroom less prepared to engage with students’ spontaneous questions or to provide deep, personalized insights.
The Myth of the Time-Saver
A common selling point for AI in education is its ability to save time. However, Sungu’s personal experience and the study’s findings challenge this notion. Sungu noted that when he uses AI to create interactive games or polls for his own university-level courses, the initial output often looks impressive but lacks the necessary accuracy or classroom-specific calibration.
"I end up spending an equal amount of time to improve the output or calibrate it to my class," Sungu stated. "It’s not a time saver."
This highlights a critical misunderstanding of how generative AI functions. Large Language Models (LLMs) operate on probabilities, not a fundamental understanding of pedagogical goals. If a teacher does not "immerse" themselves in the AI output—checking every example, verifying every number, and tailoring the tone—the resulting material can be confusing or irrelevant, ultimately requiring more human intervention to fix than if it had been created from scratch.
Broader Implications and the Path Forward
The implications of this study are profound for school districts and policymakers currently rushing to implement AI "guardrails" and platforms. The research suggests that simply providing access to the technology is not only insufficient but potentially counterproductive.
If AI is to be successfully integrated into the educational system, the focus must shift from "access" to "literacy and training." The study highlights several key areas for future development:
- Teacher Training: Educators need to be trained not just on how to prompt AI, but on how to critique and adapt its output. Training should emphasize the importance of maintaining a "personal voice" and using AI as a starting point rather than a final product.
- Interface Design: AI tools for education may need to be designed to be "less helpful" in a way that forces human engagement. Interfaces that require teachers to make choices or edit content before it can be finalized could prevent the "crutch" effect.
- Guardrails for High-Stakes Materials: Schools may need to implement policies regarding the use of AI in high-stakes areas like exam creation and core lecture notes to ensure that the quality of instruction does not dip below a certain threshold.
Conclusion: A Call for Human-Centric AI
The findings from the University of Pennsylvania team serve as a cautionary tale for the "AI-first" movement in education. While the technology holds the potential to augment human capability, its organic and unguided use currently poses a risk to the quality of teaching and the motivation of students.
As Alp Sungu concludes, the lesson is not that AI is inherently "terrible" or destined to ruin education, but rather that its benefits are not automatic. The "human in the loop" is not just a safety requirement; it is the essential ingredient that makes education meaningful. Without the creative spark and personal investment of a dedicated teacher, AI-generated education risks becoming a hollow, uniform experience that fails to inspire the next generation of learners. The challenge for the coming decade will be to ensure that AI remains a tool for empowerment rather than a substitute for the essential, difficult, and deeply human work of teaching.









Leave a Reply