A two-year school experiment testing Khan Academy's Khanmigo AI tutoring tool has wrapped up, offering one of the longest-running real-world assessments of generative AI in education. The study tracked student engagement and learning outcomes across multiple classrooms, providing early signals on how AI tutors might reshape instruction.
The Two-Year Trial
Researchers worked with schools that integrated Khanmigo into their math and science curricula. Students used the AI tutoring system during class and for homework. The study measured not just test scores but also student confidence and independence in solving problems.
Key findings from the experiment include:
Why This Matters
The implications extend beyond this single experiment. Public school districts across the United States are evaluating AI tutoring as a way to address learning gaps widened by the pandemic. If Khanmigo and similar tools can deliver consistent results at scale, they could lower the cost of personalized tutoring dramatically. However, the experiment also reveals a risk: students may become dependent on AI assistance and lose the ability to struggle productively through problems. The balance between support and independence will define whether AI tutoring becomes a classroom staple or a crutch.
Limitations and Next Steps
The experiment was limited to schools that already used Khan Academy's platform, raising questions about whether results apply to schools with less digital infrastructure. Researchers emphasize that AI tutoring should supplement rather than replace human teaching. A larger rollout across more diverse districts is planned for the coming year.
As the Year School Experiment concludes, educators are left with a clear trade-off: AI can personalize learning at scale, but only if students use it as a coach, not a shortcut.



