Ask a school district what AI it has adopted and you get a procurement list. Ask the teachers in that district what they opened this week and you get a much shorter, much more honest answer. The gap between the two is the whole story of AI in education right now.
This is an attempt at the honest version: what is genuinely in daily use, what quietly gets abandoned after a month, and what the evidence actually supports. We build a math solver, so treat the section about math tools as interested testimony and weigh it accordingly.
What teachers actually use
A general assistant, for the writing around teaching. The single most-used AI tool in schools is not an education product at all. It is ChatGPT, Claude or Gemini, used for the enormous volume of writing that surrounds instruction: worksheet variants, rubrics, exemplar answers, parent emails, IEP-adjacent phrasing, quiz questions at three difficulty levels, and the second version of a task for the student who finished early. This works because it is drafting, not deciding. The teacher already knows what good looks like and is editing rather than trusting.
Purpose-built lesson tools, mainly for the templates. MagicSchool, Diffit, Brisk and Curipod wrap the same underlying models in education-shaped forms: paste a text, get a leveled version; paste a standard, get a lesson skeleton. Their real value is not smarter AI, it is that a teacher with eleven minutes between classes does not have to write a prompt. Adoption is high where a district has bought a licence and near zero where teachers pay themselves.
Question and quiz generation. An AI question generator is genuinely good at producing twenty structurally similar practice problems from one worked example, which is a task that used to eat an evening. It is much weaker at producing good distractors for multiple choice, and it silently produces questions with no correct answer often enough that every generated set needs to be worked through before it is issued.
Grading support, cautiously. Rubric-aligned feedback drafting is in real use for essays. Automated scoring is not, in most places, and should not be. The pattern that works is AI writes the feedback paragraph, the teacher fixes the judgement and owns the grade.
What students actually use
Surveys consistently find majority use among secondary and university students, and the composition matters more than the headline number:
- Explaining, not writing. The most common student use is "explain this paragraph / this proof / this error message to me again, differently." That is a legitimate use and probably the single biggest genuine win of AI in education.
- Summarising sources. NotebookLM is the current favourite here because it is grounded in documents you upload and cites back into them, which sharply reduces invention compared with asking a chat assistant about a paper it has not read.
- Homework, photographed. Scanning a problem and getting a solution is the highest-volume behaviour in math and science, and it is where the learning benefit is most contested. See our piece on what photo solvers actually read.
- Retrieval practice. Flashcard generation from notes, mostly through Quizlet and Anki add-ons. The spaced repetition, not the AI, is what makes this work.
Where the evidence is strongest
Three uses have the best support and they are all narrow.
Immediate, specific feedback. The oldest and best-replicated finding in learning science is that fast targeted feedback beats slow generic feedback. An AI that tells you at 11pm that your integration by parts failed because you differentiated the wrong factor is doing something a graded worksheet returned on Friday cannot.
Rephrasing on demand. A textbook explains a concept once. A student who does not get that explanation is stuck. Being able to ask for the same idea four different ways, with a different worked example each time, is a real structural improvement over static material.
Volume of practice. Generating fresh problems at a fixed difficulty removes the "I have done all the odd-numbered exercises" ceiling.
Where it does not work
AI detection. Detectors are unreliable in both directions and their false positives land hardest on non-native English writers. Assessment redesign works; detection does not.
Replacing the struggle. The uncomfortable finding is that difficulty is often the mechanism, not the obstacle. A student who reads a perfect worked solution feels they understood it and frequently cannot reproduce it a day later. The instrument is not the problem — how it is used is. Attempting first, then checking, produces a different result from asking first.
Anything requiring the teacher's knowledge of the student. No model knows that this student's algebra collapses whenever a negative sign is involved, that this one has stopped attending since a family illness, or that this class needs the geometry unit slowed by a week. That judgement is why the job is not automatable.
Long unverified derivations. General assistants remain capable of a confident sign error in step six of eleven, and they will not flag it. This is the main reason math-specific tools exist: AI-Math runs the algebra through a symbolic engine so each step is checked rather than merely generated, and shows the steps so you can catch what it got wrong. That is the honest claim. It is a solver with visible work, not a tutor who knows you, and features like a conversational tutor and flashcard generation are still in development rather than available today.
A practical policy for a student
- Attempt the problem yourself first, to the point of being genuinely stuck. The stuck point is the information.
- Ask for the concept, not the answer, and ask for a different worked example of the same type.
- Check any computed result in a tool that shows steps — an equation solver or derivative calculator — rather than trusting prose arithmetic.
- Close everything and redo the problem cold. If you cannot, you have not learned it.
- Know your course policy in writing, and never submit reasoning you cannot defend out loud.
That last rule is doing most of the work. Almost every genuinely bad outcome with AI in education, for teachers and students alike, comes from submitting or issuing something nobody checked.
Related reading: