Part 6 of 7
×The Tool Was Never the Hard Part
What does a tool need before it can help a teacher?
Key Takeaways
A new classroom tool only helps when someone decides what it is for and who it will help.
Time saved by AI gets filled with other work unless the school removes another task.
AI tools work best when built on a list of common student mistakes, which exists for maths but not yet for writing.
Every few years a school is handed a new tool and told it will lighten the load or lift results. In the 2000s it was extra adults in the classroom, then computers, then something called the Scandinavian method, and now it is AI.
There are so many different methods and tools. They are often adopted because it sounds good or based on misconceptions, or just because you’re supposed to. But in reality, a tool only helps in education when someone decides three things before it arrives: what it is for, which students it will reach, and what the teacher will stop doing to make room for it.
Where nobody decides, the tool will negatively impact learning or get dusty in the corner. Where it is used to stand in for the teacher with the weakest students, those students do worse. AI is following the same pattern. The time it saves is kept only if something comes off the timetable, and its best use so far keeps the teacher deciding, gives the typing to the machine, and has only been built where someone first wrote down how students go wrong in the subject.
The clearest warning comes from the teaching assistants in English schools. Between 2003 and 2008 a team at the Institute of Education in London followed what happened as those schools took on many more of them.
They tracked 8,200 pupils in 153 schools, in two cohorts across seven year groups, and compared the progress of pupils who received a lot of support from an assistant with that of similar pupils who received little.
The pupils with the most support made less progress in English and maths. The pupils with the highest level of need did worst.
The researchers were clear that this was not the assistants’ fault. The cause was how they were used. Schools had attached assistants, more or less permanently, to the lowest-attaining pupils in each class, and over time those assistants had become the main educators of the pupils in most need, while the qualified teacher taught everyone else.
The guidance that came out of the study is short. Do not use assistants as an informal teaching resource for low-attaining pupils and give them structured programmes with evidence behind them. In short, they should be used to add to what the teacher does, not to replace it. And they should also connect what happens in the programme with what happens in the lesson.
None of that is complicated. It took a study of 8,200 children to find that adding adults to a classroom, without deciding what they would do, produced worse results than adding nobody.
And the same thing happens to time.
Teachers are overworked and underpaid, a view you’d be hard pressed to find a teacher to dispute, so we assume that new tools with modern technology will bring their workload down.
But that assumes that tools are designed properly, with the correct intentions.
In the summer term of 2024, the Education Endowment Foundation ran a randomised trial with 259 science teachers in 68 English secondary schools. Half were given access to ChatGPT and a guide to using it for lesson preparation. The other half were asked not to use generative AI.
The teachers using ChatGPT spent about 56 minutes a week preparing lessons. The teachers without it spent 81. That is a saving of nearly a third, and the lessons were no worse. A panel judged the resources without knowing which had been made with the tool and found no difference in quality. The question is what happened to the 25 minutes. Most teachers said they spent them on other planning and teaching tasks. Some said they worked less overall.
A survey two years later shows which of those two answers won. In 2026 YouGov asked 1,033 teachers across the UK about AI for the Bett education show. Four in five were using it. Of those, just over half said it had eased their workload, and 35% said they were working fewer hours. The rest worked the same hours or more. Bett’s own summary was that teachers say the tool saves them time on particular jobs, but “the time comes back, and something else fills it”.
The same observation was made about the civil service in 1955. C. Northcote Parkinson’s essay in The Economist said “Work expands so as to fill the time available for its completion”, and nothing about a school exempts it. A teacher who finishes planning early has marking waiting, and a school that notices its teachers finishing early has a new initiative waiting.
Unless someone decides what will stop, nothing does.
England has one period on record when something did stop. The Department for Education’s workload survey of 2016 found teachers and middle leaders working an average of 54.4 hours a week in term time. That spring three review groups reported on marking, planning and data management.
The marking group said that written comments on every piece of work, and policies that made teachers respond to the student’s response to their comments, cost hours and did little for learning. The target was the volume of written comments, not feedback itself, which has also been covered in this series.
Ofsted said it expected no particular type or amount of marking. Schools were free to drop it, and many did. The 2019 survey found 49.5 hours, nearly five fewer, and traced most of the fall to less time on marking, planning and supervision. Nobody had been given a faster tool for marking. The job had been made smaller.
Those hours sit on a long flat line. A 2021 analysis pooled four datasets covering more than 40,000 teachers in England between 1992 and 2017. Across those 25 years, teachers worked about 47 hours a week in term time, eight more than the average across comparable OECD countries, and no tool that arrived in those years moved the figure. The 2019 fall brought teachers back to that line, not below it.
The problem also occurred with mass introduction of computers into schools. Delivered without anyone deciding what they were for, they did nothing. In 2000 the government delivered four computers to each of a hundred municipal primary schools in the Indian city of Vadodara.
Two years later a survey found that very few were being used by children. When the charity Pratham then ran a programme in those schools, with maths games tied to the curriculum for two hours a week, maths scores rose by 0.35 standard deviations in the first year and 0.47 in the second. The computers had been there the whole time.
Once the session around the software is counted, the software itself adds little. A trial in rural China separated the two. Across 130 boarding schools and just over 4,000 pupils in grades four to six, one group got extra maths sessions on computer software, a second got identical sessions working through the same material in paper workbooks, and a third got nothing. The sessions raised maths grades a little. The workbook sessions did almost all of it, and the part of the gain that could be put down to the technology itself was small and statistically indistinguishable from zero. The authors’ phrase was that the tech in EdTech may have relatively small effects.
So the question for AI is the same one. Which design puts the tool inside work that someone has thought about, and which adds it to a class and hopes.
The best answer so far comes from a trial in five English secondary schools in 2025, run by the maths platform Eedi with Google DeepMind. Eedi’s questions are written by former maths teachers so that every wrong answer points to a named misconception, and the platform now holds more than 8,000 of them. The tutoring sat on top of that map. 165 students aged 13 to 15 were tutored online over seven weeks. For half of them, the messages were written by one of 17 human tutors. For the other half, an AI model drafted each message, and the same tutors read every draft before it was sent, approving it, editing it or writing something else.
In practice the tutors sent 76% of the drafts with no change or a change of a character or two. The students taught this way did at least as well as those taught by the tutors alone on every measure, and on the one that matters most, solving new problems on a later topic, they did a little better, 66% against 61%.
Of 3,617 messages, none was harmful and five contained a factual error. The tutors said the work felt more fluid and that they could support several students at once. The authors call the trial exploratory, it covers one subject over one term, and a larger trial with about 1,500 students in ten schools is running in 2026.
What the design did was keep the human decision in the loop, at least, change the scope of the decision. Writing the message takes much more time than deciding if it’s correct.
Tutoring’s reputation is one of the big misconceptions in education. It goes back to 1984, when Benjamin Bloom reported that students tutored one to one outperformed 98% of a class taught conventionally. A 2024 review of 265 randomised trials shows how little of that survives growth. Programmes with fewer than a hundred students averaged 0.55 standard deviations, and programmes with more than a thousand averaged 0.14.
What tutoring offers is teaching pitched at what the student knows and a correction the moment they go wrong. Both depend on someone who knows where the student is likely to go wrong. The Eedi trial suggests that the someone can be a person who checks rather than a person who types. The map of how students go wrong in a subject still comes from people who have taught it.
The limit shows when the teacher sets the tool up well and then leaves the student alone with it. A University of Washington study published in 2026 followed 16 American teachers who set up an AI assistant for their own classes, writing the prompts that told it what to do and how hard to push, and analysed 1,479 student conversations that followed.
Seven in ten stayed fully on the track the teacher set, and fewer than one in a hundred went badly off it. Even so, 38% fell short of the depth of thinking the teacher had asked for, and when the teacher aimed for the hardest level, strategic reasoning rather than recall, the share falling short approached half. Setting the tool up correctly and getting hard thinking out of it turned out to be two different problems.
Put that next to the teaching assistants and the risk needs no new name. A struggling student who is told to work with the AI while the teacher’s attention goes to students who need it less is putting the student with a teaching assistant, just with different hardware.
And how schools keep records is another version of this failure.
Every learning platform logs when a task was opened and when it was handed in, how many drafts a piece of writing went through, how often a student asked for a hint and when they gave up.
Leveraged correctly, it could be used to identify their strengths and weaknesses. But almost none of it reaches a teacher. Twelve minutes on question four, three attempts and then nothing is not yet a claim about whether the student understood the idea and froze or never understood it, and turning the first into the second is expensive.
Well-funded attempts have joined up the data and claimed to have done the rest. Purdue University’s Course Signals system flagged students at risk from their activity logs, and in 2012 the university reported that students who had taken a course using it stayed on at a rate nearly twenty points higher than those who had not.
A year later a critic pointed out that the figure ran backwards. Students who stay at university longer take more courses, and so take more courses with Signals in them. In a simulation, a colleague replaced Signals with chocolates handed out at random in some courses, and the students who stayed longest still came out with the most chocolates. Purdue conceded that more research was needed. The retention claim collapsed. The flags themselves were never shown to be wrong, only never shown to help, and nobody had turned them into something a teacher could act on.
Turning the data into a claim about what a student understands is the hard part, because it has to be built for each subject. A teacher who looks at a wrong answer in maths and a teacher who looks at a weak paragraph are using different knowledge to decide what went wrong, so a single system that promises to read every subject reads none of them. A login counter scales to every classroom in the world for nothing. A tool that can tell why a particular kind of essay error keeps recurring does not, which is why the login counter is what gets built.
Where that map has been written, the tools help. In maths, a free American homework tool called ASSISTments gives students feedback as they work and gives teachers a report of the common wrong answers in each problem set. In a randomised trial across 43 schools in Maine, published in 2016, seventh-graders whose teachers used it scored 0.18 standard deviations higher at the end of the year, and the students who started behind gained most. Eedi’s tutoring trial sat on a map of more than 8,000 misconceptions that former maths teachers had spent a decade writing. Nothing like it is open for writing. Cambridge has coded the errors in millions of words of learners’ exam writing, but that record is private, and it names the error, not the reason behind it. The best-known tool for marking English writing, Cambridge’s Write & Improve, gives a level on the European scale and feedback on spelling, grammar and vocabulary, and describes no map of how students go wrong and no trial. The knowledge of where a piece of writing goes wrong, and why that student keeps making that error, is in English teachers’ heads, and nobody has written it down where other teachers can use it.
One boundary has not moved and nothing in this evidence suggests it will. In 2000 Richard Ryan and Edward Deci set out the three things that sustain a person’s motivation, a sense of competence, a sense of choice and a sense of connection to the people around them, and nothing in the trials above shows a tool supplying any of the three.
England’s Department for Education, in its 2025 guidance on generative AI, put it as plainly as a government does. Technology should not replace the relationship between teachers and pupils. Every use of AI that has earned a place in this essay sits on the qualification side of what school is for. None of it touches the part where a child becomes a person among other people.
Sources
Teaching Assistants
- Blatchford, P., Bassett, P., Brown, P., Koutsoubou, M., Martin, C., Russell, A., Webster, R. with Rubie-Davies, C. (2009). Deployment and Impact of Support Staff in Schools: The Impact of Support Staff in Schools (Results from Strand 2, Wave 2). DCSF Research Report RR148, with research brief RB148. England; 8,200 pupils in two cohorts in seven year groups in 153 schools; "a consistent negative relationship between the amount of support a pupil received and the progress they made in English and mathematics" (RB148, p. 10); decisions about deployment "largely outside their control" (p. 12). The "16 of the 21 results were in a negative direction and there were no positive effects" line is in Webster, R. and Blatchford, P., "Teaching assistants", chapter for the International Guide to Student Achievement (UCL Discovery). See also Blatchford, P. and colleagues (2011), British Educational Research Journal, 37(3), 443–464, and Webster, R., Blatchford, P. and Russell, A. (2013), School Leadership & Management, 33(1), 78–96. The guidance: Sharples, J., Webster, R. and Blatchford, P. (2015, updated 2018 and 2021). Making Best Use of Teaching Assistants. Education Endowment Foundation. Seven recommendations; the four in the essay are recommendations 1, 2, 5 and 7, and "primary educator for pupils in most need" is the guidance's own phrase (p. 8).
Saved Planning Time
- Roy, P., Poet, H., Staunton, R., Aston, K. and Thomas, D. (2024). ChatGPT in Lesson Preparation: A Teacher Choices Trial. Evaluation Report. Education Endowment Foundation, evaluated by NFER, published 12 December 2024. 34 schools and 129 teachers in the ChatGPT group, 34 schools and 130 teachers in the comparison group; Key Stage 3 science; ten weeks in the summer term of 2024. In weeks six to ten the ChatGPT group spent about 56.2 minutes a week on lesson and resource preparation against 81.5 minutes, 69% of the comparison group's time, with a high security rating. An expert panel blind to how resources were made found no evidence of a difference in quality, from a limited sample of resources. "Teachers reported typically using the time saved to complete other lesson and resource planning or teaching tasks, or to reduce overall workload" (p. 4). The report does not measure total working hours.
The Survey
- Bett with Lenovo (2026). The State of AI in Education 2026: Working It Out. YouGov survey of 1,033 UK teachers, fieldwork June 2026, published August 2026. 80% use AI; of users, 51% say it has eased their workload, 35% say they work fewer hours, 55% the same hours and 4% more. The closing wording is from Duncan Verry of Bett, quoted in Tes, 27 August 2026: "They tell us it saves them time on particular jobs, but most say they are not working fewer hours. The time comes back, and something else fills it." A trade-show poll, not a government survey.
Parkinson
- Parkinson, C. N. (1955). "Parkinson's Law." The Economist, 19 November 1955. Source of the sentence "Work expands so as to fill the time available for its completion." Date and wording checked against secondary sources.
Workload in England
- Higton, J., Leonardi, S., Richards, N., Choudhoury, A., Sofroniou, N. and Owen, D. (2017). Teacher Workload Survey 2016. Department for Education, p. 6: 54.4 hours for classroom teachers and middle leaders in the reference week. Walker, M., Worth, J. and Van den Brande, J. (2019). Teacher Workload Survey 2019. Department for Education and NFER, p. 10: 49.5 hours, "down 4.9 hours from the 54.4 hours reported in 2016"; p. 11: most of the reduction attributable to less time on non-teaching activities, with planning, marking and pupil supervision each down by one to two hours; pp. 11–12: "quite possible" that the three review groups contributed. The three Independent Teacher Workload Review Group reports on marking, planning and resources, and data management were published 26 March 2016; the marking report (p. 5) warns against "extensive written comments on every piece of work when there is very little evidence that this improves pupil outcomes" and describes "dialogic marking, triple marking and quality marking" (p. 6). Ofsted, "Ofsted inspections: myths" (2016) and its 2018 clarification for schools: "Ofsted does not expect to see any specific frequency, type or volume of marking and feedback." The 2019 survey (p. 13) reports that most respondents said their school's marking and feedback approaches had changed to reduce workload. Allen, R., Benhenda, A., Jerrim, J. and Sims, S. (2021). "New evidence on teachers' working hours in England. An empirical analysis of four datasets." Research Papers in Education, 36(6), 657–681. More than 40,000 teachers, 1992 to 2017; about 47 hours a week in term time, "8 hours per week longer than the OECD average".
Computers in India and China
- Banerjee, A. V., Cole, S., Duflo, E. and Linden, L. (2007). "Remedying education: evidence from two randomized experiments in India." Quarterly Journal of Economics, 122(3), 1235–1264. Section II.B: in 2000 the government delivered four computers to each of the 100 municipal primary schools in Vadodara, and a 2002 Pratham survey "suggested that very few of these computers were actually used by children"; the CAL programme raised maths scores by 0.35 standard deviations in year one and 0.47 in year two. Ma, Y., Fairlie, R., Loyalka, P. and Rozelle, S. (2024). "Isolating the 'tech' from EdTech: experimental evidence on computer-assisted learning in China." Economic Development and Cultural Change, 72(4), 1923–1962 (NBER Working Paper 26953, 2020). 352 classes in 130 rural boarding schools in Shaanxi, 4,024 pupils in grades 4 to 6. The full programme raised maths grades by 1.74 percentile points (significant at the 10% level) and had no significant effect on test scores; the software arm against the workbook arm was 0.059 standard deviations on test scores and 0.21 points on grades, neither significant. "When we isolate the technology-based effect of CAL we find point estimates that are small and statistically indistinguishable from zero"; "the 'Tech' in EdTech may have relatively small effects on academic outcomes."
The Human-in-the-Loop Tutor
- LearnLM Team, Google DeepMind, and Eedi; Brazão, V., McKee, K. R. and colleagues (2025). "AI tutoring can safely and effectively support students: an exploratory RCT in UK classrooms." Report of 11 November 2025 and arXiv:2512.23633. 165 students in Years 9 and 10 in five UK secondary schools, 17 tutors, mathematics, seven weeks in May and June 2025. Tutors approved 76.4% of drafted messages with zero or minimal edits (one or two characters). Transfer to new problems on later topics 66.2% against 60.7%; immediate correction 93.0% against 91.2%; 3,617 messages, none harmful, five factual errors. Tutors' reports that the work felt "more fluid and efficient" and that they could support several students at once are self-report, and the capacity gain is modelled in an appendix rather than measured. Funded through Eedi by the Learning Engineering Virtual Institute and others, with the model and research team from Google DeepMind. The scaled trial is registered as AEARCTR-0018079, ten UK schools, 1,200 students planned and about 1,500 enrolled, March to August 2026.
The Subject Map
- Eedi (2026), Diagnostic Engine and Learning Design pages at eedi.com, and the news post "Eedi's misconception graph is now a public good": "every wrong answer is engineered by our Learning Design team to reveal a specific misconception"; "Every distractor in Eedi's diagnostic questions maps to a specific, documented misconception"; "over 8,000 misconceptions", written by "former classroom math teachers" over "eleven years" and released under CC BY 4.0. The arXiv paper itself calls it only "the Eedi mathematics platform"; the misconception design is from Eedi's own pages, checked 6 October 2026. Roschelle, J., Feng, M., Murphy, R. F. and Mason, C. A. (2016). "Online mathematics homework increases student achievement." AERA Open, 2(4). Randomised trial in Maine, 46 schools recruited and 43 analysed, 2,850 seventh-grade students, TerraNova outcome after a warm-up year; effect 0.18 standard deviations (Hedges's g), 0.29 for students at or below the median prior achievement and 0.12 above it; "Teachers receive reports on how students perform on the assigned problem sets, including information about common wrong answers." Write & Improve, Cambridge University Press & Assessment, writeandimprove.com (checked 6 October 2026): a CEFR level and "feedback in seconds, covering spelling, vocabulary, grammar and general style"; no error taxonomy and no trial evidence on the site.
Tutoring
- Bloom, B. S. (1984). "The 2 sigma problem: the search for methods of group instruction as effective as one-to-one tutoring." Educational Researcher, 13(6), 4–16. Kraft, M. A., Schueler, B. E. and Falken, G. T. (2024). "What impacts should we expect from tutoring at scale? Exploring meta-analytic generalizability." EdWorkingPaper 24-1031, Annenberg Institute, Brown University. 265 randomised trials of 340 tutoring programmes, pooled effect 0.42 standard deviations; by programme size, under 100 students 0.55, 100 to 399 0.32, 400 to 999 0.25, 1,000 and over 0.14.
The Ceiling
- Liu, A., Sun, M., Esbenshade, L., Tian, V., Zhang, Z. and He, K. (2026). "Teacher-authored prompts for configuring student–AI dialogue: K–12 classroom implementation." arXiv:2604.16738. University of Washington with Colleague AI; 16 teachers, 94 prompts, 1,479 conversations with 878 students, April to June 2025, United States. "71% were fully on-track, and fewer than 1% were substantially off-track"; "38% of conversations under-reached the teacher-targeted DOK level, approaching 50% when targeting DOK 3" (Webb's Depth of Knowledge, where level 3 is strategic thinking).
The Writing Corpus
- Nicholls, D. (2003). "The Cambridge Learner Corpus: error coding and analysis for lexicography and ELT." Proceedings of Corpus Linguistics 2003, Lancaster University, https://ucrel.lancs.ac.uk/publications/CL2003/papers/nicholls.pdf. Then 16 million words, 6 million of them error-coded, from learners with 86 first languages; "There are 88 possible codes in all." The full corpus is not public.
Purdue
- Arnold, K. E. and Pistilli, M. D. (2012). "Course Signals at Purdue: using learning analytics to increase student success." Proceedings of LAK '12, 267–270; and Pistilli, Arnold and Bethune, "Signals: using academic analytics to promote student success", EDUCAUSE Review, 17 July 2012: four-year retention 87.42% for students with at least one Signals course against 69.40% without. The critique: Caulfield, M. (2013). "Why the Course Signals math does not add up." Hapgood, 26 September 2013; Feldstein, M. (2013), e-Literate, 26 September and 3 November 2013, reporting Alfred Essa's simulation in which randomly distributed "chocolates" reproduced the retention pattern; Purdue's Matthew Pistilli told Inside Higher Ed (6 November 2013) that "more research needs to be done".
The Boundary
- Ryan, R. M. and Deci, E. L. (2000). "Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being." American Psychologist, 55(1), 68–78. Department for Education (2025). Generative artificial intelligence (AI) in education, policy paper updated 12 August 2025: "Technology, including generative AI (artificial intelligence), should not replace the valuable relationship between teachers and pupils."