Part 1 of 7
×Who Checks the Teacher?
Who checks the judgement behind a mark?
Key Takeaways
The law asks schools to develop the whole person, but schools only check exam results.
Teachers' judgements of students are often right, but they can be biased, and nobody checks them.
When a teacher's unchecked judgement decides a student's future, poorer students lose out most.
Ask a government what school is for and the answer is broad. England’s Education Act 2002 requires state schools to offer a curriculum that promotes pupils’ “spiritual, moral, cultural, mental and physical development” and prepares them for “the opportunities, responsibilities and experiences of later life”. Spain’s education law gives as its first aim the full development of students’ personality and capacities.
In 2009 the education researcher Gert Biesta gave names to what laws like these describe. He argued that education has three purposes. Qualification gives students knowledge and skills, socialisation brings them into the traditions and ways of a society, and subjectification helps them become independent people in charge of their own lives. Both laws cover all three.
Ask any teacher and they will tell you that all of this is implied, myself included. But whether they would say they themselves are measured across all three purposes in the development of a student is another question.
When two UCL researchers analysed the OECD’s 2018 survey of teachers, they found that 68% of teachers in England said being held responsible for their students’ achievement was a source of stress, against an average of about 45% across the countries surveyed. In 2017 England’s chief inspector of schools reported that the primary curriculum was narrowing in some schools because of the time given to preparing for national tests, and that ten of the 23 secondary schools her inspectors visited had cut a year from the middle phase of schooling so that GCSE courses could start early.
So of the three purposes, schools check one, and they check it with two instruments, an exam score and a teacher’s judgement.
Both are ways of measuring a student, and each can be frequent or rare, and carry high or low stakes. A weekly quiz that counts for nothing is frequent and low-stakes, and a final exam is rare and high-stakes. Daisy Christodoulou has argued that teacher assessment is biased and that frequent, low-stakes tests serve learning better than either exams or teacher grades, a view my experience has also led me to believe. But furthermore, the larger fault is high stakes resting on a judgement nobody checks, and that fault survives even when the exams are abolished.
Start with the judgement. In 1968, the infamous study by Robert Rosenthal and Lenore Jacobson told teachers at a California primary school that about a fifth of their pupils were about to bloom intellectually, even though the pupils had actually been picked at random. A year later the named pupils had gained more on an IQ test than their classmates. Later attempts to repeat it gave mixed results, but the finding is striking, and I can see it happening. What is not in doubt is that the effect was concentrated in the two youngest year groups.
Timing explains both the mixed results and the youngest groups. A 1984 review of those attempts by Stephen Raudenbush found that the effect appeared when teachers had known their pupils for two weeks or less, and disappeared once they knew them well.
Who the teacher is shapes the expectation too. A 2016 American study compared two teachers’ expectations of the same student. Non-Black teachers expected less of Black students than Black teachers did.
This raises various issues within education, but for the scope of this topic, teacher judgement is still worth having. A 2012 review of 75 studies found that teachers’ judgements of attainment matched test results reasonably well on average, with wide variation from one teacher to the next. What it lacks is a check.
This has happened to me. While I never stopped caring for the students, I have written them off, who went on to ace the exam. And more dangerously, I have been sure of students, who then went on to fail it. Either way, when you actually receive the feedback on results it’s too late. In both cases, the student has been failed.
Then the other two purposes leave almost no record at all, other than perceptions and memories.
This isn’t the teacher’s fault. With large classes of varying dependencies and levels, the teacher must push the high achievers and support the struggling, because they all have exams to pass. Then on top of that, they must find a way to support the other purposes. In preparing them for future responsibilities of later life, such as the skills needed in the workplace, it’s difficult to monitor, especially long term.
Whether a student has learned to work with people they did not choose, or to carry on with something they find hard, appears nowhere in their file unless a teacher thinks to write a line about it, and that line is one adult’s impression, which nobody checks either.
Spain shows what happens when a judgement of that kind carries high stakes. Spanish law gives regions and schools wide freedom to set their own assessment criteria, and in the schools I knew, each subject teacher wrote their own tests, marked them and turned the marks into the grade, with nobody outside the classroom seeing the paper or checking the marking. Subjects were tested about once a month, mostly on recall, with most results counting towards the final grade. For years, failing three or more subjects normally meant repeating the year.
The 2025 PISA tests found that 22% of 15-year-olds in Spain had repeated at least one year, the third-highest rate in the OECD, where the average was 9.8%. About 40% of the poorest quarter of students had repeated, against about 7% of the richest quarter. An analysis by EsadeEcPol and Save the Children compared students who scored the same in the PISA tests. Those from poor families were five times more likely to have repeated. An outside test says these students can do the same work, and the grades given inside their schools sent them different ways.
The system I have experience with, in the Basque Country, it turns out holds back fewer students than most of Spain. But it is also where the clearest evidence comes from. Basque students sit external diagnostic tests that are marked blind and count for nothing. A 2022 study of more than 31,000 of them compared those results with the grades their teachers gave. At the same test score, boys, students from immigrant families and poorer students got lower grades from their teachers.
Interestingly, Spain reformed assessment in 2020. The new law scrapped the national end-of-stage exams an earlier law had introduced, kept national diagnostic tests that have no effect on a student’s record, and said assessment should be continuous and formative. It also made repeating a year an exception, to be decided by a student’s teachers as a group. The share of 15-year-olds who had repeated was the same in 2025 as in 2022.
The reform dealt with the rare, national measurement, and brought it in line with what the studies suggest is the best educational practice. Yet in the classrooms I saw it left the monthly tests untouched, so students had more high-stakes moments than before, not fewer.
England has the opposite arrangement. Its exams are rare and marked by strangers, so the judge (and potential executioner) is checked, but nothing else is measured. When the exams were cancelled in 2020 and 2021 and grades rested on teachers’ assessments, the share of A-level entries in England given an A or A* rose from 25% in 2019 to 44% in 2021. Under high stakes and without a check, teacher judgement moved there too.
One system already puts a second pair of eyes on a teacher’s marking without taking the marking away from the teacher. In New Zealand the national qualifications authority moderates teacher-assessed work every year. Each school sends samples of marked work for each standard it assesses, and moderators check the school’s judgements against the standard and tell the school where its marking drifts.
The teacher keeps the judgement, the student keeps the grade, and the school learns each year whether its marks mean what they say. Nothing in it needs a national exam. A school’s own tests, with a sample of the marking read outside the school each year, would give Spain the check it lacks and England a measurement it could afford to repeat.
Both countries promise in law to develop the whole person, and both check one purpose out of three. England checks the marker and measures rarely. Spain measures often and routinely checks no one. A fair measurement would be frequent, carry little weight, and be tested against something outside one teacher’s head. Few systems combine all three. Until schools have a measurement like that, the law will keep promising the whole person while a student’s future turns on a single number and on the one adult who set it.
Sources
Laws and Purposes
- Education Act 2002, section 78(1), legislation.gov.uk. The duty applies to maintained schools; academies are bound through their funding agreements.
- Ley Orgánica 2/2006, de Educación (LOE), article 2.1(a), consolidated text at boe.es. Article 2.1 opens "El sistema educativo español se orientará a la consecución de los siguientes fines: a) El pleno desarrollo de la personalidad y de las capacidades de los alumnos." The wording of letter (a) is the same in the 2013 and the 2024 consolidated texts.
- Biesta, G. (2009). "Good education in an age of measurement: on the need to reconnect with the question of purpose in education." Educational Assessment, Evaluation and Accountability, 21(1), 33–46.
Accountability and Narrowing in England
- Jerrim, J. & Sims, S. (2021). "School accountability and teacher stress: international evidence from the OECD TALIS study." Educational Assessment, Evaluation and Accountability, doi 10.1007/s11092-021-09360-0. TALIS 2018 asked lower-secondary teachers how far "being held responsible for students' achievement" was a source of stress; 68% in England answered "quite a bit" or "a lot", against a cross-country average of about 45%. The England and average figures are reported in Jerrim's UCL IOE blog post of 18 March 2021 summarising the paper.
- Spielman, A. (2017). "HMCI's commentary: recent primary and secondary curriculum research." Ofsted, 11 October 2017. Reports that "the primary curriculum is narrowing in some schools as a consequence of too great a focus on preparing for key stage 2 tests", and that "ten of the 23 secondary schools visited for this current survey were reducing key stage 3 to just a 2-year period of study."
Frequency and Stakes
- Christodoulou, D. (2017). Making Good Progress? The Future of Assessment for Learning. Oxford University Press.
Teacher Expectations and Judgement
- Rosenthal, R. & Jacobson, L. (1968). Pygmalion in the Classroom. Holt, Rinehart and Winston. Gains were concentrated in grades 1 and 2; replications are mixed.
- Raudenbush, S. W. (1984). "Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction." Journal of Educational Psychology, 76(1), 85–97.
- Gershenson, S., Holt, S. B. & Papageorge, N. W. (2016). "Who believes in me? The effect of student–teacher demographic match on teacher expectations." Economics of Education Review, 52, 209–224.
- Südkamp, A., Kaiser, J. & Möller, J. (2012). "Accuracy of teachers' judgments of students' academic achievement: a meta-analysis." Journal of Educational Psychology, 104(3), 743–762. Mean correlation .63 across 75 studies.
Spain
- Cobreros, L., Gortazar, L., Gairal, M. & del Moral, C. (2026). Todo lo que debes saber de PISA 2025 sobre equidad. EsadeEcPol and Save the Children, September 2026. Executive summary: 22% of 15-year-olds have repeated, "igual que en 2022"; Spain is "el tercer país de la OCDE con más repetición, más del doble que la OCDE-35 (9,8 %)"; the top socioeconomic quartile repeats at about 7% and the bottom quartile at 40%; "a igualdad de competencias, un estudiante del cuartil bajo tiene 5 veces más probabilidades de haber repetido que uno del alto". Note that on 22 September 2026 the authors said comparisons between 2022 and 2025 by socioeconomic level are under revision because the PISA socioeconomic index is not comparable across the two rounds; the essay uses only 2025 figures by socioeconomic level, and the one 2022 comparison it makes is the overall repetition rate.
- Gortazar, L., Martínez de Lafuente, D. & Vega-Bayo, A. (2022). "Comparing teacher and external assessments: Are boys, immigrants, and poorer students undergraded?" Teaching and Teacher Education, 115, 103725. Two cohorts, 31,183 pupils in publicly funded Basque schools, in the fourth year of primary and second year of ESO; teacher grades compared with blindly marked external tests.
- Ley Orgánica 3/2020, de 29 de diciembre (LOMLOE), BOE núm. 340, 30 December 2020, amending the LOE. Its preamble states that ESO assessment "será continua, formativa e integradora", that promotion decisions "serán adoptadas de forma colegiada por el equipo docente", and that the diagnostic evaluations have "carácter informativo, formativo y orientador". The relevant articles are LOE articles 28 (ESO evaluation and promotion), 21 and 29 (diagnostic evaluations) and 144. The LOMCE end-of-stage evaluations it removed had been suspended since 2016 and never took full effect. Regional and school autonomy over assessment: Eurydice, Spain, "Assessment in general lower secondary education".
England's Teacher-Assessed Grades
- FFT Education Datalab (10 August 2021). "A-Level results 2021: the main trends in grades and entries." Analysis of JCQ data, all-age figures for England: A or A* 25.2% in 2019, 38.1% in 2020, 44.3% in 2021. In 2020 the final grade was the higher of the centre assessment grade and the calculated grade; in 2021 grades were teacher-assessed (Ofqual, "Summer 2021 results analysis and quality assurance").
New Zealand
- NZQA, "External moderation" and "External moderation outcomes" pages (NCEA for teachers and schools), read 2 October 2026. Schools submit samples of internally assessed student work for each standard in their moderation plan, normally six samples per standard spread across the grades; moderators assess each sample against the standard and rate the school's judgements as consistent, not yet consistent or not consistent, and the school receives the feedback. Moderation does not change current students' grades, and the schools choose which samples to send.