Skip to content

Conclusion

×

When a Number Starts to Count

How do we measure students without gaming the system?

Key Takeaways

  1. When a number decides something important, like a grade or a job, people will bend it.

  2. Measure students often with little riding on each result, and have someone outside check the marks that still count.

  3. When someone outside checks a teacher's marks, the gap between poorer and richer students narrows.

In 1976 the social psychologist Donald Campbell set down what he called a pessimistic law. “The more any quantitative social indicator is used for social decision-making,” he wrote, “the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.”

He applied it to schools. Achievement tests, he said, may well be valuable indicators under normal teaching aimed at general competence. “But when test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways.”

Basically, when something can be gamed, it will be. A number that decides nothing gives nobody much reason to bend it, and a number that decides something, will be. A school’s rating, a teacher’s job or a student’s year gives everyone a reason.

The series has found that law at work everywhere. It’s been a constant theme. And I have seen it happen in so many places across education. I have succumbed to it as well. It’s something every person within education can relate to. But it’s something we need to ask ourselves.

Do we keep making the easy choices of how we measure things? Or do we try to reach beyond good exam results?

Reaching beyond them means measuring two things nobody measures now: what a teacher knows about how students go wrong, and the habits that predict a student’s life as well as their scores do: conscientiousness, self-control and turning up (please see 2. What a Report Card Cannot Say). And it means measuring them in a way that does not bend.

And as it happens, the grades bent first. When England cancelled exams in 2020 and 2021 and grades rested on teachers’ assessments, the share of A-level entries given an A or A* rose from 25% to 44%. In Spain, where teacher-set grades decide who repeats a year, a student from a poor family is five times more likely to have repeated than a student from a well-off family with the same test scores. Nothing checked the judgement in either case, and in both it moved.

Judgements about teachers bent the same way. Academies hire for fluency because parents can hear it, and the knowledge of how students go wrong in the subject, which is what predicts whether they learn, is not measured at the interview or afterwards. One in four Spanish secondary teachers works in a school that never appraises anyone. Where a score for teachers does exist, it moves so much from one year to the next that it says as much about the year as about the teacher, and it still decides careers.

Governments bent too. The Treasury’s case for education is a sum of the tax a person will pay over a working life, because tax is the one thing anyone collected. So the state spends on what shows in that sum and calls the rest unmeasured.

And the thing that counts most was never measured at all. What makes a teacher effective is knowing in detail how students go wrong in a subject, and no exam tests it, no interview reveals it and no appraisal records it. The qualities that predict a student’s life as well as their scores do, conscientiousness and self-control, appear on no report card, though a teacher’s effect on them predicts finishing school better than their effect on test scores. The knowledge that matters most sits in teachers’ heads, and nobody checks it or writes it down.

The obvious answer is to take the stakes away, and the evidence is against it. The American No Child Left Behind law raised fourth-grade maths scores on a national test that carried no consequences, by about 0.23 standard deviations by 2007, though it produced no gain in reading. When Wales abolished school league tables in 2001 and England kept them, Welsh schools fell behind by about two GCSE grades per pupil per year.

In Texas, pressure on schools at risk of a low rating raised their students’ later earnings, while pressure on schools chasing a higher rating harmed their lowest scorers. The evidence for measuring with no stakes at all is thinner. An English trial of peer observation with no grades attached found no detectable effect on exam results, and the programme’s designers, analysing more students, estimated a small gain. People work for a number that counts, and some of that work is real.

So low stakes is not the answer on its own, and the two halves of the evidence fit together in one place. Stakes produce effort and distortion at once. What decides the balance is whether anyone outside the system is checking the number, and whether the record that is meant to help a teacher is kept apart from the number that ranks a child.

Campbell’s remedy is a routine check from outside, by someone who has nothing riding on the number. He took it from the case that gave him the law. In Texarkana a company was paid according to its pupils’ test gains and taught them the answers to the test. Nobody inside the scheme had a reason to notice. It took an outside evaluator, with nothing riding on the gains, to find it.

The best recent evidence is of that kind. In Italy, middle-school teachers who were shown their own score on a test of implicit bias before setting end-of-term grades narrowed the gap between immigrant and native students.

Then in Denmark, a 2016 reform that handed the marking of a ninth-grade exam to outside examiners alone cut the grade gap between richer and poorer students by 15 to 20%. The grades still counted in both cases. The judge was checked.

New Zealand shows that the check need not take the judgement away though. Its qualifications authority moderates teacher-assessed work every year: each school sends samples of marked work, moderators check the school’s judgements against the standard, and the school learns where its marking drifts. The teacher keeps the judgement and the student keeps the grade.

So four things change.

The first is to write the knowledge down. Maths has a map of more than 8,000 ways students go wrong, written by former teachers, and the tools built on it help. Writing has no open map like it. It would have to be written by people who have taught it, subject by subject, and that is slower than building one chatbot with one set of rules.

The second is to record what a student does, not what they are, and compare it only with their own past. Drafts, attempts, time on a task, requests for help: a student doing less than they did last month is a student a teacher should look at. The record holds what happened on the task and nothing about the student’s family or life outside school.

The third is to keep that record away from grades, rankings and labels. This is where the stakes come off, and only here. The moment a mark or a punishment depends on it, students will manage the record, and Campbell’s law takes over. The record should only exist to show a teacher who needs help. The student and the parents see it when they ask, and when it has done its job it is deleted.

The fourth is to put someone outside the room on the judgements that carry weight. For most of them the teacher still marks and someone else checks the marking: moderation of the kind New Zealand runs, or a bias score shown to the teacher before the grades go in. For the one mark that can cost a student a year, the check is a second marker on the work, as Denmark did with its exam. The teacher keeps the judgement where it carries less weight, and shares it where one mark decides the most.

Put together, that is a school that looks different in small ways. Its teachers keep a record of how each student goes wrong and compare it only with that student’s past, and nothing in it counts towards a grade. Its marking is moderated from outside, so a grade means the same from one classroom to the next. And because the record carries more, the exam can carry less: fewer moments where a single number decides a year, not fewer tests.

Nothing in this series shows that exams should go or change shape. What it shows is that the judgement around them should be checked, and the record beneath them kept apart from the number that counts.

And this is where we neatly loop back to the beginning. The law in England and in Spain asks three things of a school: a qualification, a place among other people, and a person able to run their own life. Schools only measure the first, because it is the only one an exam can test.

The other two never reach an exam. A teacher sees them in small things: whether a student turns up, keeps going when the work gets hard, keeps their temper, asks for help. Those small things are what the New Zealand children’s self-control and the North Carolina teachers’ effects were made of, and they predicted health, income, convictions and finishing school better than the scores did.

A record that keeps them, compared with the student’s own past and never used against them, is the first measure of the two purposes nobody measures. The cost of leaving them unmeasured is the one the Treasury already counts: the young people who leave school and neither work nor study.

For ten years the number that counted for me was an exam in June. What I wanted, and never had, was a real picture of where each student was, in real time. Not just an idea of where they were. We have the ability to build the picture. To really help our students. We don’t have to wait for the exam. Whether a picture like that improves results or reduces bias is the test still to run.

Underneath all of it is bias. Where this series found a judgement left unchecked, it fell hardest on the student with the least behind them: the poorer child in Spain, held back a year when a richer child with the same scores was not; the immigrant child in Italy, marked below the native one; the poorer child in Denmark, marked down by their own teacher. Where someone checked, the gap narrowed. Removing that bias is the reason to measure differently at all.

Sources

Most of this essay restates findings sourced in the earlier essays; those entries are not repeated here. "Read" means the original was opened.

Campbell's Law

  • Campbell, D. T. (1976). Assessing the Impact of Planned Social Change. Occasional Paper No. 8, Public Affairs Center, Dartmouth College. Read. The law is on p. 49, the achievement-test passage on pp. 51 to 52, and the Texarkana case and the "objectivity-preserving features" line are in the same section. Reprinted in Evaluation and Program Planning, 2(1), 67–90 (1979). Campbell described his evidence as mostly anecdotal.

From Who Checks the Teacher

  • A-level A or A* share 25.2% in 2019 to 44.3% in 2021 (Ofqual and JCQ figures for England, read). Repetition in Spain: Cobreros, Gortazar, Gairal and del Moral (2026), Todo lo que debes saber de PISA 2025 sobre equidad, EsadeEcPol and Save the Children, read. New Zealand moderation: NZQA's external moderation of internally assessed standards, checked for that essay.

From What Makes a Great Teacher

  • Hiring for fluency: the Catalonia language-school study and the German study of teacher knowledge. Appraisal in Spain: OECD TALIS 2018, about one in four lower secondary teachers in a school that never formally appraises. Teacher effect scores: McCaffrey, Sass, Lockwood and Mihaly (2009), Education Finance and Policy, 4(4), 572–606; year-to-year correlations of 0.2 to 0.5 for primary teachers.

From The Case for the Treasury

  • The OECD's public-returns sum (Education at a Glance 2021, Indicator A5) and the World Bank review's admission that non-money benefits are left out for want of evidence.

From What a Report Card Cannot Say

  • Conscientiousness and grades: Poropat (2009). Self-control and adult outcomes: Moffitt and colleagues (2011), Dunedin cohort. Teacher effects on behaviour versus scores: Jackson (2018), North Carolina. The conditions on the record (task only, own past only, no consequences, deleted when done) are that essay's.

From The Tool Was Never the Hard Part

  • The maths misconception map (Eedi, more than 8,000 misconceptions, former maths teachers) and the absence of an open equivalent for writing.

Stakes That Worked

  • Dee, T. and Jacob, B. (2011), Journal of Policy Analysis and Management, 30(3), 418–446. Burgess, S., Wilson, D. and Worth, J. (2013), Journal of Public Economics, 106, 57–67. Deming, D., Cohodes, S., Jennings, J. and Jencks, C. (2016), Review of Economics and Statistics, 98(5), 848–862. Dee and Jacob: fourth-grade maths "effect size = 0.23 by 2007" and "no evidence that NCLB increased fourth-grade reading achievement" (abstract). Burgess, Wilson and Worth: the fall in Welsh schools' effectiveness amounts to about two GCSE grades per pupil per year (CMPO summary). Deming and colleagues: risk of a Low Performing rating raised college attainment and earnings at 25; pressure to reach a higher rating had "large negative impacts on attainment and earnings for the lowest-scoring students" (abstract).

The English Observation Trial

  • Worth, J., Sizmur, J., Walker, M., Bradshaw, S. and Styles, B. (2017). Teacher Observation: Evaluation report. EEF and NFER. Read. Effect −0.01, pre-registered, 7,366 pupils. Burgess, S., Rawal, S. and Taylor, E. (2021). Journal of Labor Economics, 39(4), 1155–1186. Read. Same 82 schools, both cohorts and both subjects pooled, about 28,000 pupils, estimate 0.07.

Italy

  • Alesina, A., Carlana, M., La Ferrara, E. and Pinotti, P. (2024). American Economic Review, 114(7), 1916–1948. Abstract read for What a Report Card Cannot Say. The finding used here is from the first experiment.

Denmark

  • Birkelund, J. F. (2026). "Teacher Bias, Socioeconomic Inequality, and the Role of Academic Reputations: Evidence from a Grading Reform." Sociology of Education, 99(4), 330–351. Abstract read; the full paper has not been opened. Written Danish exam; no effect in mathematics.