Part 2 of 7
×What a Report Card Cannot Say
What should we track instead?
Key Takeaways
Habits like effort and self-control shape a student's life as much as grades do.
Schools should record what students do in class and compare each student only with their own past.
A record of student behaviour must never count towards a grade, or students will start gaming it.
Everything a school writes down about a student fits on a page, and nearly all of it is grades. Whether they try again after a wrong answer, ask when they are lost, or carry their share of a group task appears nowhere, although the evidence says those habits shape the rest of their life. What little a school knows about them sits in one teacher’s head, and nobody checks it. Schools should write down what each student does on their work, the attempts, the requests for help, the gaps between tries, and use that record for two things only: to show which students need help, and to check the teacher’s own picture of each student. Compared only with the student’s own past, and never touching a grade, it is fair. Used any other way, it becomes the thing it was meant to fix.
One condition comes before everything else, which is that whatever is recorded must never be used against the student. It should exist only to show a teacher who needs help. The moment a grade, ranking or punishment depends on such a system, it stops working.
The case for looking beyond test scores is strong. A 2009 analysis of studies covering more than 70,000 students found that conscientiousness predicts grades about as well as intelligence does at secondary school and university, though not in primary school. A study that followed just over a thousand New Zealand children to the age of 32 found that self-control in childhood predicted adult health, wealth and criminal convictions, after accounting for IQ and family background.
Teachers affect these qualities. When the economist Kirabo Jackson studied ninth-grade teachers in North Carolina, he measured their effect on four things schools already record, which were attendance, suspensions, grades and moving up a year on time. A teacher’s effect on those behaviours predicted whether students finished school better than the teacher’s effect on test scores did, and because the two effects were only weakly related, a teacher could be good at one and ordinary at the other.
The labour market has noticed the same qualities. In the United States, jobs that depend on working closely with other people grew by nearly twelve percentage points as a share of all jobs between 1980 and 2012, and the pay premium for social skills rose with them. In Sweden, the pay return to a psychologist’s rating of teamwork and leadership roughly doubled between 1992 and 2013, while the return to cognitive skill stayed flat.
So the qualities matter and schools shape them, but measuring them is difficult.
Asking students to rate themselves fails because the standard they judge themselves by moves. In Boston, students who won a lottery for a place at a high-achieving charter school made larger test-score gains than those who lost, yet rated themselves lower on self-control and grit. The schools, the researchers concluded, had raised the bar the students measured themselves against.
Angela Duckworth gave the name grit to perseverance towards long-term goals in a 2007 paper, and a TED talk and a bestselling book made it a word every school knew. Yet in 2015 she warned with David Yeager that no available measure of these qualities, questionnaire or teacher rating, is fit for judging schools against one another. Two years later a review of 88 samples found that grit overlaps almost entirely with conscientiousness and adds little to predicting performance. The one part that added something was perseverance of effort.
That points to the alternative, which is to record what students do.
Some evidence for this already exists. One study took six long-running American surveys of teenagers and counted how many of the questions each teenager bothered to answer, and that share predicted how many years of schooling they went on to complete. Research on maths tutoring software has separated students who keep working productively from those who repeat the same failed attempt, and productive persistence in middle school predicted who later enrolled in college. The same body of research found students gaming the software by clicking through hints for the answer, and those students learned less.
That evidence shows behaviour predicts outcomes. It does not show that writing it down helps anyone.
A classroom record could hold similar things, such as how many attempts a piece of work took, whether the errors are careless slips or signs of a misunderstanding, whether a student asks for help when stuck or guesses and moves on, and how long they leave between a failed attempt and the next one. Much of this already sits inside tools schools use.
The record costs nothing to collect, because the counts already sit in the platforms, but it costs time to read, and nothing in a teacher’s week gives that time up on its own. Added on top of marking and planning, it becomes a dashboard nobody opens. A school has to decide what the teacher will stop doing before it switches the record on.
Two rules decide whether such a record is fair.
The first is that each student is compared only with their own past. A record showing that this student asks for help less often than they did last term needs no comparison with anyone else.
The second is that the record describes behaviour on a task and never turns into a score for a trait. A 2023 American study found that calling a child a “visual learner” or a “hands-on learner” changed how parents, teachers and other children judged that child’s intelligence, and adults expected the “hands-on” child to get worse grades in maths and language. Learning styles have no evidence behind them, and a persistence score would be a label of the same kind. “Took one attempt at each of the last five tasks and stopped” tells a teacher something to act on, while “low persistence” tells them what kind of child to expect.
A record like this has two uses, and the evidence for each is different.
The first is spotting who needs help. American schools have used early warning systems for years, built on attendance, behaviour and course grades, and these indicators predict well. Whether flagging students changes anything is less clear.
In a randomised trial across 73 American high schools, schools given an early warning system had fewer chronically absent students after a year, 10% against 14%, and fewer failing a course, 21% against 26%. Suspensions and credits earned did not change, and grades barely did. A trial across 42 Norwegian schools found no effect at all on attendance, completion or results after two years. A flag helps when a teacher knows what to do next.
The second use is checking the teacher’s judgement of each student. Every teacher carries a private ranking of the class, who is able, who is trying, who has given up, and for most of the year nothing tests it. Mine was tested once a year through examination. And more than once my ideas of a student were wrong in both directions. A record of what each student did gives a teacher something to set against that ranking in November rather than June. That ranking leaks into grades, and two studies show that an outside check on it makes grading fairer.
In Denmark, a 2016 reform took teachers out of marking a ninth-grade written Danish exam and left it to outside examiners alone. A study of the reform found the grade gap between students from more and less advantaged homes was 15 to 20% wider when the students’ own teachers had a hand in the marking. Most of that gap came from teachers leaning on what they already knew of each student, so a child with a strong record got the benefit of the doubt on a weak paper.
In Italy, researchers tested middle-school teachers for implicit bias against immigrant students. Teachers shown their own result before setting end-of-term grades narrowed the gap between immigrant and native students. A general talk about stereotypes did as well on average, but the most biased teachers changed only when they saw their own result. Teachers grade more fairly when their judgement meets an outside check.
The Danish study is also the strongest objection to writing any of this down. The teachers who widened the gap were leaning on what they knew of each student’s past work, and a written record gives a teacher more to lean on.
That is why the two rules matter. Compared only with the student’s own past, the record holds no ranking to lean on, and kept out of grading, it cannot be used against anyone. One risk remains, since the teacher who reads the record also sets the mark, and a perception with more to stand on can be a perception corrected or a perception confirmed. Which happens is the thing to test, and the two rules make the test safe to run.
The problem is that every measure here is easy to game, since a student can pad out time on a task or add drafts, and the tutoring research found students gaming software even when nothing rode on it. Once a grade depends on the record, students will manage the record. A flag can also harden into a label, which is how teacher expectations often go wrong in the first place, and the record can carry bias of its own, since a teacher decides who gets a second attempt and whose request for help is noticed.
Each of these is a reason to keep the record away from grades and labels, not a reason to go on keeping nothing. So the record holds what happened on the task and nothing about the student’s family or life outside school. The student’s teacher sees it, the student and the parents see it when they ask, and nobody whose interest is in ranking the child sees it at all. When it has done its job, it is deleted. Spanish law asks the same of a school anyway. Hopefully, this will keep what teaching needs and nothing more.
With all this in mind, what I have tried to show is that a record of what a student does, compared with their own past and kept free of consequences, can show a teacher who is struggling and give them something outside their own head to check their read against. Bias is a problem everywhere. Education is no different. I believe that tools should be built to combat it. Whether these tools can improve results or reduce bias is what I intend to find out.
Sources
The order follows the essay. "Read" means the original report or article was checked; "abstract read" means the abstract was.
Conscientiousness and Grades
- Poropat, A. E. (2009). "A meta-analysis of the five-factor model of personality and academic performance." Psychological Bulletin, 135(2), 322–338. Read. Cumulative samples "ranged to over 70,000". Conscientiousness and intelligence had similar validity at secondary and tertiary level; at primary level intelligence was the stronger predictor.
Childhood Self-Control
- Moffitt, T. E. et al. (2011). "A gradient of childhood self-control predicts health, wealth, and public safety." PNAS, 108(7), 2693–2698. Abstract read. Cohort "of 1,000 children from birth to the age of 32"; effects "could be disentangled from their intelligence and social class". The Dunedin study has 1,037 members. Correlational.
Teacher Effects on Behaviour
- Jackson, C. K. (2018). "What Do Test Scores Miss? The Importance of Teacher Effects on Non-Test Score Outcomes." Journal of Political Economy, 126(5), 2072–2107. Abstract read: the behaviours are "absences, suspensions, course grades, and grade repetition in ninth grade"; teacher effects on them "predict larger impacts on high school completion and other longer-run outcomes than their effects on test scores", and the two kinds of effect "are weakly correlated". North Carolina is named in the earlier working-paper title (NBER WP 18624). Course grades are in the measure.
Labour Market
- Deming, D. J. (2017). "The Growing Importance of Social Skills in the Labor Market." Quarterly Journal of Economics, 132(4), 1593–1640. Abstract read (NBER WP 21473 version): jobs requiring high social interaction "grew by nearly 12 percentage points as a share of the U.S. labor force" between 1980 and 2012, and "the labor market return to social skills was much greater in the 2000s than in the mid 1980s and 1990s". Edin, P.-A., Fredriksson, P., Nybom, M. & Öckert, B. (2022). "The Rising Return to Noncognitive Skill." American Economic Journal: Applied Economics, 14(2), 78–100. Abstract read: between 1992 and 2013 the return to "a psychologist-assessed measure of teamwork and leadership skill" roughly doubled, while the return to cognitive skill "was relatively stable and decreased modestly during the 2000s". Swedish administrative data; the sample is men.
The Boston Paradox
- West, M. R., Kraft, M. A., Finn, A. S., Martin, R. E., Duckworth, A. L., Gabrieli, C. F. O. & Gabrieli, J. D. E. (2016). "Promise and Paradox: Measuring Students' Non-cognitive Skills and the Impact of Schooling." Educational Evaluation and Policy Analysis, 38(1), 148–170. Read (open access). 1,368 Boston eighth graders; "exploiting admissions lotteries, we find positive impacts of charter school attendance on achievement and attendance but negative impacts on these non-cognitive skills", with "suggestive evidence" that reference bias explains it.
Grit
- Duckworth, A. L., Peterson, C., Matthews, M. D. & Kelly, D. R. (2007). "Grit: Perseverance and Passion for Long-Term Goals." Journal of Personality and Social Psychology, 92(6), 1087–1101. Duckworth's TED talk "Grit: the power of passion and perseverance" (April 2013) and her book Grit: The Power of Passion and Perseverance (Scribner, 2016, a New York Times bestseller) carried the term into schools.
Measurement Warning
- Duckworth, A. L. & Yeager, D. S. (2015). "Measurement Matters: Assessing Personal Qualities Other Than Cognitive Ability for Educational Purposes." Educational Researcher, 44(4), 237–251. Abstract read: self-report questionnaires, teacher-report questionnaires and performance tasks are each "imperfect in its own way", and "we do not believe any available measure is suitable for between-school accountability judgments".
Grit and Conscientiousness
- Credé, M., Tynan, M. C. & Harms, P. D. (2017). "Much Ado About Grit: A Meta-Analytic Synthesis of the Grit Literature." Journal of Personality and Social Psychology, 113(3), 492–511. Read. 584 effect sizes from 88 samples and 66,807 people; grit "very strongly correlated with conscientiousness" (ρ = .84), and "the perseverance of effort facet has significantly stronger criterion validities than the consistency of interest facet".
Survey Effort
- Hitt, C., Trivitt, J. & Cheng, A. (2016). "When you say nothing at all: The predictive power of student effort on surveys." Economics of Education Review, 52, 105–119. Abstract read: "six nationally-representative, longitudinal surveys of American youth"; the share of questions skipped as an adolescent "is a significant predictor of later-life educational attainment, net of cognitive ability".
Tutoring Software
- Adjei, S. A., Baker, R. S. & Bahel, V. (2021). "Seven-Year Longitudinal Implications of Wheel Spinning and Productive Persistence." AIED 2021, LNCS 12748, 16–28. Abstract read: "productive persistence during middle school mathematics is associated with a higher probability of college enrollment", while wheel-spinning "is not statistically significantly associated with college enrollment in either direction". Beck, J. E. & Gong, Y. (2013), AIED, define wheel-spinning. Baker, R. S., Corbett, A. T., Koedinger, K. R. & Wagner, A. Z. (2004). "Off-task behavior in the Cognitive Tutor classroom: when students 'game the system'." CHI 2004, 383–390. The original could not be opened; Baker (2007, CHI) summarises it as showing gaming ("systematic guessing and persistent overuse of hints") "significantly associated with poorer learning".
Labels
- Sun, X., Norton, O. & Nancekivell, S. E. (2023). "Beware the myth: learning styles affect parents', children's, and teachers' thinking about children's academic potential." npj Science of Learning, 8, 46. Three US experiments; children, parents and teachers rated "visual learners" as more intelligent than "hands-on learners", and adults expected visual learners to get better grades in maths, language arts and social studies. Read.
Early Warning Trial, US
- Faria, A.-M., Sorensen, N., Heppen, J., Bowdon, J., Taylor, S., Eisner, R. & Foster, S. (2017). Getting Students on Track for Graduation: Impacts of the Early Warning Intervention and Monitoring System after One Year (REL 2017–272). IES/REL Midwest. Read (ERIC record ED573814). 73 high schools in three states, 2014/15, grades 9 and 10. Chronic absence 10% v 14%, failing a course 21% v 26%, both significant. No significant effect on the share with a low GPA (17% v 19%), on suspensions (9% in both) or on insufficient credits (14% in both); a sensitivity analysis found a small significant rise in mean GPA (2.98 v 2.87), hence "grades barely did". The authors note "overall implementation of the EWIMS seven-step process was low".
Early Warning Trial, Norway
- Sletten, M. A., Tøge, A. G. & Malmberg-Heimonen, I. (2023). "Effects of an early warning system on student absence and completion in Norwegian upper secondary schools: a cluster-randomised study." Scandinavian Journal of Educational Research, 67(7), 1151–1165 (online 2022). Abstract read (ERIC EJ1401081): the IKO model; 7,677 first-year students in 42 schools, 20 randomised to the model and 22 to control; "after two school years, there were no significant effects on absence from lectures, completion rates, or academic results".
Denmark
- Birkelund, J. F. (2026). "Teacher Bias, Socioeconomic Inequality, and the Role of Academic Reputations: Evidence from a Grading Reform." Sociology of Education, 99(4), 330–351. Abstract read. A 2016 reform replaced joint marking by the student's teacher and an external examiner with a single external examiner for the ninth-grade written Danish exam. Difference-in-differences; "teacher involvement increases the socioeconomic grading gap by 15 percent to 20 percent", and the effect "is largely explained by objective measures of prior academic performance", which the author reads as reputational rather than socioeconomic bias. No effect in mathematics.
Children's Data
- Ley Orgánica 2/2006 de Educación, disposición adicional vigesimotercera, "Datos personales de los alumnos": schools may collect the data needed for their educational function; the data must be limited to what the teaching and guidance function needs and may not be processed for other purposes without express consent.
Italy
- Alesina, A., Carlana, M., La Ferrara, E. & Pinotti, P. (2024). "Revealing Stereotypes: Evidence from Immigrants in Schools." American Economic Review, 114(7), 1916–1948. Abstract read (AER and NBER WP 25333). Middle schools; immigrant students get lower grades than natives with the same scores on blind standardised tests; teachers told their Implicit Association Test score before term grading reduced the gap. In a second experiment generic debiasing had a similar average effect, but "teachers with stronger negative stereotypes do not respond to generic debiasing but change their behavior when informed about their own IAT".