Skip to content

Part 4 of 7

×

Good Design, Not Learning Styles

What makes a tool help a student think?

Key Takeaways

  1. Learning styles have no evidence behind them, and labelling children "visual" or "hands-on" does harm.

  2. A good learning tool adapts to what a student already knows about the topic.

  3. A tool like that has to be built with teachers who know how students go wrong in the subject.

Almost every EdTech product now describes itself as adaptive or personalised, and some people talk about AI agents that could one day replace teachers. But few of them say what the product is adapting to.

Everyone’s natural response would be, “Errr… the students? Duh.” But that still misses the point.

Most teachers would fill that gap with learning styles. A review of 37 studies, covering 18 countries and more than 15,000 educators, found that 89% believe teaching should be matched to a student’s learning style. Something I’ve seen in teacher training and from day to day conversations with fellow educators. One study put the figure at almost 98%, and the review found no sign that the belief had declined between 2009 and 2020. So when a product promises to personalise learning, this is what many of the people buying and using it will hear.

It is easy to see why the idea took hold. The students in a classroom sit at different stages on any topic, one explanation lands differently from one student to the next, and no teacher can adapt a lesson to every person in the room. Learning styles offers a tidy answer, which is to sort the students, match the method, and consider the problem solved.

But the problem is that the research rejected that answer almost twenty years ago.

In 2008 the Association for Psychological Science commissioned a group of cognitive psychologists led by Hal Pashler to settle the question. They tested what they called the meshing hypothesis, which is what most people mean by learning styles, and which says that a visual learner remembers more from a diagram and an auditory learner remembers more from hearing the same thing aloud. Proving it takes a particular kind of experiment, in which you sort students by their stated style, assign them at random to different teaching methods, and then show that the method that works best for one group works worse for the other.

Pashler’s team found that very few studies had ever used that design, and several of those that did flatly contradicted the hypothesis. They concluded that there was no adequate evidence base for teaching students in their preferred style.

They also suggested why the idea survives.

It is more comfortable to believe the teaching didn’t suit your style than to put your lack of success down to ability or effort. Myers-Briggs has the same appeal.

The belief does harm. A 2023 American study found that calling a child a “visual learner” or a “hands-on learner” changes how intelligent parents, teachers and other children judge them to be. Adults expected visual learners to get better grades in maths and language, and hands-on learners to do better in art, PE and music, and children as young as six showed the same bias about intelligence.

So what should an adaptive product adapt to?

The best-evidenced answer comes from cognitive load research, which studies how much the brain can hold and process at once, and it is called the expertise reversal effect. Someone new to a topic learns better with heavy structure, worked examples and step-by-step guidance, but the same support can stop helping, and can even hinder, someone who already knows the topic well. A learning style is supposed to be a fixed trait of a person, whereas this is a state tied to a topic, so the same student might need every bit of scaffolding in one subject and almost none in another.

An adaptive product should adapt to what a student knows about the topic in front of them.

Adapting to knowledge has worked across many products. A 2016 meta-analysis pooled fifty controlled evaluations of intelligent tutoring systems and found a median gain of 0.66 standard deviations over conventional teaching, roughly the jump from the 50th percentile to the 75th. The gain depended heavily on the test, and was much smaller on standardised tests than on tests built around the tutoring content, which is the figure closer to what a school would see.

Research on teaching points the same way. One of the best-evidenced factors in great teaching is pedagogical content knowledge, which means understanding how students think about a topic well enough to catch their common misconceptions before they harden into habit. A teacher builds a picture of where each student’s understanding is wrong and why, and is limited by how many students they can hold in mind at once. A well-designed AI system can apply that picture to every student at once, but it cannot draw the picture, because the map of how students go wrong in a subject has to come from people who have taught it.

Khan Academy’s Khanmigo shows the approach in a working product. It runs on the same kind of model as ChatGPT, but with instructions that stop it handing over answers, so when a student asks for one it asks a guiding question back, and it works from the Khan Academy lesson the student is on. Questioning of this kind only works once a learner has some footing in the topic, because someone with no foundation who is asked to reason their way there will guess. Any developer can build a chatbot that answers homework questions. The hard part is building one that won’t.

Its guardrails have limits, and the clearest one is the subject. A question-first design suits problems with one right answer, where every question can steer towards it, better than open writing. Ask it for help with an essay and the line between guiding a student and doing the work for them is harder to keep in line.

The strongest single test comes from a Harvard physics course, where researchers built an AI tutor on the same principles as good active learning, so that it managed how much a student had to hold in mind at once and gave personalised feedback.

In a randomised crossover trial, 194 undergraduates studied two physics units, one with the AI tutor and one in an active learning class taught by highly rated instructors, so each student served as their own comparison. Their learning gains with the tutor were more than double those in the class, most finished in less time, and they reported feeling more engaged. The comparison was with some of the best classroom teaching available, for motivated students at a selective university.

Adaptivity has good evidence behind it once you can say what is being adapted to. That might be a student’s current knowledge, a specific misconception, or which explanation worked last time. A product that cannot name its target is probably not doing much for learning.

Naming the target has to be done again for each subject and each task, which makes it slower and harder than building one chatbot with one set of rules. But it is worth it. Teachers who have taught the subject are the ones who know in detail how students go wrong in it, and naming the target depends on that knowledge.

Much of the talk about AI in education asks whether teachers will be replaced, but the better the design gets, the more of their expertise it needs.

And unless something dramatic happens, AI is coming and is only going to be a bigger part of our lives in all spaces. Learning how to make better use of it now is more important than ever.

Sources

  • Newton, P. M. and Salvi, A. (2020). How common is belief in the learning styles neuromyth, and does it matter? A pragmatic systematic review. Frontiers in Education, 5, 602451. The 37 studies, 18 countries, 15,405 educators and the 89.1% weighted mean. The highest single figure was 97.6%, among pre-service teachers in Turkey.
  • Pashler, H., McDaniel, M., Rohrer, D. and Bjork, R. (2008). Learning styles: concepts and evidence. Psychological Science in the Public Interest, 9(3), 105–119. The meshing hypothesis, the crossover design it requires, the finding that very few studies used it, and the suggestion about why the belief survives. The Myers-Briggs comparison is mine, not theirs.
  • Sun, X., Norton, O. and Nancekivell, S. E. (2023). Beware the myth: learning styles affect parents', children's, and teachers' thinking about children's academic potential. npj Science of Learning, 8, 46. Three experiments in the United States with children aged six to twelve, parents and teachers. Children were asked about intelligence only; adults were also asked to predict grades by subject, and science showed no difference.
  • Kalyuga, S., Ayres, P., Chandler, P. and Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31.
  • Kulik, J. A. and Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: a meta-analytic review. Review of Educational Research, 86(1), 42–78. Fifty controlled evaluations, median effect 0.66 standard deviations, "from the 50th to the 75th percentile". The abstract says the size of the gain "depended to a great extent on whether improvement was measured on locally developed or standardized tests". The full paper is paywalled, so the exact figures by test type are not given here.
  • Khan, S. (2023). Harnessing GPT-4 so that all students benefit. Khan Academy blog, 14 March 2023. Khanmigo's question-first design. Khan Academy has also said the tutor is anchored on its own instructional and practice content and runs on OpenAI models. The point about open writing is my own reasoning; I have found no study comparing Khanmigo's results by subject.
  • Kestin, G., Miller, K., Klales, A., Milbourne, T. and Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. 194 students, randomised crossover over two lessons, comparison lessons led by instructors rated above the department average. Median post-test scores were 4.5 with the tutor against 3.5 in class, an effect of 0.63 standard deviations by linear regression, which the authors call an underestimate because of a ceiling effect (their quantile regression gives 0.73 to 1.3). Median time with the tutor was 49 minutes, and 70% of students spent under an hour, against a class learning time the authors assume to be 60 minutes rather than measure. Engagement and motivation were higher with the tutor; enjoyment did not differ.
  • Coe, R., Aloisi, C., Higgins, S. and Major, L. E. (2014). What makes great teaching? Review of the underpinning research. Sutton Trust. Rates pedagogical content knowledge as one of two factors with strong evidence of impact on student outcomes. See also Baumert, J. and colleagues (2010), Teachers' mathematical knowledge, cognitive activation in the classroom, and student progress. American Educational Research Journal, 47(1), 133–180, the German study that found this knowledge predicted pupils' gains in maths.