AI tools are (mostly) great for retrieval practice

Share

This is the third post in a brief series that contextualizes AI in education within prominent theories and phenomena of educational psychology. The previous posts discussed Flow Theory and Expectancy-Value Theory. Today’s focus is on the testing effect.

The testing effect – also known as retrieval practice – is a well-established phenomenon in educational and cognitive psychology demonstrating that people learn content better if they are “tested” on it frequently. “Testing” here doesn’t mean high-stakes, final-exam-type testing that school conditions us to associate with the term. Rather, it refers to any task that requires someone to retrieve information from long-term memory, including answering practice problems, taking brief quizzes, or teaching others. One of the most-common versions of this, particularly for those of us who grew up before computers and the internet were ubiquitous in schools, is studying with flash cards. I can still vividly remember making periodic table flash cards for chemistry in Mr. Stietzel's class…

But I digress! Taking these “tests” requires the brain to actively search for and retrieve information from long-term memory, which strengthens neural pathways and improves long-term retention of the content being tested on. Plenty of research indicates that this approach to studying is more effective any simply rereading material (see e.g. Roediger & Karpicke, 2006). A particularly effective approach is distributed practice – low-stakes, repeated, spaced testing (see e.g. Cepeda et al., 2006).

AI’s Potential Role

Ok, so, regularly taking brief, low-stakes practice “tests” is an effective way to learn content. It seems like this is something that AI would be great at facilitating. And it mostly is. Teachers (or students) can ask Gemini or ChatGPT or Claude to throw together a 5-question mini-quiz on factoring polynomials. Students can then either type their answers into AI and receive feedback, or the AI tool can pre-generate an answer key that students check their answers against. It seems like a straightforward, slam-dunk application of AI. And it’s one of the applications I encounter the most whenever I read about AI in education. It’s such an obvious use case that Gemini even has a built-in “Guided Learning” mode.

There are a few small concerns I’ll raise here, but overall this does feel like a pretty great way for teachers and students to use AI.

The first concern has to do with hallucinations. At this point, we’ve all heard the warnings about how AI models can hallucinate answers and we should be skeptical of their output because they might tell us it's safe to put metal in the microwave or whatever. Sure sure, fine. This largely isn’t true anymore, at least not for common content. Models are good enough now that if you ask them for a practice quiz on the periodic table, it’s going to be able to generate a quiz and a correct answer key. Maybe you should be wary of hallucinations if you’re asking about uber-niche topics or cutting-edge content, but that’s obviously not going to be the case for the vast majority of uses, especially in K-12 education.

The second concern has to do with asking students to develop their own practice tests. This requires some level of metacognition in that students will need to be aware of what they know and what they don’t know about a given topic to get the AI tool to produce a test that’s tailored to their needs. On the one hand, this is a potential strength, since using metacognitive strategies is another profoundly effective way to learn. On the other hand, students, and particularly younger students, aren't just automatically good at enacting metacognitive strategies. Just like anything else, successfully using them requires direct instruction and practice. If we don't teach students how to enact these strategies, then they probably won't do so well.

It’s worth mentioning that plenty of AI-powered ed-tech apps shortcut the whole student-practicing-metacognitive-strategies stuff by feeding students’ data and past responses into AI tools that will just generate differentiated practice problems automatically, sans metacognition.

The final concern is the most fundamental. The whole mechanism of the testing effect is predicated on a learner actively retrieving information from long-term memory and this retrieval strengthening neural pathways. The search-and-retrieval part is literally what makes it effective. If we give students a brief “test” and they use Gemini – rather than their own brains – to answer the questions in the test, then they’re not doing the search and retrieval. Which means they won't learn much. This feels very obvious to write, but it's also probably the biggest threat.

If you’re enjoying reading these posts, please consider subscribing to the newsletter by entering your email in the box below. It’s free, and you’ll get new posts to your email whenever they're published.

Read more