Different popular LLMs 2025 12 13

When students get an answer handed to them, something interesting happens in the brain. They start to feel like they understand the material, even when they don’t. This is the core problem at the heart of how we use answer keys and AI tools in education today. According to Nelson and Narens’ (1990) foundational framework on metacognition, our brain operates on two levels. This includes one that actually processes and stores information, and a “supervisor” level that watches and evaluates how well we’re learning. The problem is that these supervisors can be fooled pretty easily.

What usually fools them is something called the fluency heuristic. When information feels easy to process, the brain reads that ease as a sign of real learning. Bjork and Bjork describe this as an “illusion of competence,” and it shows up constantly in how students study. Someone reviews material that feels familiar, assumes they know it, and then struggles to recall it on a test days later. The feeling of understanding and the actual storage of information are two different things, but they’re hard to tell apart from the inside.

Traditional answer keys can feed into this problem, but they don’t have to. A study by Tomanek and colleagues looked at what happened when answer keys were redesigned to include reflection questions, asking students to explain their reasoning or connect concepts together. Students who used these enhanced answer keys reported better understanding of what they actually knew versus what they still needed to study. The reflection added a small delay between getting the answer and judging their own learning, which gave the metacognitive supervisor a chance to do its job.

AI assistance creates a different pattern. A 2024 study by Fernandes and colleagues had participants solve LSAT-style logic problems, some with ChatGPT and some without. The AI group scored about 3.5 points higher on average, but they also overestimated their performance by around 4 points. The gap between how well they did and how well they thought they did was actually larger than the improvement AI gave them in the first place. The study also found that the usual Dunning-Kruger pattern disappeared with AI use. Normally, lower performers overestimate their abilities while higher performers underestimate theirs, but with AI, students at every skill level became similarly overconfident.

The implication isn’t that AI shouldn’t be used in classrooms, but rather that  we’ve been evaluating educational tools by the wrong standard. Performance and metacognitive accuracy aren’t the same thing, and a tool that improves one while quietly damaging the other isn’t a clean win. The Tomanek research suggests reflection prompts can preserve metacognitive accuracy with traditional answer keys, and something similar might work with AI, like asking students to predict their score before checking, or to explain an AI-generated answer in their own words. The source of an answer isn’t a neutral variable. It shapes what students learn and what they think they’ve learned, and those two things need to be measured separately.

Leave a Reply

Previous Post
Next Post

Quote of the week

"People ask me what I do in the winter when there's no baseball. I'll tell you what I do. I stare out the window and wait for spring."

~ Rogers Hornsby

Designed with WordPress

Discover more from Phoenix Journal

Subscribe now to keep reading and get access to the full archive.

Continue reading