Test Intelligence Review

Explanatory Feedback Versus Answer-Only Feedback in AP Prep

Elaborated feedback explaining why answers are correct beats answer-only feedback for AP test prep.

Editor at Large · · 11 min read
Cover illustration for “Explanatory Feedback Versus Answer-Only Feedback in AP Prep”
Feedback Loops · September 26, 2026 · 11 min read · 2,537 words

More than 1.3 million students sat for over 4.8 million AP Exams in a single year. That number is climbing, up 7% from 2024 to 2025, with growth spread across most AP courses. This is a growing program straining under its own size, not a shrinking one looking for relevance. It is a growing one straining under its own size.

The strain appears in what happens after the exam. Many of those millions of tests include a free-response section, where a student has to build an argument instead of picking from four choices. The 2025 AP Reading, the largest grading event in the program's history, needed more than 31,000 educators to read and score over 20 million student responses. Twenty million individual pieces of writing, each one needing a human to judge whether the reasoning holds up.

And after all that grading, all that practice leading up to exam day, pass rates still vary a lot by subject. In some subjects, fewer than two out of three test takers manage to earn a 3 or better. A huge share of students are doing the work, sitting the exam, and still coming up short.

That gap between effort and outcome is the real story. Something breaks down between practicing a skill and actually learning it, and the most likely suspect is feedback: specifically, what kind of feedback students get while they practice, and whether it teaches them anything.

What the research means by "feedback" (and why the type matters enormously)

Not all feedback is the same animal. Researchers split it into three distinct types, and most AP practice tools quietly pick the weakest one by default.

Knowledge of Results (KR) tells a student right or wrong. Nothing else. It is the most minimal form of feedback a practice tool can deliver.

Knowledge of Correct Response (KCR) shows the correct answer, but never explains why it is correct or why the student's own answer missed.

Elaborated Feedback (EF) goes beyond the correct answer by including some form of explanation or guidance meant to support a student's understanding of the material.

These are not three levels of the same thing, stacked like bronze, silver, and gold. They ask the brain to do three different jobs. KR asks a student to notice they were wrong. KCR asks them to notice what right looks like. EF asks them to engage with an explanation aimed at improving their reasoning process.

Take a thesis statement on an extended written response. Elaborated feedback does not just say "incomplete thesis." It names which piece of the argument is missing, which piece of evidence should have been used, and where the reasoning stopped short of what the rubric wants.

Without that explanation, a predictable thing happens. Call it the correct answer trap. A student sees the right answer, copies it down, and moves on. No new thinking happens. That is especially damaging in problem-solving subjects, where the whole point of practice is the thinking behind the answer.

Meta-Analytic Evidence on Elaborated Feedback and Learning

Diagram: Why Feedback Type Is Doing Most of the Work. Visualizes: Visualize the three distinct feedback types and their relative impact on learning outcomes.

The numbers back this up, and the gap between feedback types is bigger than most people expect. Research synthesizing multiple studies, including work by Van der Kleij and colleagues, has found that elaborated feedback produces substantially larger effect sizes than simpler feedback types, with Knowledge of Results showing minimal impact and Knowledge of Correct Response falling in between.

Telling a student they got it wrong barely moves the needle. Showing them the right answer helps some. Explaining why moves the needle nearly ten times further than KR alone.

The advantage for elaborated feedback was largest for higher order thinking: analysis, synthesis, building an argument from evidence. That is exactly the kind of skill AP exams are built to test, and explanation matters here because analysis and argument-building require unpacking multiple decision points rather than a single right-or-wrong check. Recall does not need much unpacking; right or wrong tells a student most of what they need to know. Building an argument is a process with dozens of small decision points, and a student cannot fix a decision point they do not know exists.

The advantage for elaborated feedback appears particularly meaningful for students who are newer to the kind of reasoning AP courses demand. A huge share of AP students are meeting college-level analytical work for the first time. They are the ones who need the "why" the most, and they are the ones most likely to get only the "right answer."

Zoom out, and feedback overall, across all types lumped together, still produces a meaningful effect on learning in broader reviews. But that number hides a wide spread. Some feedback barely helps. Some helps enormously. The type is doing most of the work. Treating "feedback" as one interchangeable feature misses the point.

Complications in the research and their implications for AP practice design

Not every study lines up neatly behind elaborated feedback, and pretending otherwise would be dishonest.

Some research on digitally delivered feedback has found patterns that complicate the picture, with simpler feedback types outperforming more elaborate ones in certain contexts.

A 2025 study in Frontiers in Education, run by Lu, Yang, Wei, and Heffernan out of Brandeis and Worcester Polytechnic, tested this on semi-open-ended foreign language questions. Students who got correct-response feedback alone finished faster, learned more, and reported feeling more confident than students who got correct-response feedback plus elaboration. Adding explanation did not help. It just added more to read and more time spent reading it.

A study from the EDM conference, led by Fleischer and colleagues at Leibniz Universität Hannover, tracked 61 first-year students across more than 1,169 individual feedback messages. Correctness feedback helped every student's learning rate. But explanatory feedback describing the next step, delivered before a student had actually made an error, hurt learning rates. For students with weak prior knowledge, the explanation was too vague to latch onto. For students with strong prior knowledge, it told them something they already knew. Either way, it landed wrong.

A 2025 randomized controlled trial in a general biology course, with 1,002 students, found no meaningful difference between AI-generated elaborated feedback and a plain right-or-wrong control. What moved the needle was a separate ingredient built into the study: a requirement for structured self-reflection, students asked to think about their own thinking. Not the feedback content. The reflection habit.

Putting these together, the lesson for anyone building an AP practice tool is not "add more explanation."" Timing matters. Whether the student has actually struggled with the problem first matters. Explanation dumped on someone who has not wrestled with the question yet, or piled onto a task where the answer alone already made things clear, can do nothing or actively backfire. Elaborated feedback is a tool that has to be aimed, not a feature to bolt on. It is a tool that has to be aimed.

The interaction of feedback timing, student metacognition, and explanation quality

Timing predicts learning outcomes more reliably than how many practice problems a student grinds through. Feedback delivered before a student has genuinely wrestled with a problem tends to fall flat, no matter how well the explanation is written.

Structured reflection, prompting a student to think about how they arrived at an answer, may matter as much as the explanation itself. Research on structured self-reflection suggests that explanation alone does little if a student does not actively engage with it.

Two long-standing frameworks explain why. Hattie and Timperley describe effective feedback as answering three questions: where is the learner going, how is the learner doing, and where does the learner go next. That feedback is supposed to work across four levels: the task itself, the process behind it, the student's own self-regulation, and the student as a person. Most answer-only feedback only ever touches the first level. It says whether the specific answer was right. It says nothing about the process that produced it and nothing about how to regulate that process next time.

Sadler's framework adds a second angle. For feedback to help, a learner needs a clear picture of what the target standard looks like, an honest comparison between their own work and that standard, and a real plan for closing the gap. Simple right-or-wrong feedback does little to support these demands. It confirms an outcome. It confirms an outcome but builds nothing.

The capabilities and limits of AI-generated feedback compared to human feedback

A meta-analysis in Educational Psychology, covering 41 studies and 4,813 students, found no statistically significant difference in learning outcomes between students given AI-generated feedback and students given feedback from a human. The pooled effect size was small and did not clear the bar for significance. On average, well-built AI feedback and human feedback land in roughly the same place.

Averages hide the interesting cases, though. The FeedbackWriter study, a randomized controlled trial set for CHI '26 in Barcelona, tracked 354 students across 1,366 graded essays. Teaching assistants reviewed AI-generated feedback suggestions before sending them, sometimes adopting them as-is, sometimes editing them. Students who received this AI-mediated, human-reviewed feedback produced meaningfully stronger revisions: a Cohen's d of 0.50, roughly landing at the 70th percentile instead of the 50th. The more suggestions the TAs adopted from the AI, the bigger the gain.

That AI-mediated feedback also proved more actionable and did more to build independent learning skills than feedback written entirely by hand. Actionable and independence-building line up almost exactly with what AP free-response prep needs. A student writing an extended written response needs to know what to fix, and needs to learn to catch the same mistake unprompted next time.

Fully automated AI feedback, with no human reviewing it before it reaches the student, is a different story. Research on these systems documents real accuracy problems: in some systems, a meaningful share of AI-generated hints have been flagged and rejected for low quality. When the AI gets something wrong, that error does not just fail to help. It actively hurts learning. Unreviewed AI feedback carries real risk, and any system worth trusting needs some mechanism, human or otherwise, for keeping errors out before they reach a student.

The tools AP students and teachers use for explanatory feedback on practice work

A handful of tools have built themselves specifically around closing this gap instead of settling for right-or-wrong grading.

Class Companion gives students instant, rubric-aligned feedback that explains mistakes, offers concrete suggestions, and calls out what a student did well, measured against actual AP standards. Teachers using it report it saves significant grading time and, just as important, produces feedback specific enough that students actually read it instead of skipping straight to the score.

Khanmigo sits on top of a free content library already covering a lot of AP material, for $4 a month or $44 a year. It is built to hold back the answer and guide a student toward it through scaffolded questions and hints, support that spans the whole writing process while leaving teachers in control of what students see. New AI tools built on Google's Gemini models moved out of pilot testing and into classrooms for the 2026 school year, after roughly six months of integration work, adding real-time adaptive diagrams and practice material teachers can customize.

DeAP Learning builds its AI tutors to mirror well-known subject-specific AP educators (names like Heimler's History, Mr. Sinn, and Jacob Clifford) across more than 17 AP subjects. It grades short-answer questions, DBQs, LEQs, and FRQs instantly, using the same rubrics College Board scorers rely on, turning what normally takes days of waiting into feedback delivered in seconds.

AP Exam Tutor, from Jenova AI, grades free-response answers against the rubric, generates adaptive practice questions targeting a student's weak spots, and builds study plans tied to an actual exam date. Students using it have improved FRQ scores by an average of 1.5 rubric points across subjects including AP Chemistry and AP US History. That number matters because of where AP points actually get lost. Most students are not losing points because they lack content knowledge. They lose points on structure: a missing thesis element, evidence with no commentary explaining its relevance, a three-part prompt answered in two parts.

The digital SAT's feedback gap and the growing importance of explanatory tools

The SAT went fully digital in 2024, and it now runs on an adaptive format: two modules per section, where performance on the first module decides how hard the second module gets. That adaptivity personalizes the test experience for every student. It also opens a blind spot around feedback that did not exist with the old paper test.

Because each student ends up seeing a partly unique set of questions, and because those questions get reused across future administrations, the full test can no longer be released afterward the way it once was. What students get back instead is broad skill-band feedback, categories like "Expression of Ideas" or "Advanced Math," with no visibility into which specific questions were missed or why.

A skill band tells a student where the weakness lives in general terms, and that vagueness is a real problem. A skill band tells a student where the weakness lives in general terms. It says nothing about the actual failure. Did the student misread the question? Make an arithmetic slip? Fundamentally misunderstand the underlying concept? Those are three different problems needing three different fixes, and a skill band, which is really just Knowledge of Results wearing a category label, cannot tell them apart.

That gap raises the value of practice tools that adapt too. A static question bank, set at one fixed difficulty, cannot recreate what the digital SAT does, and cannot target the specific gaps that format is designed to expose.

What knowledge gap detection adds beyond single-question feedback

Elaborated feedback on a single question fixes the error sitting right in front of the student. It does not answer the bigger question: is this a one-off slip, or a sign of a deeper hole in what the student actually understands?

Newer approaches try to answer that by building curriculum-aligned knowledge graphs, maps of how concepts depend on each other, so a system can trace a wrong answer back to the specific upstream idea that is actually broken. That is a real shift in approach. Instead of reacting to each mistake as it happens, the system tries to spot which foundational concept is causing a string of downstream errors before they pile up.

The Fleischer study fits directly into this. Students with weak prior knowledge benefited from explanatory feedback aimed at heading off a misunderstanding before it took root. Simply showing them the correct answer, with no explanation, produced no learning gain at all for that group. Knowing exactly where a student's understanding breaks down has to come before feedback can actually fix it.

Apply that to AP prep directly. A student who keeps losing points on LEQ evidence commentary might have a problem that has nothing to do with formatting. The student might not actually understand what corroboration means, how one piece of evidence is supposed to support or challenge another. Feedback that treats this as a formatting fix, some version of "add more analysis here," will keep missing the real issue, no matter how many times it repeats the same instruction.

Sources

  1. Does Student Learning Rate Depend on Feedback Type and Prior Knowledge?
  2. Frontiers | Students
  3. AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course
  4. Assessing the Impact and Underlying Pathways of Sequenced AI feedback on Student Learning
  5. Detecting Knowledge Gaps from Conversational AI Interactions Using Curriculum Prerequisite Graphs
  6. LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
  7. sciencedirect.com
  8. jenova.ai
Filed underFeedback Loops

More in Feedback Loops