Test Intelligence Review
FeaturesLong read

Why Mixing Problem Types Beats Topic Drills in AP Prep

Interleaved practice builds the recognition skill blocked drills never train.

Correspondent · · 10 min read
Cover illustration for “Why Mixing Problem Types Beats Topic Drills in AP Prep”
Features · September 15, 2026 · 10 min read · 2,262 words

Blocked practice feels like studying. It looks like progress. Solve twenty derivative problems in a row, get nineteen right, and it seems like calculus is clicking. But on the AP exam, no problem tells you what kind of problem it is. Nobody labels question 14 "related rates" before you read it. You have to figure that out yourself, on the spot, with no warning. That's the skill blocked practice never trains, and it's the exact skill interleaving builds.

The distinction sounds small. It isn't. A group of researchers at a public university studying learning put it bluntly: with blocked practice, students can solve word problems without reading the words. The surrounding context (the fact that every problem in this set is a "convert to fractions" problem) hands them the method before they've read a single sentence. That's not a knock on students or teachers. It's just what happens when every problem in a row calls for the same strategy. The fix isn't more effort. It's a different structure.

What the AP exam actually asks students to do at the moment of each question

Every AP exam mixes problem types in a sequence with zero labels. No headers. No "this section is derivatives." Students have to do two things at once: figure out what kind of problem is sitting in front of them, then apply the right strategy to solve it.

The exam demands a two-step skill: choose the right strategy, then execute it. Blocked practice only ever trains step two. Students walk in already knowing the strategy because the drill set told them. The exam never does that favor.

The shift to digital testing has made this more obvious than ever. As of 2025, 28 of 36 AP subjects with end-of-course exams run on Bluebook, where students scroll through a mixed sequence of question formats with no topic headers in sight. A student who spent months drilling unit by unit has never actually practiced the context-switching the exam demands at every single question.

The scale here isn't small, either. More than 1.3 million students in the class of 2025 sat for over 4.8 million AP Exams in public high schools nationwide, up from 4.3 million the year before. And the results show where prep is falling short: AP Statistics had only 60.3% of students score a 3 or higher in 2025; AP Computer Science Principles landed at 61.9%. Those aren't small subjects with weird edge cases. Those are mainstream courses with a real gap between what students studied and what the test asked them to do.

So the question becomes concrete: what kind of practice trains both steps, recognition and execution, at the same time?

How interleaving trains the missing skill

Interleaved practice mixes problem types within the same assignment, so a student never knows what's coming next. That uncertainty is the entire point. Because the problem type isn't given away by context, students have to read each question, figure out its category, pull up the right strategy, and then execute it. That's the full two-step sequence the exam actually demands, done for every problem instead of skipped for two-thirds of them.

Rohrer and colleagues describe two benefits working together here. Putting different problem types side by side forces strategy selection. Spacing out same-type problems across sessions strengthens retention of each individual strategy. Two separate mechanisms, one practice structure.

This lines up with something cognitive psychology has studied for decades: the contextual interference effect. Mixed, unpredictable practice produces better long-term retention than sorted, predictable practice. The difficulty isn't a side effect to tolerate. It's the mechanism doing the work.

That's why researchers call this kind of interleaving a "desirable difficulty." It makes students perform worse in the moment while they're practicing, and better months later when it counts. The struggle during practice isn't a warning sign. It's the thing building durable memory.

And what matters most for exam day specifically is transfer. Students who train interleaved don't care what order the real exam throws problems at them, because they've never relied on order as a clue. Students who train blocked are stuck hoping the exam matches their practice conditions. Research from Schorn and Knowlton at UCLA backs this up directly: sequence-dependence is a vulnerability blocked practice builds in, whether anyone intends it or not.

What the research actually shows about interleaving's effect on test scores

Diagram: Blocked vs. Interleaved: How the Score Gap Grows Over Time. Visualizes: Show how the performance advantage of interleaved practice over blocked practice widens as time passes before testing.

Rohrer, Dedrick, and Stershic ran a classroom study with 126 seventh-grade students, using identical practice problems over three months, split into interleaved and blocked groups. At a one-day delay, the interleaved group already outperformed the blocked group (Cohen's d = 0.42, a solid effect). At a 30-day delay, that gap nearly doubled (d = 0.79).

Read that again: the longer the wait before testing, the bigger interleaving's advantage got. For graph problems specifically, the interleaved group scored 89% versus 73% for the blocked group at the one-day mark. Now stretch that timeline out to the length of actual AP prep, which runs for months before a single exam day in May. The 30-day finding isn't a curiosity. It's the exact scenario every AP student is living through.

The pattern appears outside seventh-grade math, too. A university physics study published in npj Science of Learning in 2021 tracked interleaved homework over eight weeks and found median improvements of 50% on a first surprise test and 125% on a second, using novel, harder problems the students hadn't drilled directly.

One science-concepts retention study found something almost counterintuitive: the blocked group actually outperformed the interleaved group during practice. But a week later, the interleaved group came out ahead, with an effect size of 1.34, the largest number in the whole set of findings. That reversal is the pattern to hold onto. Blocked practice looks better while it's happening. That's exactly why students and teachers keep choosing it. But the advantage evaporates, or flips entirely, once real testing conditions (a delay, a mixed sequence) occur. A broader research summary backs this up across the board: interleaved math practice produces better test scores across lab experiments and classroom studies, across different ages, topics, and delay lengths.

Why students resist interleaving, and why that resistance is itself a signal

Students hate interleaving, and they're wrong to hate it, and the reason they hate it is actually useful information.

A survey of 233 seventh-grade math students found they rated interleaved practice as less effective, less preferable, more time-consuming, and more difficult than blocked practice. Every one of those judgments is a misread of their own learning.

Why does this happen? Blocked drilling produces smooth, fast, confident retrieval. That smoothness feels like mastery. But smooth retrieval during practice is often just the absence of struggle, and struggle is what encodes a strategy so it survives 30 days. Fluency in the moment is a bad predictor of memory later. It's just the only signal students have access to while they're sitting there doing the work, so they trust it.

This isn't a discipline problem or a motivation problem. Distributed and interleaved practice are both largely unfamiliar territory for most students, and both get judged with mixed, often skeptical, reviews, according to research on how students perceive their own study strategies. The discomfort isn't a bug. It's contextual interference doing its job.

Look at how students talk about this after the fact. On forums, students who drilled one question type at a time, session after session, describe hitting a wall on full practice tests where "the sheer variety of question types" throws them off completely. They didn't fail because they didn't know the material. They trained for a context the real test never gives them.

So if interleaving feels harder and less confidence-inspiring while doing it, that's not a reason to bail back to blocked sets. That discomfort is the expected experience of a method that actually works.

When blocked practice still makes sense, and when to switch

None of this means blocked practice is useless. It has a real job, just not the whole job.

For a genuine beginner, someone with zero prior exposure to a concept, interleaving too early adds load before there's anything to retrieve. Mixing in strategy A when a student hasn't even learned strategy A yet doesn't build judgment. It just builds confusion. Some initial blocked exposure lays the foundation that interleaving later exercises.

That points to a staged approach: get oriented with blocked practice, then switch to interleaving once there's baseline familiarity. Practitioner guidance backs this staged model, and a 2025 Language Learning study found that a hybrid approach (blocked first, then interleaved) may outperform either method used alone on delayed tests.

There's a refinement worth noting, too. A 2025 ScienceDirect study on adaptive interleaving suggests working memory capacity might shape how much mixing a student can actually handle productively. Adaptive sequencing, where frequently confused categories appear close together rather than randomly scattered, may reduce overload for students who struggle most. Interleaving isn't one-size-fits-all chaos. It can be tuned.

One more exception: the day of the exam itself. Evidence suggests blocking may actually help on exam day, since the goal at that point is fluency with material already learned, not new encoding. But that only applies after the real prep work is done. It's a last-minute exception, not a strategy.

This is also where the standard tutoring-industry advice (focus on weak areas, drill in short timed blocks, repeat until fluent) actually fits. It's a fair description of early-stage work. The problem occurs when it's the entire plan instead of the first third of it.

A phased structure makes the roles clear:

Phase 1: Re-learn forgotten units, build a base of knowledge. Blocked practice is appropriate here. Phase 2: Targeted training with unit-specific multiple-choice and free-response questions, used to find weak spots. Interleaved sets start entering the mix. Phase 3: Build stamina with timed, full-length simulations. Interleaving is the main method here, because this phase is meant to mirror exam conditions directly.

A diagnostic test, taken after covering all units, is the bridge between these phases. It shows which topics need heavy review and which need a light touch. Its job is to shape what goes into the interleaved mix, not to send a student back into pure topic blocking.

How the digital SAT's adaptive format makes the same demand interleaving trains for

The digital SAT runs on the same logic, with an extra layer of complexity stacked on top. Each section (Reading and Writing, Math) splits into two separately timed modules. Module 1 mixes easy, medium, and hard questions across topics in no predictable order.

The digital format is no longer a niche option for SAT takers. It has become the standard experience.

The adaptive routing raises the stakes on getting Module 1 right. Strong performance there routes a student into a harder Module 2, which is the only path to the top of the score range for that section. A student who trained purely by topic has never once practiced the mixed, unlabeled question environment that Module 1 actually is, and that module is the gatekeeper for the score ceiling.

The digital SAT adds something AP doesn't: not just unlabeled problem types, but a difficulty level that shifts based on how the student is already performing. That demands flexible retrieval under conditions that keep moving, which is precisely what interleaved practice trains for. A student who only ever practiced topic-by-topic might genuinely know the material and still underperform on Module 1, not from a knowledge gap, but from never having trained for context-switching under a countdown clock.

What a practical interleaved practice session looks like for an AP or SAT student

What does this look like on an actual worksheet, on an actual Tuesday night?

Rohrer and colleagues' original model gives a clean template: after a lesson on proportions, an assignment might include four new proportion problems, plus one problem each from eight different skills learned earlier. New material gets introduced. Older material gets spaced back in, on purpose, so it doesn't fade.

Apply that structure across subjects:

AP Calculus: mix derivative identification, integral setup, series convergence tests, and related rates in a single session, not because the student is weak in all four, but because the real exam will present them in exactly that jumbled order. AP History or English: alternate between document-analysis questions, long-essay prompts, and multiple-choice inference questions within one sitting, instead of finishing every document-based question before touching a single long-essay question. Digital SAT Math: deliberately mix linear equations, quadratic applications, data analysis, and geometry without grouping by type, matching how Module 1 is actually built.

The starting point is still a diagnostic. Once it flags the worst gaps, blocked review is the right tool to patch those specific holes first. After that, restructure sessions to fold those repaired skills back in alongside stronger ones, mixed together. The diagnostic decides what goes into the interleaved set. It doesn't decide whether to interleave at all.

Some practice tools now use AI to sequence problems based on what a student has already worked through, deciding which skills need to be spaced back in and which need first-time exposure, which is really just the Rohrer et al. model running at scale instead of on a printed worksheet. Passionfruit's approach fits this pattern directly: tracking not just which answers were wrong, but where a student's reasoning actually breaks down, then sequencing practice around that.

The test for any tool claiming to "interleave" is simple. Does it genuinely mix problem types across material a student has already seen? Or is it just another blocked set, sorted differently, wearing a new label?

Sources

  1. files.eric.ed.gov
  2. Students’ Perceptions of Effective Math Learning Strategies
  3. The role of executive function abilities in interleaved vs. blocked learning of science concepts
  4. researchgate.net
  5. researchgate.net
  6. sciencedirect.com
  7. onlinelibrary.wiley.com

More in Features