A meta-analysis published in The Clinical Teacher in 2026 pooled 14 studies and 21,415 medical learners and found that those who studied with spaced repetition outperformed those using standard study techniques on objective tests by a standardized mean difference of 0.78, with a 95% confidence interval of 0.56 to 0.99 and a p-value below 0.0001. That is a large gap by the usual conventions of education research, and the interval never comes near zero. If the question on your mind is whether the flashcard habit is worth the hours it eats, this evidence says yes.
What the paper does not do is tell you when to review. It compared spacing against not spacing. It did not run three-day intervals against seven-day intervals, or morning sessions against evening ones, or one app against another. Read the number as a verdict on the principle, not as an instruction for Tuesday.
This page is for people memorizing high-volume factual material on a deadline: gross anatomy, pharmacology, the science sections of the MCAT, board review. It covers what Maye and Hurley's review actually measured, where it stops, and how to turn a pooled effect size into a week you can run. The mistake most people make with a finding like this is to treat the effect size as a schedule. The schedule has to come from somewhere else, and it is worth being honest about where.
What a 0.78 effect size is saying
A standardized mean difference is a way of averaging results from studies that used different tests. One study scores out of 40 on an anatomy quiz, another out of 100 on a pharmacology exam. You cannot add those raw scores together, so each study's gap between groups gets converted into standard deviations, a measure of how spread out the scores were. A standardized mean difference of 0.78 means the spaced group sat roughly three quarters of a standard deviation above the comparison group.
If scores in the comparison group are spread in a normal curve, that puts the typical spaced learner near the 78th percentile of the group that studied the standard way. That translation assumes the curve behaves, which real exam data does not always do, but it gives you the scale of the thing. It is not a rounding error.
The confidence interval matters as much as the point estimate. The 95% interval of 0.56 to 0.99 is the range of true effects compatible with the data. Even the pessimistic end, 0.56, is a substantial difference. Nothing in the interval suggests spacing is a wash.
A few method details are worth knowing, because they define what the number can bear. The review followed PRISMA reporting guidelines, ran its electronic database searches in February 2025, and assessed the quality of included studies with the Medical Education Research Study Quality Instrument, a scoring tool built specifically for medical education research. Effect sizes, heterogeneity and publication bias were calculated in R. The paper is indexed by the National Library of Medicine under PMID 41601436, which confirms its peer-reviewed status and its placement in Volume 23, Issue 2, article e70353.
Two honest gaps. First, the February 2025 search date means anything published after that is not in the pool. Second, heterogeneity and publication bias were computed, but the values are not in the bibliographic record, so a reader working from the indexed abstract cannot see how much the individual studies disagreed with each other. That is the figure you would want before treating 0.78 as a promise about any single classroom.
A second 2026 systematic review, indexed as PMID 42468294, restricted itself to randomized controlled trials and looked at health professions education more broadly, including nursing and allied health. It examined three outcomes: knowledge acquisition, knowledge retention, and learning-related anxiety. It found spaced learning influenced all three positively, and still described the method as a "promising educational strategy" rather than a settled one. When two reviews published in the same year land on the same direction with different inclusion rules, the direction is worth trusting. The exact size is not.
The pooled review compared spaced repetition against standard study techniques on objective test performance. It did not, in the record available, break its results down by how the spacing was delivered. That leaves the reader with a choice the evidence does not make for them.
| Approach | App-scheduled flashcards | Fixed-calendar review blocks | Spaced digital modules or emails |
|---|---|---|---|
| How the interval gets set | The algorithm decides: well-known cards are shown less often, struggled cards more often (Dedicated Prep, USMLE Anki Guide 2026) | You decide, and the rule of thumb is to leave enough time for some forgetting to occur (Lecturio, updated February 5, 2026) | The provider's platform pushes material back on a preset spacing (JMIR spaced digital education review, 2024) |
| Daily load suggested by practitioner guides | 100 to 150 cards per day as a USMLE baseline (Dedicated Prep, 2026); 15 to 20 minutes per day of pure review to hold existing cards (iatrox, 2026) | Roughly 1 to 2 hours every other day rather than one long session (Lecturio, February 2026) | Whatever the module sends; no per-learner volume figure appears in the review (JMIR, 2024) |
| Format-specific meta-analytic support | Not broken out by delivery format in the indexed record of the 2026 pooled review (Maye and Hurley, The Clinical Teacher, 2026) | Not broken out by delivery format in the indexed record of the 2026 pooled review (Maye and Hurley, 2026) | Spaced online education beat massed online education on post-intervention knowledge, and beat both massed education and no intervention on clinical behavior change (JMIR, 2024) |
| What it does not do | Does not teach novel concepts (PMC10842980) and does not train the clinical reasoning that single-best-answer questions test (iatrox, 2026) | Same limits, plus nothing records which specific items you are forgetting unless you keep the log yourself | Optimal design and delivery remain unsettled, and the authors call for further research on both (JMIR, 2024) |
| Setup work before day one | Pick a deck, then set learning steps, graduating interval and maximum interval (iatrox, 2026) | A calendar and a topic list | None on the learner's side; the provider schedules it (JMIR, 2024) |
Building the week, step by step
The numbers below come from practitioner guides written for medical exam prep, not from the meta-analysis. That distinction is the whole point of this section. The research supports spacing; it does not supply a card count. Where a figure has a source and a date, it is a recommendation someone published, and you should treat it as a starting position to adjust, not a finding.
-
A fixed exam date, or at least a horizon
Every scheduling decision below depends on how far out you are. Without a date you cannot cap intervals sensibly.
-
Material you have already understood once
Spaced repetition consolidates what you have met before. It is not a first encounter with a concept.
-
One card source, chosen and closed
A faculty deck, a third-party deck, or your own cards. Switching sources mid-semester resets your intervals.
-
Two protected slots in your day
Dedicated Prep's 2026 USMLE guide recommends a fixed morning session and a fixed evening session rather than binge sittings.
-
A question bank or problem set
Both practitioner guides insist cards alone are not a study plan. You need something that tests application.
-
A place to record what you got wrong
An error log turns a missed card into a topic to revisit, which the algorithm alone will not do for you.
1. Cap the maximum interval to your horizon. Anki's default maximum interval is 36,500 days, a setting built for lifelong learning. The 2026 iatrox guide to medical exams recommends capping it at 180 days for exam preparation, so that nothing gets scheduled past the point where it still has to be in your head. It worked when no card in your deck has a due date beyond the exam. Cost: about two minutes in the deck options.
2. Learn the concept before you card it. The PMC article arguing for spaced repetition in medical curricula is otherwise enthusiastic about the method, and it still draws one hard line: "While spaced repetition is highly effective for consolidating information, it is not appropriate for learning novel concepts, which typically requires more in-depth discussion." Dedicated Prep puts the same rule in study terms: comprehension before memorization. It worked when you can explain the mechanism in a sentence before you ever see the card. Cost: this is your lecture and reading time, not extra time.
3. Set the new-card number by your distance to the exam. The iatrox guide suggests 20 to 40 new cards per day when you are more than three months out, rising to 40 to 60 per day in the final weeks. Dedicated Prep's baseline of 100 to 150 cards per day for USMLE prep counts new cards and reviews together, adjusted up or down based on your backlog and retention rate. It worked when your review queue empties most days without dread. Cost: this is the number that sets everything else.
4. Configure the front end of the schedule. The iatrox settings are specific: learning steps of 1 minute then 10 minutes before a card graduates, a graduating interval of 1 day, an easy interval of 4 days, and the review cap set to unlimited (9999) so that finished reviews are never artificially held back. If you are using the bundled scheduler, note that Anki shipped a new FSRS version alongside two security patches in September 2026, which changes how those intervals get calculated. It worked when a card you failed comes back within the same session and a card you nailed disappears for days. Cost: five minutes, once.
5. Anchor two sessions a day. Dedicated Prep recommends a fixed morning session for activation and a fixed evening session for reinforcement, and argues that consistency beats intensity. The iatrox guide puts pure review maintenance at 15 to 20 minutes a day, separate from new-card work. It worked when the sessions happen without a decision being made about them. Cost: 15 to 20 minutes daily for maintenance, plus whatever your new cards demand.
6. Put the non-card work in every-other-day blocks. Lecturio's medical education team recommends distributing study as roughly 1 to 2 hours every other day across a week rather than one long session. That is where the material that does not fit on a card goes: pathways, mechanisms, worked problems, the anatomy you need to draw rather than name. The gap between blocks is deliberate. Lecturio's authors write that the interval "should be long enough for some forgetting (and therefore difficulty) to occur," which is the principle researchers call desirable difficulty: recall that costs you something sticks better than recall that comes free. It worked when the start of each block feels slightly harder than you expected. Cost: 1 to 2 hours, three or four days a week.
7. Add questions that test reasoning, not recall. "Anki alone does not prepare you for medical exams. SBA questions test clinical reasoning, not recall," the iatrox guide states, referring to single-best-answer questions, the format used across UK and US medical exams. Dedicated Prep makes the same point about Step 2 CK and Step 3, where spaced repetition supports but does not replace clinical reasoning. It worked when you are missing questions for reasoning errors rather than blank recall. Cost: one question block on the days you are not doing a long study block.
8. Grade honestly. The whole system depends on you distinguishing true recall from recognition, meaning the difference between producing the answer and merely nodding at it when you see it. If you press Good because the answer looked familiar, the algorithm pushes the card out and you have quietly deleted it from your study plan. It worked when your retention rate drops for a week after you start grading strictly, then recovers. Cost: nothing but discomfort.
9. Audit once a week. Look at three things: whether the review queue grew across the week, what your retention rate did, and how many cards you made instead of studied. Adjust the new-card number first, because it is the only lever that controls the other two. It worked when next week's queue is flat or shrinking. Cost: ten minutes on a Sunday.
Where these schedules break
The queue outruns you. The iatrox guide warns directly that adding too many new cards daily creates unsustainable review backlogs. You can recognize it by the shape of the week: Monday's queue is 180 cards, Friday's is 340, and you have started skipping days. The fix is to cut new cards to zero for several days and clear the backlog before adding anything, then restart at a lower number. Do not fix it by suspending cards at random, which destroys the schedule for the material you have already paid for.
The default maximum interval is still set. If you never changed it, cards are being scheduled up to 36,500 days out. Symptom: a card you found genuinely hard three weeks ago has a due date in 2029. Cap it at 180 days for exam prep and the problem disappears for the cards scheduled afterward.
Card-making becomes the work. The same guide flags excessive custom card creation as a form of procrastination. It feels productive because it involves the material, and it produces no retrieval at all. Recognize it by the ratio: if you spent 90 minutes on your deck and reviewed 40 cards, you were writing, not studying. A premade faculty or third-party deck removes the temptation entirely.
Cards meet the concept first. This is the failure the PMC authors warn about. Symptom: you have been reviewing a card for two weeks, you can produce the answer, and you cannot say why it is the answer. The fix is not more reviews. It is going back to the lecture or the chapter, then letting the card consolidate something you understand.
The deck is treated as the plan. Dedicated Prep's guide states plainly that Anki is not a standalone strategy and must be paired with practice questions and conceptual understanding. Symptom: high card retention, mediocre question-bank scores. The gap between those two numbers is the size of your reasoning problem, and no flashcard schedule closes it.
Reviews get batched. Saving five days of reviews for Saturday converts a spaced schedule into a massed one, which is the comparison condition the 2026 review found losing by 0.78 of a standard deviation. It is the same reason cramming the week before a test tends to disappoint even when the hours are real. Lecturio's recommendation to distribute study across the week rather than concentrate it is the countermeasure, and it costs nothing but scheduling.
Cases where the answer changes
You are more than three months out versus inside the final weeks. The volume guidance moves: 20 to 40 new cards a day at distance, 40 to 60 in the last stretch, per the 2026 iatrox guide. The intervals should compress toward the end too, which is what capping the maximum interval accomplishes.
You want the knowledge after the exam. The 180-day cap exists because board prep is time-boxed. If you are learning a subject to keep, the default long intervals are appropriate and the cap is counterproductive. Decide which you are doing before you touch the settings, because the two goals want opposite configurations.
The exam tests reasoning more than recall. For single-best-answer and clinical-vignette formats, both practitioner guides treat spaced repetition as supporting infrastructure rather than the method. The card schedule stays; the proportion of your week spent on questions goes up.
You are outside medicine. The 2026 pooled review sampled medical learners, and the companion randomized-trial review covered health professions education broadly. Neither speaks directly to law, languages or undergraduate humanities. The principle has evidence elsewhere; this particular effect size does not transfer, and saying otherwise would be borrowing authority the study did not lend.
You are already practicing. The 2024 JMIR review of spaced digital education for health professionals found benefits for knowledge, skills, confidence and clinical behavior change among both students and practicing clinicians, and its authors note an explicit gap in evidence on long-term and continuing medical education impacts. For CME, the direction is supported and the durability is not yet established.
Test anxiety is your actual bottleneck. The 2026 randomized-trial review is the only one of these sources that measured learning-related anxiety, and it found spaced learning influenced it positively. That is one review, and anxiety is a less studied outcome in this literature than test scores, so treat it as a reason to try rather than a reason to expect.
The time cost, and what the evidence buys
In time, the floor is modest: 15 to 20 minutes of daily review to maintain a deck, plus the new-card load, plus 1 to 2 hours every other day on the conceptual work per Lecturio's February 2026 guidance. At the upper practitioner recommendation of 100 to 150 cards a day, the card work alone becomes a real block of the morning and evening rather than a filler task.
In money, the software side can be nothing. The costs that show up are decks, question banks and courses, and none of the sources here price those, so no figure belongs in this section.
What the evidence buys is confidence in the bet, not precision in the execution. A pooled difference of 0.78 standard deviations across 21,415 learners is about as strong a signal as medical education research produces, and the 0.56 floor of the confidence interval means even a conservative reading leaves a substantial gap. What it does not buy is an interval. As Lecturio's authors put it, "Scientists have yet to clearly define the time periods for interstudy space, but it should be long enough for some forgetting (and therefore difficulty) to occur." The JMIR reviewers said the same thing from the research side, calling for further work on optimal design and delivery.
So the honest summary of the state of play, as of this review's February 2025 search cutoff: spacing beats not spacing by a lot, in medical education specifically, on objective tests within the training period. The best interval, the best delivery format, and how much survives years later are open. What would settle them is a trial that randomizes learners to different interval lengths with the same content and follows them past graduation. Until someone runs it, the schedule above is a reasonable, sourced default rather than a proven optimum, and the number you should tune first is how many new cards you add tomorrow.
Comments
No comments yet. Be the first to comment!
Leave a Comment