BlogWhy reviewing flashcards too early wastes the reviewMemory

Why reviewing flashcards too early wastes the review

In short

  • Reviewing at 98% recall buys 4.2 days of extra durability.
  • Waiting until 90% buys 19.9 days — from the identical review.
  • The gain comes from difficulty, not from effort or repetition count.
  • This is why re-reading a deck the next day feels productive and is not.

There is a specific kind of studying that feels excellent and does almost nothing. You finish a lecture, you re-read the deck that evening, you re-read it again the next morning, and each pass is smooth — you recognise everything, nothing trips you up, and you close the laptop with the sense of a job done.

That smoothness is the problem. It is the sensation of reviewing something you had not yet begun to forget.

We wanted to know the size of the penalty, so we isolated it.

The setup

The previous experiments in this series simulated whole cohorts of concepts over months. This one does the opposite: a single card, held fixed, reviewed once.

We took a concept with a stability of 10 days and a mid-range difficulty, then asked what happens if the review lands at different points on its decay curve. For each point we recorded the new stability if recall succeeds, the new stability if it fails, and the expected value across both weighted by how likely each outcome is at that moment.

Holding the card constant is what makes the comparison clean. Same concept, same student, same amount of effort. The only thing that varies is when.

Bar chart showing days of durability bought by a single review, plotted against the recall probability at the moment of review. Reviewing at 98% buys about four days; reviewing at 60% buys over sixty.
Expected durability gained from one review, by recall probability at the time of review. Starting stability 10 days, difficulty 5.

What one review buys

Reviewed when recall is You waited Durability gained
98% 1.8 days 4.2 days
95% 4.6 days 10.2 days
90% 10.0 days 19.9 days
85% 16.4 days 28.9 days
80% 24.0 days 37.2 days
70% 44.4 days 51.6 days

The review done at 98% recall — the next-morning re-read — adds 4.2 days of durability. The identical review, delayed until recall has drifted to 90%, adds 19.9 days.

Nearly five times the return, for the same work. The only difference is that you waited eight more days before doing it.

Where the gain comes from

FSRS models the strengthening effect as proportional to how much retrieval difficulty was actually overcome. If recall probability was 98%, retrieving the item was nearly free, so there was nothing to strengthen: the memory was not under any strain. If it was 90%, the retrieval was genuinely effortful and the model rewards it accordingly.

This is the desirable difficulty effect that cognitive psychology has described for decades, made numerically explicit. Struggle is the active ingredient. Reviews that feel easy are, by construction, reviews that did very little.

Fluency is not a sign that the review worked. It is a sign that the review was not needed.

It also explains why cramming underperforms per unit of effort despite involving enormous effort. A crammer’s reviews sit a day apart, at 95–99% recall, deep in the low-return zone. They did seven passes over every concept and got 16.9 days of stability. The spaced student did 4.2 passes, timed further out, and got 55.4.

So why not wait forever?

The table keeps climbing as you wait, which raises the obvious question. We ran it further out and the expected gain does keep rising — reviewing at 50% recall buys about 69 days.

There are two reasons that is not the recommendation.

The first is that the failure branch gets ugly. At 90% recall, one review in ten fails and the card drops to roughly 2.1 days of stability — a substantial reset. At 60%, four in ten fail. Each failure costs a relearning cycle, so the expected gain per review flatters a distribution that is increasingly bimodal: usually great, sometimes back to the start.

The second is that retrievability is not a number in a database. It is your actual ability to recall the fact. Deliberately running your entire knowledge base at 60% means walking around unable to reliably retrieve four in ten things you have studied — fine for a card in a simulation, not fine on a ward or three weeks before an exam.

FSRS’s default target of 90% is a considered compromise between per-review efficiency and staying functional. What our numbers argue against is not the 90% target but the very common practice of reviewing at 98%, which is what re-reading the deck the next day amounts to.

What to do with this

The practical version is short: stop re-reading recent material and start reviewing older material.

The instinct runs the other way. Recent lectures feel urgent and older ones feel settled, so attention flows toward the material that needs it least. Every pass over yesterday’s deck is a 4.2-day purchase; the concept from five weeks ago that you have not touched is sitting in the 20-to-30-day range, waiting.

Acting on that requires knowing where each concept currently sits on its curve, which is not something intuition provides — the feeling of “I know this well” is generated by recent exposure, which is exactly the signal you want to ignore. This is what Forgetting computes inside KoiSwarm: it tracks the estimated stability of each concept and surfaces the ones near threshold, so the material you are shown is the material where a review is worth doing.

One important exception, from the first article in this series: the very first review, done within a day of the lecture, is worth doing even though it lands in the low-return zone. A brand-new memory has a stability of about three days, so it is genuinely at risk of being gone within the week. After that first consolidation, waiting is the better move.

Limits of the model

One card, one starting stability, one difficulty value. The absolute numbers shift with both — an easier concept gains more from every review, a harder one less — but the shape of the curve is a structural property of the model rather than an artefact of the parameters we chose.

The bigger caveat is that FSRS models a discrete recall event: you saw a prompt, you retrieved or you did not. Much of medical study is not shaped like that. Working through a clinical vignette exercises several linked concepts at once and builds the connections between them, which this model does not represent at all. Read this as a result about factual recall, which is a large part of preclinical study but not the whole of it.

Model: FSRS-5, default published parameters. Single card, initial stability 10 days, difficulty 5, expected values weighted by recall probability at each interval.