BlogWhy weak multiple choice distractors make questions easierAssessment

Why weak multiple choice distractors make questions easier

In short

  • Two dead options raise the guessing floor from 20% to 33%.
  • That removes 25% of the question's information.
  • You need 89 such questions to learn what 67 clean ones would tell you.
  • Distractor quality, not stem quality, is where generated questions usually fail.

Everyone has met this question. Five options, and two of them are not answers at all — a drug from the wrong class entirely, a structure from a different organ system. You discard them without reading properly and choose between the three that remain.

It still counts as a five-option question. It is not one, and the difference is measurable.

What a dead option does

An implausible distractor does not function as an option. If nobody plausibly picks it, it is decoration, and the effective number of choices is the number a reasonable student would actually consider.

That changes the guessing floor directly. Five live options give a blind examinee a 20% chance. Three live options give 33%. And because the guessing floor is the parameter that determines how much a question tells you about the person answering it, everything follows from there.

We ran the same three-parameter IRT model we used to compare question formats, varying only how many of the five options are real.

Bar chart showing information per item falling as live options decrease from five to two, with the number of items needed rising from 67 to 134.
Information per item as implausible options are removed from consideration, with items required for the same measurement precision.

The cost

Dead options Live options Effective guessing Information Items needed
0 5 20.0% 0.167 67
1 4 25.0% 0.150 75
2 3 33.3% 0.125 89
3 2 50.0% 0.083 134

One dead option costs 10% of the item’s information. Two costs 25%. Three — a question with a single serious competitor — costs half, at which point the item is worth no more than a true/false.

The right-hand column is the one that matters for a student. Two dead options per question means you need 89 questions to learn what 67 well-built ones would tell you about your own readiness. Twenty-two extra questions, and the only thing you get for them is the information you should have had already.

Where generated questions actually fail

It is easy to write a plausible clinical stem. A 58-year-old presents with crushing chest pain radiating to the jaw — the vocabulary and rhythm are well represented in any medical corpus, and the result reads convincingly.

Writing four wrong answers that are all genuinely tempting is a different job. It requires knowing which confusions a student at this stage actually makes: which two drugs get mixed up because they sound alike, which pathway is commonly misattributed, which condition presents similarly enough to be a real candidate.

The failure mode is systematic. A generated item pulls its distractors from things that are topically related rather than diagnostically confusable, which produces options a knowledgeable person recognises as wrong instantly. The question looks fine. It is measuring less than it appears to.

The stem is what a question looks like. The distractors are what a question is.

This is a bad deal in both directions. Your practice score drifts up because guessing carries more of it, and the practice itself is doing less work, because eliminating an obviously wrong option exercises nothing.

The check we run

This is the reason USMLE puts every generated item through a separate verification pass rather than shipping the first draft. A second model solves the question independently and reviews it for medical soundness, and items that do not hold up do not reach you.

Verification is not free — it doubles the model work per question, and that cost is real. The alternative is to hand students items that quietly need 89 of themselves to do the job of 67, which is a worse trade.

MCQ draws its distractors against the concepts in your own uploaded material rather than from generic association, so the wrong answers are things from your actual course that could plausibly be confused with the right one. That is what makes a distractor live.

A test you can run on your own bank

Take twenty questions from whatever you are practising with and, before answering, count how many options you would seriously consider.

If the average is four or five, the bank is well built and your scores mean roughly what they appear to. If it is consistently three, your effective guessing floor is 33% rather than 20%, and a practice score of 70% corresponds to genuinely knowing about 55% of the material — the arithmetic is in our piece on score inflation.

That is a large enough gap to change how prepared you think you are, three weeks before an exam.

Limits of the model

We treat a dead option as contributing nothing, which is the clean version of a messier reality. In practice an implausible option still absorbs a little attention and occasionally catches someone genuinely lost, so the true effect sits slightly below what we report.

The model also assumes the remaining live options are equally attractive. Real distractors vary — usually one is nearly right and the others are weaker — and a full analysis would use a nominal response model that scores each option separately rather than collapsing to right/wrong.

Both simplifications are conservative in the same direction. And the headline does not depend on the details: options nobody picks are not options, and a question is only as good as its worst-considered alternative.

Model: 3PL item response theory, discrimination 1.0, difficulty 0.0, evaluated at θ = 0, effective guessing 1/(live options). Code in research/question-quality.