← Back to blog
Research

Building distractors from misconceptions, not guesswork

Adam··4 min read

In most multiple-choice items, three of the four options exist only to be wrong. They’re plausible-looking numbers, chosen so the correct answer doesn’t stand out. That’s a wasted opportunity: a well-built distractor is a diagnostic instrument, and the item costs exactly the same to write.

The research tradition here is older than edtech. Item response work going back decades treats distractor selection patterns as a signal about which misconception a student holds, not merely that they erred. Diagnostic assessment programs — from concept inventories in physics to misconception-mapped math inventories — are built entirely on this idea.

What makes a distractor useful

A distractor earns its place when it corresponds to a specific, nameable error in student thinking. That gives it three properties:

  • It’s attractive. A student with that misconception will actively choose it, not eliminate it.
  • It’s diagnosable. When 40% of your class picks option C, you know what to reteach on Monday — not just that they “struggled.”
  • It’s on-standard. The error it represents is one the standard is actually about, so the item stays aligned even in how it fails.

A distractor that no one picks tells you nothing. A distractor that half the class picks tells you what to teach tomorrow.

Compare that to the default approach: take the correct answer, add one, subtract one, double it. Those options are eliminable by estimation, so strong students discard them without reasoning and weak students scatter randomly. The item measures test-wiseness.

A working misconception model

For any standard, you can usually enumerate the errors before you write a single item. For ratio reasoning (6.RP.A.3), the recurring ones are well documented:

  • Additive thinking. Treating the relationship as a difference rather than a multiplicative comparison — going from 3 : 1 to 6 : 4 by adding 3 to both.
  • Reversal. Inverting the ratio, or mapping the wrong quantity to the wrong role.
  • Part–whole confusion. Reading 3 : 1 as “3 out of every 1” or as the fraction 3/4 of the total.
  • Unit-rate slip. Computing a correct rate, then applying it to the wrong quantity.

Each of those maps cleanly onto a numeric option. Build the item once, and every wrong answer becomes labeled.

Item spec · distractor map

stem: 3 : 1 water to concentrate, 12 cups of water
key: B · 4 cups — scales the ratio correctly
distractor A: 3 cups ← reversal (uses the water term)
distractor C: 6 cups ← additive thinking
distractor D: 36 cups ← multiplies both quantities
report: selection % per misconception

Rules that keep distractors clean

Keep them parallel. Same grammatical form, similar length, same units. An option that’s noticeably longer or more specific than the others is a giveaway — and in practice, the longest option is correct often enough that students learn to pick it.

One error per distractor. If an option requires two mistakes to arrive at, almost no one will land on it and you’ve learned nothing.

No “all of the above.” It rewards partial knowledge in ways you can’t interpret, and it collapses your diagnostic signal.

Avoid absolutes. “Always” and “never” options are eliminated on grammar, not content.

Order options logically. Numeric options should run in ascending order. Ordering by “hide the key in position C” is a habit students detect faster than you’d like.

Reading the results

The payoff comes the morning after. Instead of a percentage correct, you get a distribution across named errors:

  • Heavy on the additive distractor → reteach ratio as a multiplicative comparison, with a double number line.
  • Heavy on the reversal distractor → the concept is intact; the labeling isn’t. That’s a fifteen-minute fix, not a reteach.
  • Evenly scattered → students are guessing. The item was too hard, or the concept wasn’t taught yet.

That third case is worth pausing on. A flat distribution across all four options is itself a finding: it means the item isn’t measuring anything, and you should discard the data rather than act on it.

Where generation fits

Writing four honest distractors for every item is real work — perhaps two or three minutes per item once you know the misconception set, and much longer when you don’t. That’s precisely the sort of skilled-but-repetitive task worth automating: hold a misconception model per standard, generate options from it, label each one, and let the teacher accept or rewrite.

The teacher still owns the call. But the default becomes a diagnosable item instead of three throwaway numbers.