ExplainerResearch-backed

Desirable Difficulties: When Harder Learning Helps—and When It Does Not

Effort can strengthen retention and transfer, but difficulty is useful only when it engages the target process without overwhelming the learner.

The desirable-difficulty gate. A mechanism, capacity, feedback, and outcome gate for distinguishing productive challenge from friction, overload, and unsupported struggle. Download the SVG asset.
Direct answer

A difficulty is desirable when it makes the learner perform useful cognitive work—such as retrieving, discriminating, generating, or adapting—and improves delayed retention or transfer. Difficulty is undesirable when it consumes capacity without serving the target, arrives before prerequisites, prevents corrective feedback, or causes avoidable exclusion and disengagement. Effects depend on prior knowledge, element interactivity, feedback, motivation, accessibility, and the criterion used to define learning.

The decision behind the phrase

“Make learning harder” is memorable advice and a dangerous design rule. This article is for learners and designers deciding whether to remove friction or preserve challenge. It distinguishes effort that exercises the target capability from effort spent navigating the lesson.

The paradox behind desirable difficulties is real: conditions that depress performance during practice can sometimes improve performance later. Spacing creates some forgetting between sessions; retrieval withholds the answer; interleaving makes the next procedure less predictable. Immediate ease can therefore mislead.

But the inverse does not follow. Confusing instructions, inaccessible design, irrelevant complexity, sleep deprivation, humiliation, and unsupported search are not beneficial merely because they are hard.

What the evidence establishes for the desirable-difficulty gate

Evidence snapshotHigh confidence

Reviews distinguish temporary practice performance from relatively durable learning and identify retrieval, spacing, generation, and varied practice as conditions that can introduce productive challenge. Research also documents boundary conditions: high element interactivity and insufficient prior knowledge can turn a proposed desirable difficulty into overload. The name describes an empirical result, not an intrinsic property of a task.

bjork-difficulties, soderstrom-performance, chen-undesirable

Claim sources: bjork-difficulties, soderstrom-performance, chen-undesirable

Difficulty needs a causal story

For every added obstacle, finish this sentence:

This difficulty is expected to improve [later capability] because it requires [learning process], and we will detect the benefit with [criterion and delay].

“No hints” might require retrieval, or it might strand a novice who lacks the representation needed to begin. “Mix problem types” might strengthen discrimination, or it might prevent initial comprehension. The same surface feature can have different mechanisms at different stages.

The desirable-difficulty gate

| Gate | Pass question | Warning signal | |---|---|---| | Mechanism | What learning process does the effort require? | “Struggle builds character” is the only account | | Capacity | Does the learner possess the prerequisites to engage it? | Errors are random rather than diagnostic | | Feedback | Can misconceptions be detected and corrected? | Wrong responses repeat without correction | | Outcome | Is benefit measured later or in transfer? | Only time-on-task or immediate fluency is reported | | Sustainability | Can the learner continue without avoidable harm? | Access barriers and exhaustion are reframed as rigor |

Failing one gate does not always require removing the challenge. It may require a worked example, a smaller step, a cue, better feedback, or a different moment.

Two learners, one problem

A novice and an experienced analyst receive the same unfamiliar forecasting case with no guidance. The analyst recognizes relevant variables and can use the open problem to expose assumptions. The novice spends most of the session decoding vocabulary and guessing what a finished answer looks like.

The activity is not simply “productive struggle.” For the analyst, it may exercise model selection. For the novice, it may impose search without a usable schema. Give the novice a worked case and a completion problem, then fade the support. Preserve the hard judgment; remove the avoidable search.

Productive discomfort versus bad friction

Productive discomfort often has a precise target:

  • recalling before seeing;
  • distinguishing close alternatives;
  • producing an explanation rather than copying one;
  • varying context so a rule cannot be tied to one example;
  • delaying feedback briefly enough to complete a genuine attempt.

Bad friction has no defensible connection to the target:

  • hunting through a broken interface;
  • deciphering undefined jargon;
  • guessing hidden scoring rules;
  • enduring irrelevant time pressure;
  • overcoming a format that conflicts with sensory or motor access.

Accessibility and rigor are not opposites. A caption can remove an auditory barrier while the conceptual inference remains difficult. A calculator can offload arithmetic when the target is statistical reasoning. The design question is always: which work must the learner own?

Run the difficulty gate

Select one difficult activity already in your learning system.

  1. Name the delayed capability it is meant to improve.
  2. Identify the cognitive action the difficulty forces.
  3. Check prerequisites with one short diagnostic.
  4. Define an error pattern that would indicate overload.
  5. Add the smallest support that preserves the target action.
  6. Compare delayed performance with a less difficult condition.
  7. Retain, tune, or remove the difficulty based on the result.

Do not compare only how much learners enjoyed the conditions, but do not ignore experience either. A theoretically effective activity that people cannot or will not sustain has limited practical value.

Ways difficulty becomes theatre

A useful transfer check changes both the surface cues and the timing. If a learner succeeds only on the practised format five minutes later, the result may be supported performance rather than retention. Ask for an unaided explanation after a delay and an application in a new context. For novices, preserve enough guidance to make the relevant relationship visible; for learners with prior knowledge, remove support selectively. Review evidence on desirable difficulties is strongest as a conditional design principle, not a demand that every learner experience maximum struggle.

  • Adding arbitrary limits because rigor must look severe.
  • Calling poor explanation “discovery.”
  • Removing all aids when real performance allows tools.
  • Celebrating errors without providing correction.
  • Generalizing a result from simple paired associates to a complex profession.
  • Using effort as a moral judgment about the learner.
  • Measuring immediate performance and claiming durable learning.

Limits of the evidence

Limits and counterevidence

Desirable difficulties are a family resemblance among effects, not one intervention with a universal effect size. Laboratory tasks, educational settings, age groups, and outcome measures vary. Challenge can interact with prior knowledge, disability, language, anxiety, motivation, and available time. The safest conclusion is conditional: specify the mechanism, learner, support, and later test before preserving difficulty.

Difficulty earns its place by producing stronger capability, not by demanding more suffering.

Use cognitive load to assess capacity, compare worked examples and productive struggle, and verify the outcome with evidence of actual learning.

Named sources

Evidence and further reading

  1. Desirable Difficulties in Theory and Practiceresearch · accessed 2026-07-28
  2. Learning Versus Performance: An Integrative Reviewresearch · accessed 2026-07-28
  3. Undesirable Difficulty Effects in High-Element Interactivity Materialsresearch · accessed 2026-07-28
Publication record

Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.