The Error Taxonomy Method: Learn From Patterns, Not Anecdotes
Define observable error classes, code representative cases with reliability, quantify patterns, route interventions, and retest the taxonomy on new work.
Collect a representative set of errors, define observable categories with inclusion and exclusion rules, code a sample independently with another reviewer, reconcile ambiguity, then quantify frequency, consequence, and context separately. Route each high-value class to a different intervention and test whether its rate falls on new cases. Categories simplify continuous and interacting causes, and observed output does not by itself reveal the cognitive or system mechanism.
Use this when errors recur across inspectable cases
Use this method when a learner, team, model, or process produces enough comparable outputs to distinguish a pattern from a memorable miss: writing revisions, diagnostic decisions, citation checks, code reviews, support tickets, or task failures.
Do not use a taxonomy to diagnose a person, rank moral worth, infer hidden motives, or avoid repairing an urgent severe error. It is also a poor fit for one-off creative exploration where variation is not meaningfully “wrong.” Formal safety or clinical classification requires validated domain methods.
Research reviews describe conditions under which errors and corrective feedback can support learning. metcalfe, reason Human-error scholarship distinguishes among forms and system levels rather than treating every unwanted outcome as one category. reason Neither source licenses a universal taxonomy for all work.
What survives the comparison: the Error Coding Manual and Intervention Router
The coding manual
Every code has:
| Field | Requirement | |---|---| | Name | Neutral description of observable failure | | Definition | What must be present | | Exclusion | Similar cases that do not belong | | Anchor | One positive and one negative example | | Context | Task, stage, support, and condition | | Consequence | Effect independent of frequency | | Repair route | Knowledge, selection, execution, feedback, or system | | Confidence | Clear, probable, or unresolved |
Allow multi-code cases when mechanisms interact, but define how they are counted.
Build code calibrate route and retest
Step 1 — Define the unit. One claim, decision, response, transaction, or task attempt. Do not mix units silently.
Step 2 — Sample representatively. Include success, routine misses, rare severe cases, different users, and different conditions.
Step 3 — Describe before explaining. Write what is observable: omitted source, wrong method selected, calculation mismatch, unauthorized action.
Step 4 — Draft categories. Aim for distinctions that imply different interventions. Merge categories with identical responses.
Step 5 — Write anchors and exclusions. Resolve boundary cases before full coding.
Step 6 — Calibrate reviewers. Two people independently code a sample. Measure agreement, inspect disagreements, and revise the manual.
Step 7 — Code and stratify. Report frequency, consequence, context, and uncertainty. Do not collapse them into one severity number.
Step 8 — Route interventions. Missing knowledge gets instruction; selection errors get contrasting cases; execution slips get interface or checklist analysis; system errors get workflow redesign.
Step 9 — Retest new cases. Preserve the baseline and determine whether the targeted class falls without another class rising.
Worked taxonomy: source verification
“Citation error” divides into fabricated source, inaccessible source, wrong passage, scope inflation, outdated authority, duplicate secondary sourcing, and missing citation. Each demands a different response.
A blanket instruction to “be more careful” cannot repair a generator that invents identifiers, a workflow that never opens pages, and an editor who broadens the source’s population. Coding reveals whether the dominant intervention is tool validation, primary-source tracing, claim rewriting, or current-source verification.
The negative finding matters: an intervention may reduce missing citations while increasing irrelevant citations. Report displacement.
Adaptations by volume and consequence
- Low volume: use a small case log and qualitative pattern threshold.
- High volume: sample by task and subgroup before automated coding.
- AI evaluation: calibrate automated labels against blinded human review and inspect drift.
- Team learning: code the artifact, not the person, and protect confidential cases.
- High consequence: prioritize severity and containment before frequency.
Adaptation should preserve observable definitions and a new-case retest.
Before automation, run a drift challenge. Add new cases from a later period, a different contributor, and an adjacent task. Review which codes become ambiguous or disappear. A stable taxonomy should preserve decision-relevant distinctions without forcing novel failures into obsolete boxes; record every code revision and recode the baseline when comparisons require it.
Failure modes that turn categories into labels
- Creating categories from one vivid example.
- Mixing cause, output, and consequence in the same level.
- Coding only failures and losing the success baseline.
- Letting one reviewer redefine codes during the run.
- Treating agreement as truth.
- Counting correlated errors as independent.
- Optimizing the most frequent class while ignoring severe rare cases.
- Hiding unresolved or negative findings.
When a category cannot be recognized consistently or route a distinct action, revise or retire it.
Where the Error Coding Manual and Intervention Router travels—and where it does not
Give the practitioner a new set from another task. Without the codebook template, they must define the unit, sample successes and failures, create operational categories and exclusions, calibrate a second coder, route at least two classes to different interventions, and specify a new-case retest. Score reliability and actionability, not taxonomy size.
Turn the dominant class into Deliberate Practice, recover event context through The After-Action Review, and investigate a recurring causal branch with The Five Whys, With Evidence.
A code describes the output not the whole cause
The same visible error can arise from different knowledge, attention, interface, incentive, or environmental conditions. Reviewer agreement does not prove causal validity. Samples can omit rare harms, and automated classifiers can reproduce the biases of the codebook. Use domain experts and formal investigation when error classification affects safety, rights, or professional judgment.
A useful taxonomy does not make failure tidy. It makes the next repair testable.
Named sources
Evidence and further reading
Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.