Field LabPractitioner-tested

Which Citation Errors Can a Checklist Catch? A 100-Case Boundary Test

Across one hundred frozen citation cases, which faults are observable from structured records and which still require semantic judgment? Executed on frozen inputs with inspectable

Which Citation Errors Can a Checklist Catch? A 100-Case Boundary Test. A visible map of the frozen unit, baseline, principal result fields, and interpretation boundary for citation boundary. Download the SVG asset.
Direct answer

The frozen cases draw a sharp boundary: all sixty structural faults were flagged, none of the twenty semantic overclaims was detected, and no clean control was falsely flagged. Use structural validation before semantic review, then open the source and test entailment, population, date, and causal language. The rule can clear only the fields it inspects; meaning inside one citation case remains a separate judgment.

The boundary a green check cannot cross

A citation boundary checklist becomes dangerous at the moment its green light is mistaken for understanding. The predetermined research question is: Across one hundred frozen citation cases, which faults are observable from structured records and which still require semantic judgment? The unit of observation is one citation case; neither truth nor human judgment appears as a hidden outcome.

A citation that resolves to a plausible source can still overstate causality or population scope, so link resolution alone is an inadequate baseline. Against that baseline, citation boundary separates a machine-readable defect from an interpretive dispute. Moving the structured citation checks line after seeing the flags would make the test easier, but it would also erase the boundary under examination.

Frozen citation boundary inputs contain cases: 100; structural: 60; semantic: 20; clean controls: 20. Designed structured citation checks faults and clean controls appear together, so the rule can be penalized for both silence and overreach.

One hundred cases, three kinds of decision

The 3-stage procedure gives the checklist no access to undeclared meaning:

  1. Freeze sixty structural faults, twenty semantic faults, and twenty clean controls before running the checklist.
  2. Allow the checklist to inspect identifiers, locators, ledger use, and dates but not source meaning.
  3. Score structural recall, semantic recall, and clean-control false positives separately.

Across the citation boundary diagram and JSON, the same unit, sample, and result fields remain visible. If those structured citation checks representations disagree, the visual is wrong; visual polish cannot override the canonical executed record.

Results: where the citation boundary rule succeeded—and went silent

| Recorded result | Value | |---|---| | structural Flags | 60 | | structural Recall | 1 | | semantic Flags | 0 | | semantic Recall | 0 | | clean False Positives | 0 |

The decisive citation boundary finding is asymmetric: all sixty structural faults were flagged, none of the twenty semantic overclaims was detected, and no clean control was falsely flagged. Its rule is reliable only inside the structured citation checks fields it can inspect. A clean citation boundary result means “no declared structural fault was found,” not “the underlying claim is sound.”

The prespecified negative finding for citation boundary is equally important: The checklist caught none of the semantic overclaim cases because every identifier and locator was structurally valid. It marks the point at which this structured citation checks method becomes silent, a condition a reader needs before deciding whether to use the rule.

Inspect the cases the citation boundary rule could see

The complete sanitized raw data is the canonical record for this run. The first twenty-eight citation boundary cases stay in source order. They include quiet structured citation checks controls as well as failures, preventing a conclusion assembled only from conspicuous examples.

  • Row 1 — case Id: C001; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 2 — case Id: C002; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 3 — case Id: C003; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 4 — case Id: C004; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 5 — case Id: C005; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 6 — case Id: C006; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 7 — case Id: C007; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 8 — case Id: C008; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 9 — case Id: C009; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 10 — case Id: C010; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 11 — case Id: C011; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 12 — case Id: C012; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 13 — case Id: C013; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 14 — case Id: C014; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 15 — case Id: C015; fault Type: unknown-source-id; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 16 — case Id: C016; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 17 — case Id: C017; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 18 — case Id: C018; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 19 — case Id: C019; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 20 — case Id: C020; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 21 — case Id: C021; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 22 — case Id: C022; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 23 — case Id: C023; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 24 — case Id: C024; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 25 — case Id: C025; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 26 — case Id: C026; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 27 — case Id: C027; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
  • Row 28 — case Id: C028; fault Type: missing-source-locator; decision Class: structural; checklist Flagged: true; human Semantic Review Required: false.
Evidence snapshotHigh confidence

The executed citation boundary record shows that all sixty structural faults were flagged, none of the twenty semantic overclaims was detected, and no clean control was falsely flagged. Row-level structured citation checks fields support that bounded finding, while no field represents human learning, reader comprehension, or real-world deployment.

lab-record

Claim sources: lab-record

Evidence snapshotModerate confidence

The reviewed method source supplies a relevant standard for context, traceability, or explicit evaluation of structured citation checks. It disciplines interpretation of structured citation checks; it does not generate or independently confirm this local aggregate.

method-source

Claim sources: method-source

The objection: the citation boundary rule is too narrow

A critic might say the zero semantic recall simply proves that the checklist was badly designed. That criticism mistakes a declared boundary for an implementation defect: population scope and causal force are not present in the inspected fields.

Taken seriously, the citation boundary objection shows that a deliberately narrow checker can still be badly specified. Its recommendation reverses if a missing structured citation checks field can be represented deterministically and changes a consequential decision without creating false assurance.

That reversal condition keeps citation boundary from becoming either technological maximalism or ritual caution. The structured citation checks procedure earns its place only when it makes a consequential uncertainty, tradeoff, or failure more visible.

Run the checklist before reading for meaning

Use structural validation before semantic review, then open the source and test entailment, population, date, and causal language. Reproduce the citation boundary flags, then inspect cases nearest the boundary between an observable structured citation checks field and semantic judgment. A new field or rule constitutes a new citation boundary test with its own controls.

The working sequence for citation boundary is specific to this study: lock the question and baseline, freeze the unit, execute the declared transformation, retain negative findings, and separate the local result from any transfer claim.

False assurance is the central failure mode

The most serious citation boundary misuse is false clearance: treating silence outside the rule's inputs as evidence that no problem exists. Selective deletion of structured citation checks controls and borrowed method authority create the same illusion by different means.

Reproducibility in structured citation checks also fails when a download cannot regenerate the claim in the prose. This citation boundary record keeps protocol, sample, aggregates, limitations, negative findings, and row-level output in one parseable object so that disagreement can reach the actual computation.

Where this citation boundary result stops

Limits and counterevidence

The cases are designed boundary probes, not an estimate of citation-error prevalence in published scholarship. The designed controls expose rule behavior; they do not estimate how often structured citation checks fails in published or operational work. No inspected field contains the missing semantic judgment.

This citation boundary limit specifies the next experiment. Transfer of this structured citation checks result requires records from the target context, the same visible denominator, and a fresh execution—not stronger adjectives attached to the present run.

Related reading:

The citation boundary rule earns trust by making its silence visible, not by painting that silence green.

Named sources

Evidence and further reading

  1. Which Citation Errors Can a Checklist Catch? A 100-Case Boundary Test — Sanitized Raw Recordpractitioner · accessed 2026-07-28
  2. Framework for Information Literacy for Higher Educationofficial · accessed 2026-07-28
Publication record

Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.