Critical Thinking: How to Evaluate Claims and Evidence
A practical method for separating claims, evidence, assumptions, inference, and uncertainty before deciding what to believe or do.
Reason under uncertainty
Tools for evaluating claims, framing problems, seeing systems, generating options, and making better decisions.
Topic map
Starting sequence
Start with the first article, then choose by goal—not by whatever is newest.
A practical method for separating claims, evidence, assumptions, inference, and uncertainty before deciding what to believe or do.
First-principles reasoning rebuilds from constraints; analogy transfers patterns from prior cases. Strong problem solvers alternate between them.
See behavior as the result of relationships, stocks, flows, delays, and feedback—not isolated events—and design safer interventions.
Make better uncertain decisions by defining options, using base rates and ranges, separating value from information, and favoring reversible tests.
Separate what was observed, what the evidence implies, and what should be done. AI can assist each layer, but it must not silently collapse them.
Association describes a pattern; causation asks what an intervention would change. Judge explanations with counterfactuals, diagrams, and rival mechanisms.
A belief becomes accountable when you state what would lower confidence, distinguish it from rivals, and trigger a changed decision.
Read the claim, denominator, scale, uncertainty, and missing comparison before trusting a chart. Then reconstruct the decision it invites.
Reframe a vague complaint into a decision question by separating symptoms, mechanisms, constraints, stakeholders, and evidence that could change the frame.
Decide faster without becoming careless by separating reversible experiments from choices that create durable harm, lock-in, or lost option value.
Make opportunity cost concrete by naming the best displaced alternative, testing scarcity, and comparing portfolios instead of isolated benefits.
Start forecasts with a relevant reference class, then update for case-specific evidence instead of letting a coherent story erase historical frequencies.
Use expected value as a transparent comparison, then stress-test probability, utility, tail risk, dependence, uncertainty, and the value of more information.
Turn verbal certainty into testable forecasts, score repeated predictions, inspect calibration by category, and improve without confusing confidence with truth.
Use forecasts for resolvable uncertainty and scenarios for structurally different futures, then connect both to decisions, signposts, and invalidation.
Prevent research volume from hardening a preferred belief by separating discovery from testing, budgeting disconfirmation, and scoring rival explanations.
Bias literacy becomes useful only when it changes the decision environment: records, defaults, independent estimates, feedback, incentives, and accountability.
Trace reactions, feedback, delays, substitution, and displaced costs before a promising intervention turns its first success into a later failure.
Separate option generation from evaluation without romanticizing creativity: use purposeful constraints, independent variation, criteria, tests, and reopening rules.
Treat every mental model as a purpose-built compression: expose its entities, assumptions, omissions, scale, predictions, rivals, and failure boundary.
Distinguish observed AI convergence from broader systemic risk, then preserve source, model, prompt, reviewer, and non-AI diversity at consequential bottlenecks.
Imagine a committed plan has failed, generate independent causes, convert them into evidence and controls, and define stop conditions before launch.
Record the decision state before outcomes arrive, then score predictions, separate process from luck, and update recurring judgment patterns.
Reconstruct the strongest supportable argument, secure fidelity, then test its premises, evidence, alternatives, and reversal conditions adversarially.
Trace a conclusion back through selected data, meaning, assumptions, and beliefs, then seek missing observations and test a rival interpretation.
Branch each why into evidence-backed causal hypotheses, test interventions, and stop before a neat single-root story replaces a complex system.
Keep hypotheses append-only, define rival predictions before searching, and update confidence from dated evidence without rewriting intellectual history.
Test Jonathan Rauch’s institutional account of truth against reproducibility, exclusion, concentrated power, platform incentives, and AI-mediated disagreement.
Test Julia Galef’s truth-seeking ideal against motivated reasoning, incentives, identity, power, debiasing transfer, and the institutional conditions for honest updating.
Test Donella Meadows’s systems framework against contested boundaries, institutional power, strategic actors, unequal harms, and the politics of intervention.
Test David Deutsch’s optimism about explanatory knowledge against social institutions, measurement, tacit skill, power, implementation, and irreversible harm.
Use James C. Scott’s critique of high-modernist planning to examine AI classification, metrics, local knowledge, administrative scale, and coercive simplification.
Test Gary Klein’s recognition-primed decision model against low-validity environments, feedback quality, automation, overconfidence, and the need for analysis.
Test Amartya Sen’s capability approach against GDP, preference, measurement, paternalism, inequality, conversion factors, and the difference between resources and freedom.
Test Yuval Noah Harari’s history of information networks against information theory, institutional epistemology, propaganda, bureaucracy, power, and AI-generated scale.
Why cheaper generation can move cost into verification, choice, accountability, and correction—and how to distinguish abundance from value.
A falsifiable ledger of capability, adoption, work, trust, learning, multilingual access, and infrastructure signals—without pretending to forecast.
Which combinations of observability, direction, threshold, time window, and scenario exclusivity make an indicator discriminating rather than decorative? Executed on frozen inputs
Across one hundred quality decisions in five knowledge-work domains, which decisions reduce to declared observables and which require contextual interpretation? Executed on frozen