The Alignment Problem: Whose Values Enter the System?
Test Brian Christian’s account of machine-learning alignment against contested objectives, data politics, abstraction, plural values, institutional power, and recourse.
Values enter AI before and after model training: in the problem definition, who is represented, labels, data exclusions, loss function, thresholds, interface, deployment incentives, monitoring, and appeal. Ask who chose each translation, whose error counts, and who can reverse the decision. A technically optimized target is not evidence of legitimate human alignment.
Reconstructing alignment as a translation problem
Christian’s book traces machine learning through examples of reward, imitation, bias, and human values. Its central argument is that powerful learning systems do not automatically inherit what humans mean. They optimize signals, objectives, and data proxies. Misalignment appears when the formal target diverges from the human purpose.
The familiar thought experiment is a system that pursues a stated objective in an unintended way. The harder real-world case begins earlier: humans disagree about the objective, institutions translate some values and exclude others, and the people bearing error may not control the specification.
value-translation chain: the claim-bearing evidence
Christian reconstructs the historical and technical development of alignment problems across machine learning. NIST’s risk framework treats AI as contextual and life-cycle based, with governance, mapping, measurement, and management. Sociotechnical-fairness research shows how abstraction can hide institutions, feedback, and social meaning. Alignment cannot be confined to a loss function.
christian-alignment, nist-ai-rmf, selbst-abstractionClaim sources: christian-alignment, nist-ai-rmf, selbst-abstraction
Audit the value-translation chain
Document each handoff:
| Stage | Question | |---|---| | Purpose | What human condition should improve? | | Standing | Who can define or contest that purpose? | | Construct | What concept represents success? | | Data | Whose histories and absences train the system? | | Label | Who judged examples, under which instruction? | | Objective | What mathematical signal is optimized? | | Threshold | Which trade-off between errors is selected? | | Interface | What does the user see, trust, or misunderstand? | | Deployment | Which incentive governs actual use? | | Recourse | Can an affected person obtain explanation and remedy? |
No stage is value-free. Some are empirically testable; others require legitimate deliberation.
Counterevidence to a single alignment target
Plural values do not always aggregate. Privacy can conflict with personalization; individual autonomy with collective safety; consistency with accommodation; and present users with future populations. “Human values” is not one latent variable waiting to be learned.
Preference learning also faces adaptive and manipulated preferences. A person can click compulsively without endorsing an engagement-maximizing life. Observed behavior is evidence, not sovereign moral truth.
The counterargument to pluralism is operational: a system must act. The answer is not to avoid choice but to expose thresholds, preserve vetoes and rights, separate contexts, and assign accountable authority.
The hardest objection to The value-translation chain
Expanding alignment to society can make the term mean everything and therefore nothing. Engineers need bounded technical problems: reward misspecification, robustness, distribution shift, interpretability, and control.
Keep levels distinct. A technical alignment claim should specify its target and test. A sociotechnical claim should identify institutions and stakeholders. A political claim should identify legitimate authority. Connecting levels is necessary; collapsing them prevents evaluation.
Intellectual inheritance behind the Alignment Problem
The intellectual genealogy includes cybernetics and control, Norbert Wiener’s warnings about literal machines, reinforcement learning, principal–agent problems, safety engineering, moral philosophy, and algorithmic-fairness research. It also inherits older political questions: who defines the public good, who represents affected people, and how authority can be contested.
This genealogy prevents a category mistake. Engineering can improve a system’s fidelity to a target. Political and ethical reasoning determine whether the target and authority are defensible.
A hiring model
A firm builds a model to predict “high performer” from historical employees. Past performance ratings become labels. The objective appears neutral: classification accuracy.
The chain exposes choices. Ratings reflect managers and roles; people denied entry never appear; the firm values short-term output over mentoring; false negatives affect applicants while false positives cost the company; and an automated recommendation changes recruiter attention. Improving accuracy against the historical label can deepen the original institutional preference.
A responsible response may change the target, collect task-relevant evidence, use structured human review, monitor subgroup outcomes, and provide appeal. The decision may also be not to automate.
Alignment after deployment
Even a well-specified model enters a workflow. Users overtrust or ignore it; managers change targets; affected people adapt; data drift; vendors update components; and metrics become goals.
Build monitoring around material failure, not only model performance:
- who is affected and how;
- where human override occurs;
- whether appeal changes outcomes;
- new dependence or skill loss;
- feedback that reaches redesign;
- conditions requiring suspension.
Alignment is maintained governance, not a one-time training achievement. It must be re-established as purposes, actors, data, and consequences change.
Reversal conditions for the value-translation chain
Confidence rises when the purpose is bounded, affected stakeholders helped shape it, labels and outcomes are validated, error trade-offs are public, monitoring captures real harm, and recourse works. It falls when the system optimizes a proxy detached from the purpose, value conflict is hidden as technical tuning, or the deploying institution escapes accountability.
For low-stakes reversible tasks, lighter governance may be proportionate. For rights, safety, employment, health, or mass-scale dependency, average accuracy is an insufficient warrant.
Alignment category errors
- Treating a benchmark as a human value.
- Equating historical preference with legitimate purpose.
- Assuming one stakeholder speaks for humanity.
- Solving objective fidelity while ignoring objective choice.
- Counting human review without measuring its authority or quality.
- Calling appeal available when it cannot change the outcome.
- Confusing model safety with deployment safety.
- Expanding “alignment” until no claim can be tested.
The boundary of this reading of the Alignment Problem
The book spans technical history, ethics, and social consequences but predates many later systems and policies. This article does not resolve long-term control questions or select a universal value theory. Specific deployments require current technical evaluation, legal analysis, stakeholder participation, and verified product facts at publication.
Begin with legibility and abstraction, scale the question through containment, and test claimed public purpose in The Technological Republic.
Named sources
Evidence and further reading
Published July 29, 2026. No substantive revision has been recorded. Evidence last verified July 28, 2026.