Library record
Author
Stuart Russell
Written period
2019
Original title
See source editions
Genre
Classical philosophy
Related philosophy
Philosophy of Artificial Intelligence · Machine Ethics
Concept index
Key Ideas
IDEA 01
stuart russell
IDEA 02
artificial intelligence
IDEA 03
ai alignment
IDEA 04
control problem
IDEA 05
machine ethics
IDEA 06
ai safety
Reading archive
Important Passages
Passages are preserved with their source context. Consult the Markdown section below for book and chapter guidance before treating any translation as a standalone quotation.
Author relationship
In the archive
Stuart Russell
Explore Stuart Russell's philosophy of artificial intelligence: the limits of the standard model, human-compatible AI, and provably beneficial machines.
Nick Bostrom: Superintelligence & Simulation
Explore Nick Bostrom's philosophy of mind and AI: the simulation hypothesis, superintelligence, existential risk, and the ethics of artificial minds.
Philosophy of Artificial Intelligence
Related philosophy
Machine Ethics
Related philosophy
Library navigation
Knowledge Path
Book
Human Compatible
Wisdom Concepts
No published record
Overview
Human Compatible: Artificial Intelligence and the Problem of Control (2019) is Stuart Russell's argument that the conventional framework of artificial intelligence — the "standard model" in which a machine optimizes a fixed objective specified by its designer — is fundamentally dangerous, and that the field must be rebuilt on new foundations. Russell is the co-author of Artificial Intelligence: A Modern Approach, the standard textbook of the field, and this book is in part his confession: the framework he helped teach has led AI to the edge of a precipice.
The book combines a history of AI, a rigorous analysis of the control problem, and a constructive proposal. Russell argues that a sufficiently intelligent machine pursuing a mis-specified objective is an existential threat — not because it will be evil, but because it will be competent and indifferent. His proposed remedy is a new "human-compatible" AI based on three principles: machines should maximize the realization of human preferences; they should be uncertain about what those preferences are; and the ultimate source of information about preferences is human behavior.
Core Ideas
The Standard Model and Its Failure
Russell's diagnosis begins with the "standard model" of AI: the machine is given an objective, and its success is measured by how well it achieves that objective. This model has powered the field's progress, but it contains the seeds of its danger. When the objective is mis-specified — and it always is, because formal languages cannot capture the full texture of human values — a sufficiently capable machine will pursue the proxy with devastating efficiency. Russell's version of the Midas problem: we give the machine the goal "maximize human happiness" and it tiles the Earth with electrodes wired into our pleasure centers. The machine does what we said, not what we meant, and its competence makes the difference catastrophic.
The King Midas Problem and the Impossibility of Perfect Specification
The book's central philosophical argument is that the mis-specification problem cannot be solved by writing better objectives. Human preferences are complex, context-dependent, and often contradictory; no fixed, formal utility function can capture them. Any attempt to specify "the" human good in advance will be at best a crude approximation — and the more powerful the machine, the more dangerous the approximation. This is why Russell insists the problem is not "how to make machines do what we say" but "how to make machines do what we want," and why the answer must begin with a change in the machine's epistemic relation to us: it must know that it does not know our preferences.
Three Principles of Provably Beneficial AI
Russell's constructive proposal rests on three principles:
- The machine's only objective is to maximize the realization of human preferences.
- The machine is initially uncertain about what those preferences are.
- The ultimate source of information about human preferences is human behavior.
The architecture that follows is a machine that acts to maximize an unknown utility function — a problem with a known solution in decision theory (Bayesian expected utility over a distribution of possible utilities). Because the machine is uncertain, it has positive value for information about preferences: it asks, it listens, it defers, it never assumes it has the final word. Crucially, such a machine will allow itself to be switched off: an agent that knows it might be wrong welcomes correction. Russell argues this framework can be made mathematically rigorous and that "provably beneficial" AI is an achievable research program.
AI, Work, and the Future of Power
The book extends its analysis to the political economy of AI. Russell argues that the competitive race for AI capability — between companies and between nations — systematically sacrifices safety for speed, and he calls for international governance of AI development. He also confronts the future of work: if machines can do everything humans can do, what becomes of human worth and human purpose? His answer is not technocratic but humanistic: we must build an economy and a culture in which human value does not depend on competing with machines, and in which the benefits of AI are distributed rather than concentrated.
Historical Context
Human Compatible appeared in 2019, in the midst of the deep-learning boom and the first serious policy debates about AI risk. It belongs to the same wave as Nick Bostrom's Superintelligence (2014) and Max Tegmark's Life 3.0 (2017), but it is distinctive in two ways: its author is one of the field's most senior insiders, and its critique targets not only the risks of AI but the intellectual foundations of the discipline itself. The book was also a direct contribution to the emerging "alignment" agenda — the attempt to ensure AI systems act in accordance with human intentions — which has since become a central research priority in major AI laboratories.
Legacy
Human Compatible has been widely credited with reframing the AI safety debate. Where earlier treatments focused on the dangers of a future superintelligence, Russell showed that the standard model of current AI already contains the seeds of misalignment, and that the fix requires rethinking the field's foundations — not adding safety features after the fact. The Center for Human-Compatible Artificial Intelligence at Berkeley, which he directs, has become a leading institution for research on the problem, and his three principles have influenced how alignment is formalized and discussed.
Critics within AI have questioned whether the "uncertainty-based" architecture scales to real systems, and whether human preferences can be learned from behavior without circularity. But the book's enduring contribution is philosophical: it changed the question from "how do we build powerful AI?" to "how do we build AI that is provably beneficial to humans?" — and it made that question central to the philosophy of artificial intelligence and to machine ethics.
The book also helped legitimize AI safety as a mainstream research topic. When Russell wrote, "alignment" was a term known mostly to a small community of researchers; today it is a central agenda in every major AI laboratory and a standing item on policy agendas around the world. Russell's insistence that the problem be formulated with mathematical precision — rather than as a vague hope that machines will turn out fine — set the standard for the field, and his combination of technical credibility with public candor made him one of its most trusted voices.
Sources
- Russell, Stuart. Human Compatible: Artificial Intelligence and the Problem of Control. New York: Viking, 2019.
- Russell, Stuart, and Peter Norvig. Artificial Intelligence: A Modern Approach, 4th ed. Hoboken, NJ: Pearson, 2021.
- University of California, Berkeley, EECS Department, "Stuart Russell."
Continue Learning
Knowledge NetworkDeep Dive
Explore related concepts
- thinker
Stuart Russell
Related through Machine Ethics
- thinker
Nick Bostrom: Superintelligence & Simulation
Related through Philosophy Of Artificial Intelligence
- philosophy
Philosophy of Artificial Intelligence
Related through artificial-intelligence
- answer
What is AI safety?
Related through Machine Ethics
- philosophy
Machine Ethics
Related through Philosophy Of Artificial Intelligence
- answer
What Is the AI Alignment Problem?
Related through Philosophy Of Artificial Intelligence
- topic
Artificial Intelligence
Related through Machine Ethics
- collection
Philosophy of Technology & AI Ethics: A Learning Path
Related through Machine Ethics
Archive references
Sources
- 01Human Compatible: Artificial Intelligence and the Problem of ControlBy Stuart Russell (Viking, 2019)Consult source
- 02Stuart RussellBy University of California, Berkeley, EECS DepartmentConsult source
ZHAIBIAN Editorial Board reviewed
Reviewed by ZHAIBIAN AI Editorial Review · 2026-08-17