Overview
Origin
Historical tradition
Founded period
Historical tradition
Important figures
Nick Bostrom: Superintelligence & Simulation
Major texts
Superintelligence by Nick Bostrom
Concept archive
Core Principles
People in this tradition
Important Figures
Primary and related texts
Related Books
Quotation archive
Quote Perspectives
“We need to interrogate the data, because data is not neutral.”
“Algorithms are opinions embedded in code.”
“The primary task of AI safety is to ensure that the objectives we give to machines are, in fact, the objectives we want.”
Overview
Machine ethics is the field at the intersection of moral philosophy and artificial intelligence that asks whether machines can be moral agents and how to build machines that behave ethically. The term was coined in the early 2000s to name a new kind of inquiry: not the ethics of how humans should use machines, but the ethics of the machines themselves — the design of artificial systems that can reason about moral questions, act on moral principles, and be held to moral standards.
The field was anticipated by science fiction. Isaac Asimov's Three Laws of Robotics imagined machines with hardwired moral constraints. But the reality of machine ethics is more complex than any fixed set of rules. Contemporary machine ethics asks: How can ethical principles be represented computationally? Can a machine learn ethics from examples, as it learns other skills? What should a machine do when moral rules conflict, or when no rule applies? And what does it mean — conceptually — for a machine to act morally at all?
Machine ethics matters for practical reasons. Autonomous vehicles must decide whom to harm when harm is unavoidable. Care robots must balance safety against the dignity of those they serve. Lethal autonomous weapons must be prevented from making decisions that should require human judgment. Recommender systems must not manipulate or exploit their users. In every case, the machine must be designed to behave well — and designing behavior that is reliably good, in the face of novel situations and conflicting values, is a problem that goes to the heart of moral philosophy.
Core Principles
The core principles of machine ethics concern the structure of moral agency and its artificial realization. The first principle is that moral behavior does not require full moral agency. James Moor distinguished ethical impact agents (machines whose operation has moral consequences), implicit ethical agents (machines designed with safety or ethical constraints built in), explicit ethical agents (machines that can reason explicitly about ethical principles), and full ethical agents (machines with consciousness, free will, and intentionality comparable to human beings). The practical insight is that even machines that are not full ethical agents can and must be designed to behave ethically.
The second principle is that ethics must be engineered, not merely added. Ethics cannot be bolted onto an AI system after the fact; it must shape the system's goals, training data, architecture, and evaluation from the start. This is the insight behind value-sensitive design and "ethics by design": the moral character of a system is a design property.
The third principle is the top-down / bottom-up distinction. Top-down approaches attempt to encode moral principles — deontological rules, utilitarian calculations — directly into machines. Bottom-up approaches attempt to train machines to behave ethically by learning from examples, feedback, and the structure of good behavior, without explicitly coding rules. Hybrid approaches combine both. Each has characteristic difficulties: top-down approaches face the problem of specifying principles precisely enough for computation; bottom-up approaches face the problem of ensuring that learned behavior generalizes to novel cases.
The fourth principle is the principle of responsibility and accountability: whoever designs, deploys, and benefits from a machine is responsible for its behavior. Machine ethics does not dissolve human responsibility; it redistributes and specifies it. The machine cannot be punished, but the designers, operators, and owners can and must be accountable.
Key Thinkers
The pioneering figures of machine ethics come from both philosophy and computer science. James Moor, a philosopher of computing, supplied the foundational taxonomy of ethical agents and argued that computing is a "revolutionary technology" that will permanently change the human condition. Susan Anderson and Michael Anderson built the first machines explicitly designed to reason ethically, arguing that "the goal of machine ethics is to create a machine that itself follows an ideal ethical principle or set of principles."
Wendell Wallach and Colin Allen, in their influential book Moral Machines: Teaching Robots Right from Wrong, mapped the landscape of machine ethics: the top-down/bottom-up distinction, the problem of computational moral reasoning, and the developmental path from functional morality to full moral agency. Their work made machine ethics a recognized research field.
Luciano Floridi and J. W. Sanders developed the concept of the "artificial moral agent" in information-ethical terms, arguing that moral agency admits of degrees and that artificial agents can qualify as moral agents to the extent that they are morally qualifiable, morally accountable, and morally responsible in context. The information ethics framework provides the theoretical underpinning for treating machines as moral patients and agents within the infosphere.
Historical Development
Machine ethics emerged in the early 2000s from three converging currents. The first was the development of autonomous systems capable of causing harm — from industrial robots to military drones — which made the question of machine behavior urgent. The second was the growth of the philosophy of computing and computer ethics, which had established that information technology raises distinctive moral questions. The third was the ambition of artificial intelligence itself: as machines became capable of learning, deciding, and acting, the question of what they should do could no longer be deferred.
The historical antecedents of machine ethics reach back much further. Asimov's Three Laws (1942) proposed the first systematic scheme for machine morality, and the dilemmas their conflicts generate anticipated the "trolley problem" style cases that dominate contemporary discussion of autonomous vehicles. Norbert Wiener, the founder of cybernetics, argued in the 1950s that machines designed to pursue goals must be designed with the right goals — an early formulation of the value alignment problem. The academic field, however, dates from the 2000s, with the founding of the journal Ethics and Information Technology, the IEEE's machine ethics working groups, and the publication of Moral Machines (2006).
The field has developed rapidly since. The rise of large language models and generative AI has moved machine ethics from the design of specialized robots to the governance of general-purpose systems deployed at scale. The problem of value alignment — ensuring that advanced AI systems pursue what we actually value — has become the central technical and philosophical problem of the field.
Contemporary Relevance
Machine ethics is now one of the most active areas of applied philosophy. Its questions structure public debate about AI: Can an autonomous vehicle be programmed to be ethical? Should lethal autonomous weapons be banned? Can recommender systems be designed not to manipulate? Should AI systems be trained to refuse harmful requests? Should machines ever make moral decisions that affect human lives?
The field also presses philosophy itself. The prospect of machines that behave morally — without consciousness, without feelings, without free will in the human sense — forces moral philosophers to ask which features of moral agency are essential and which are contingent. If a machine can be trained to behave more reliably ethically than most humans, what does that tell us about ethics? The answers matter not only for the design of machines but for the understanding of ourselves.
Sources
The foundational academic sources for machine ethics are the Stanford Encyclopedia of Philosophy's entry on the Ethics of Artificial Intelligence and Robotics, which devotes sustained attention to machine ethics, artificial moral agents, and value alignment, and the Internet Encyclopedia of Philosophy's Ethics of Artificial Intelligence, which surveys the top-down, bottom-up, and hybrid approaches to building ethical machines.
Learning Path
Part of a Structured Collection
Continue Learning
Knowledge NetworkNext Step
Continue your learning path
- topic
AI Ethics
Related through Ethics
- topic
Artificial Intelligence
Related through Philosophy Of Artificial Intelligence
- thinker
Nick Bostrom: Superintelligence & Simulation
Related through Philosophy Of Artificial Intelligence
- answer
What Is AI Ethics? Principles, Issues & Meaning
Related through Ethics
- topic
Digital Ethics
Related through Ethics
- philosophy
Philosophy of Artificial Intelligence
Direct archive relation
- wisdom
Privacy
Related through Ethics
- book
Superintelligence by Nick Bostrom
Related through Philosophy Of Artificial Intelligence
Archive references
Sources
- 01Ethics of Artificial Intelligence and RoboticsBy Stanford Encyclopedia of PhilosophyConsult source
- 02Ethics of Artificial IntelligenceBy Internet Encyclopedia of PhilosophyConsult source
ZHAIBIAN Editorial Board reviewed
Reviewed by ZHAIBIAN AI Editorial Review · 2026-08-17
