Quotation archive
“We need to interrogate the data, because data is not neutral.”
Timnit Gebru · On the Dangers of Stochastic Parrots
Quote record
Author
Timnit Gebru
Source
On the Dangers of Stochastic Parrots
Chapter / location
2021
Tradition
gebru · data · datasets · ai-ethics · criticality
Source information
From On the Dangers of Stochastic Parrots, 2021.
Original language: English
Translation
Translated from English into English using a named scholarly edition.
Context
Read the contextual commentary in this archive entry.
Interpretation
Gebru urges us to interrogate the data, since data is not neutral. On the dangers of stochastic parrots and critical approaches to AI datasets.
Knowledge network
Related Archive Records
We need to interrogate the data, because data is not neutral. — Timnit Gebru, on the dangers of unexamined datasets (2021)
Timnit Gebru, a computer scientist and co-lead of Google's Ethical AI team until her departure in 2020, has been one of the most influential voices arguing that the datasets on which artificial intelligence is built are not innocent raw material. The sentence captures her core demand: data must be examined, questioned, and understood as a product of history, power, and labor.
Context
The claim belongs to the context of Gebru's research and public advocacy, most prominently the paper "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" (2021), co-authored with Emily M. Bender, Angelina McMillan-Major, and Shmargaret Shmitchell. The paper warned that very large language models — trained on massive corpora scraped from the internet — absorb and amplify the values, biases, and harms embedded in those corpora: stereotypes, hate speech, and the disproportionate representation of some voices over others. The paper also raised the environmental costs and the unexamined limits of scale as a research paradigm.
Gebru's broader research career, including her work on the Gender Shades project with Joy Buolamwini, established her as a leading critic of the assumption that datasets objectively mirror the world. Facial recognition trained on unrepresentative data performs worse on dark-skinned women; language models trained on the internet learn the internet's prejudices. The problem is not an accidental bug but a structural feature: data is collected, curated, labeled, and selected by people, in specific conditions, with specific purposes — and it carries those conditions with it.
Meaning
The sentence means that data is never given in the raw. Every dataset is a selection from reality, and selection is already interpretation: what was included and excluded, what was labeled and how, who was counted and who was erased, which sources dominate and which are absent. To treat data as neutral is to treat the history that produced it as irrelevant — and to allow the biases of that history to be reproduced at scale by machines that appear objective precisely because they are automated.
"Interrogating the data" names the required practice: asking where the data came from, who collected it, under what incentives, with what exclusions; testing what it represents and what it omits; auditing the models built on it for disparate impact. The demand is both technical and political. Technical, because dataset documentation, provenance, and evaluation are engineering tasks that can be done well or badly; political, because the answers to the questions reveal structures of power — whose language dominates the internet, whose faces are in the image archives, whose labor produced the annotations.
The deeper meaning is epistemic. A model trained on data inherits the world-view of the data, and a world-view is not a fact. If the data says that certain groups are overrepresented in certain outcomes, the model will "learn" the correlation and reproduce it as prediction, converting a contingent social condition into a statistical destiny. The appearance of objectivity is precisely the danger: data is not neutral, but it can be made to look neutral, and that appearance is itself a political achievement.
Philosophical Significance
The sentence is a cornerstone of critical data studies and the ethics of AI. It extends the long philosophical critique of "the given" — the idea that experience or observation presents us with pure, uninterpreted facts — to the domain of machine learning. Data, like perception, is theory-laden: it is shaped by the concepts, interests, and conditions under which it is gathered. Gebru's claim operationalizes this insight for an age in which data is the raw material of automated judgment.
The claim also grounds a program of responsible AI research: dataset documentation, participatory data collection, impact assessment, and the inclusion of the communities most affected by AI systems in decisions about what counts as data. Gebru's insistence on interrogating data connects the ethics of AI to the politics of knowledge — to questions of who gets to define problems, who is represented, and who bears the costs. It is an argument against the deferral of responsibility to "the data," which is always, in the end, a deferral to someone's choices.
The philosophical stakes are therefore about accountability. If data is not neutral, then neither are the systems built on it, and neither are the institutions that deploy them. The demand to interrogate the data is a demand to keep the choices visible, to keep the history in view, and to refuse the comfort of technological inevitability. In an era of large language models trained on the entire internet, Gebru's warning has become a research program and a movement: before we trust what the machine has learned, we must ask what the data has taught it — and who wrote the lesson.
Sources
- Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), 610–623. arXiv:2101.05783.
Continue Learning
Knowledge NetworkDeep Dive
Explore related concepts
- thinker
Timnit Gebru
Related through Machine Ethics
- wisdom
Criticality: Meaning, Philosophy & Wisdom
Related through criticality
- answer
What Is Algorithmic Bias?
Related through ai-ethics
- philosophy
Machine Ethics
Direct archive relation
- wisdom
Objectivity: Meaning, Philosophy & Wisdom
Direct archive relation
- answer
Do machines have rights?
Related through Machine Ethics
- answer
What are robot rights?
Related through Machine Ethics
- answer
What is AI governance?
Related through Machine Ethics
Archive references
Sources
- 01On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?By Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret ShmitchellConsult source
ZHAIBIAN Editorial Board reviewed
Reviewed by ZHAIBIAN AI Editorial Review · 2026-08-17