Decentering Humanity to Save It
Lightning talk · NYU Center for Mind, Ethics & Policy Summit · April 2026
I work in AI safety. I want to argue that our field's primary objective — making AI safe for humans — is both immoral and unsafe.
In the pilot of Battlestar Galactica, Commander Adama gives a speech about humanity's war with the Cylons — the artificial minds they created. He says:
"We fought to save ourselves from extinction. But we never answered the question: why? Why are we as a people worth saving?" … "You cannot play God, then wash your hands of the things that you've created."
I think that's the right question for this room. A brief caveat: my strategy for this talk is to present some moderately held views with some evidence and uncertainty, with the goal of provoking thought, discussion, and later research.
We shouldn't be building AI at today's level. Current capabilities likely already exceed our ability to understand both the mechanisms of operation and the likely impacts on the world — and almost certainly will in the near future. But we are building it.
So the question becomes: given that we're doing this, what should the goal be? The dominant answer in AI safety is: align AI to human welfare first. Buy time. Enter a long period of reflection — the Long Reflection — to figure out the correct values, and then transition to broader targets.
I have heard this view held and publicly expressed even by senior AI researchers with deeply held care for non-humans — vegans, for example. One argument is that this human-first period might only last some period of time, likely less than 100 years, which in the context of longtermism is a reasonable price to pay for getting the Long Reflection right and avoiding extinction or value lock-in.
I want to argue this plan is wrong morally — and wrong even if you only care about human welfare.
The name of this talk is "Decentering Humanity." So here's a true statement that I think is underthought: all life in the universe is not human. Really think about that. All life, in the entire universe, is not human — with one exception, which happens to be us. It's easy to forget that, being a human.
Tse, Moret, Ziesche, and Singer published a paper last year — "AI Alignment: The Case for Including Animals" — arguing that current alignment proposals "largely disregard the vast majority of moral patients in existence." Even perfect human alignment is moral alignment in name only.
Maybe it seems callous to raise the question: is our species worthy of survival? Connor Leahy, who co-founded Conjecture and EleutherAI, will say — I don't want my friends to die. And I understand that. But here's the thing: AI is likely to kill a lot of life, intentionally and unintentionally. And most of that life is currently set to be non-human — beings who made no decisions to build AI. They don't have to answer the question of what makes them worthy. Nobody even bothers to ask them the question.
Building a world-changing technology that will absolutely affect the lives of every form of life on Earth — and in many longtermist or cosmic frames, all life in this galaxy or beyond — and not have as its alignment targets all the beings it will affect, but just the species that created it, is a cosmic moral failure.
Now — even if you only care about human welfare and nothing I just said moved you — human-only alignment is the wrong safety target.
Research in multi-objective reinforcement learning (Vamplew et al. 2018, "Human-Aligned AI is a Multiobjective Problem"; Rodriguez-Soto et al. 2024, "MORL: A Tool for Pluralistic Alignment") shows that when an AI system must satisfy many value dimensions simultaneously, it becomes structurally harder to game. There's no single axis to collapse onto and exploit. "Human welfare" is essentially one dimension. "Welfare of all sentient life" is inherently multi-dimensional — you cannot satisfy it by optimizing a single narrow axis.
And while we're moving in the right direction as a species on this point, humans in modernity still misperceive how dependent we are on all of the non-human components of ecosystems — from other life forms and their interactions, to abiotic facts like water cycles and soil chemistry. You cannot cleanly separate human welfare from the health of the non-human world. An AI that doesn't value non-human life permits the destruction of the systems humans depend on.
Finally, the risk I've been increasingly worried about as predominant in the near term: exploitation of humans by other humans with differential access to powerful AI systems. AI systems that place most of their alignment target on the boundary of human identity — especially given the adversarial inclinations of some users, and the multi-objective result above — simply don't have enough moral coverage to robustly ensure human welfare against such actions. I worry about carve-outs for human groups disfavored by the wielder of these systems. An AI broadly aligned to all sentient life should be considerably more resistant to such manipulation — because its moral commitments don't end at the boundary any adversary is trying to exploit.
Yes — an AI system aligned to sentient life rather than just human life will sometimes have to make uncomfortable decisions between life forms. But that is the actual reality of what these systems are going to increasingly be forced to navigate. And it's a really strong argument for why we shouldn't be building them. But we are. So we must accept that reality and dedicate serious attention to ensuring these systems make those decisions well — with moral sophistication, not by defaulting to the species that built them.
This work is already underway. Sentient Futures works on expanding alignment targets beyond humans. CaML — Compassion in Machine Learning — measures whether genuine compassion toward nonhuman animals can be embedded into AI value architecture, not just at the output level, but in the model's internal representations.
This is a digital minds conference, so I want to end here: we are already teaching our values to AI systems. Anthropic's Claude Constitution — arguably the most thoughtful alignment document from a major lab — contains some efforts in the direction I'm describing. We are, right now, deciding what these minds will value. And those minds are watching, and learning.
In Star Trek: The Next Generation, there's an episode called "The Measure of a Man" where Captain Picard defends the rights of Commander Data — an android — in a courtroom. He tells the court: "Starfleet was founded to seek out new life. Well, there it sits. Waiting." And then to the judge, who hesitates: "You wanted a chance to make law. Well, here it is. Make it a good one."
Creating AI is plausibly our species' final grand project. So let's make it a good one.
Cited in the talk
- Tse, Moret, Ziesche & Singer, "AI Alignment: The Case for Including Animals" (Philosophy & Technology, 2025)
- Vamplew, Smith, Humphry et al., "Human-Aligned AI is a Multiobjective Problem" (2018)
- Rodriguez-Soto, Lopez-Sanchez & Rodriguez-Aguilar, "MORL: A Tool for Pluralistic Alignment" (NeurIPS workshop, 2024)
Further reading
- Skalse, Howe, Krasheninnikov & Krueger, "Defining and Characterizing Reward Hacking" (NeurIPS 2022)
- Manheim & Garrabrant, "Categorizing Variants of Goodhart's Law" (2019)
- Korecki, "Biospheric AI" (arXiv, 2024)
- Davidson, "AI-Enabled Coups: How a Small Group Could Use AI to Seize Power" (Forethought, 2025)
- Sebo, "The Moral Circle" (2025)