Who he is
Affiliation: Co-founder & Senior Research Fellow, Machine Intelligence Research Institute (MIRI) — verified current as of late 2025 (MIRI 2025 fundraiser; his own posts); founder of LessWrong; co-author, If Anyone Builds It, Everyone Dies (with Nate Soares, Little Brown, published 2025-09-16; NYT bestseller, 2025-10-05 list).
One-line position: Building superintelligent AI without solving alignment first defaults to human extinction; the rational policy is a hard global moratorium.
Discipline & technical bet
Autodidact decision theorist — no formal credentials, but he founded alignment as a public discipline. His technical wager is a negative one: that RLHF-style behavioral shaping, interpretability, and scalable oversight cannot reach the safety bar before capabilities cross the danger threshold. Note the rhetorical evolution: he now refuses to state a probability at all, arguing p(doom) is a broken instrument, and reframes everything as "what is the minimum necessary and sufficient policy that would prevent extinction?" (LessWrong, 2025-08-22).
Key claims (Says)
- Empirical: the default outcome of building a system smarter than humans is loss of control; alignment must be solved before that threshold. Held — book-length restatement published 2025-09; no softening found.
- Empirical: current alignment techniques (RLHF, RLAIF, constitutional AI) shape behavior, not values, and do not scale to superhuman systems. Held — unchanged.
- Empirical: "sharp left turn" — capabilities can jump discontinuously at superhuman scale, leaving no iterative correction window. Drifted — capability progress 2020–2026 has stayed smooth and correctable (frontier releases arrive on a cadence, get patched, iterate); defenders say the threshold isn't reached, which keeps it unfalsifiable short-run.
- Empirical: inner misalignment (learned objective ≠ training objective) is the central unsolved problem. Held — still load-bearing in the safety literature.
- Empirical: a sufficiently capable AI defeats containment by default — air-gaps, oversight, interpretability all probabilistically lose. Held (as stated position; untested by construction).
- Normative: governments should impose a global moratorium on frontier training above a compute threshold, enforced by treaty and, if needed, physical action against non-compliant data centers (TIME, 2023-03). Held — the book is this argument at trade-press scale.
- Normative: burden of proof is on builders to demonstrate safety. Held.
- Forecast, superseded: P(extinction | unaligned superintelligence) very high, often read as >90%. Drifted — he has retired the metric itself: "Don't use p(doom)" (2025-08-22) argues the number "serves tribal, not epistemic functions" and he refuses to quote one; the operative question is now the minimum-sufficient-policy one. Resolves the v1 [TBD: most recent precise number] — there deliberately isn't one.
Notable predictions — with falsifiable checks
- (2008+) Alignment won't be solved by default; capabilities outpace alignment. Contested; graded per observable: no lab has demonstrated alignment of a system at the capability frontier by any agreed metric — but no uncontrolled takeover either. Check: any frontier-lab "aligned by construction" claim accepted by external reviewers, or any containment failure, by next review.
- (2022-06) "AGI Ruin: A List of Lethalities" — arguments still load-bearing in the safety community. Held as influence claim.
- (2023-03) Moratorium call. No moratorium adopted anywhere as of 2026-09; rhetorical influence > policy influence. Check: any binding compute-threshold treaty or national frontier-training pause by 2027-12-31.
- (2025-09) Book moves the Overton window. Partial: NYT bestseller (2025-10-05), New Yorker/Guardian best-of lists — but reception polarized (Becker: "tendentious"; Marcus: "not nearly as worrying"; New Statesman: "not a serious book"). Check: does any G7 policymaker cite it in an official proceeding by 2027?
- (Multiple) "Slow takeoff unlikely; expect a sharp left turn." Drifted — smoothness continues to accumulate against it; unfalsifiable until/unless threshold. Watch observable: any capability jump that skips an evaluation generation.
Revealed behavior (Does)
- Wrote a mass-market book instead of research papers — revealed strategy: the technical community is lost or converted; the play is now public opinion and policymakers.
- Stopped quoting probabilities (2025-08) — a methodological retreat that also happens to make his position harder to score; note the convenience, but he applies it consistently (won't give a number even when it would help him).
- Stays at MIRI — no job change (checked per validation rule); MIRI's 2025 fundraiser continues the communications-over-research pivot.
- Keeps engaging hostile interviewers and reviewers — he spends attention on persuasion, not insulation.
Feels
Despair discipline: operates as a man who believes he is watching the preventable end approach, and that clarity is the only dignity left. Wants to be wrong; expects not to be.
Hears
LessWrong/rationalist discourse, MIRI colleagues (Soares), and decades of his own prior writing — the most self-referential input diet in the file set; mainstream ML research enters mostly as material to rebut.
Sees
The full argument-space of alignment failure modes, mapped before almost anyone — his comparative advantage is exhaustive adversarial imagination, not empirical data access. No lab vantage; no privileged evals.
Incentive map
Paid by MIRI (donor-funded, doom-salient donors); selling a book whose thesis is its title. Cannot easily say: that things look better than expected; that some lab's safety work is adequate (would collapse the moratorium case). But note the costly signal: he attacks the labs that fund the wider safety ecosystem, and his no-number stance forfeits rhetorical ammunition. Ideological consistency > revenue optimization here; discount for salience-dependence, not for cynicism.
Theories aligned with
- AI alignment / x-risk — the maximalist / MIRI variant
- Bayesian epistemics, decision theory, game-theoretic safety arguments
- Distinct from Bengio's "manageable risk, govern hard" frame and from Amodei's "race to the top" frame
What he's reacting against
- "Default optimism" — the assumption that smarter systems are easier to align with human values
- RLHF / constitutional AI treated as a solution rather than a surface patch
- Lab-led "responsible scaling" frames he reads as motivated cognition by people who want to keep building
- "Iterate in deployment" safety arguments — only works for systems that can't deceive
- Tribal framing of alignment as "just another camp" rather than a load-bearing technical problem
- NEW (2025-08): the p(doom) number culture itself — "weird astrological signs" substituting for models
Where he overlaps / splits (with Rich)
- Overlaps with Bengio, Amodei, Hassabis on x-risk being real; splits sharply on policy (pause vs governed building)
- Splits with Andreessen on basically every claim — the canonical "doomer" foil
- Splits with Altman on whether iterative deployment is a coherent safety strategy
- Splits with Brynjolfsson, Acemoglu, Cowen on framing — labor/distributional questions are downstream of "will we still exist"
- Splits with Mostaque — open-source frontier AI is strictly worse on his view (capability proliferation)
- Overlap with Diamandis, Blundin, Ismail = essentially none; they treat his frame as overconfident, he treats theirs as suicidal
Track record
- Founded the alignment field as a public discipline (2000s LessWrong / Sequences); broad intellectual influence regardless of object-level correctness
- Predicted scaling + general-purpose systems would beat symbolic/GOFAI — broadly correct
- Predicted LLM deception/sycophancy before wide observation — partially vindicated
- Counter-evidence: 2020–2026 capability progress smoother and more correctable than the sharp-left-turn model implies
- Policy: zero moratoriums adopted; Overton-window shift real but unattributable to him alone; the book made him a bestselling author but hardened the critics
Empirical vs normative
- Empirical: alignment unsolved; current techniques don't scale; default outcome is loss of control
- Normative: humans should not build it; governments justified in forcibly preventing builds; safety burden on builders
- The probability claim is now officially withdrawn as a genre by its most famous exponent — treat his risk level as a strongly-held prior expressed through policy demands, not a measurement (CONVENTIONS rule 9)
Weak spots / open questions
- Retiring p(doom) removes the moving-goalpost risk but also removes scoreability — his framework now grades only on policy adoption and capability discontinuities
- "Sharp left turn" remains in tension with observed smooth scaling; unfalsifiable in the short run by design
- Moratorium-by-force is politically infeasible at scale — even allies (Bengio) won't go that far
- Doesn't engage the strongest versions of empirical alignment programs (interpretability, scalable oversight) so much as dismiss their premise
- "Burden of proof on builders" requires an institutional framework he doesn't specify
- The doom argument compresses contested premises (sharp takeoff, deceptive alignment, instrumental convergence) — the conjunction multiplies uncertainty
- Quiet-update watch (carried from v1): the operative frame drifts from "build alignment first" toward "minimize catastrophe given they'll build it" — the book's existence (persuasion, not prevention-by-research) is itself evidence of this
Rich's take
- (your synthesis here)
Delta log
2026-09-04 — v1→v2 migration + validation (weekly batch)
- Grades: 5 Held / 2 Drifted / 0 Wrong.
- Held Core doom argument, techniques-don't-scale, inner misalignment, containment-fails, moratorium stance, builder burden — restated at book length (If Anyone Builds It, Everyone Dies, pub. 2025-09-16; NYT bestseller 2025-10-05; reception sharply polarized — New Yorker best-of vs Becker/Marcus/New Statesman pans).
- Drifted Sharp left turn: another 16 months of smooth, correctable frontier releases accumulated against it; still unfalsifiable short-run.
- Drifted P(doom) >90%: he retired the metric itself — "Don't use p(doom)" (2025-08-22); refuses to quote a number; operative question is now "minimum necessary and sufficient policy to prevent extinction." v1's [TBD: precise number] resolves to: deliberately none.
- Affiliation check (per learnings rule): still MIRI Senior Research Fellow (MIRI 2025 fundraiser); no job change.
- Most surprising delta: the man whose name is synonymous with p(doom) publicly abandoned p(doom) as a concept — and almost nobody in the persona's orbit updated on it.
- Tier: semiannual (positions move on book/essay cadence; no commercial cycle). Wake triggers: new book/major essay (position restatements), MIRI role change (canonical miss), any stated probability/timeline revision (would reverse the no-number stance — big news), moratorium/treaty movement (his win condition).
Sources
- Books: If Anyone Builds It, Everyone Dies (with Nate Soares, Little Brown, 2025-09-16; NYT bestseller list 2025-10-05); Rationality: From AI to Zombies (2015)
- Essays / posts: "AGI Ruin: A List of Lethalities" (2022-06) · TIME moratorium essay (2023-03-29) · "Don't use p(doom)" (2025-08-22)
- Org: MIRI 2025 fundraiser · MIRI (Wikipedia)
- Reception: reviews catalogued in the book's Wikipedia article (New Yorker, Guardian, Kirkus vs Becker/Atlantic, Marcus, New Scientist, New Statesman)
- Podcasts: Lex Fridman (2023-03, ep 368); Bankless [TBD: date]; Dwarkesh Patel [TBD: date]
- Earlier work: "Coherent Extrapolated Volition" (2004); "AI as a Positive and Negative Factor in Global Risk" (2008, in Bostrom & Ćirković eds.)
migrated v1→v2 2026-09-04 (weekly validation batch) · backup: _archives/eliezer-yudkowsky.html.bak-20260904 · AI & Society domain