Eliezer Yudkowsky

persona · active · confidence: medium · reviewed 2026-09-04 · tier: semiannual

Who he is

Affiliation: Co-founder & Senior Research Fellow, Machine Intelligence Research Institute (MIRI) — verified current as of late 2025 (MIRI 2025 fundraiser; his own posts); founder of LessWrong; co-author, If Anyone Builds It, Everyone Dies (with Nate Soares, Little Brown, published 2025-09-16; NYT bestseller, 2025-10-05 list).
One-line position: Building superintelligent AI without solving alignment first defaults to human extinction; the rational policy is a hard global moratorium.

Discipline & technical bet

Autodidact decision theorist — no formal credentials, but he founded alignment as a public discipline. His technical wager is a negative one: that RLHF-style behavioral shaping, interpretability, and scalable oversight cannot reach the safety bar before capabilities cross the danger threshold. Note the rhetorical evolution: he now refuses to state a probability at all, arguing p(doom) is a broken instrument, and reframes everything as "what is the minimum necessary and sufficient policy that would prevent extinction?" (LessWrong, 2025-08-22).

Key claims (Says)

Notable predictions — with falsifiable checks

Revealed behavior (Does)

Feels

Despair discipline: operates as a man who believes he is watching the preventable end approach, and that clarity is the only dignity left. Wants to be wrong; expects not to be.

Hears

LessWrong/rationalist discourse, MIRI colleagues (Soares), and decades of his own prior writing — the most self-referential input diet in the file set; mainstream ML research enters mostly as material to rebut.

Sees

The full argument-space of alignment failure modes, mapped before almost anyone — his comparative advantage is exhaustive adversarial imagination, not empirical data access. No lab vantage; no privileged evals.

Incentive map

Paid by MIRI (donor-funded, doom-salient donors); selling a book whose thesis is its title. Cannot easily say: that things look better than expected; that some lab's safety work is adequate (would collapse the moratorium case). But note the costly signal: he attacks the labs that fund the wider safety ecosystem, and his no-number stance forfeits rhetorical ammunition. Ideological consistency > revenue optimization here; discount for salience-dependence, not for cynicism.

Theories aligned with

What he's reacting against

Where he overlaps / splits (with Rich)

Track record

Empirical vs normative

Weak spots / open questions

Rich's take

Delta log

2026-09-04 — v1→v2 migration + validation (weekly batch)

Sources

migrated v1→v2 2026-09-04 (weekly validation batch) · backup: _archives/eliezer-yudkowsky.html.bak-20260904 · AI & Society domain