Affiliation: Ex-Anthropic alignment researcher (departed 2026); ex-OpenAI superalignment co-lead (resigned May 2024); PhD in machine learning One-line position: Alignment is the central technical challenge of our time, frontier labs are systematically under-investing in it, and the people closest to the problem keep leaving because they can't fix it from inside.
What he's reacting against
- Lab leadership treating safety as a PR function rather than a core engineering constraint
- OpenAI's dissolution of the superalignment team (2024) — triggered his resignation
- The assumption that scaling alone will solve alignment — argues it requires dedicated research at scale
- Complacency about the gap between capability progress and alignment progress
Key claims
- "Safety has taken a backseat to shiny products": his resignation statement from OpenAI (May 2024) — the single most-cited insider critique of frontier-lab safety culture
- Superalignment is tractable but under-resourced: believes the technical problem (aligning systems smarter than humans) is solvable with sufficient dedicated effort
- Scalable oversight is the key research agenda: humans can't directly supervise superhuman systems, so we need AI-assisted evaluation chains
- Weak-to-strong generalization: can we use weaker models to align stronger ones? Central research question he brought to Anthropic
- Serial departure pattern is the signal: left OpenAI, then left Anthropic — the pattern itself is his strongest argument that labs aren't doing enough
Theories aligned with
- AI alignment / superalignment (technical, optimistic-if-resourced variant)
- Scalable oversight and iterated amplification (Christiano lineage)
- Adjacent to interpretability research but focused on oversight rather than understanding
Where he overlaps / splits
- Overlaps with Christiano on scalable oversight as the core research agenda; splits on institutional approach — Christiano built an independent org (ARC), Leike tried to work inside labs
- Overlaps with Kokotajlo on labs deprioritizing safety; splits on response — Leike stayed constructive longer, Kokotajlo went adversarial earlier
- Overlaps with Dario Amodei on alignment being tractable; splits on whether Anthropic is actually doing enough
- Splits with Yudkowsky on tractability — Leike believes alignment is solvable, Yudkowsky is more pessimistic
Notable predictions
- (2024) OpenAI will not maintain serious safety investment without external pressure — outcome: supported by subsequent trajectory
- (Implicit) Anthropic would be different from OpenAI on safety prioritization — outcome: his own departure in 2026 suggests the answer was "not enough"
- (Ongoing) Scalable oversight is the most promising alignment research direction — outcome: TBD; active research area
Track record
- Co-led OpenAI's superalignment team — had direct visibility into frontier safety work
- Joined Anthropic as the highest-profile safety hire of 2024 — lent significant credibility
- Departure from Anthropic (2026) was arguably more significant than the OpenAI exit — suggests the problem is structural, not org-specific
- Technical publications on RLHF, reward modeling, and oversight are widely cited
Empirical vs normative
- Empirical: technical claims about alignment difficulty, scalable oversight feasibility
- Normative: labs have a moral obligation to invest seriously in safety; commercial incentives are insufficient
- His serial departures function as empirical evidence for his normative claim
Sources
- OpenAI resignation statement (May 2024): "safety has taken a backseat to shiny products"
- Anthropic announcement (May 2024): joining to continue superalignment mission
- Technical work: RLHF papers, reward modeling, scalable oversight research
- Anthropic departure (early 2026): warning that the world is "in peril"
Weak spots / open questions
- Serial departure could indicate unrealistic expectations rather than lab failure — how much safety investment is "enough"?
- Technical research agenda (scalable oversight) is promising but unproven at the scale that matters
- Less public communication than other safety voices — impact is through institutional moves, not arguments
- Hasn't built an independent institution (unlike Christiano/ARC) — critique without an alternative organizational model
Rich's take
- (your synthesis here)
converts-from: personas/jan-leike.md · schema v1 · AI & Society domain