Affiliation: Founder, Alignment Research Center (ARC); head of AI safety, US AI Safety Institute (NIST, as of 2024); ex-OpenAI researcher; UC Berkeley PhD One-line position: AI alignment is a solvable but urgent technical problem; the right approach is rigorous worst-case analysis plus practical evals of frontier models, and ~10–20% existential risk this century is the honest number.
What he's reacting against
- Alignment nihilism (Yudkowsky-style "we're doomed") — believes the problem is tractable with sufficient effort
- Lab self-assessment as sufficient — built ARC specifically to provide independent capability evaluations
- Vague safety claims without technical substance — insists on formal, worst-case reasoning
- The gap between academic alignment theory and frontier-model reality
Key claims
- ~10–20% existential risk from AI this century: one of the most widely cited calibrated risk estimates from a technical insider
- Iterated amplification / scalable oversight: humans can align superhuman systems by building chains of AI-assisted evaluation — his core technical contribution
- RLHF as a foundation: co-developed reinforcement learning from human feedback at OpenAI — the technique underlying ChatGPT-era alignment
- Independent evals are essential: labs cannot credibly evaluate their own models for dangerous capabilities; third-party assessment is a governance requirement
- Alignment tax should be low: if done right, making models safe shouldn't cost much in capability — this is an optimistic structural claim
- AGI timelines: 2030s–2040s median: longer than Kokotajlo/Aschenbrenner but still within policy-relevant window
Theories aligned with
- AI alignment (technical, tractable variant)
- Scalable oversight / iterated amplification
- Independent evaluation as governance infrastructure
Where he overlaps / splits
- Overlaps with Leike on scalable oversight research agenda; splits on institutional approach — Christiano built an independent org, Leike worked inside labs
- Overlaps with Yudkowsky on risk being real and large; splits sharply on tractability — Christiano believes we can solve it, Yudkowsky doubts it
- Overlaps with Dario Amodei on alignment being the right technical bet; splits on whether labs can self-govern (Christiano says no, hence ARC)
- Splits with Cowen/Brynjolfsson on risk magnitude — they treat x-risk as small tail, Christiano puts it at 10–20%
Notable predictions
- (Ongoing) ~10–20% existential risk from AI this century — outcome: not directly testable; track whether estimate moves
- (Ongoing) AGI median in 2030s–2040s — outcome: TBD
- (2022) Independent evals will become a governance requirement — outcome: partially vindicated (NIST involvement, executive orders)
- (Ongoing) Alignment research can stay ahead of capability if resourced — outcome: TBD
Track record
- Co-developed RLHF at OpenAI — now the dominant alignment technique in production, widely vindicated
- Founded ARC (2021) — established the most credible independent evals organization
- Appointed to lead AI safety at NIST (2024) — institutional influence is real
- Iterated amplification framework is widely cited but not yet proven at superhuman scale
Empirical vs normative
- Empirical: risk estimates, capability forecasts, technical alignment feasibility
- Normative: independent evaluation is a governance necessity; labs should not self-assess; alignment research deserves major investment
- Unusually calibrated for the field — gives specific probability estimates and updates them
Sources
- ARC website: alignment.org
- Key papers: "AI Alignment Landscape" (2020); RLHF work at OpenAI; iterated amplification papers
- LessWrong / Alignment Forum: extensive technical writing
- NIST role: appointed head of AI safety, US AI Safety Institute (2024)
- Podcasts/interviews: 80,000 Hours, Bankless, various EA forums
Weak spots / open questions
- RLHF is widely used but widely criticized as shallow — does behavioral alignment scale to superhuman systems?
- ARC evals have influence but limited enforcement power — what happens when a lab disagrees with the assessment?
- "Alignment tax is low" is an optimistic assumption that may not hold for truly superhuman systems
- Government role (NIST) may constrain his public communication and independence
- 10–20% x-risk is a big number to hold while also believing the problem is tractable — tension worth tracking
Rich's take
- (your synthesis here)
converts-from: personas/paul-christiano.md · schema v1 · AI & Society domain