Chris Olah

persona · active · confidence: medium · reviewed 2026-06-04

Affiliation: Co-founder, Anthropic (interpretability lead); previously Google Brain and OpenAI (Clarity team); co-founder of the Distill journal One-line position: We are deploying systems we do not understand; mechanistic interpretability — reverse-engineering the internal computations of neural networks — is the load-bearing path to making them safe.

What he's reacting against

Key claims

Theories aligned with

Where he overlaps / splits

Notable predictions

Track record

Empirical vs normative

Sources

Weak spots / open questions

Rich's take

converts-from: personas/chris-olah.md · schema v1 · AI & Society domain