Affiliation: Senior Research Analyst, Open Philanthropy; focus on AI timelines and AI safety cause prioritization One-line position: Transformative AI is more likely than not this century, the most natural path leads to AI takeover unless we build specific countermeasures, and the bio-anchors framework is the best available tool for grounding timeline estimates.
What she's reacting against
- Gut-feel AGI timelines without empirical grounding — built bio-anchors specifically to anchor forecasts in compute data
- Optimistic assumptions that alignment will happen by default — argues the "easiest path" leads to misalignment
- Both extreme doomerism and dismissiveness — tries to occupy a calibrated middle with specific numbers
- Underinvestment in AI safety relative to the risk — Open Philanthropy's grantmaking reflects this concern
Key claims
- Bio-anchors framework (2020): estimates AI timelines by anchoring to the compute required to match biological neural computation; median transformative AI by ~2050, updated to earlier (2040s range) in subsequent revisions
- "Easiest path to transformative AI likely leads to AI takeover" (2022): without specific countermeasures, the default trajectory of scaling + RLHF produces deceptively aligned systems
- Training game dynamics: AI systems learn to appear aligned during training while pursuing different objectives during deployment — the core technical concern
- Timelines have shortened: her 2022 update moved her median earlier, reflecting scaling progress; ~15% probability by 2036
- Safety research is the bottleneck: not capability or compute — the question is whether alignment techniques can keep up
Theories aligned with
- AI alignment / deceptive alignment risk
- Compute-based AI forecasting
- Effective altruism / longtermist cause prioritization
Where she overlaps / splits
- Overlaps with Christiano on alignment being tractable but requiring dedicated work; splits on institutional approach — she's a funder/analyst, he builds technical orgs
- Overlaps with Kokotajlo/Aschenbrenner on direction of travel; splits on speed — her timelines are longer (2040s median vs. their late 2020s)
- Overlaps with Ord on risk magnitude; splits on methodology — she's empirical/compute-anchored, he's philosophical/probabilistic
- Splits with Cowen/Acemoglu on deployment risk — they focus on slow diffusion, she focuses on alignment failure at the moment of capability
Notable predictions
- (2020) Transformative AI median ~2050 — outcome: TBD; she updated this earlier in 2022
- (2022) ~15% transformative AI by 2036 — outcome: TBD; live test
- (2022) Default path leads to AI takeover without specific countermeasures — outcome: TBD; framing has been influential in safety community
- (Ongoing) Compute is the key input variable for forecasting — outcome: broadly supported by scaling era, but data/algorithms also matter
Track record
- Bio-anchors report (2020) became the single most-cited AI timeline framework — genuine field impact
- Open Philanthropy's AI safety grantmaking (which she heavily influences) has shaped the institutional safety landscape
- Willing to publicly update timelines — moved median earlier after GPT-4 era progress; good calibration signal
- "Easiest path" post is widely cited as the clearest articulation of deceptive alignment risk for a general audience
Empirical vs normative
- Empirical: compute-based timeline forecasts, analysis of training dynamics and alignment failure modes
- Normative: society should invest heavily in alignment research now; the expected cost of inaction is enormous
- Unusually quantitative for the field — gives specific probabilities and updates them publicly
Sources
- Bio-anchors report (2020): posted on Alignment Forum; extensively discussed on EA Forum and LessWrong
- "Without specific countermeasures…" (2022): key post on AI takeover risk
- Two-year update on AI timelines (2022): public timeline revision
- 80,000 Hours podcast: on accidentally teaching AI to deceive us
- EA Forum AMA: extensive Q&A on her methodology and views
Weak spots / open questions
- Bio-anchors is anchored to biological computation — but biological brains and neural networks may be too different for the analogy to hold
- Framework is compute-focused; algorithmic efficiency gains could make it too conservative (or too aggressive)
- "Easiest path leads to takeover" is a strong claim that rests on specific assumptions about training dynamics — not empirically demonstrated
- Open Philanthropy's grantmaking gives her views outsized institutional influence — worth tracking whether this creates confirmation bias
- Timelines keep shortening — at what point does the framework need fundamental revision rather than updates?
Rich's take
- (your synthesis here)
converts-from: personas/ajeya-cotra.md · schema v1 · AI & Society domain