Affiliation: Director of Strategy and Foundational Research Grants, Center for Security and Emerging Technology (CSET), Georgetown University (2019+); former OpenAI board member (2021-09 to 2023-11); previously Open Philanthropy (AI policy); Australian background, Melbourne University One-line position: Frontier AI labs cannot be trusted to self-regulate; "trust me" is not a governance strategy; the institutional design problem is at least as load-bearing as the technical alignment problem.
What she's reacting against
- "Move fast and figure out governance later" framing common in frontier labs — argues the OpenAI 2023-11 episode is direct evidence this fails
- Lab-led safety frameworks (RSPs, voluntary commitments) presented as sufficient — argues they're necessary but not sufficient
- "Trust me, I'm a good actor" framing from lab leadership — argues the question is institutional, not personal
- Pure-technical framings of AI safety that ignore organizational and political dynamics
- Doomer framings that treat governance as either irrelevant or hopeless — argues there is concrete institutional work to do
- US-vs-China zero-sum framings that crowd out the possibility of bilateral coordination
Key claims
- Empirical: frontier labs face structural commercial pressure that erodes safety commitments over time — corporate governance has to be designed against this, not assumed away
- Empirical: the OpenAI 2023-11 firing/return episode demonstrated that mission-aligned governance structures (nonprofit board with capability to fire CEO) can fail under commercial pressure when employee equity, customer continuity, and investor leverage all push the same direction
- Normative: external accountability — third-party audits, government evals, mandatory incident reporting — is required because internal accountability has repeatedly failed
- Empirical: compute governance (export controls, training-run reporting thresholds, chip-level provenance) is the most tractable lever for frontier AI policy because compute is concentrated, measurable, and bottlenecked by a small number of fabs
- Empirical: China's frontier AI capabilities are real but often overstated in US discourse; bilateral track-2 dialogue on catastrophic risk is feasible and underway
- Mixed: licensing regimes for frontier models are reasonable in principle but vulnerable to incumbent capture; design matters more than the binary "regulate / don't regulate"
- Normative: AI governance scholarship needs more practitioners with insider knowledge of how labs actually operate, not just academics observing from outside
- Empirical: lab-funded "safety" research is structurally compromised — funding source shapes what questions get asked even when individual researchers are honest
Theories aligned with
- AI alignment / x-risk — institutional / governance variant, distinct from MIRI (Yudkowsky) and CHAI (Russell) technical variants
- AI safety governance / institutional design — adjacent to Bengio but more focused on US domestic and lab-level institutions
- Compute governance as a tractable policy lever — distinct enough to be its own thread
Where she overlaps / splits
- Overlaps with Bengio on need for external governance and on lab self-regulation being insufficient; splits on emphasis — Bengio more focused on international institutions and Scientist AI research direction, Toner more focused on US domestic accountability and corporate governance
- Overlaps with Russell on x-risk being real and on the inadequacy of current technical safety; splits on the locus of the fix — Russell wants architectural redesign (CIRL); Toner wants institutional/governance redesign
- Overlaps with Hinton on alarm about pace; splits on programmatic specificity — Toner provides concrete institutional proposals where Hinton is more declarative
- Splits sharply with Andreessen on whether AI governance is mostly incumbent capture — Toner argues the OpenAI episode is direct evidence that lab self-governance fails, not that external governance fails
- Splits with Altman publicly and personally — was on the board that fired him 2023-11; her post-firing writing is among the sharpest insider critiques of OpenAI's drift from nonprofit mission
- Splits with Amodei, Hassabis, Suleyman on whether RSPs / containment frameworks are sufficient without external accountability mechanisms with teeth
- Splits with Yudkowsky on framing — Toner thinks governance work is tractable and worth doing; Yudkowsky treats it as insufficient and argues for moratorium
- Less engaged with Brynjolfsson, Acemoglu, Cowen on labor questions — sees them as parallel rather than primary
- Splits with Gebru on framing — Toner takes catastrophic / x-risk framing seriously; Gebru argues it crowds out present-harm accountability. Both agree lab self-regulation is inadequate but disagree on what "AI safety" centrally means
Notable predictions
- (2023-10) Co-authored "Decoding Intentions" CSET paper that drew Altman's reported objection — the paper argued frontier-lab "voluntary commitments" were weak signals of safety effort. Outcome: directionally vindicated by 2023-11 governance episode and by subsequent evolution of frontier-lab safety commitments [TBD: cite specific paper section]
- (2023-11) OpenAI board exercised firing authority over CEO citing "not consistently candid" — outcome: occurred; reversal within 5 days; subsequent governance restructuring confirms structural fragility of the original nonprofit board design
- (2024-04, "What Went Wrong at OpenAI" / TED Talk) Frontier labs will continue to weaken safety commitments under competitive pressure absent external mechanisms — outcome: TBD; consistent with 2024–25 trajectory of lab safety team departures and capability releases [TBD: catalog specific 2024–26 examples]
- (Multiple, 2023+) Compute governance will become a major policy lever — outcome: directionally correct; US BIS export controls (2022-10, expanded 2023-10), entity-list expansions, training-run reporting in EO 14110 (2023-10) [TBD: 2026 status given EO repeal and replacement]
- (2024+) US-China track-2 dialogue on catastrophic AI risk is possible despite headline tensions — outcome: TBD; some confirmed track-2 meetings (Shanghai, Bletchley follow-ups) [TBD: catalog]
Track record
- CSET strategy lead since 2019 — institution has become one of the most-cited US AI policy shops; durable institutional contribution
- OpenAI board 2021-09 to 2023-11 — was inside the room for the highest-profile lab governance episode of the period
- Co-authored "Decoding Intentions" (2023-10) — paper used frontier-lab voluntary safety commitments as a case study; reportedly part of the friction that preceded the November firing
- Post-firing public writing (Economist op-ed 2024-05; TED Talk 2024-04) — among the sharpest insider critiques of nonprofit-to-capped-profit drift
- Pattern: rare combination of insider experience + scholarly framing + willingness to publish critically; less prolific than Bengio on technical research but more institutionally embedded than most external critics
Empirical vs normative
- Empirical: lab self-regulation has empirically failed (2023-11 case); compute governance is concentrated and measurable; track-2 US-China dialogue is feasible
- Normative: external accountability is required; governance scholarship should include practitioners with insider experience; lab-funded safety research is structurally compromised regardless of researcher integrity
- Less commercial conflict than frontier-lab CEOs (CONVENTIONS rule 25) — academic / think-tank position; CSET is funded primarily by Open Philanthropy and other AI-safety-aligned funders [TBD: confirm 2026 funding mix], which is its own (smaller) bias to flag
Sources
- Peer-reviewed / institutional reports: CSET reports under her authorship and editorship (2019+) [TBD: catalog 5–10 most-cited]; "Decoding Intentions: Artificial Intelligence and Costly Signals" (Toner & Imbrie & others, CSET, 2023-10) [TBD: confirm exact authorship]
- Op-eds: "What Went Wrong at OpenAI" The Economist (2024-05) [TBD: confirm exact title / date]; multiple TIME, Foreign Affairs pieces [TBD: catalog]
- Talks: TED Talk 2024 — "How to govern AI — even if it's hard to predict" [TBD: confirm exact title]; CSET public events (ongoing)
- Testimony: US Senate (multiple, 2022+) [TBD: specific dates and committees]; UK AI Safety Summit (2023-11)
- Podcasts: TED Interview with Bilawal Sidhu (2024) [TBD: confirm]; The Logan Bartlett Show, Lawfare, multiple AI-policy podcasts (2023–24) [TBD: dates]
- Earlier work: Open Philanthropy AI policy research (~2018–19); studies of China's AI ecosystem in collaboration with Jeffrey Ding [TBD: specific publications]
Weak spots / open questions
- Insider-experience-as-credibility cuts both ways — some critics frame the OpenAI board episode as evidence that the board was the failure mode, not Altman; her post-firing writing has not fully engaged with the strongest version of that counter-narrative
- "Compute governance is the tractable lever" thesis is vulnerable to (a) algorithmic efficiency gains making compute thresholds obsolete and (b) China's domestic chip ecosystem (SMIC, Huawei Ascend) reducing leverage of US export controls — track both
- Less specific than Russell or Bengio on the technical alignment side; her framing is institutional-first, which is a feature for governance work but a gap if foundational technical assumptions are wrong
- CSET's funding base (heavily Open-Philanthropy-aligned) creates a less-flagged bias: research portfolio over-indexes on catastrophic / x-risk framings vs present-harm framings (Gebru critique applies)
- Rhetorical position post-firing is necessarily shaped by the personal stakes of that episode — track whether her policy positions today are downstream of the substantive analysis or of the board fight
- Doesn't fully engage with Mostaque-style argument that compute governance is itself a concentration mechanism that benefits incumbents
- Track record on dated falsifiable predictions is thinner than e.g. Cowen — she writes more in the framework / institutional-design register than the forecast register
Rich's take
- (your synthesis here)
converts-from: personas/helen-toner.md · schema v1 · AI & Society domain