Stuart Russell

persona · active · confidence: medium · reviewed 2026-09-14 · tier: semiannual

Who he is

Affiliation: Professor of Computer Science, UC Berkeley (faculty since 1986); founder & director, Center for Human-Compatible AI (CHAI, 2016+); co-author with Peter Norvig of Artificial Intelligence: A Modern Approach (1995; 4th ed. 2020); UN, UK, OECD AI policy advisor. No role change found — still speaking as Berkeley professor at the AI Impact Summit, New Delhi (2026-02); employer check run 2026-09-14.
One-line position: AI risk is real but tractable through architectural redesign — build systems that are uncertain about their objective and learn human preferences rather than optimize a fixed spec. New since v1: the sharpest anti-race language of his career — private companies betting humanity in "Russian roulette" is "a serious government failure," and CEOs can't stop unilaterally because "they would be ousted by investors."

Discipline & technical bet

Academic AI (probabilistic reasoning, planning) turned safety architect. Standing technical wager: assistance games / CIRL — machines uncertain about objectives, deferring to human preference evidence — as the replacement for fixed-objective optimization. That bet remains unadopted at the frontier: RLHF-plus-agents is still the shipped paradigm, and no lab has trained a flagship on assistance-game architecture. His live levers are governance (LAWs process, red-lines advocacy) rather than code.

Key claims (Says)

Notable predictions — with falsifiable checks

Revealed behavior (Does)

Feels

Frustration of the engineer whose fix is on the shelf while the building burns: he believes a tractable redesign exists and watches the field scale the broken architecture instead. The "less than a human being" language signals the worry has widened from control failure to human diminishment.

Hears

CHAI researchers and Berkeley colleagues, the UN/CCW diplomatic circuit, FLI and safety-statement coalitions, and textbook-generation academics worldwide; thin input from frontier-lab product teams and from the abundance camp.

Sees

Three decades of the field's conceptual foundations (he wrote the textbook), plus the inside of the LAWs diplomatic process — he sees how slowly binding governance actually moves, which grounds his "government failure" charge in procedural experience rather than rhetoric.

Incentive map

Cleanest map in this batch: salaried academic; CHAI philanthropic funding, book royalties, and speaking fees are the residual incentives — none reprice on capability optimism. Cannot easily say: "the standard model turned out fine" (30 years of framing invested) or "assistance games failed" (his program). Rule-25 weight: high credibility on structural claims; discount only for paradigm-defense bias on CIRL's prospects.

Theories aligned with

What he's reacting against

Where he overlaps / splits (with Rich)

Track record

Empirical vs normative

Weak spots / open questions

Rich's take

Delta log

2026-09-14 — v1→v2 migration + validation (weekly batch)

Sources

migrated v1→v2 2026-09-14 · backup: _archives/stuart-russell.html.bak-20260914 · converts-from: personas/stuart-russell.md · AI & Society domain