Affiliation: Co-founder, Anthropic; leads research / science-of-scaling [TBD: exact current title]; previously OpenAI (co-led scaling-laws and large-batch-training work); theoretical physics PhD [TBD: institution — likely Stanford, verify] One-line position: The behavior of large models under training is a lawful, predictable physical system — and that predictability is itself the foundation of safety, because you can forecast a model's capabilities (and risks) before you build it.
What he's reacting against
- The view that deep-learning training is unpredictable "alchemy" — his career argues the macro dynamics are lawful even when the micro is opaque
- Trial-and-error scaling — argues you should be able to predict the right batch size, compute budget, and resulting loss before a run, not discover them empirically each time
- Treating safety and capabilities as separable workstreams — predictability of training is offered as the bridge between them
- Scaling skepticism (LeCun / Marcus lane) — the smoothness of the curves is the empirical counter-bet
- Surprise as an acceptable operating mode for frontier labs — frames "no surprises" forecasting as a precondition for responsible scaling
Key claims
- Large-batch training has a predictable optimum: the "gradient noise scale" predicts the largest useful batch size across tasks, from supervised learning to RL (empirical; lead author, "An Empirical Model of Large-Batch Training," 2018-12)
- Scaling laws are empirical regularities: model loss falls as a smooth power law in compute, data, and parameters over many orders of magnitude (empirical; co-author, "Scaling Laws for Neural Language Models," 2020-01)
- Training dynamics are a science, not a craft: the field can build predictive models of how networks learn, not just that they do — the same epistemic move physics makes (interpretive frame)
- Predictability is a safety property: if capabilities can be forecast from compute budgets, then dangerous capabilities can be anticipated and gated before deployment — this underwrites Anthropic's Responsible Scaling Policy (normative + empirical bridge)
- Architecture matters weakly, scale matters strongly: shared scaling-laws finding — total scale dominates structural details (empirical)
- Safety-motivated labs should hold the frontier: shared Anthropic founding premise — if scale is the path, careful actors should lead it rather than cede it (normative)
- The curves have not bent off-trend: continued frontier gains through 2024–2025 are consistent with extrapolation, including the inference-time-compute extension [TBD: direct public statement] (empirical, contested)
- Forecasting capabilities is tractable enough to act on: capability prediction is good enough to inform governance decisions, not merely an academic exercise (normative claim about decision-relevance)
Theories aligned with
- AI Alignment / X-Risk — moderate, technical-tractability variant; predictability-as-safety lens
- Adjacent to Techno-Optimism — conditional, safety-gated; optimism located in the trend lines, not the rhetoric
- White-Collar Labor End — short-timeline knowledge-work automation follows from the same curves (implicit; thinner public articulation than Kaplan)
- Adjacent to Jevons' Paradox — cheaper, more predictable compute drives more total compute use
Where he overlaps / splits
- Overlaps with Jared Kaplan almost completely — co-equal authors on scaling laws; the split is emphasis: Kaplan anchored the loss-vs-scale law and Constitutional AI; McCandlish anchored the training-dynamics side (batch size, optimization, the science of the run itself)
- Overlaps with Tom Brown on scaling-as-predictable-engineering — Brown supplies the compute/infrastructure, McCandlish the predictive science that says what that compute will buy
- Overlaps with Dario Amodei on safety-gated frontier strategy; splits in register — Dario makes the civilizational case, McCandlish the quantitative/forecasting one
- Complementary with Chris Olah — Olah works backward from a trained model (interpretability); McCandlish works forward to predict it before training (forecasting); two halves of "understand the model"
- Complementary with Noam Brown — McCandlish's laws govern pretraining dynamics; Brown argues inference-time search is a separate scaling axis — agreement on lawfulness, open question on which axis dominates
- Overlaps with Leopold Aschenbrenner on straight-line extrapolation; splits on the geopolitical race framing — McCandlish stays near the curves
- Splits with Yann LeCun on whether scaled autoregressive models reach general capability — the cleanest empirical crux in the roster
Notable predictions
- (2018-12) Gradient noise scale predicts the critical batch size across domains — outcome: largely correct; the diagnostic was adopted as a practical tool for planning large runs
- (2020-01) Loss follows a power law across further orders of magnitude — outcome: largely correct through at least 2024; Chinchilla (2022-03) revised the optimal data/parameter ratio but kept the power-law framework intact
- (2023–2025) Frontier gains stay on trend; no scaling wall — outcome: largely correct to date; reasoning-model paradigm (2024-09 onward) arguably extended rather than broke the trend [TBD: verify a directly attributed statement]
- (Ongoing) Capabilities are forecastable enough to gate via RSP thresholds — outcome: TBD; the live institutional test of predictability-as-safety
Track record
- Lead author, "An Empirical Model of Large-Batch Training" (2018-12) — introduced the gradient noise scale; a genuinely distinct technical contribution from his scaling-laws co-authors
- Co-author, "Scaling Laws for Neural Language Models" (2020-01) — among the most consequential empirical results of the deep-learning era; set frontier compute strategy
- Co-founded Anthropic (2021) and built/leads its research science function — Claude staying frontier-competitive is the execution evidence
- Caveat: Chinchilla (2022-03) showed the original scaling exponents were not final — the framework held, the constants moved; relevant when weighing current extrapolations
Empirical vs normative
- Empirical: gradient noise scale, scaling laws, training-dynamics predictability, no-wall claims — his strongest ground
- Normative: predictability should gate scaling; safety-focused labs should lead the frontier — asserted partly via Anthropic's institutional choices
- Anthropic co-founder with equity at stake — capability and timeline claims carry commercial interest per CONVENTIONS rule 25; pre-Anthropic academic physics work partially mitigates but does not remove this
Sources
- Papers: "An Empirical Model of Large-Batch Training" (2018-12, arXiv:1812.06162, lead author); "Scaling Laws for Neural Language Models" (2020-01, arXiv:2001.08361, co-author with Kaplan) — preprint/peer tier
- Anthropic materials: company research posts and Responsible Scaling Policy documentation [TBD: links]
- Talks / interviews: public footprint unusually thin for a co-founder [TBD: locate podcast or conference appearances]
- Background: theoretical physics publications [TBD: confirm field and institution — context for the epistemic style, not AI claims]
- Caveat: persona leans on papers and institutional behavior more than direct personal statements; lower confidence on stated personal views than for public-facing peers (same limitation as the Tom Brown file)
Weak spots / open questions
- Scaling laws and training-dynamics laws predict loss and optimization behavior, not capabilities — the loss-to-capability mapping remains the weak link (the emergence debate sits exactly here)
- "Predictability-as-safety" is contested: critics (Yudkowsky lane) argue that if danger is forecastable, that is a reason to pause, not a license to keep scaling
- Chinchilla correction is a caution — the published scaling constants were wrong once; current extrapolations could be similarly revisable
- Thin public record — many positions are inferred from papers and Anthropic's institutional voice; risk of projecting the company onto the individual
- Largely silent on economics and distribution — no known position on labor, displacement, or who captures the gains
- Does the science-of-scaling frame survive the shift to inference-time compute and data constraints, or does the predictive model need rebuilding for the new regime? No clear public answer found
Rich's take
- (your synthesis here)
converts-from: personas/sam-mccandlish.md · schema v1 · AI & Society domain