Affiliation: Davis Professor, Santa Fe Institute; author of Artificial Intelligence: A Guide for Thinking Humans (2019); research on analogy, abstraction, and conceptual understanding in AI One-line position: Current LLMs are impressively capable but do not understand in the way humans do — they lack genuine conceptual abstraction and analogical reasoning, and confusing fluency with comprehension is the central error of the current AI discourse.
What she's reacting against
- The conflation of language fluency with understanding — argues LLMs produce coherent text without forming genuine concepts
- Both hype ("LLMs understand") and dismissal ("LLMs are just stochastic parrots") — argues reality is more nuanced and scientifically interesting
- Benchmark-driven capability claims — shows that LLMs fail on tasks requiring genuine abstraction and analogy even when they pass surface-level tests
- The metaphor problem: the way we talk about AI shapes what we believe about AI — and most metaphors are misleading
Key claims
- LLMs don't do genuine analogical reasoning: her empirical testing (2024–2025) shows that LLMs fail at counterfactual and novel analogy tasks that humans handle easily
- Abstraction is the key gap: human cognition builds abstract concepts from experience; LLMs approximate this through statistical patterns but break down at the edges
- "The Metaphors of Artificial Intelligence" (Science, 2024-11): argues that the metaphors we use for LLMs (individual minds, stochastic parrots, alien intelligence, etc.) shape research directions and policy — none is adequate alone
- Understanding is not binary: the debate over whether LLMs "understand" is poorly framed; there may be degrees and types of understanding, and LLMs occupy a novel position
- Benchmarks don't measure understanding: LLMs can pass benchmarks through pattern-matching without the underlying competence the benchmarks were designed to test
Theories aligned with
- Cognitive science of analogy and abstraction (Hofstadter lineage — she was his PhD student)
- Complexity science / Santa Fe Institute approach to intelligence
- Adjacent to AI capability skepticism but empirically grounded, not ideological
Where she overlaps / splits
- Overlaps with Marcus on LLMs lacking genuine understanding; splits on tone — Mitchell is measured and empirical, Marcus is more polemical
- Overlaps with Kambhampati on "LLMs don't reason"; splits on focus — Mitchell emphasizes analogy/abstraction, Kambhampati emphasizes planning/reasoning
- Overlaps with Chalmers on LLMs lacking key cognitive features; splits on the question asked — Chalmers asks about consciousness, Mitchell asks about understanding
- Splits with scaling optimists (Altman, Amodei) who expect more compute to close the understanding gap — Mitchell argues the architecture may be insufficient
- Splits with Bubeck ("Sparks of AGI") — Mitchell would argue the "sparks" are pattern-matching, not genuine intelligence
Notable predictions
- (Ongoing) Scaling alone will not produce genuine understanding — outcome: TBD; the central empirical question of the scaling era
- (2024–2025) LLMs will continue to fail at novel analogical reasoning tasks — outcome: supported by her testing; frontier models still fail at counterfactual analogies
- (Implicit) The AI field needs better evaluation methodology — outcome: growing consensus (see Liang/HELM, Kapoor/Narayanan)
Track record
- Artificial Intelligence: A Guide for Thinking Humans (2019) — widely praised as the most accessible and balanced overview of AI for general audiences
- Published in Science, PNAS, ICML, CogSci on LLM understanding — peer-reviewed empirical work, not just commentary
- Santa Fe Institute affiliation brings genuine complexity-science rigor — she's not a pundit
- Hofstadter lineage (PhD advisor: Douglas Hofstadter) gives deep theoretical grounding in analogy and cognition
Empirical vs normative
- Primarily empirical: runs experiments, publishes peer-reviewed results on what LLMs can and cannot do
- Normative content is subtle: we should be honest about what AI can do; we should resist metaphor-driven overconfidence; evaluation methodology matters
- Source quality: high — peer-reviewed publications in top venues; this is serious science, not op-ed commentary
Sources
- Book: Artificial Intelligence: A Guide for Thinking Humans (2019)
- Papers: "The Debate Over Understanding in AI's Large Language Models" (PNAS, with Krakauer); "The Metaphors of Artificial Intelligence" (Science, 2024-11); multiple 2024–2025 papers on analogy and abstraction in LLMs
- Santa Fe Institute: santafe.edu/people/profile/melanie-mitchell
- Personal site: melaniemitchell.me
Weak spots / open questions
- "Understanding" is itself a contested concept — if the definition keeps shifting, the critique can't be falsified
- Her tests focus on analogy and abstraction — but AI impact doesn't require human-style understanding; economic displacement can happen without genuine comprehension
- The gap between "doesn't understand like humans" and "can still do the job" is the key question she doesn't fully address
- Santa Fe Institute approach is rigorous but niche — influence on mainstream AI discourse is smaller than Marcus or Cowen
- If future architectures (not just scaled LLMs) achieve genuine abstraction, the critique applies to a generation, not to AI in general
Rich's take
- (your synthesis here)
converts-from: personas/melanie-mitchell.md · schema v1 · AI & Society domain