Tao: Math in the Age of AI (2026 ICM Lecture Review and Extensions)
Review: Terence Tao’s 2026 ICM lecture on AI’s impact on mathematical practice, values, and future challenges, including problem-solving, proof generation, and community goals.
Tao: Math in the Age of AI (2026 ICM Lecture Review and Extensions)
1.
Mathematics in theAge of AI
A Summary and Critical Review of Terence Tao’s 2026 ICM Public Lecture
August 2026
Abstract
Terence Tao’s 2026 International Congress of Mathematicians public lecture, Mathemat-
ics in the Age of AI, argues that the prospect of capable mathematical AI is forcing a crisis —
not in the logical foundations of mathematics, but in the foundations of mathematical values
and practice. Tao’s method is explicitly conditional: he separates the empirical question
of what AI systems can do from what he calls its “orthogonal complement,” the question
of what the mathematical community actually values. Conditioning on a Working Hypoth-
esis of substantial AI capability, he traces the problem-solving component of mathematics
through a six-stage pipeline running from proof generation to canonicalization, and argues
that acceleration of the earliest stage produces “proof indigestion” at every later one. This
paper summarizes that argument, assesses its strengths and limitations, and — in a clearly
separated second part — develops several extensions the lecture does not itself undertake.
Sources are limited to the published lecture slides; the reviewer did not have access to a
recording or transcript, so remarks made only in delivery are not represented here.
Contents
1 Introduction 2
2 The Structure of the Argument 2
2.1 The motivating question . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2
2.2 The first subquestion: capability . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.3 Orthogonal decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
2.4 The one capability data point: the First Proof challenge . . . . . . . . . . . . . . 4
3 The Goals and Values Question 4
4 Case Study: Problem Solving 5
4.1 Generation is not verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
4.2 Verification is not understanding . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
4.3 Exposition, and the hazard of over-polish . . . . . . . . . . . . . . . . . . . . . . 5
4.4 Publication, digestion, and canonicalization . . . . . . . . . . . . . . . . . . . . . 6
4.5 Proof indigestion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
5 Recommendations 7
6 Strengths 8
1
2.
7 Limitations andOpen Questions 8
7.1 The taxonomy of systems could be sharper . . . . . . . . . . . . . . . . . . . . . . 8
7.2 Community acceptance deserves its own scrutiny . . . . . . . . . . . . . . . . . . 9
7.3 The publication standard needs implementation detail . . . . . . . . . . . . . . . 9
7.4 The economics are deferred rather than addressed . . . . . . . . . . . . . . . . . . 9
8 Extension: Proof Generation Versus Theory Generation 10
9 Extension: A Staged Model for Education 10
10 Extension: Redistribution of Mathematical Labor 11
11 Assessment 11
12 Conclusion 12
Note on scope and attribution. Sections 1 through 7 summarize and assess Tao’s lecture
[1]. Sections 8 through 10 are the reviewer’s own extensions; they are not claims about
what Tao argued, and are marked as such at the head of each section. Section 11 returns to
assessment of the lecture.
1 Introduction
Tao opens with a historical prologue. For centuries mathematics operated on “naive” foundations,
delegating questions such as what is a set? and what are the axioms of mathematics? largely
to philosophers. Russell’s paradox (1901) and Gödel’s incompleteness theorems (1931) [3] forced
practicing mathematicians to re-examine those implicit assumptions directly.
The important feature of this analogy, for Tao’s purposes, is not the turbulence of the period
from roughly 1900 to 1930 but its outcome. The crisis produced an explicit, rigorous, standard-
ized foundational framework which has since survived strenuous testing and now constitutes a
trusted environment for mathematical work. Tao’s claim is that mathematics is entering a simi-
larly turbulent period — a crisis in the foundations of mathematical values and practices — and
his stated expectation is that once those foundations are thoroughly examined and codified, the
community will emerge stronger and more resilient than before. The register of the lecture is
therefore closer to constructive urgency than to alarm.
The lecture’s slides are available at
https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf .
2 The Structure of the Argument
2.1 The motivating question
Tao organizes the lecture around a single question:
2
3.
Community Response Question.How should the mathematical community respond to the
advent of modern AI technologies, and their real and/or claimed capabilities to perform mathe-
matical tasks?
He is explicit that this is not a mathematical question. It is a metamathematical one, and
also a political, ethical, and cultural one. His method is to borrow the precise, familiar language
of mathematics — conjecture templates, hypotheses, flow networks, orthogonality — as a device
for clarity, not because the subject matter admits mathematical treatment. He also disclaims
any presumption of having the answers; the question is posed to the community.
2.2 The first subquestion: capability
The answer depends on a subquestion Tao states pseudomathematically as a family of conjectures:
AI Capability Conjecture (template). At some point in the near future, some AI tools will,
at some expense, and with some level of human supervision, be able to correctly accomplish some
research-level mathematical tasks in some fields of mathematics, with some non-trivial success
rate, and at some level of correctness and quality.
Every occurrence of “some” is a placeholder. Different instantiations yield different conjec-
tures, and Tao deliberately declines to work through the finer distinctions, drawing only a coarse
line between weak and strong forms. The consequence of that line is sharp: if even weak forms
of the conjecture are false, the community could safely dismiss AI tools as being of no long-term
significance and continue largely as before. If the strongest forms are true, it becomes challenging
to maintain current culture and practice — particularly if the profession continues to prioritize
obtaining as many solutions to unsolved problems as possible.
Two further qualifications deserve emphasis, because they are easily lost in summary. First,
Tao distinguishes the truth value of a given form of the conjecture from its desirability; these are
separate questions and are frequently conflated in public debate. Second, he asks the audience
only to condition on the hypothesis, not to want, believe, or accept it. The analysis that follows
is conditional throughout.
2.3 Orthogonal decomposition
Tao notes that debate about AI in mathematics has concentrated almost entirely on which ver-
sions of the capability conjecture hold, and that he has himself contributed extensively to that
debate elsewhere. He then sets it aside. The lecture is about what he calls the “orthogonal com-
plement” to the AI Capability Conjecture within the Community Response Question. Evidence
for or against the hypothesis is, in his framing, orthogonal to everything that follows.
The operative assumption for the remainder of the lecture is:
Working Hypothesis. AI tools will, reasonably soon, become capable of performing a rea-
sonable fraction of research-level mathematical tasks, with reasonable levels of success, quality,
supervision, and cost.
The precise meaning of “reasonable” is stated to be non-critical.
3
4.
2.4 The onecapability data point: the First Proof challenge
Consistent with this decomposition, Tao offers exactly one item of capability evidence, and offers
it briefly. He observes that many data points now exist for and against various forms of the
conjecture, but that most have not been gathered under controlled scientific conditions, and that
much publicly available evidence is subject to reporting bias and non-scientific incentives, with
important costs and variables undisclosed.
Against that background he cites the First Proof challenge [7], described as an independent
assessment of frontier models and harnesses. Each batch consists of ten novel research-level
problems across various fields. The second batch was tested under controlled scientific conditions
against four AI harnesses on 28 May 2026. Results were refereed by experts for both correctness
and exposition. Seven of the ten problems were solved at publication-level quality by at least
one team, with compute costs ranging from ten to one thousand US dollars per problem. Further
batches are planned.
This is stated as fact rather than report, and the controlled conditions and expert refereeing
are load-bearing features of why Tao selects this data point in preference to the wider public
record.
3 The Goals and Values Question
Once the Working Hypothesis is granted, a second subquestion becomes unavoidable:
Goals and Values Question. What are the precise goals, objectives, and values of our math-
ematical community, and the enterprise of mathematical research?
Tao stresses that the relevant goals include not only the explicit ones communicated to the
public and to funding agencies, but the implicit ones actually pursued in practice. He observes
that mathematicians have largely delegated this question to the humanities while attending
to the technical aspects of the profession, and argues that under the Working Hypothesis this
luxury disappears — though he adds that critical examination of the question would be valuable
regardless of the hypothesis’s status.
The goals he lists include solving unsolved problems, pure and applied; developing new theo-
ries and techniques; understanding the world; building a community of mathematicians; training
the next generation to guide the field’s future directions; contributing to the shared network of
mathematical knowledge; and creating enduring works of aesthetic value.
His structural claim about these goals is the pivot of the lecture. They have historically been
positively correlated: progress on one typically produced progress on the others. Consequently
one or two could serve as proxies for the rest, and most could be left implicit. Tao then invokes
Goodhart’s law [4] — when a measure becomes a target, it ceases to be a good measure1 — and
identifies two specific reasons AI tools are unusually vulnerable to it: the inherently ungrounded
nature of generative AI, and the financial incentives of AI companies. The second reason is a
claim about institutional structure rather than about technology, and it does substantial work
in his argument.
The predicted consequence is divergence. Under excessive optimization, goals that were previ-
ously aligned — solving research problems, solving Erdős problems, solving Olympiad problems,
1
Tao dates the law to 1975, following Goodhart’s original paper on UK monetary management [4]; the compact
slogan formulation is usually attributed to Strathern [5].
4
5.
building theory, applyingknowledge, building community, training the next generation, creating
enduring aesthetic works — may pull apart from one another and from mathematical knowledge
itself. Tao presents this as a deliberately oversimplified diagram and notes that the true picture
should be much higher dimensional.
4 Case Study: Problem Solving
Tao selects problem solving as his worked example, with two explicit caveats: it is not the only
aspect of the profession, and theory building is a complementary aspect requiring its own separate
analysis. Problem solving is chosen because it appears particularly susceptible to impact under
the Working Hypothesis.
The exposition proceeds by successively refining a goal statement, each refinement adding a
stage to a flow network.
4.1 Generation is not verification
The first attempt is simply: solve as many unsolved problems as possible, optimizing flow across a
single edge from open problems to solutions. Tao points out that the failure mode here predates
AI entirely — optimizing this metric already produces a large volume of incorrect solutions to
major open problems, the Riemann hypothesis being the standing example. The second attempt
therefore adds verification, and the network acquires an intermediate node: open problems,
unverified solutions, verified solutions.
Tao notes that advances in AI and autoformalization, using proof assistants such as Rocq,
HOL, or Lean, have already significantly accelerated both generation and verification, and that
under the Working Hypothesis this acceleration continues. He sets formalization aside as requir-
ing a lecture of its own.
4.2 Verification is not understanding
The question that motivates the next refinement is posed concretely rather than abstractly. Tao
observes that erdosproblems.com [8] already contains dozens of AI-generated proof submissions;
that many are likely correct; that no human expert has yet volunteered to verify and vouch
for them; and that in several cases the human submitters have themselves declared they are
unqualified to do so. From this he draws the question directly: could we have a verified proof of
a major result that no human understands well enough to explain it?
The third attempt therefore adds exposition.
4.3 Exposition, and the hazard of over-polish
Tao’s assessment of current AI exposition is mixed in a specific way. Spelling, grammar, and
formatting are close to flawless — he remarks, in a footnote, that one can argue they are too
flawless. But the writing often dwells at length on trivialities while passing briefly through, or
actively obscuring, the most interesting and novel portions of an argument, and frequently fails
to note connections with prior literature or to supply a high-level overview.
He then makes a subtler point. Exposition is a fuzzier optimization target than verification,
and the Working Hypothesis predicts AI exposition will improve — but exposition can also be
over-optimized. A proof may be too slickly written, presenting routine and difficult parts of an
5
6.
argument as equallyeasy to digest. In human-written proofs, the passages the author found
hard typically retain natural friction that slows the reader down. Excessive AI polish removes
both artificial and natural friction without encouraging the reader to learn the key ideas. Tao’s
formulation is that “mistakes” in human exposition can paradoxically help the reader, and he
illustrates it with a slide showing a 1991 paper of Bourgain annotated by his younger self.2
The stated standard is Thurston’s [2]: the measure of success in mathematics is not meeting
a production quota of definitions, theorems, and proofs, but whether the work enables people to
understand and think more clearly and effectively about mathematics. This is the intellectual
anchor of the lecture’s central thesis and is credited as such.
4.4 Publication, digestion, and canonicalization
Correctness and readability are still not sufficient. A result must be accepted and valued by the
community; other mathematicians must digest it and incorporate it into their own work. Tao
notes that authors assist digestion by describing their own insights and the story of how they
worked on the problem, and that current AI tools are quite opaque about their problem-solving
process — particularly proprietary models whose inner workings are corporate secrets. This is a
specific transparency deficit tied to a specific stage of the pipeline, not a general complaint.
Community acceptance is by its nature slow and human. Good exposition and careful writing
encourage it, but it is an external process that cannot be optimized purely by authors and their
tools. The current publication infrastructure depends on human editors and referees providing
this service voluntarily, work that is regarded as less prestigious than generating proofs but
which is how individual achievements are converted into collective progress. Tao allows that AI
evaluation tools may serve as useful filters — journals could automatically reject papers flagged
for inadequate verification or exposition — while insisting they cannot substitute for community
acceptance.
Publication is still not the terminal state. Key results should become part of the definitive
textbooks and reference material of a subject and be taught to the next generation. Canonical-
ization is the slowest stage of all, requires broad deliberative consensus, and is the stage least
amenable to AI optimization. Tao argues it is also the most valuable: many applications become
feasible only once the underlying mathematics has been fully digested, and the success of AI
tools in mathematics itself relies crucially on the canonical theories human mathematicians built
over centuries.
The final goal statement is: solve unsolved problems, verify them to be correct, ensure they
are clearly communicated, and have them digested, accepted, and incorporated into the definitive
theory of the field. The corresponding network is shown in Figure 1.
4.5 Proof indigestion
If the Working Hypothesis holds, then absent suitable policy and cultural change, significant
impedance mismatches emerge throughout this network. AI-generated proofs accumulate await-
ing verification. Verified proofs await readable writeups. Correct, well-written proofs overwhelm
traditional peer review. Even published proofs prove too numerous for the community to work
into definitive form.
Tao’s summary formulation is that mathematics transitions from an era of proof scarcity to
an era of proof abundance. He notes that some signs of indigestion are already visible, and — a
2
The slide does not identify the paper, and it is not identified here.
6
7.
open
problems
unverified
solutions
verified
solutions
well-written
solutions
accepted
solutions
definitive
solutions
proof
generation
proof
verification
proof
exposition
proof
publication
proof
digestion
Figure 1: Theproof pipeline, reconstructed from the lecture’s final goal slide: six states, with
each successive goal statement in the lecture adding one node to the previous network. The slide
names six transitions — proof generation, verification, exposition, publication, digestion, and
canonicalization — of which the last describes the passage of an accepted result into the definitive
literature of a field; the exact assignment of the final two labels to edges is not recoverable from
the slide text alone.
point worth retaining — that strains were emerging even before modern AI.
5 Recommendations
Tao’s recommendations are not presented as freestanding proposals. He introduces the Leiden
Declaration [6] as an excellent starting point, referring the audience to a lecture by Jim Portegies
at the same congress on 26 July, and then offers three commentaries.
Disclosure. The worst case to avoid is one in which authors use AI tools covertly and conceal
that usage to avoid peer criticism. Responsible disclosure should be normalized. Tao practices
this in the lecture itself: a footnote discloses that AI tools were used to autocomplete text and
generate diagrams in the slides, and a second footnote notes that all em-dashes in the slides were
human-generated.
Reweighting prestige. Emphasis should decrease on proof generation and on being first to
solve a problem, and increase on “proof digestion”: exposition, publication, and canonicalization.
A publication standard. Tao’s suggested rule of thumb is that if authors cannot convincingly
demonstrate that they can give a clear, expert-level talk on their results, correct and properly
attributed, then the result should not be published. The standard as stated is a demonstrable
capacity to give such a talk, not a requirement that one be delivered.
Closing thoughts. Tao notes that problem solving is only one aspect where the Working
Hypothesis forces inspection of goals and values; teaching, mentoring, hiring, grant applications,
and public outreach require similar analyses. In education and training particularly, he argues it
will be crucial to emphasize the human aspect of the work and tightly restrict AI use; in other
areas the community should take the initiative and define best practices on its own terms. He
states that new workflows and infrastructures will be needed to complement traditional ones,
explicitly defers that topic as a different talk, and closes by pointing to five existing efforts:
Mathlib [9], Mathematical Discourse [10], the Erdős problems database [8], an optimization
constants database [11], and the SAIR Foundation competitions [12].
7
8.
6 Strengths
The orthogonaldecomposition is the substantive methodological move. Almost all
public argument about AI and mathematics is argument about capability. By construing the
Community Response Question as a direct sum of a capability subquestion and a values subques-
tion, and then addressing only the second, Tao produces an analysis whose conclusions do not
depend on winning a forecasting argument. Even a perfect theorem prover would leave questions
of education, authorship, credit, publication, and problem selection unresolved.
The pipeline decomposes a word that was doing too much work. “Proof” is routinely
used to cover generation, verification, explanation, publication, digestion, and canonicalization.
Current discussion overweights the first two because they have crisp success criteria: a theorem
is proved or not, a formal proof passes the kernel or does not. Explanation quality, definitional
usefulness, and theoretical importance resist quantification. Tao’s argument that the less mea-
surable stages become more important precisely as the measurable ones become cheap is the
lecture’s most transferable claim.
The over-polish argument is non-obvious and correct. The observation that fluency can
destroy informative friction cuts against the intuitive view that better writing is monotonically
better. It also identifies a failure mode that will worsen as AI exposition improves, which is the
opposite of the usual pattern in AI criticism.
The forecasting is restrained. Tao does not claim autonomous mathematical research has
arrived. He cites one controlled result and declines to argue the capability question. The institu-
tional problems he identifies arise under partial automation; complete automation is not required
for review capacity to be exceeded.
7 Limitations and Open Questions
7.1 The taxonomy of systems could be sharper
Tao does distinguish system types where it matters to his argument: proof assistants such as
Rocq, HOL, and Lean; natural-language generative models; AI evaluation filters; proprietary
systems with opaque internals. He also defers formalization as a separate topic. The criticism
available is therefore not that he treats AI as unified, but that the distinctions appear locally
where needed rather than as a systematic framework.
A more explicit classification would be useful for institutional design, since the relevant policy
questions turn on properties that vary sharply across system types: autonomy, transparency,
verifiability, reproducibility, cost, and position in the pipeline. A Lean proof checked by a small
trusted kernel and a natural-language proof from a proprietary model present different reliability
problems, and an autonomous agent that selects conjectures and writes papers raises different
authorship questions from a retrieval system that locates lemmas. These differences bear directly
on the disclosure recommendation, which asks which tasks were delegated and how output was
verified.
8
9.
7.2 Community acceptancedeserves its own scrutiny
Tao’s account of acceptance and canonicalization is descriptively accurate and his valuation of
them is persuasive. The gap is that community consensus is treated as the terminal criterion
without being subjected to the same critical examination he applies to proof generation.
Mathematical communities can be conservative, status-sensitive, and slow to absorb unfamil-
iar methods. Work that lies between established fields, or uses unfamiliar language, or arrives
from outside the usual networks, can be ignored for reasons unrelated to importance. If canoni-
calization is the most valuable stage, its own failure modes warrant analysis — particularly under
the Working Hypothesis, where the volume of candidate results makes attention allocation more
consequential and plausibly more status-driven than before.
Tao himself supplies the beginning of an answer, in that he allows AI evaluation tools as filters.
That suggestion could be extended: systems that translate terminology across subfields, identify
structural analogies between proofs, or surface overlooked prior work would act on exactly the
failure modes described. The point is not that consensus should be replaced but that it should
be an object of scrutiny rather than a backstop.
7.3 The publication standard needs implementation detail
The expert-talk rule is stated as a demonstrable capacity, which is a more flexible standard
than a delivery requirement, and the criticism sometimes made of it — that it disadvantages
researchers with differing language ability, communication style, disability status, or collaborative
role — applies to a stricter reading than Tao’s text supports.
What the rule does need is implementation detail, because the demonstration is where the
work happens. If the underlying principle is that a submitting team must show substantive com-
mand of a result’s claims, methods, provenance, limitations, and relation to existing work, and
must take responsibility for its correctness, then several instruments could serve: oral presenta-
tion, structured written responses to referee interrogation, or a required provenance statement
of the kind the disclosure recommendation already implies. Specifying the acceptable instru-
ments matters more than specifying the format, and the lecture, being a public lecture, does not
attempt it.
7.4 The economics are deferred rather than addressed
This is the most consequential gap, and it is a deferral Tao makes explicitly: new workflows and
infrastructures are, he says, an entirely different talk. He does point to five concrete efforts, so
the deferral is not a blind spot.
Nevertheless the recommendations as given are cultural — disclose, reweight prestige, raise
the publication bar — while the bottleneck he diagnoses is one of labor. Verification, exposi-
tion, formalization, refereeing, database maintenance, and textbook synthesis all require skilled
human time. Prestige reallocation changes incentives at the margin; it does not by itself create
capacity. If output grows substantially without corresponding support for downstream work, the
publication system can fail through volume alone regardless of what the profession claims to
value. Plausible instruments include funding lines for synthesis and formalization, salaried veri-
fication roles, curated knowledge bases, journal formats designed for verified-but-unassimilated
results, and infrastructure for tracking dependencies among results. The transition from proof
scarcity to proof abundance is an economic and organizational problem as much as a cultural
one, and the cultural recommendations are unlikely to be sufficient without it.
9
10.
The three sectionsthat follow are the reviewer’s own extensions. They develop lines
of argument the lecture does not pursue and should not be read as summary or interpretation
of Tao’s claims.
8 Extension: Proof Generation Versus Theory Generation
Tao states that theory building is a complementary aspect of the profession requiring separate
analysis, and does not undertake that analysis. The following is an attempt to indicate why it
may not reduce to the problem-solving case.
Solving isolated problems and constructing theories are not different quantities of the same
activity. Theory construction involves selecting definitions, finding invariants, identifying repre-
sentative examples, choosing abstractions, and determining which questions are worth asking at
all. These are choices about what to attend to, made before any statement exists to be proved.
A system could in principle become exceptionally strong at proving supplied statements while
remaining weak at reorganizing a field conceptually, because the two draw on different kinds of
judgment.
The converse possibility is equally open. Clustering results, identifying analogous proof struc-
tures across subfields, proposing intermediate concepts, detecting hidden commonalities, and
suggesting reusable abstractions are all tasks with some computational purchase, and progress
on them would contribute to theory formation directly.
If so, the governing question is not how many proofs a system can produce but whether it
can generate compressive understanding: frameworks from which many results follow, and which
can be remembered, taught, and applied. One informal way to express the intuition is
mathematical value ≈
important phenomena explained or controlled
conceptual and inferential complexity required
.
This is not a metric and is not proposed as one; neither quantity admits measurement. It
records only the observation that a powerful theory compresses many facts into few concepts and
principles, and that a long list of isolated machine-generated theorems may therefore represent
less progress than a short theory explaining why those theorems hold.
9 Extension: A Staged Model for Education
Tao argues that education and training require tight restriction of AI use, and does not specify
what the resulting curriculum looks like. Restriction alone is incomplete as policy, since students
will work in an environment where these tools are ubiquitous; the objective is to preserve unaided
competence while teaching disciplined use. A three-stage model is one way to reconcile these.
In the first stage, students develop core skills without assistance: understanding definitions,
constructing examples and counterexamples, performing symbolic manipulation, estimating plau-
sibility, detecting errors, and writing elementary proofs.
In the second, AI functions as a constrained assistant — proposing examples, generating
alternative strategies, criticizing a draft, translating informal argument into formal notation,
locating background material — with the student retaining the decisions.
10
11.
In the third,students are assessed on their ability to verify, explain, modify, and defend
AI-assisted work: identifying which parts of an argument are routine, which are essential, and
where the proof depends on external results.
The operative rule is: generate with AI, criticize independently, verify mathematically or
formally, then explain and take responsibility. Human judgment is placed at the final and decisive
stage rather than removed. This is compatible with Tao’s position but is a specification of it,
not a report of it. The rationale for retaining unaided problem solving is instrumental as well as
intrinsic: students cannot evaluate machine-generated mathematics without first developing the
competence against which to evaluate it.
10 Extension: Redistribution of Mathematical Labor
The likely near-term effect of the Working Hypothesis is not the disappearance of mathemati-
cians but a change in the relative price of mathematical activities. Routine derivation, symbolic
manipulation, literature retrieval, proof search, and formal proof completion become cheaper.
Verification, interpretation, problem selection, exposition, synthesis, education, and theory con-
struction become relatively more valuable — this is the labor-market restatement of Tao’s pipeline
argument.
The redistribution is not self-executing. Institutions currently reward novelty, priority, and
publication volume more strongly than verification or synthesis. If incentives do not move, the
predictable outcome is large volumes of low-value output alongside neglected assimilation work:
proof abundance producing noise and congestion rather than accelerated understanding.
Several roles that currently exist informally would need to become recognized positions with
associated funding and career structure:
• verifiers who check machine-generated arguments;
• formalization specialists who translate results into proof assistants;
• synthesizers who organize related results into coherent theories;
• expositors who make complex machine-assisted work intelligible;
• curators who maintain trusted mathematical knowledge bases;
• evaluators who identify important results within large volumes of output.
The five infrastructure efforts Tao lists at the close of the lecture are plausibly read as early
instances of the curation and formalization roles acquiring institutional form.
11 Assessment
The lecture avoids both technological triumphalism and reflexive dismissal, and its central thesis
is convincing: AI does not merely present mathematicians with a new class of tools, it obliges
the profession to state explicitly which parts of its activity it values.
The deepest point is that a proof is not yet mathematical progress. Progress occurs when a
result is verified, explained, connected to prior knowledge, evaluated by competent readers, and
incorporated into a durable conceptual structure. Artifact production and collective understand-
ing are not the same thing, and the pipeline is Tao’s device for keeping them apart. The debt
to Thurston is acknowledged in the lecture and is real.
11
12.
Two limitations boundthe argument’s reach. The lecture is a public lecture and does not
attempt institutional design; the recommendations are accordingly cultural where the diagnosed
bottleneck is economic. And the analysis is conditional on a hypothesis Tao declines to defend
— which is a deliberate and defensible structural choice, but means the lecture supplies no
independent purchase on how urgent the problem is.
Within those bounds the argument holds. If the Working Hypothesis is even approximately
right, proof generation becomes less scarce while trustworthy evaluation and conceptual synthe-
sis become more scarce, and mathematicians spend less time on routine derivation and more on
deciding which questions matter, which results are reliable, how results fit together, and what
they teach. That outcome is not automatic. Absent changes in incentives, infrastructure, pub-
lication, education, and professional recognition, abundance produces congestion and declining
trust instead.
The lecture is best read as what its historical prologue implies: a call to examine and codify
the foundations of mathematical practice before the volume of machine-generated output over-
whelms institutions built for an era of scarcity — and, following the precedent of the earlier
foundational crisis, an expectation that doing so leaves the community stronger.
12 Conclusion
Artificial intelligence may change mathematics most profoundly not by replacing mathematicians
but by changing the relative scarcity of mathematical activities. When proofs are scarce, pro-
ducing one is naturally the central achievement. When proofs are abundant, the central tasks
shift toward verification, explanation, selection, synthesis, and canonicalization.
The community must therefore distinguish theorem production from mathematical under-
standing, and preserve the educational processes by which humans acquire the competence needed
to evaluate machine-generated work. The appropriate response is neither unrestricted adoption
nor blanket prohibition, but disciplined research and educational practice in which AI assists
generation and criticism, formal and mathematical tools support verification, and humans retain
responsibility for interpretation, judgment, and the direction of the field.
References
[1] T. Tao, Mathematics in the Age of AI, Public lecture, International Congress of Math-
ematicians 2026, University of California, Los Angeles, July 24, 2026. Slides: https:
//teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf.
[2] W. P. Thurston, On proof and progress in mathematics, Bulletin of the American Mathe-
matical Society 30 (1994), no. 2, 161–177. Also available as arXiv:math/9404236.
[3] K. Gödel, Über formal unentscheidbare Sätze der Principia Mathematica und verwandter
Systeme I, Monatshefte für Mathematik und Physik 38 (1931), 173–198.
[4] C. A. E. Goodhart, Problems of Monetary Management: The U.K. Experience, in Papers
in Monetary Economics, Volume I, Reserve Bank of Australia, 1975.
[5] M. Strathern, “Improving ratings”: audit in the British University system, European Review
5 (1997), no. 3, 305–321.
12
13.
[6] The LeidenDeclaration. https://leidendeclaration.ai. Cited in [1]; a lecture on the
declaration by J. Portegies was given at ICM 2026 on July 26, 2026.
[7] First Proof. https://1stproof.org. Independent assessment of frontier AI models and
harnesses on novel research-level mathematical problems; second batch tested May 28, 2026,
as reported in [1].
[8] Erdős Problems. https://www.erdosproblems.com.
[9] Mathlib: the Lean mathematical library. https://mathlib.org.
[10] Mathematical Discourse. https://www.mathematicaldiscourse.org.
[11] T. Tao et al., Optimization constants database. https://github.com/teorth/
optimizationproblems.
[12] SAIR Foundation Competitions. https://competition.sair.foundation/competitions.
Sourcing note. All statements attributed to the lecture derive from the published slide
deck [1]; no recording or transcript was consulted, and remarks made only in oral delivery
are not represented. References [6]–[12] are resources cited within the slides and are listed
here as pointers rather than as independently assessed sources.
13