数学の終焉
原題: The End of Mathematics
日本語訳
# 数学の終焉
私は今、OpenAIで開催された数学の未来に関するサミットからトロントへ戻っているところだ。セバスチャン・ブベックから、人間が数学的な力を失ってしまうという、誰もが避けたい未来について少し話してほしいと頼まれた。ジェイコブ・ツィマーマンからは「正確さよりも詳細を優先せよ」と助言されたが、私は正確さを二の次にすることには間違いなく成功したようだ。
あまり大げさすぎないタイトルを探した。
ワークショップの前提は(建設的な議論のために、議論の対象ではなく出発点として扱ったが)、AIが数学において極めて強力な超人類(superhuman)になるということだ。私は、それにもかかわらず数学の進歩が停滞するという物語を語りたい。はっきりさせておくが、これは予測ではない。私は本質的に楽観主義者であり、私たちは適応する方法を見つけられると考えている。しかし、現在の特定の傾向が続いた場合にどのような未来になるかを想像しようとしているのだ。
### 2026年
明らかなのは、数学的なアウトプットが爆発的に増加し始めているということだ。例えば、以下は2021年後半以降、arXivに毎週投稿されている組合せ論の論文数である。他の分野でも、劇的ではないにせよ同様の上昇が見られる。数学の結果に関するツイートの時系列データも、似たような形になるだろう。
それよりも不透明なのは、この余剰分がいかに興味深く、あるいは正確であるかということだ。ましてや、そのどれほどが意味を持って取り組まれているかについてはなおさらである。とはいえ、そこには驚くべき、かつ重要な新しい結果がいくつも含まれている。
同時に、数学コミュニティの特定の機能が萎縮しつつある。以下はMathOverflowの月ごとの質問数と回答数のグラフである。MathOverflowの機能がDiscordなどに取って代わられてきたため、これらの数値はしばらくの間緩やかに減少していたが、2025年初頭からの減少は、その大部分がAIによるものと考えられる。ここで驚くべきは、質問数も回答数も共に減っていることだ。例えば、統計データの中に、古い質問への回答が増えているといった兆候は見られなかった。ポジティブな方向に解釈できる統計は、実のところ一つも見当たらない。
AIによって生み出された最も興味深い結果の間でさえ、奇妙なことが起き始めている。例えば、3つのグループがほぼ同時に、フェイゲの1/e予想に対して非常に似た証明を独立して導き出した。そのうち2つのグループは、その結果がAIによって発見されたことを公表している。@__alpoge__が次元 $\geq 3$ におけるヤコビ予想へのFableの反例を投稿した後、OpenAIの内部モデルがそれを再現した。同様に、AnthropicもOpenAIが発表した最近の多くの結果を再現している。モデルも、それを使う人間も、同じ問題を解いているようだ。
実務的には、これは生身の人間とシリコン(AI)の両方において、膨大な重複作業が費やされていることを意味する。その数学的な限界価値は、本質的にはトークンコストと、おそらく「その問題は既存のモデルで解ける」ということを示すわずかな情報に過ぎない。
### 2027年
もちろん、こうした仕事は、それを発表する人々にとっては価値があるかもしれない(実績やPRなどの意味で)。
現在、私たちは論文を作成し、定理を証明し、予想を解決する人々に報いることで、高品質な科学の生産を動機づけようとしている。しかし、これらのアウトプットは現在、誤った価格付けがなされており、それらにインセンティブを与えることは、高品質な科学の生産にとって明らかに最適とは言えない。もし今後数年間、このまま続いたらどうなるだろうか?
もし続ければ、キャリア形成における支配的な戦略は、「予想というスロットマシンを回すこと」になるだろう。実際、予想を選ぶ必要さえなく、Codexに選ばせ、解かせ、検証させるだけでいい。もし正確な論文を作成したいのであれば、この方法で1日に複数の短い論文を量産できる(実際にそうしている人々もいる)。正確さを気にしないのであれば、さらに多くの論文を量産できる(これも既に行われている)。
付加価値はどこにあるのか? トークンコストか? 習得される専門知識については、間違いなく「ない」と言える。著者でさえ、こうした仕事の多くを読んでいない。数学者は、その根底にある数学から切り離されている。モデルがより信頼できるようになるにつれ、人間による検証の価値さえも低下していく可能性がある。
さらに、これは数学コミュニティに深刻な悪影響を及ぼす。モデルがいくつかの主要なアイデアを与えられれば論文を再構成できる段階に、私たちは近づいている。私たちが目にし始めている自律的なAIによる成果のいくつかは、「ラストワンマイル」的な性質を持っている。つまり、他者による深い研究の後に、その問題の仕上げを行うものだ。このような世界では、自身の進行中の研究について話すこと、あるいは特定のモデルがその問題を解けることを示すことさえ、ますます危険なことになっていく(少なくとも、そのような研究に名声や職などの報酬が与えられ続けるのであれば)。
最近、複数の同僚から、そのような理由で進行中の研究について議論したくないと言われた。
### 2028年
とはいえ、明るい兆しもある。自動形式化(Autoformalization)が安価で効果的になる。文献における多くの欠落や誤りが発見され、修正されるようになる。
ここでは、命題や定義が正しく形式化されているかを確認するために、人間の判断が必要だという議論が盛んになされている。私はこれに懐疑的だ。モデルがこれを効果的に行えない理由はないと考えている。
一方で、形式化された内容が、元の英語のテキストとは、読者には分かりにくい形で異なっているケース(下のスライドの2つの例など)をすでに目にし始めている。ここでも、数学者は数学から切り離されつつある。過去の研究における命題は信頼できても、その背後にあるアイデアを信頼することは難しくなっているのだ。非形式化(Informalization)はこれを多少助けるが、コストと時間がかかる。
### 2029年以降
それにもかかわらず、職業としての数学は依然として論文の生産を動機づけている。モデルは、現在の数学者が行っているすべての機能(理論構築、予想、解決、反復など)を果たし始める。人間の数学者は、エージェントを用いた「実験科学」を行っており、おそらく自分が興味のある問いに対して計算リソースを割り当てる役割を担う。
では、誰がこの仕事に関わっているのだろうか? 私たちは次世代をどのように育てていくのだろうか? 現在の制度がこの新しい体制に適応しなければ、高品質な数学者を輩出し続けることはできないだろう。実際、既存のインセンティブ構造は、数学に深く関わらない人々、あるいは数学を全く気にかけない人々を報いるようになると私は考えている。
これは持続可能な数学の実践につながるだろうか? おそらく、そうではないだろう。数学に携わるエージェントに対して、なぜそのような人々がリソースを割き続ける必要があるのだろうか? おそらく、この世界における数学研究の長期的な姿とは、このようなものなのだろう。
**起こりうるリスクの要約:**
これらの問題は、数学という職業自体がすでに様々な点で不完全であることの反映であることを指摘しておきたい。これは驚くべきことではない。AIによる変化が私たちの制度に負荷をかけるとき、制度はすでに欠陥のある箇所から当然のように崩れていく。おそらく、この外生的なショックは、これらの欠陥を修正する機会を与えてくれるだろう。
私たちの制度には、特定の価値(高品質な科学の生産、人的資本、人間の理解など)があり、それらに貢献する人々に楽しみや名声などを与えることで、その達成を図ろうとしている。高度なAIが存在する世界でもこれらの価値は存続するが、それらを達成するために用いている仕組みの多くは、高度なAIに対して脆弱である。
**最後に:**
参考までに、私は数学が生き残り、繁栄することについて、大筋では楽観的である。私たちには、信じられないようなことを学び、理解する機会がある。私たちは適応していくと思う。
高度なモデルが抽象数学の世界を超えて大規模な社会変動を引き起こすにつれ、これらの懸念の多くは、数年後には古臭い、あるいは些末なものに見えるかもしれない。私の願いは、ここで提起した問いが、建設的に検討されるのに十分なほど限定的なものであり、私たちの答えが、同様の影響を受ける他の人々にとってのモデルとなることだ。
原文(英語)を表示
I'm currently returning to Toronto from a summit on the future of mathematics, at OpenAI. Sebastian Bubeck asked me to talk a bit about the future we'd all like to avoid, where humans are mathematically disempowered. Jacob Tsimerman advised us to try to prioritize detail over correctness, and I have no doubt that I succeeded in deprioritizing correctness.
I tried to find a title that wasn't too bombastic:
The premise of the workshop (which we took as a starting point, rather than subject to debate, for the sake of productive discussion) was that AI will become robustly superhuman at mathematics. I want to tell a story in which, despite this, mathematical progress stalls. To be clear this is not a prediction--I'm optimistic by nature and think we'll find a way to adapt--but I am trying to imagine what a future in which certain existing trends continue might look like.
2026
What's clear is that we are at the start of a massive explosion of mathematical outputs; for example, below is the number of combinatorics papers posted per week to arXiv since late 2021. Other areas show a similar, but not quite dramatic, rise. I imagine a time series of tweets about math results would look similar.
What's less clear is how interesting or correct this surplus is, let alone how much of it is being meaningfully engaged with. Nonetheless it contains a number of striking and significant new results.
At the same time certain organs of the mathematical community are atrophying. Below is a graph of MathOverflow questions and answers by month; these numbers have been in slow decline for some time as MathOverflow's function has been cannibalized by Discord etc., but the decline since the beginning of 2025 is likely due in large part to AI. What I find striking here is that there are both fewer questions and fewer answers. For example, I was not able to find an increase in answers to older questions in the statistics here, or really any other statistic I could spin as positive.
Even among the most interesting results produced by AI, something odd is starting to happen. For example, three groups independently produced very similar proofs of Feige's 1/e conjecture almost simultaneously; two groups disclosed that the result was found by AI. After @__alpoge__ posted Fable's counterexample to the Jacobian conjecture in dimension \geq 3, an internal model at OAI replicated it; likewise Anthropic replicated many of the recent results OpenAI has announced. The models, and the people using them, seem to be solving the same problems.
In practice this means that a huge amount of duplicative labor, both in flesh and in silico, is being devoted to work whose marginal value to mathematics is, essentially, the cost of the tokens and perhaps a few bits of information indicating that the problem can be solved by existing models.
2027
Of course this work might have value to the people announcing it (credit, PR, etc.).
Right now we try to incentivize the production of high quality science by rewarding people who produce papers, prove theorems and resolve conjectures, etc. But these outputs are now mispriced, and incentivizing them is not obviously optimal for the production of high quality science. What happens if we continue to do so in the next years?
I think if we do, the dominant strategy for career success (at least in the medium term) is playing the slot machine for conjectures. In fact one does not even have to choose the conjectures--you can just ask codex to pick them and resolve them and check the work. If you care about producing correct papers you can produce multiple short papers per day this way (and people who are doing so); if you don't care about correctness you can produce far more (and people are doing this too).
What's the value-add? The cost of the tokens? Certainly not the expertise developed--there is none. No one, not even the author, is reading much of this work. Mathematicians are no longer connected to the underlying mathematics. Even human verification is arguably less valuable as the models become more reliable.
Moreover this has seriously negative effects on the math community. We are near the point where the models can reconstruct a paper given a few key ideas. Some of the autonomous AI results we are starting to see have a "last mile" flavor, where they finish off a problem after deep recent work by others. In this world talking about one's work in progress--or even indicating that the models can solve a given problem--is increasingly dangerous (at least if we still reward such work with prestige, jobs, etc.).
I've recently been told by multiple colleagues that they are unwilling to discuss work in progress for this reason.
2028
Nonetheless there are some bright spots. Autoformalization becomes cheap and effective. Many gaps or errors in the literature are discovered and repaired.
Much hay has been made of the necessity of human judgment here, to check that statements and definitions are formalized correctly. I am skeptical of this--I see no reason the models will not be able to do this effectively.
On the other hand, we are already starting to see cases (e.g. the two examples in the slide below) where formalizations differ from the English text they are formalizing in ways that may not be obvious to the readers. Again mathematicians are becoming disconnected from mathematics--while they might be able to trust the statements in past work, it is harder to trust the ideas. Informalization helps with this a bit but it is costly and time-consuming.
2029-
Despite this, the profession still incentivizes the production of papers. Models start to fulfill all the functions human mathematicians do now: theory-building, conjecturing, resolving conjectures, iterating, etc. Human mathematicians are doing "lab science" with agents, perhaps directing compute to questions they find interesting.
Who is engaging with this work? How are we training the next generation? It's not clear to me that our current institutions, if they do not adapt to this new regime, continue to produce high-quality mathematicians. Indeed it seems to me that our existing incentive structures will start to reward people who do not engage deeply with the mathematics, or, arguably, care about it at all.
Will this lead to a sustainable mathematical practice? I think plausibly not. Why would such people continue to devote resources to agents doing mathematics at all? Perhaps this is what the long term of mathematics research looks like, in this world:
A summary of some possible risks:
I want to point out that these problems are reflections of the fact that the profession itself is already imperfect in various ways. This isn't surprising--as AI-induced change puts stress on our institutions, they will of course crack in the places where they are already flawed. Perhaps this exogenous shock will give us a chance to fix some of these flaws.
Our institutions have certain values (production of high quality science, human capital, human understanding, etc.) that we try to achieve by rewarding people who contribute to them, with fun, prestige, etc. These values persist in a world with highly capable AI, but the mechanisms we use to achieve them are in many cases not robust to highly capable AI.
Some final questions:
For what it's worth, I'm broadly optimistic that mathematics will survive and flourish. We have the opportunity to learn and understand incredible things. I think we'll adapt.
I think many of these concerns may seem quaint or parochial in the next few years, as highly capable models cause massive social upheaval beyond the world of abstract mathematics. My hope is that the questions I raise here are narrow enough to be considered productively, though, and that our answers might serve as a model for others as they too are impacted.