GPT-6 Astra and Fermat's Last Theorem: two announcements, no independent verification
OpenAI has rolled out its new model to Daybreak program enterprise clients; Anthropic states that Claude formalized Fermat's Last Theorem in Lean. In both cases, the figures come from the companies themselves.
On September 3, OpenAI announced the release of GPT-6 Astra to enterprise clients of the Daybreak program. The company stated that a version with additional restrictions on cybersecurity capabilities will be rolled out to paying ChatGPT users, and that both versions block access to the most advanced computer-use functions. The company presents the model as a milestone on the path toward artificial general intelligence and claims state-of-the-art results on benchmarks such as Agents’ Last Exam, AutomationBench and ScreenSpot Pro.
On this last point, the clarification is warranted and not a mere nuance: these are claims made by the company, not verifications carried out by third parties. Al Jazeera notes that in the announcement OpenAI devoted considerable space to the model’s risks and safety measures; Bloomberg reports the release and the additional restrictions. These are two separate sources documenting the announcement, not its technical content.
Among the outside voices gathered by Al Jazeera is that of Toby Walsh, professor and artificial intelligence expert at the University of New South Wales in Sydney, according to whom “the intelligence in artificial intelligence is still very jagged today” (transl. from English): the capability of these systems remains uneven depending on the task.
A model that claims to surpass a benchmark is asserting, not proving.
Eleven days and thirteen million lines. The second announcement comes from Anthropic, which claims to have completed, with the Claude model, the first fully computer-verified formalization of Fermat’s Last Theorem. Formalization consists of translating a mathematical proof into code checkable by a proof assistant — in this case Lean — so that the correctness of every step can be verified mechanically.
According to the company, the model worked largely autonomously for eleven days, generating roughly 13 million lines of Lean code; the final formalization reportedly contains nearly 29,500 intermediate theorems and was produced by dozens of Claude agents coordinated on the Prove2Me platform.
Here the documentation gap is twofold. First: the news currently comes from a single source (Anthropic’s announcement, picked up by the specialized outlet AIdapted); no independent confirmation is available. Second: the figures are company claims, and no independent verification has been published at this time. It is worth noting that, in principle, formalization work is among the results most readily checkable by third parties — the code either gets accepted by the proof assistant or it does not — but this check, as things stand, does not appear in the available materials.
What we don’t know. We don’t know whether the formalization’s code is public and inspectable, nor whether the mathematical community has examined the result. We don’t know the computational cost of the operation. On the OpenAI side, we don’t know what methodology was used to measure the cited benchmarks, nor who has access to the version without additional restrictions, nor when the model will become available beyond the Daybreak program and paid subscriptions.
The two announcements share the same informational structure: a company publishes a figure, the press records it, verification comes later — if it comes at all. For this reason we keep separate, including graphically, what is documented (the release date, the existence of the announcement) from what is claimed (the milestones and performance figures). The next verifiable element will be the rollout of the restricted version of GPT-6 Astra to paying ChatGPT users, announced by the company without a date.
Sources: Al Jazeera; Bloomberg; AIdapted.
← Archive · Front page · Past editorials · Report an error · Original article (in Italian)