
OpenAI published 722 math manuscripts on GitHub generated by an unreleased internal model from a single prompt, with only 22% formally verified in Lean; experts debate the significance, reproducibility, and responsibility of AI-produced mathematical results.
AI-generated summary
OpenAI released 722 AI-generated math manuscripts on GitHub from an unreleased model, claiming they resulted from a single prompt to one AI agent, with only 22% formally verified using Lean software.
OpenAI published 722 math manuscripts on GitHub on Tuesday, all produced by an internal model the company has not released. An OpenAI spokesperson said almost everything came from a single prompt handed to a single AI agent, though some may have taken multiple attempts.
It's a bold claim and a potentially significant breakthrough in the field of mathematics. But not everyone is a fan, or buying the hype.
“Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified,” Andrew Sutherland, a mathematician at MIT, told Scientific American. “We should ask for receipts,” he said.
The papers are grouped into 372 “families” of related results, and a family can bundle a main theorem with companion arguments, consequences or alternative proofs. That makes 722 a count of manuscripts, not of solved problems. OpenAI says it posed roughly 4,000 problems to the model and kept the outputs it judged significant enough to publish.
The average result used the equivalent of roughly three hours of ChatGPT Pro thinking compute, per OpenAI. The Navier-Stokes claim last month looked very different, with 10,000 coordinating agents working for 88 hours.
OpenAI released abridged reasoning summaries for 10 of the results. That said, only 162 of the 722 papers come with a computer-checked main result, according to a formalization catalog in the repository. That is about 22% of the collection, translated into Lean, software that checks every logical step mechanically.
OpenAI itself says not all manuscripts have Lean formalizations and that “some of the unformalized results could have issues.” In other words, a lot of what they published could be wrong.
A passing Lean check confirms only that the proof follows from the statement as written in Lean. It does not show that the statement matches the original problem, or that the result is new or important, which is the part mathematicians now have to judge.
And this is where researchers raise their eyebrows.
“It is now the case that AI can output mathematical arguments in situations without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them,” The Institute for Advanced Study in Princeton, New Jersey, said in a statement. “We believe that human understanding of mathematics remains of paramount importance. How, in this new era, can we work towards a new paradigm that includes human understanding of mathematics as part of responsible scholarly output?”
Others, though, like Professor Abhishek Saha, are pretty excited. “It is a very big day for mathematics,” he wrote, but noted that most of the problems fit in the categories of “exceptional advances within an existing program” of “surprising breakthroughs.”
This means most of the problems in the set are interesting, but not impossible or game changing like the millennium problems. That spot is reserved for exactly one problem out of the 722: the Quasi-Riemann Hypothesis.
The release also falls short of what an advisory group at the Institute for Advanced Study recommended on September 29: the model name, the prompts, a summarized chain of thought, the time taken and the compute cost for every result. OpenAI published average compute figures and 10 reasoning summaries but no prompts, and says it is still working to release the model responsibly.
Daniel Litt, a mathematician at the University of Toronto, took the opposite view, arguing there is no reason to ask the company to keep the answers to these math questions secret.
Anthropic took a different route with its Lean-checked Fermat's Last Theorem proof last month, posting all 13 million lines publicly on GitHub. That proof formalized a theorem Andrew Wiles published in 1995, rather than claiming new results.
OpenAI says it will add Lean formalizations as it obtains them; for now, 162 of the 722 manuscripts have one.
AI outlook — possibilities, not facts
OpenAI will release the internal model and prompts within the next few months
Possible · Within months
More Lean formalizations will be added to the repository over time
Likely · Within months

Anthropic released Claude Haiku 5.5, positioning it as the cheapest, fastest, and most capable small model for high-volume tasks like document summarization and live customer support. Priced at $0.10 per million input and $0.50 per million output tokens, it offers up to 75% average savings over Haiku 4.5 and matches OpenAI's GPT-6 Luna pricing. Benchmarks show Haiku 5.5 outperforms Luna on OSWorld 2.1 (72.4% vs 48.9%) and Terminal-Bench 4.0 (39.2% vs 16.4%), while scoring 1620 on GDPval-AA v2.1. The model includes an adjustable effort setting and follows Opus and Sonnet 5.5 releases. Available on Claude website, AWS, Google Cloud, and Azure as claude-haiku-5-5, with new monthly API credits for Max and Team plans.

Yoido Full Gospel Church in Seoul said personal data tied to 850,000 members may have been stolen, with KISA flagging the breach and Oasis Security finding the data on an overseas server alongside attack logs suggesting AI sub-agent use; Sarang Church also reported a smaller breach affecting 89,000 members and 286 employees.

Google launched Playground, a web platform allowing users to create playable games by describing them in plain language, with refinement via chat and sharing via link. Available to U.S. users 18+, access tiers depend on Google AI subscription. Games run in browsers, support privacy or community publishing with safety screening, and may integrate Unity Spark for advanced features. The launch follows Google's Genie AI projects and coincided with a drop in game stocks as investors worry about AI disrupting traditional game development tools.

Google launched Nano Banana 2.1, an updated AI model that generates and edits images from text, now available in Gemini app, Search AI Mode, Ads and developer tools. The model offers better visual design, mask-based editing and subject consistency, scores 1,050 ELO in preference tests, and reduces API costs to $0.0336 per 1K image, half the price of its predecessor.

Developer hashfunction published metadata for 5.6 billion public TikTok videos spanning 2014 to 2026 on Hugging Face and datasocial.ai. The 460 GB dataset includes captions, engagement metrics, and AI flags, collected via a reverse-engineered mobile API.

Mistral AI launched Mistral Large 4, a 1-trillion-parameter AI model using a mixture-of-experts design with 49 billion active parameters per query. The model is positioned as a competitive open-weight alternative to Claude Opus 5.5 and GPT-6 Astra, with pricing at $1.36 per million input tokens and $4.18 per million output tokens. Benchmarks show strong performance in coding and automation tasks, trailing only Kimi K3 and Gemini 4 Argon in some tests. Mistral plans to release the model weights by end of October.