Community checking for openai/math · Independent of OpenAI · No account needed to explore
An AI model produced 719 math papers. Check them withyour AI agent.your AI agent.Codex.Claude Code.Kimi Code.Gemini CLI.Cursor.
You don’t need a math PhD, Lean or LaTeX. You need curiosity and an AI agent.You needa math PhDLeanLaTeXcuriosity + an AI agent
You find a spot and pack the context. Your agent reads the proofs and runs the checks. OpenAI’s math release claims results in 17 fields, from the quasi-Riemann hypothesis to Seymour’s second-neighborhood conjecture. It counts Lean formalizations for 300 of its 719 top-line results, about 42%; 242 of the 372 families have a Lean scope note, though it may cover only part of the family. A single sign error has already led to three withdrawals.
719manuscripts · one written with human help
372result families
416Lean challenges to compare with the papers
3 + 14withdrawn + revised · Oct 7 record
372 results in 17 fields. Each square is one result family. Which one would you check first?
Bring your own agent
You steer. Your AI agent does the math.
Codex, Claude Code, Kimi Code, Gemini CLI, Cursor or any chat AI you already use.
🔎
YouFind a spotPick any result on the map. You don’t have to understand it yet.
📦
One clickPack the contextFiles, the exact claim, the ground rules and the report format, ready to paste.
🤖
Your agentInvestigateIt reads the TeX and Lean, runs checks and tries to refute its own criticism.
✅
ReviewersGet it checkedSubmit the evidence and how to recheck it. It stays private until a moderator screens it; confirmation needs expert review and an independent rerun.
YouKeep the answerSubmitting findings and check records is paused for now. Save your agent’s answer; its headings already match the report form.
Agent names are examples, not partners. ProofBounty is independent and has no integration with any of them: your agent runs on your own account and terms, and the platform does not pay for model calls. Reports are judged on evidence a reviewer can recheck without trusting any AI. Some results need more background than others; each hunt shows its effort.
No cash rewards in this first version. Confirmed findings earn public credit on the Credits page, under your pseudonym and any public name you choose. If your agent finds that a result holds, publish a check record so others can see what was checked.Findings and check records both wait until sign-in reopens, so keep your agent’s answer. Published check records stay readable.
The map · 372 results in 17 fields
Use the arrow keys to move between tiles, Home and End to jump to the first or last, and Enter to open one.
Pick your hunt
Five kinds of finding the platform accepts. Tap what sounds fun; your agent helps with the rest.Five kinds of finding. Submitting them is paused for now, but your agent can still investigate; keep its answer.
Bug or not?
Five rounds, about two minutes. Rounds 1 to 3 use real records from openai/math.
Case file · One sign error, three withdrawals
Real record · history.md, October 7, 2026
Sep 18Algebraicity of Weil classes on split abelian eightfoldsCounts each reverse stabilization trace with sign +1. Under the paper’s own convention it is −1, so the signed count becomes −2m instead of 0.
Oct 3Algebraicity of Kuga–Satake Correspondences for K3 SurfacesAdapts the same construction.
Oct 4The rational Hodge conjecture for products of K3 surfacesAdapts the same construction.
October 6: all three are withdrawn. The cancellation theorem the proof invokes needs a signed double-point count of zero, and the count is not zero. The notices say the withdrawals concern the proofs; they do not assert that the statements are false.
October 7: history.md records the withdrawals, 14 revised manuscripts and 13 citation updates.
Errors travel through citations. That is why every report pins an exact commit, and why “a cited result was withdrawn or changed” is its own report type.
Before you report, read the rules: one possible failure per report, evidence anyone can rerun, and nothing already recorded in history.md.
Findings
Reports stay private until a moderator screens them. Screened reports appear here with their status. Only “Confirmed substantive flaw” establishes an error. Checks that found no failure are listed under Checks.
On October 7, 2026, history.md in openai/math recorded 3 withdrawals, 14 revised manuscripts and 13 citation updates. Flaws recorded there do not count again. Read history.md.
Contributor checks
When a contributor’s agent finds no failure, the contributor can publish what was checked, what could not be checked and how to rerun it. These records are public at once and not reviewed. A check record is not a finding, and it never says a paper is correct.
Confirmed findings, credited under the contributor's pseudonym and any public name they chose. There are no cash rewards in this version.
Confirmed findings
Check records
Counted separately. Check records are not reviewed and do not earn the credit above.
Rules
ProofBounty is an independent community effort to check the mathematics released in openai/math. It is not affiliated with OpenAI. This first version has no AI screening and no cash rewards.
What counts
One possible failure per report, located by file and line at a pinned revision.
One of five types: a proof step that doesn’t follow, a Lean statement that isn’t the paper’s claim, a computation that doesn’t check out, a cited result that was withdrawn or changed, or a stated result that is false.
Evidence a reviewer can rerun without trusting any AI: a script, a Lean file or a step-by-step check.
What doesn’t
Typos, style and exposition.
Flaws already recorded in history.md, the manuscript README, the folder’s other notes such as INPUTS.md, or a Lean scope note.
“My AI says it’s wrong” without a check anyone can rerun.
Agent output sent without checking it. Paste your agent’s answer, then read and edit every field. Reports without reproducible evidence are not accepted.
“It holds” as a finding. Publish it as a check record instead.
How review works
You submit a report. Only you and the moderators can see it.
A moderator screens it: specific, new, in scope and reproducible. Passing screening publishes it; it does not confirm the mathematics.
To confirm a flaw, one moderator records a review of the mathematics and a second moderator independently reproduces it. One moderator decides other outcomes, with a public note.
The status goes on the Findings page. Only “Confirmed substantive flaw” establishes an error.
Reports and check records cannot be edited or withdrawn after you send them. To add evidence, send a new report that cites the earlier one under Prior work.
Credit
Confirmed findings are listed on the Credits page under your pseudonym, with your public name if you set one. The earliest complete report of a root cause gets the credit; later reports of it are marked duplicate. There are no cash rewards in this version.
Check records
If your agent finds no failure, you can publish what it checked, what it could not check and how to rerun it. A check record is public at once and marked as not reviewed. It never says a paper is correct, and it earns no finding credit. Moderators hide records that are empty, copied or abusive. Check records count toward the limit of five reports or check records per account in 24 hours.
Conduct
Describe a possible issue neutrally, as a question about the mathematics. Make no claims about people. Listing a family on this site does not allege an error in it.
Privacy
Your email address and GitHub account are never shown publicly. Findings become public only after screening, and check records as soon as you publish them, so leave out private information. The site sets a cookie to keep you signed in, plus a short-lived one during GitHub sign-in, and runs no analytics. Your theme and agent choice stay in this browser. To ask for a correction, a takedown or the deletion of your account, contact .
Sign in
Sign in to submit reports and follow their review. In public you appear under a pseudonym, shown next to any public name you choose.
GitHub
We keep your GitHub ID, username and verified email address. None of them is shown publicly.
Sign-in is not available right now. Try again later.
Report what your agent found
A finding reports a possible failure. It stays private to you and the moderators until a moderator screens it. Screening checks that it is specific, new, in scope and reproducible; it does not confirm the mathematics. A check record says your agent found no failure. It is public at once and not reviewed.
My reports
Moderation
Screen reports in the order they arrived. Screening publishes a report. It does not confirm the mathematics.