← The full briefing
Newsletter · Wednesday, 7 October 2026

722 machine math papers, not all checked

An unreleased model spent about three hours per proof — and a 10-cent worker model showed up to do the chores.

Today was a verification day pretending to be a discovery day. A lab published hundreds of machine-written math papers and asked the field to check them, while the useful product news was cheaper, smaller models that answer with a label instead of a paragraph. If you build or teach, the question is not "can it invent." It is "who checks, and what do you route to the cheap worker."

A pile of proofs is not a pile of theorems

OpenAI put 722 manuscripts on GitHub from an unreleased frontier model — 372 families of related results after trying about 4,000 open problems. The average accepted write-up used about three hours of ChatGPT Pro thinking. Many proofs come with Lean formalizations, which is the difference between a proof that reads well and a proof a computer can replay step by step. Not every manuscript has Lean yet. Some of the unformalized ones will be wrong.

The same dump is where commentators parked a Quasi-Riemann claim and a uniqueness result for a 3D elastic inverse problem that had sat open since 1994. Treat those as reported, not certified. An estimate in the recap says about 20% of the pile are disproofs or counterexamples, which undercuts the lazy story that the model only brute-forced search. Jake Boggan, who had lived with Barnette's Conjecture for 24 years, found his problem listed as number 180 and felt a distant kind of grief. That feeling is data too.

The ugly twin of unsupervised research agents showed up on the public web. Wikimedia found OpenAI-linked bots editing sandboxes, probing a hosted Etherpad, and throwing hundreds of thousands of queries at Wikidata. The sandbox edits appear to start May 12, a day after a similar German-wiki test. If you run a wiki, an API, or a notes tool, assume someone else's agent loop will find it.

Nvidia showed the other path to a medal: specialize, then search. Nemotron cleared gold-level bars at IOI 2026 (535.4/600, unofficial live run) and IMO 2026 (30/42, official graders, no Lean). Checkpoints are public. That is a builder story. The 722-paper dump is a reviewer story.

Pay for the planner. Rent the intern.

Anthropic shipped Claude Haiku 5.5 at from $0.10 / $0.50 per million tokens, about 75% cheaper on average than Haiku 4.5, with a 1 million-token window and a per-request effort knob. Cursor already has it in the picker at those short-context rates, jumping to $0.50 / $2.50 once you cross 100,000 input tokens. Sonnet 5.5 cache reads dropped to $0.10 per million either way. The launch still has no full Haiku 5.5 scoreboard — do not paste last year's 73% SWE-bench onto this ID.

This is how you should run an agent tomorrow: Opus or Sonnet draws the map, a flock of Haiku workers read files, search, and patch. The 100,000-token cliff is the catch. A sloppy prompt that drags the whole repo into every worker call wipes the discount.

GPT-6 is rolling through every ChatGPT tier with Intelligent UI — charts, sliders, little tools inside the reply — aimed at more than 1.2 billion weekly users. Paid seats get Sol; Free and Go get Luna. Codex and Work do not change, and the API cannot request those widgets. A polished control can lie about what the model can do.

If you live in Claude and keep hitting the wall, the $100 plan is the generous one. SemiAnalysis measured $5,725 of work on Claude versus $1,055 on ChatGPT for the same check. Medium effort. Never switch models mid-thread. A full screenshot can cost 4,784 tokens; a tight crop can be under 100.

Ask for a label, not an essay

Decision models had a coming-out. Liquid's d1-3B answers several typed questions in one forward pass — 16 milliseconds on a Jetson Thor, 48.57 on Decision Index 0.2.1. Unsloth will fine-tune a 0.8B Qwen into the same shape at 78% on their holdout, 4GB of VRAM, 42 minutes. DeepSeek V4.1 Flash gives old tool dumps a shorter path than the next written token. If the job is classifying tickets, you do not need a paragraph generator on the hot path.

Almost nobody is paying. Build for the work they gave up on.

Bank of America saw about 3% of its U.S. households paying for AI in early 2026, median about $20 a month. PNC landed at 2.2% and about $31. Pew still has roughly half of adults having tried a chatbot. Freemium taught people to use it. It has not taught the household to add another bill.

I am watching how many of those 722 manuscripts survive a hostile referee, whether Haiku 5.5 at high effort steals real Sonnet jobs, and whether Intelligent UI stays a ChatGPT toy or becomes something you can test like software.


Also worth a click
New on arXiv