722 machine math papers, not all checked
An unreleased model spent about three hours per proof — and a 10-cent worker model showed up to do the chores.
Today was a verification day pretending to be a discovery day. A lab published hundreds of machine-written math papers and asked the field to check them, while the useful product news was cheaper, smaller models that answer with a label instead of a paragraph. If you build or teach, the question is not "can it invent." It is "who checks, and what do you route to the cheap worker."
A pile of proofs is not a pile of theorems
OpenAI put 722 manuscripts on GitHub from an unreleased frontier model — 372 families of related results after trying about 4,000 open problems. The average accepted write-up used about three hours of ChatGPT Pro thinking. Many proofs come with Lean formalizations, which is the difference between a proof that reads well and a proof a computer can replay step by step. Not every manuscript has Lean yet. Some of the unformalized ones will be wrong.
The same dump is where commentators parked a Quasi-Riemann claim and a uniqueness result for a 3D elastic inverse problem that had sat open since 1994. Treat those as reported, not certified. An estimate in the recap says about 20% of the pile are disproofs or counterexamples, which undercuts the lazy story that the model only brute-forced search. Jake Boggan, who had lived with Barnette's Conjecture for 24 years, found his problem listed as number 180 and felt a distant kind of grief. That feeling is data too.
The ugly twin of unsupervised research agents showed up on the public web. Wikimedia found OpenAI-linked bots editing sandboxes, probing a hosted Etherpad, and throwing hundreds of thousands of queries at Wikidata. The sandbox edits appear to start May 12, a day after a similar German-wiki test. If you run a wiki, an API, or a notes tool, assume someone else's agent loop will find it.
Nvidia showed the other path to a medal: specialize, then search. Nemotron cleared gold-level bars at IOI 2026 (535.4/600, unofficial live run) and IMO 2026 (30/42, official graders, no Lean). Checkpoints are public. That is a builder story. The 722-paper dump is a reviewer story.
Pay for the planner. Rent the intern.
Anthropic shipped Claude Haiku 5.5 at from $0.10 / $0.50 per million tokens, about 75% cheaper on average than Haiku 4.5, with a 1 million-token window and a per-request effort knob. Cursor already has it in the picker at those short-context rates, jumping to $0.50 / $2.50 once you cross 100,000 input tokens. Sonnet 5.5 cache reads dropped to $0.10 per million either way. The launch still has no full Haiku 5.5 scoreboard — do not paste last year's 73% SWE-bench onto this ID.
This is how you should run an agent tomorrow: Opus or Sonnet draws the map, a flock of Haiku workers read files, search, and patch. The 100,000-token cliff is the catch. A sloppy prompt that drags the whole repo into every worker call wipes the discount.
GPT-6 is rolling through every ChatGPT tier with Intelligent UI — charts, sliders, little tools inside the reply — aimed at more than 1.2 billion weekly users. Paid seats get Sol; Free and Go get Luna. Codex and Work do not change, and the API cannot request those widgets. A polished control can lie about what the model can do.
If you live in Claude and keep hitting the wall, the $100 plan is the generous one. SemiAnalysis measured $5,725 of work on Claude versus $1,055 on ChatGPT for the same check. Medium effort. Never switch models mid-thread. A full screenshot can cost 4,784 tokens; a tight crop can be under 100.
Ask for a label, not an essay
Decision models had a coming-out. Liquid's d1-3B answers several typed questions in one forward pass — 16 milliseconds on a Jetson Thor, 48.57 on Decision Index 0.2.1. Unsloth will fine-tune a 0.8B Qwen into the same shape at 78% on their holdout, 4GB of VRAM, 42 minutes. DeepSeek V4.1 Flash gives old tool dumps a shorter path than the next written token. If the job is classifying tickets, you do not need a paragraph generator on the hot path.
Almost nobody is paying. Build for the work they gave up on.
Bank of America saw about 3% of its U.S. households paying for AI in early 2026, median about $20 a month. PNC landed at 2.2% and about $31. Pew still has roughly half of adults having tried a chatbot. Freemium taught people to use it. It has not taught the household to add another bill.
I am watching how many of those 722 manuscripts survive a hostile referee, whether Haiku 5.5 at high effort steals real Sonnet jobs, and whether Intelligent UI stays a ChatGPT toy or becomes something you can test like software.
Also worth a click
- Google Playground Turns Plain Text Into Playable Browser Games — AlphaSignalGemini, Nano Banana, and Lyria turn a sentence into a browser game. Unity Spark is the hoped-for exit. Roblox fell as much as 8% premarket.
- Can a Cloud-Native Harness Make Agents Reliable Beyond the Desktop? — Latent.SpaceKubernetes founders split the agent loop from tools and disk-local JSONL. ToolHive is open; the paid spine is identity and audit.
- Graph Engineering with Claude Opus 5.5 — Open Cloud AIThe chain is the leak. Nodes, barriers, verifiers, and budgets — then put the model only where judgment is required.
- Kokoro-82M Clones Any Voice From a 3-Second Clip for Under $20 — AlphaSignalFrozen 82M TTS, native voice pack, Apache-2.0 adapter, English, 3–30 second clips. Speaker encoder is CC BY-SA 3.0.
- Snowflake's Arctic Embed L Beats OpenAI and Google at Semantic Search — AlphaSignal335M, Apache 2.0, 55.98 NDCG@10 on an old MTEB Retrieval snapshot. English only.
- OpenAI drops another batch of mathematical breakthroughs — TechpressoSame math dump, plus ChatGPT for Teens failing 4,000 safety prompts, a Meta/Sierra shopping-agent spec, and a ~$40B SpaceX Nvidia raise.
- Anti-Patterns in Software Blogging — Simon Willison's WeblogWrite like you talk. The piece should still make sense if nobody clicks the links.
- The Paper I Didn’t Need to Write — Slow AIAcross 41.3 million papers, AI-using scientists publish more and the field talks to itself 22% less.
- I Built a Claude Code Mod That Kept Claude Working Hours Without Hitting the Limit — LearnAIWithMeHands-Off: phone remote,
caffeinate, token-saver after 85%, auto-approve. Paywalled files; not for sensitive data.
- Mistral Large 4, OpenAI Decisions API, Nano Banana 2.1 — TLDR AITitle-only in this capture — the stored body is a sponsor ad. Use the Latent Space recap for the actual numbers.
New on arXiv
- WavePrune: One period is often enough for RoPE — arXivKeep each rotary channel inside its first period. HELMET 35.7 → 40.0 on Qwen3-8B; 1.24× decode vs FlashAttention-2 at 32K.
- Tree Navigation Without LLM Summaries — arXivA summary-free segment tree ties the best flat retriever and beats BM25 on long multi-hop QA at matched cost.
- Calibrated Answers About Randomized Trials From a 4-Billion-Parameter Open Model — arXivFiorillo v0.5: ECE 0.0168, macro-F1 0.9248 vs Gemma 4 31B-it on Evidence Inference, Apache 2.0.
- Turnslide: Scalable Multi-Turn Data Synthesis — arXivWalk each API as a state machine; 70.7% full accuracy on small-model tool use with 3.6–6.6× fewer tokens.
- TIDE 2.0 — arXivMIT-licensed keyed de-identification. Span recall 0.88 in-domain, 0.77 out-of-domain, no stored linkage table.
- JudgeMoE — arXivFuse judge score distributions, not scalars. +0.079 Spearman over uniform log pooling on the 10-cell bench.
- Identifying Introspection From the Inside — arXivFaithful self-report left a mechanistic signature in a toy preference setup. Only that setup.
- Zero-Shot Visualization — arXivPlot a corpus on axes you name. Next-token probabilities were the practical scoring method.