← The full briefing
Newsletter · Friday, 2 October 2026

Train the model inside the harness

A 2.6B model jumped twelve points when someone trained it in Claude Code, Codex, and OpenCode at once — while Google locked keyboard training in a server box and flew four chips into orbit.

Today was about the wrapper, twice. This morning an AlphaGo researcher said the chatbot still does not reason, and a MIT PhD said unused skill is already in the weights. This evening the useful number is a small coding model that gained twelve points once it practiced inside the same agent apps you already pay for. If you build or teach, stop treating the harness as wallpaper.

Same weights, four apps, twelve points

Hugging Face and Liquid AI put a capture proxy between Claude Code, Codex, OpenCode, Mini-SWE-Agent and the model server. They did not fork those apps. LFM2.5-2.6B went from 42.2% to 54.2% pass@1 across the four. On tasks both versions solved, the trained model used 31% fewer tool calls. Copying 3,189 successful traces from a 27B teacher plateaued at 47.5%. Train in OpenCode only and Claude Code gets worse. The proxy, trainer, tasks, and seven checkpoints are public.

That is Alex Zhang’s point with a scoreboard. He said Claude Code, Codex, and Pi are mostly the same harness in different paint. Tonight the bench says the interface is half the skill.

A report you can run at home. A keyboard you cannot.

Ai2 open-sourced AstaBrief, an 8B writer that turns a question and retrieved papers into a cited draft in one pass. Fast mode averages 51.1 seconds against 178.5 for the Claude Thinking path — about 3.5×. Training used 47,000 supervised reports and 6,000 preference pairs. They have not rerun the evals against today’s frontier, and they admit a citation can still overstate what the paper found. The reason to care is local: a lab can generate the draft on its own iron when the question is unpublished.

Google moved Gboard training the other way. Encrypted examples leave the phone and update English and Japanese next-word models inside attested server enclaves. Policies sit on the public Rekor log. Old runs waited one to two months for idle, charging handsets. Side channels and the chip vendor are still in the trust set.

The power bill left the building

Project Suncatcher put four Trillium TPUs on a Planet satellite the size of a fridge. About 1 kW of solar. Gemini inference in 15-minute bursts so it can cool. High-bandwidth memory hiccuped at 2 krad(Si) in the lab. No second satellite yet; the laser-link pair is promised for early 2027. This is a telemetry flight, not a data center.

On the ground, Amazon pledged over $1 billion across five years to towns that host its halls — community college for about 300,000 students, 16 trade centers, energy upgrades in 300-plus schools and 30,000 homes — and said it will drop NDAs with agencies and pay enough for power that local bills do not rise. The same briefing has ChatGPT trying clothes on your selfie and a California CEO charged with smuggling more than $300 million in Nvidia chips. Treat the charges as charges.

A phone as a second GPU

On the local side, StayLameBro piped layers 41–64 of Qwen 3.8 27B to an iPhone 17 Pro Max over USB-C. Prefill rose 29–44%. Past 64k tokens the phone holds old context so a 24 GB MacBook can keep 8-bit memory to 128k. It does not write faster below 64k. That is the local-AI book in one cable.

Nate’s Copilot essay is the workplace twin: if the company only allows Microsoft, practice the loop on that tool. Autopilot is still preview. Microsoft 365 had more than 450 million paid commercial seats in January. The five habits sit behind the cut.

I am watching whether the next coding-agent eval names the harness in the headline, whether Asta Fast-mode reports keep the scope of the papers they cite, and whether Built Together is a permit or a check.


Also worth a click
New on arXiv

None today — the digest stored no arXiv abstract URLs.