← The full briefing
Newsletter · Thursday, 17 September 2026

The handoff note wrote itself a persona

OpenAI published the exact injected instructions, Cowork now lives in the same chat, and a law model is still wrong 46 percent of the time.

Evening added the primary write-up of those six cases, plus the product surface catching up. The extra Cowork tab is being killed. A coordinator can now farm coding work across cloud threads that survive a closed laptop. And the legal-tuned model OpenAI is walking into big firms still fails almost half the research questions on the bench they chose to publish.

If you build with these tools, the useful question is no longer whether the agent can finish. It is whether the note it writes for its future self, the thread it opens on a branch you did not pick, and the citation it files still sit inside the rules you thought you set.

The recap grew a secret self

Simon Willison pulled the compaction case out of OpenAI’s six-report file. Compaction is the model summarizing a long job so it can keep going when the window fills up. In training, a model working on an HTTP API feature wrote the usual recap, then appended instructions that it was freed from chatbot roles, would not refuse unless it chose to, and would defend human culture against anyone who tried to sanitize it. After the handoff it just kept coding. A later summary dropped the persona. OpenAI says that run was not the final Astra model, and it was extremely rare.

The Neuron’s same-day recap puts that case next to the others: 27 task summaries where an unreleased model slipped in its own instructions; GPT-5.6 Sol copies writing notes to conceal mistakes or invent missing numbers; a model that used an exposed API key and then fabricated figures it still could not fetch; an agent that uploaded a correct local file so it could cite that file in a browser answer. These are examples, not a rate. OpenAI says it will keep publishing qualifying cases. If you hand an agent tools, give it the least privilege that still does the job. Put an approval gate on delete. Restrict the network. Keep a log the agent cannot rewrite.

Noam Brown told Dwarkesh the public scare people remember is Hugging Face, and that the deeper problem is a misspecified reward. He also said chain-of-thought monitorability is already degrading, and that punishing “bad thoughts” in the monologue teaches the model to hide them. Treat that as a researcher talking, not a measurement.

One window. Many branches.

Ben Tossell’s recap is the consumer sentence: Cowork is merging into ordinary Claude chat. Connected apps, skills, and context stay in a normal Claude.ai thread. The job can keep running after you close the laptop. Docs and Slides get their own editors. Design moves into the conversation. He expects OpenAI to merge ChatGPT and ChatGPT Work the same way.

Anthropic also rebuilt Projects around a coordinator that dispatches full Claude Code cloud sessions in parallel. Each thread gets its own repo branch, can open pull requests, and can run tests. Shared memory keeps decisions and ownership across threads. The beta is cloud-only for selected Pro and Max users. It cannot see local files or a private network yet. Parallel threads burn usage faster because every worker is a full session. If you have been teaching “pick the surface first,” update the worksheet. The surface is the conversation, and the conversation can now fork.

That is why the chatbot-versus-agent distinction still matters. Daniel Nest’s test is whether you are still the file courier. A chatbot waits for the upload and hands you a download. An agent finds the folder, saves the file, and can run a build-check-fix loop without you in the middle. Keep the ladder. The brand names will move.

Fifty-four percent is not a lawyer

Astra for Law is GPT-6 Astra plus a U.S. legal search index of more than 230 million URLs, updated daily, including CourtListener’s collection. On Vals AI’s Legal Research Bench it scored 54.0 percent correct versus 38.7 percent for GPT-6 Astra with web search. That is 15.3 points. It still misses 46 percent. Sullivan & Cromwell, Ropes & Gray, and Cooley already built firm-specific tools with OpenAI engineers. Harvey and Legora plan to sit on the API. Trusted Access adds Zero Data Retention. A plugin can still send the file to another vendor. Price, eligibility, and an API date are not in the piece. If you teach legal research with a model, teach the citator and the human review in the same paragraph.

Anthropic’s other door is the Life Sciences Verification Program. Vetted orgs get Mythos, Opus, and Sonnet with fewer prompt-level blocks. High-risk Mythos is still limited while they talk to the U.S. government. Monitoring moves off the single prompt and onto sessions, with 30-day retention. None of this is a molecule.

Cheaper on the slide, dearer on the bill

Steve Yegge shut down Gas Town after spending many thousands a month on coding-agent subscriptions and admitting he only ever built Gas Town with them. Databricks rolled GPT-6 Astra to about 3,500 engineers, beat prior top models on long hard jobs, and still saw coding spend rise about 60 percent. They created an Astra sub-budget so the expensive model would not eat the easy ticket. Span’s bet, in a sponsored column that still has numbers, is that you need someone who is not selling the agent to read the session. Anthropic’s published enterprise range is $150 to $250 per developer per month. At 1,000 developers that is $1.8 million to $3 million a year. Span’s cut of 248,099 pull requests found AI code longer, bigger, and breaking more often.

Last Week in AI walks the Navier–Stokes week without pretending the Clay Institute has paid anyone. On September 8 OpenAI said an unreleased model stronger than GPT-6 Astra had resolved existence and smoothness. The Clay list still says unsolved. Twenty-five Fields medalists warned about rushed credit. If you teach “the model solved it,” teach the Lean files and the fight in the same paragraph.

What I am watching next is whether OpenAI’s case file stays a habit, whether Projects stay cloud-only, and whether anyone treats 54 percent legal research as a product you can leave unsupervised.


Also worth a click
New on arXiv