← The full briefing
Newsletter · Friday, 9 October 2026

A thousand helpers still miss four bugs

Claude will now farm a repo to a crowd. The chart from this morning still needs a query you can check.

This morning the chatbot became a live chart. This evening it became a crowd. Anthropic will let one Claude farm a repository to as many as a thousand helpers and stitch what they find. The demo number is 66 of 70 planted bugs, three times in a row. The four it missed, and the ones it might invent, are still your job. If you build, the question did not change with the clock: can you check the claim, and does “done” mean the database matches.

The swarm is a job runner, not a conscience

Claude Managed Agents dynamic workflows are in public beta. You turn them on with multiagent_20261001. A lead agent writes a JavaScript plan, splits the work, runs phases, and merges the report. One run can orchestrate up to 1,000 subagents. That is a ceiling on agents planned for the job, not 1,000 running at once. The earlier local preview sat around 16 concurrent helpers.

The test is synthetic and worth stealing anyway. Anthropic planted 70 bugs in a 116,000-line codebase. A single agent found 14, then 15, then 27. The workflow found 66 every time — 94.3% recall against 20.0% to 38.6%. They do not publish precision, duplicates, token burn, or how long a human spends checking the proposed fixes. File edits could auto-apply in the preview. Start scoped. The onboarding command is /claude-api managed-agents-onboard bug-hunter. Measure your own repo before you believe the 66.

That is the same lesson as Microsoft’s ThinkingBox-Bench from this morning. 507 workflows, twenty attempts each, graded on the final database. Kimi-K3 hit 93.89% at least once and 13.41% all twenty. Claude Opus 5 discovered fewer (79.09%) and repeated more (47.53%). 67.24% of the failures still looked finished. Pass-once is a demo. All-twenty is a product. 66 of 70 is still a recall number.

The chart still has to show the query

Anthropic also launched Claude Dashboards and Claude Motion. Dashboards is beta on paid plans and hooks to company data — Salesforce and Snowflake are the named ones — so the chart updates when the numbers change. Motion is beta on Team and Enterprise. It writes code, not generated video, so a typo is a fix. Docs, Slides, and Design left beta on every plan, including Free. Adobe will take a Motion file into Firefly. Pretty charts are easy. Trustworthy ones your boss can use are the test. Ask for the definition, the query, and the threshold that would change the decision.

The same Neuron edition says OpenAI told investors it hit roughly $50 billion in annualized revenue at the end of September, below a circulated $68 billion that counted partner gross. The product change is still the dashboard.

Saying no is not a wall

Arthur Holland Michel’s Technology Review essay is the one to read if you still treat a refusal as a safety plan. Labs train models to say no, then stack classifiers around them. Anthropic said one classifier type added 24 percent to chatbot compute. Verse broke two dozen models. Amazon unlocked Fable 5 hacking capabilities in less than three days. A Canadian shooter got shotgun advice after adding the word hypothetically. OpenAI’s country program includes the UAE. The Meta Oversight Board found five widely used models more willing to refuse a pamphlet about Thailand’s king than about Charles III.

OpenAI published 719 manuscripts covering 372 topic families from an unreleased frontier model — the same one behind the earlier Navier-Stokes write-up. David Robinson resigned after three and a half years. He led the current Preparedness Framework and oversaw safety reports on twelve frontier launches. The same week, OpenAI called three safety staff a “significant breach of trust”; they dispute the misconduct claims.

Thirty edits on one real-estate photo make the same point in pictures. Ideogram 4.5 and FLUX 3 kept at least 95% of the pixels on a small local change. GPT Image 2.5 Sunburst left about a fifth of the frame alone and rewrote the rest. Sunburst still leads the single-edit board. If you teach image prompts, score the chain, not the first pretty frame.

I am watching whether the swarm publishes precision next to recall, whether Dashboards show the SQL, and whether anyone prices cost per verified completed task instead of a green check.


Also worth a click
New on arXiv