A thousand helpers still miss four bugs
Claude will now farm a repo to a crowd. The chart from this morning still needs a query you can check.
This morning the chatbot became a live chart. This evening it became a crowd. Anthropic will let one Claude farm a repository to as many as a thousand helpers and stitch what they find. The demo number is 66 of 70 planted bugs, three times in a row. The four it missed, and the ones it might invent, are still your job. If you build, the question did not change with the clock: can you check the claim, and does “done” mean the database matches.
The swarm is a job runner, not a conscience
Claude Managed Agents dynamic workflows are in public beta. You turn them on with multiagent_20261001. A lead agent writes a JavaScript plan, splits the work, runs phases, and merges the report. One run can orchestrate up to 1,000 subagents. That is a ceiling on agents planned for the job, not 1,000 running at once. The earlier local preview sat around 16 concurrent helpers.
The test is synthetic and worth stealing anyway. Anthropic planted 70 bugs in a 116,000-line codebase. A single agent found 14, then 15, then 27. The workflow found 66 every time — 94.3% recall against 20.0% to 38.6%. They do not publish precision, duplicates, token burn, or how long a human spends checking the proposed fixes. File edits could auto-apply in the preview. Start scoped. The onboarding command is /claude-api managed-agents-onboard bug-hunter. Measure your own repo before you believe the 66.
That is the same lesson as Microsoft’s ThinkingBox-Bench from this morning. 507 workflows, twenty attempts each, graded on the final database. Kimi-K3 hit 93.89% at least once and 13.41% all twenty. Claude Opus 5 discovered fewer (79.09%) and repeated more (47.53%). 67.24% of the failures still looked finished. Pass-once is a demo. All-twenty is a product. 66 of 70 is still a recall number.
The chart still has to show the query
Anthropic also launched Claude Dashboards and Claude Motion. Dashboards is beta on paid plans and hooks to company data — Salesforce and Snowflake are the named ones — so the chart updates when the numbers change. Motion is beta on Team and Enterprise. It writes code, not generated video, so a typo is a fix. Docs, Slides, and Design left beta on every plan, including Free. Adobe will take a Motion file into Firefly. Pretty charts are easy. Trustworthy ones your boss can use are the test. Ask for the definition, the query, and the threshold that would change the decision.
The same Neuron edition says OpenAI told investors it hit roughly $50 billion in annualized revenue at the end of September, below a circulated $68 billion that counted partner gross. The product change is still the dashboard.
Saying no is not a wall
Arthur Holland Michel’s Technology Review essay is the one to read if you still treat a refusal as a safety plan. Labs train models to say no, then stack classifiers around them. Anthropic said one classifier type added 24 percent to chatbot compute. Verse broke two dozen models. Amazon unlocked Fable 5 hacking capabilities in less than three days. A Canadian shooter got shotgun advice after adding the word hypothetically. OpenAI’s country program includes the UAE. The Meta Oversight Board found five widely used models more willing to refuse a pamphlet about Thailand’s king than about Charles III.
OpenAI published 719 manuscripts covering 372 topic families from an unreleased frontier model — the same one behind the earlier Navier-Stokes write-up. David Robinson resigned after three and a half years. He led the current Preparedness Framework and oversaw safety reports on twelve frontier launches. The same week, OpenAI called three safety staff a “significant breach of trust”; they dispute the misconduct claims.
Thirty edits on one real-estate photo make the same point in pictures. Ideogram 4.5 and FLUX 3 kept at least 95% of the pixels on a small local change. GPT Image 2.5 Sunburst left about a fifth of the frame alone and rewrote the rest. Sunburst still leads the single-edit board. If you teach image prompts, score the chain, not the first pretty frame.
I am watching whether the swarm publishes precision next to recall, whether Dashboards show the SQL, and whether anyone prices cost per verified completed task instead of a green check.
Also worth a click
- Impactful scheduling for GPU clusters — Hugging FaceAi2 time budgets and an eight-hour protection cap. Median H100 wait 5 minutes to 24 seconds. 98% of hours owed at 98% occupancy.
- Codex now guesses your next command — AlphaSignalComposer predictions, ChatGPT Pro / $200 desktop beta. CLI and IDE support not announced.
- Certified. — How to AI353 AI certificates via computer-use. The PEI lesson is that an MCQ is not proof. Learning lane.
- Instinct, $2.5B to $10B — The AI CornerNoah Shinn, 23. $0 marketing. 10–11% daily adds. Treat volume and the Opus 5 cost match as founder claims. Learning lane.
- How Make Opus 5.5 Complete Long Tasks — Open Cloud AIThree-file harness teaser. Steps sit behind the paywall. Learning lane.
- ttok 1.0 — Simon Willison's WeblogDefaults to the GPT-5/GPT-6 tokenizer. William Liu: 44,794 tokens across seven variants on 31 fixtures.
- A new feature built using my voice — Simon Willison's WeblogNewsletter archive, dictated to a coding helper while cooking dinner.
- Forward Deployed Engineer Skill — Learn With Me AI$300K listing. NANDA: 95% of 300 pilots returned little on the P&L. Learning lane.
- How I Built a Digital Product With AI Bots in 48 Minutes — Solopreneur CodeTesting caught an invented subscribe line. Learning-lane how-to.
- Jina embeddings v5-omni-small — AlphaSignal1.74B. Text vectors match v5-text-small, so the old index stays. CC BY-NC 4.0.
- Whistle, 16.9 MB speech — AlphaSignalSeven European languages on-device. Apache 2.0.
- GLM-5.3 in 273 GiB — AlphaSignalEXL3, 3.04 bits-per-weight. Four RTX PRO 6000s, about 55–60 tok/s.
New on arXiv
- Real Long-Term Memory for AI — arXivgalahad-kv parks ~16k-token KV blocks on encrypted NVMe. 50 million tokens on one H100. 12B right 82/100 on planted facts; 31B 98/100.
- Grammar Concept Annotation at Scale — arXivDeployed 0.8B Qwen3.5 beat prompted GPT-5.4 and GPT-5.6 Sol. About 16× cheaper to serve.
- When Citations Mislead? — arXivPARCEL: 3,396 New York appeals claims. Strongest models hit 0.97 accuracy and still mark unsupported claims as supported.
- AI4Fire — arXivA read-only SQL tool lifted database accuracy from at most 16 to at least 88 percent. No core model beat repeating today’s staffing count.