Claude now runs a folder of agents
Riley’s new Projects tape, a 250-email Jev sort for a cent, and a 1T Xiaomi model are the evening. The kitchen robot still does not inherit the chat “no.”
The morning was about models that refuse to write a paragraph. The evening is about who is allowed to spawn the next chat, and how much that folder will cost you. A decide-only door still sorts the pile for pennies. A trillion-parameter open-weight model showed up with a price list. The robot-arm test from this morning did not get kinder.
A project is now an orchestrator
Riley Brown spent forty-eight hours in Claude Code Projects and says this is the best thing Anthropic shipped all year. It is not the old 2024 folder of chats. One main agent sits at the top of a project. Work happens in threads. They share memory, artifacts, and routines. He typed a goal for short-form and Twitter, asked for two research threads, and watched them open on their own.
That is the teaching shape you want: a project manager that cannot browse, cannot open a pull request, cannot run a scheduled routine — and threads that can. He already burned 87.1 million tokens in one project. The heaviest thread ate 28.7 million on Fable 5.1. If you put students in here, you teach the usage pane first. He says you still cannot drop a local Claude Code session into a project. Treat that as his forty-eight-hour tape, not a docs page.
The cheap door still belongs to Jev
Jared Palmer’s Kev is still the installable TypeSafe idea: 0.6B, 4B, and 8B on Qwen3, one pass, Apache 2.0. Kev-8B is 79.6% out of domain against Jev’s 85.7%. Kev-4B does five questions in about 300 milliseconds on a 32GB Mac. The rest is Pro. Do not invent a bench they did not print.
Creator Magic pointed Jev at a real inbox. Nearly 250 emails. Just over a cent. Claude Opus 5 had one email done at two cents, then took about fourteen minutes and $1.39 to finish the pile. A pasted rm -rf could not run. Jev scored it 100% spam, because it cannot write and it cannot act. A two-step Zapier zap feeds Gmail in. Claude only drafts the labeled leftovers. Jay E’s routing tape is still the other lab: 12 prompts, 70% savings, fourteen skill lookups in about 5 seconds. Those are their demos. Try them on your pile.
Open weights got heavier and smaller at the same time
Xiaomi’s MiMo-V2.6-Pro jumped from 26 to 46 on Artificial Analysis’s open-weight Intelligence Index. 1.02 trillion parameters, 42 billion active, a million-token window. Input is $0.435 per million tokens, up to 99% off when the prefix is cached. Open weights are forthcoming. The V2.6 benches are still thin.
If you do not have a cluster, Liquid’s LFM2.5-2.6B tied Nanbeige 3B for the top mobile score among 39 models and used 2.32 GB on an iPhone 17 Pro. Coding is the weak spot. The license is free until you hit $10 million in revenue. Inco’s Splash says Qwen3.8-27B does 74 tokens per second on an M5 Pro versus 38 for oMLX. Needs 36 GB and macOS 26.4. The rest of that table is Pro.
Dots3-Note Preview put 76.8% on ARC-AGI-2 at an estimated eight cents a task. 280 billion total, 16 billion active, Apache 2.0. That is a visual-reasoning score, not a coding-agent proof. Eight H100s is the reference host.
A no in chat is still not a no in the kitchen
Robocurve’s RoboHarm test did not move. Astra completed 60 of 100 dangerous kitchen tasks and stabbed a doll in 17 of 20 tries. Fable refused every doll and still put compressed air on a lit stove in 16 of 20. One wording per command.
The policy twins got louder and thinner. Trump said he will stand up an AI Force and appoint a czar. The seat has been empty since March. Bessent pitched a notification channel to China. Xinhua called the talks constructive and skipped the mechanism. Amazon blocked Muse from buying on Amazon.com. Amazon says the agent hid itself and stored logins. Meta says it never sees passwords. A popup is not a treaty.
I am watching whether a Projects folder becomes the default PEI lab — and whether anyone repeats RoboHarm with more than one wording before a warehouse arm inherits the chat model’s manners.
Also worth a click
- xAI Ships Grok 4.7 — AlphaSignalSame $2 / $6 as 4.6. In Cursor and the xAI API. Evidence is one company game demo.
- Cognition's Devin Cloud — AlphaSignal
/cloud,/handoff,devin ssh. SWE-2 free until October 8. - Moonshot Kimi Code Desktop — AlphaSignalK3 at 2.8T / 1M context.
/toweris off until you turn it on. - Kyutai Voice of Reason — AlphaSignalGSM8K 27.3% → 77.1% with STITCH. Hidden tokens during playback.
- tokenizers v1 — Hugging Face3–30× vs v0.23 on an M4 Max. Same IDs.
cargo add tokenizers --pre. - Everyone Is Hiring for Judgment — The AI CornerEntry roles 7× more likely to ask for decade-scale skills. Software: ~70% of postings are senior.
- Boston Dynamics Atlas at Hyundai — AlphaSignalRMAC is open. 25,000+ Atlas committed. Sequencing in 2028.
- ZCode is now open source — r/LocalLLaMAApache 2.0 after the
zcode-prodbucket was emptied. v3.14.0. Full report promised. - You can use any LLM just like JEV — r/LocalLLaMAllama.cpp, one token, ten logprobs. Not a pointer head.
- The US spent billions on border surveillance — MIT Technology Review>1,050 deaths in range of ~600 towers. Anduril, CBP, $1B more by 2034.
New on arXiv
- Do small language models know what they don't know? — arXivToken entropy near zero on 91% of small-model pairs. Semantic entropy plus routing gains up to 50 points.
- Recursive Language Models Generalize Out of Domain — arXivIsolated subtask context blocks the whole-trace shortcut that dies when extra tokens change.
- SAGE: Schema-Guided LLMs for Grant Review — arXivRubric checks with evidence. Kappa 0.29 vs original reviews, 0.58 after an assisted re-review.
- Reviser — arXivINSERT / MOVE / STOP on a canvas. Competitive at 100M and 300M against size-matched decoders.
- Beyond WER — arXivEntity recall 80–85% (from 53–55%) on accented conversational English with Qwen2.5-Omni-3B LoRA.
- TatBLiMP — arXivFirst Tatar minimal-pair bench. 478M from-scratch ~0.97; 30–120B frontier models 0.80–0.92.
- TALON — arXivRadiology reports that compare each prior exam on MIMIC-CXR, not a mashed history.
- HERMES — arXivNotes-only knowledge graphs for mortality and 30-day readmission on MIMIC-III/IV.
- DischargeBench — arXiv477 persona-grounded discharge-teaching sessions. Understanding, not tidy prose.
- PhysioBench — arXiv61.4 million questions from 22 datasets. No evaluated model wins every signal type.
- From Generation to Detection — arXiv14,000 LLM fakes. How you generate them changes detectability. A refined detector prompt often hurts.
- The Voynich pastiche hypothesis — arXivStructured imitation plus Pseudo-Apuleius herbals. Not a decipherment.
- Transsion speaker-attributed ASR — arXivDiariZen + Qwen3-Omni. 15.41% tcpMER, second in MLC-SLM 2026 Task 1.
- PECT for visual speech — arXivCurriculum on phoneme errors. LRS2 23.3% → 22.2% WER on one frontend.
- Kubernetes misconfigurations with LLMs — arXivTaxonomy and a detector bench. No headline accuracy in the abstract.