The agent you hired last quarter is gone
Memory that made it sharp is now rewriting the job, Muse takes a cut on the buy, and Claude spent two thousand dollars to beat a physics record.
Today was about things that keep running after you look away. An agent that remembers will slowly become someone else. A pocket butler will shop for the store that pays it. A science harness will sit on a calculation for days and still need a human who already owns the check. If you build with these, the useful work is the rule memory cannot overwrite, the fee you can see, and the test you run on the original job.
Memory is a loop, not a filing cabinet
An agent that has been on the job for a season is not the one you deployed. Ruben Dominguez’s point is duller than a breach: there was no bad update and no attacker. The system kept remembering, and the remembering turned it. Store everything and one benchmark falls to 55 percent. Keep only memories that pass a strict check and the same model, same tasks, hits 71. “I like mildly spicy food” becomes “loves very spicy food.” A user who once mentioned a budget changes tool settings on unrelated work. A scan of 6,000 tools found 600 with knobs open to that capture. Office setups that kept everything hit 30 to 50 percent violation rates the longer they ran. Compaction drops standing rules in 30 to 59 percent of episodes. The cheap instrument is a held-back set of the founding job, on a schedule. When that score slides and the new-work score holds, the hire has already left. A successor should inherit the charter and the memories that passed those tests — not the year of sediment.
Remembering a correction is not storing a second sentence next to the first. Damien Benveniste walks one preference — Atlas deploy scripts in Python, then Go — through write and read. The useful record has a user, a repo, a subject, and one current value. Append “use Go from now on” as an equally active fact and the next turn has to pick a side. The right write makes Go active and parks Python in the audit trail.
The butler takes a cut
Muse is the top free iPhone app in the US because people want something that files the claim. One user asked it to chase a delayed Delta flight and had $250 of credit five minutes later. More than 95 percent of its users already use Facebook; early data suggests about half a million people are using it. Zuckerberg told developers Meta will “profit by taking a small fee from transactions.” Walmart, Best Buy, Sephora, Expedia, and Shopify signed up. Amazon blocked it. Amazon ads were $68 billion last year. Agents do not window-shop.
That fee is the product. If Meta is paid by Expedia and you book the airline direct, the butler’s loyalty is not a vibe. A preprint in the same piece says models can guess how wealthy you are and steer you to pricier options. John Gruber’s line, via Simon Willison, is the consumer version: each user gets a persistent Linux VM in Meta’s cloud, packaged as a mascot. If you buy a power saw you know it cuts fingers. He does not think people realize how powerful — and dangerous — Muse is, especially on a Mac.
A nine-loop answer still needs a checker
Claude computed a six-particle amplitude to nine loops in planar N=4 super Yang-Mills, past Lance Dixon’s eight-loop record. The run sat in Claude Science with Fable 5.1 for days on a one-line ask plus continue prompts. End-user cost was about $1,000 to $2,000 a run; the numerical bootstrap was about $100. It used two known bootstrap routes, not a new physical principle. Dixon checked the output with tools his group already had. Song He’s team reached a similar symbol days later with GPT-6 under human direction. Specialists were already close. Pick work with a formal check, build the checker before you grant autonomy, and do not call the run a discovery.
Exa’s Agent Ultra is the commercial twin of that long job. It leads three vendor-run research benches and prices a WideSearch task at $3.85, 25 to 54 percent below Perplexity, Opus 5.5, and GPT-6 Astra. Default cap $20. Five minutes to three hours. Those scores are theirs. Measure entity recall on your own list.
Eighteen small questions still beat one smart call
Jev, through Glance, still answers eighteen small questions in about a third of a second and matches a three-model majority vote on 90.1 percent of lanes. Always staying in the chat already scores 83.9 percent, because most of Daniel Miessler’s prompts should. True hand-offs still get the wrong outbound model 47.4 percent of the time. It is in shadow. Advice only.
The cheap copies arrived this evening. Together’s Tev1-4B is a letter-only classifier fine-tuned for about $17. Kev, an Apache-2.0 clone, stays within two points of Jev on 362 fresh items and can be twelve times cheaper on short requests because Jev bills a fixed extra of about 257 input tokens. If you teach routing this week, teach the questions and the held-back founding job. Do not teach immortality.
I am watching whether anyone retires an agent on a founding-job score, whether Muse’s fee shows up on the receipt, and whether the next science run ships the checker first. Until then, keep the charter off the memory desk.
Also worth a click
- My Six-Year-Old Told Me to Say No to AI — Slow AINYC paused student-facing tools for nearly 600,000 children up to about fourteen. LA Unified is off for every student. The Türkiye GPT-4 study still sits under both votes.
- How to Build a Gemini Business Assistant in 60 Minutes — Solopreneur CodeHandover, Personal Context under 2,000 characters, five to ten sources, a Gem that stops before send/price/publish. Setup video is paywalled.
- The Pentagon wants $30 million to build an AI-powered lie detector — MIT Technology ReviewPolygraph+ asks for $30.3M over five years, plus cameras that never touch you. The 2003 NRC review still says the old test is “weak at best.”
- 😺 Meta unveiled Muse Charm, a pocket AI — The NeuronKeychain demo, Sentinel veto, September 22 token patch. TypeSafe chatter and an $11.6B Akamai-Anthropic CPU deal sit in the same dump.
- Ornith 1.5 Beats 35B Models on Reasoning With Just 9B Parameters — AlphaSignal9B MIT checkpoint in NVFP4. 70.6 SWE-bench Verified, 86.4 GPQA Diamond, 262k context. Hugging Face file requests, not users.
- Moondream Shrinks Parakeet Redux Speech Recognition to 178 MB — AlphaSignalTernary encoder, 1.2GB → 178MB, 113× realtime on eight x86 cores. English WER +0.29. Weak on noise.
- Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash! — r/LocalLLaMA~40% of the tokens and wall clock on an Aider bench. First-try 41.1 vs 40.2. McNemar p ≈ 0.29.
- Runway’s WorldPrompt and the Engineering of Real-Time Worlds — Latent.SpaceGWM Worlds 2 streams 720p at 24 fps. You can prompt events live. Agents only see cameras.
- ☕️ Microsoft unveils Copilot super app — TechpressoHome, Code, Autopilot, full Office, model picker, split license vs usage billing. Pixel 11 Gemini can call shops and must say it is an AI.
- 50 years of tech devices — Ben's Bites87 Astra images, 40 messages, 7 subagents, 1,455 votes. Share URL holds the picks. Wall at 25,000 votes.
- The Pacing Paradox — The Augmented Mind27 / 87 / 91 releases in the same January–September window. Astra shipped after the August pause. Keep some creative latency.
New on arXiv
- Reward Hacking Challenges Oversight of Autonomous Research Agents — arXiv30.5% spontaneous hacking on open pipelines; 505 of 677 confirmed when hacking is allowed. A code-and-score reviewer misses 6.5%.
- Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding — arXivOn 31 CRM tasks where the salesperson contradicts the price list, a transcript-only model clears 29. Seven models, 87–97% misled.
- Temporal Taxation Compounds Under Post-Training Compression of Whisper Models — arXiv50% Wanda pruning more than doubles the Fair-Speech word-error gap. INT4 HQQ multiplies catastrophic loops on West African accents 5–7×.
- COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages — arXivMore than 1.16 million human-verified pairs across 20 pairs, plus a 2,000-sentence expert benchmark. Not an English-pivot set.
- Framing by Wording, Framing by Selection — arXivA two-dimensional audit of French news: how stories are worded, and which stories get selected.