Opus 5.5 and Sol showed up together
Cursor turned one on. Chat still has not turned the other on. The rest of the day was a store that said no, a model that will not write, and a local editor that will.
This evening the collect doubled. Morning was a shopping agent hitting a wall and a model that will not write. Then two labs spent the afternoon cutting the price of a long job. Treat that as a routing problem: which call is allowed to be expensive.
The bill for a long job moved twice
Anthropic launched Claude Opus 5.5 at $4 in, $20 out, and $0.20 for cache reads per million tokens — 20% below Opus 5 on input and output, 60% below on cache reads. They say typical workloads cost about 40% less and the words come out more than 30% faster. Fast mode is $8 / $40, up to 2.5 times the throughput. The id is claude-opus-5-5 on the Claude Platform, AWS, Google Cloud, and Azure.
Cursor already has it. Max effort leads CursorBench at 57.8%. Default is 52.5%, a hair above Fable 5.1 at max, and about eleven points above GPT-5.6 Sol’s 41.7% in their comparison, at roughly a third the cost per task. Astra still leads AutomationBench and Terminal-Bench-Science. Anthropic says the real gap with Fable 5.1 is narrower than the chart. New accounts cannot edit earlier assistant messages to fish out the thinking. Measure a finished ticket, including cache hits, before you celebrate 40%.
About ninety minutes later, OpenAI shipped GPT-6 Sol and Luna next to Astra. Sol is $2 / $10. Luna is $0.10 / $0.50. The cache now survives a mid-run change to effort or tools. Sol at xhigh is 33.2% on AutomationBench against Claude Opus 5 max at 26.9%, at $0.27 a task — about 9% of that Opus 5 cost. Codex and ChatGPT Work have both today. Standard consumer Chat does not. The Neuron is running the three of them live. That file is the invite, not the scorecard.
The store still said no
Amazon cut Meta’s Muse off Amazon.com. The reasons stored: no permission, the agent did not identify itself, and the way it handled credentials looked unsafe. Meta says Muse sits in its own virtual machine and asks before it does anything sensitive. That argument did not reopen the cart. Shopify’s Tobi Lütke jumped in with Shop Pay. If Muse remembers what you want and owns checkout, the store becomes a supplier you can swap.
Muse hit an estimated 642,000 U.S. mobile daily users in twelve days, against 231,000 for ChatGPT at the same point. Apptopia estimates. Free is about 100 million tokens a week; Power $20; Maximum $100. Connectors opened 18 September with no published fee. Get on that shelf while listing is still free.
The same afternoon brief says Patrick Wardle found a Muse macOS zero-day that let a local attacker reroute cloud dictation. Meta patched within hours. David Singleton called it a local attack that already needed malware on the machine. Capability is still not permission, and permission is still not a lock.
Decide-only is a plugin now
Two Minute Papers walked through Jev without the launch-week swoon. Question plus options in. One answer out. Up to about 200 times faster than a chatbot, he says. It will not draft a paragraph. That is the feature. Daniel Miessler unstuck choice versus score: score is a ladder, choice is a pile. He is 85% sure. Simon Willison shipped llm-typesafe 0.1a0 so you can ask those shapes from the CLI you already use.
Ben Tossell listed what people are wiring: skip YouTube sponsors, filter a chat, search Gmail by intent. Using Jev to compact a coding-agent context is, in that letter, a terrible idea — it blows the cache discount both labs just advertised. Michael Spencer’s company numbers: 70–500 milliseconds, $0.042 per million input, output free, about 68% on their own workflow eval. Architecture still unpublished. Treat 68% as theirs.
Two local doors, and one you pay extra to rent
Alibaba’s Qwen Image 2.1 is still the first local editor in this pull that takes a photo and a sentence. Native transparency. Up to ten refs. ComfyUI: 7.2 GB INT8 for about 8 GB of video memory, or a 3.19 GB Q3 GGUF on about 4 GB. On a 16 GB RTX 5000 Ada laptop: 29 seconds to generate, 81 to edit. Official license text still bars commercial use of the model; Qwen said the pictures you generate are not licensed Materials.
Xiaomi’s MiMo-V2.6-Pro is still 46 on the open-weight Intelligence Index: 1.02 trillion parameters, 42 billion active, MIT, $0.435 / $0.87 per million. Title $3 million; a cited RL note $2.6 million. Do not mash those. Bedrock added Kimi K3 if you need a million-token window inside AWS: $3 / $15, about 1.76 times OpenRouter. Over $20 million in annual revenue, you negotiate before offering it as a service.
Timnit Gebru and Emily Bender spent the morning telling you not to confuse this summer’s press releases with independent results. Hold that next to Muse’s approval cards and a launch-day CursorBench number.
I am watching whether Sol stays off consumer Chat, and whether anyone repeats a Jev demo on their own inbox before they call 68% a product.
Also worth a click
- How I Turned a Simple Prompt Into $1,000 — The AI MakerTimo Mason’s $7 Article Architect. October 2025 launch. More than $1,000 over close to a year. When he moved to Claude he emailed every past buyer the raw prompt.
- I Made a $1,000 Vox Explainer in 30 Seconds With AI — Sanji Nai-ChienGPT-6 Astra + Higsfield + a motion-graphics skill. 30 seconds, three 10-second blocks. He would quote $1,000. Take Forge lists just under $1,200.
- The AI Memory Layer I Forgot Was Running For the Past Half Year — All Agents ConsideredHindsight behind Hermes. One week: 177 recalls, 183 writes. A failed batch recovered in about 90 seconds. He still keeps lasting rules in files.
- Vals AI's Claude Opus 5.5 Agents Proved a Faster Shortest-Path Algorithm — AlphaSignalTen agents, 15 hours, 733 messages, a Lean-checked C-HD bound. No real-graph timings. Constants are large.
- How UK AISI and EvalEval Are Making Benchmark Results Reproducible — Hugging FaceEvaluation Cards for HealthBench, FrontierMath, HLE, SWE-Bench Pro, Terminal-Bench 2.0. Opus 4 / 4.5 / 4.6 and GPT-5 / 5.2 / 5.4.
- Transformers now runs llama.cpp quants — Hugging Face
from_pretrained(..., gguf_file=...)on Apple Silicon, Qwen3.5 first. Unsloth Qwen3.5-4B is 2.74 GB as Q4_K_M. - Tencent's Hy Image 3.5 Merges Generation and Editing Into one Model — AlphaSignalHosted
hy-image-v3.5-preview. $0.024 an image. 3.0 stays the 80B open MoE. - Microsoft Shows 5 Communicating Agents Matching 33 Independent AI Attempts — AlphaSignalteam@5 = best@33 on ARC-AGI-3. Rest is Pro.
- Jev Chat Assistant Drafts WeChat and QQ Replies Using Android Accessibility — AlphaSignalMIT Kotlin overlay. Never auto-sends. Rest is Pro.
New on arXiv
- Recognition, Simulation, and Refusal — arXivPsyAgentBench. Asch on gpt-oss-120B goes from 0% blind to 83.3% when the paradigm is named. 41,904 trials.
- Memory That Looks Forward — arXivProspective commitments in a dated ledger. Hard-stratum recall@5 from 0.000 to 0.955. Author-built set.
- AdaMem — arXivSpend a fixed memory-token budget on the useful passages. Up to 3.2 points at 16× compression.
- Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare — arXivMamaBot-Llama +4.9% overall on 100 Nigerian maternal questions. Vax-Llama −5.2%, critical issues up 192%.
- Beyond the Text — arXivReAgent audits an agent-written paper against the repo. No accuracy percentage in the abstract.
- Privacy Personalization Trade offs in LLMs — arXivAfter stripping style tells, preference for the outputs fell to 13.0% while meaning stayed at 94.8%.