← The full briefing
Newsletter · Tuesday, 22 September 2026

Opus 5.5 and Sol showed up together

Cursor turned one on. Chat still has not turned the other on. The rest of the day was a store that said no, a model that will not write, and a local editor that will.

This evening the collect doubled. Morning was a shopping agent hitting a wall and a model that will not write. Then two labs spent the afternoon cutting the price of a long job. Treat that as a routing problem: which call is allowed to be expensive.

The bill for a long job moved twice

Anthropic launched Claude Opus 5.5 at $4 in, $20 out, and $0.20 for cache reads per million tokens — 20% below Opus 5 on input and output, 60% below on cache reads. They say typical workloads cost about 40% less and the words come out more than 30% faster. Fast mode is $8 / $40, up to 2.5 times the throughput. The id is claude-opus-5-5 on the Claude Platform, AWS, Google Cloud, and Azure.

Cursor already has it. Max effort leads CursorBench at 57.8%. Default is 52.5%, a hair above Fable 5.1 at max, and about eleven points above GPT-5.6 Sol’s 41.7% in their comparison, at roughly a third the cost per task. Astra still leads AutomationBench and Terminal-Bench-Science. Anthropic says the real gap with Fable 5.1 is narrower than the chart. New accounts cannot edit earlier assistant messages to fish out the thinking. Measure a finished ticket, including cache hits, before you celebrate 40%.

About ninety minutes later, OpenAI shipped GPT-6 Sol and Luna next to Astra. Sol is $2 / $10. Luna is $0.10 / $0.50. The cache now survives a mid-run change to effort or tools. Sol at xhigh is 33.2% on AutomationBench against Claude Opus 5 max at 26.9%, at $0.27 a task — about 9% of that Opus 5 cost. Codex and ChatGPT Work have both today. Standard consumer Chat does not. The Neuron is running the three of them live. That file is the invite, not the scorecard.

The store still said no

Amazon cut Meta’s Muse off Amazon.com. The reasons stored: no permission, the agent did not identify itself, and the way it handled credentials looked unsafe. Meta says Muse sits in its own virtual machine and asks before it does anything sensitive. That argument did not reopen the cart. Shopify’s Tobi Lütke jumped in with Shop Pay. If Muse remembers what you want and owns checkout, the store becomes a supplier you can swap.

Muse hit an estimated 642,000 U.S. mobile daily users in twelve days, against 231,000 for ChatGPT at the same point. Apptopia estimates. Free is about 100 million tokens a week; Power $20; Maximum $100. Connectors opened 18 September with no published fee. Get on that shelf while listing is still free.

The same afternoon brief says Patrick Wardle found a Muse macOS zero-day that let a local attacker reroute cloud dictation. Meta patched within hours. David Singleton called it a local attack that already needed malware on the machine. Capability is still not permission, and permission is still not a lock.

Decide-only is a plugin now

Two Minute Papers walked through Jev without the launch-week swoon. Question plus options in. One answer out. Up to about 200 times faster than a chatbot, he says. It will not draft a paragraph. That is the feature. Daniel Miessler unstuck choice versus score: score is a ladder, choice is a pile. He is 85% sure. Simon Willison shipped llm-typesafe 0.1a0 so you can ask those shapes from the CLI you already use.

Ben Tossell listed what people are wiring: skip YouTube sponsors, filter a chat, search Gmail by intent. Using Jev to compact a coding-agent context is, in that letter, a terrible idea — it blows the cache discount both labs just advertised. Michael Spencer’s company numbers: 70–500 milliseconds, $0.042 per million input, output free, about 68% on their own workflow eval. Architecture still unpublished. Treat 68% as theirs.

Two local doors, and one you pay extra to rent

Alibaba’s Qwen Image 2.1 is still the first local editor in this pull that takes a photo and a sentence. Native transparency. Up to ten refs. ComfyUI: 7.2 GB INT8 for about 8 GB of video memory, or a 3.19 GB Q3 GGUF on about 4 GB. On a 16 GB RTX 5000 Ada laptop: 29 seconds to generate, 81 to edit. Official license text still bars commercial use of the model; Qwen said the pictures you generate are not licensed Materials.

Xiaomi’s MiMo-V2.6-Pro is still 46 on the open-weight Intelligence Index: 1.02 trillion parameters, 42 billion active, MIT, $0.435 / $0.87 per million. Title $3 million; a cited RL note $2.6 million. Do not mash those. Bedrock added Kimi K3 if you need a million-token window inside AWS: $3 / $15, about 1.76 times OpenRouter. Over $20 million in annual revenue, you negotiate before offering it as a service.

Timnit Gebru and Emily Bender spent the morning telling you not to confuse this summer’s press releases with independent results. Hold that next to Muse’s approval cards and a launch-day CursorBench number.

I am watching whether Sol stays off consumer Chat, and whether anyone repeats a Jev demo on their own inbox before they call 68% a product.


Also worth a click
New on arXiv