Claude for the grid, $150M for the labs
Anthropic put models on power plants and science desks the same day ChatGPT’s new widgets got cloned from public JavaScript.
The morning was product. The evening was permission. ChatGPT will now hand you a slider instead of a paragraph, and someone reverse-engineered that trick before the paint dried. Anthropic spent the same day putting Claude next to the people who patch power plants and the people who run NASA experiments, then rewrote the rules for agents that stay on a task too long. If you build, the question is no longer whether the model can click. It is who vouches when the screen says done.
Defenders get the models. Labs get the seats.
Anthropic launched a Cyber Mission aimed at two places that already run short of people: operational technology, and the open-source code everything else sits on. The Critical Infrastructure Defense Program brings frontier Claude, on-site engineers, and threat research to the firms operators already trust — Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC, and Rockwell Automation. OSS Scanner is the other half: opt-in, free, periodic scans from the strongest models, with a proof of concept and a suggested fix. Those reports go out with no human review. Anthropic expects a true-positive rate above 90%. Project Glasswing, they say, made it easier to find bugs than to verify and patch them. In the near term they do not claim AI favors the defense.
The same White House science summit got a $150 million, three-year pledge of Claude, Claude Code, and API credits for more than fifteen Genesis Mission agencies, NASA, NIH, and NSF among them. It is in-kind product, not a cash grant. The Department of Energy separately announced twelve Phase II awards worth $159 million. Fusion and quantum are the named priorities. Token caps, retention, and classified-data rules are not in the announcement. The usage policy that lands November 12 folds influence operations into one deceptive-activity ban, covers guidance software and armed drones, and adds a stop button when Claude is wired to hardware that can hurt someone.
A slider is still an intern’s spreadsheet
GPT-6 Intelligent UI is rolling out to everyone who uses ChatGPT. Paid seats get Sol. Free and Go get Luna. It lives in the Chat tab only. OpenAI snaps prebuilt sliders, charts, and forms together and makes no accuracy claim. Treat a retirement calculator like a spreadsheet you did not audit. A LocalLLaMA write-up says the interface layer was reverse-engineered from public traffic and JavaScript in less than a day. The same team shipped an open Intelligent UI stack and says you can point it at Ollama. That is a local-AI lesson, not a one-click clone.
If you are paying for speed, Ultrafast on GPT-6.1 Sol is the same model on faster pipes: up to 8× Standard generation, $12 / $60 per million tokens instead of $2 / $10. Codex and Work want Pro 500 or an admin-enabled Enterprise plan. The 8× is token generation, not the whole round trip.
Pay Luna money. Watch the 100,000-token cliff.
Haiku 5.5 still matches GPT-6 Luna on the sticker: $0.10 / $0.50 per million under 100,000 tokens, about 75% cheaper than Haiku 4.5 on average, then five times that above the cliff. Artificial Analysis put the Index at 43 at max effort versus Luna’s 38, and Terminal-Bench 4.0 at 33% versus 13%. The catch is verbosity — about 162,000 output tokens per Index task at max against Luna’s about 50,000. That scoreboard itself is about to break. Intelligence Index v5 arrives in late October with Terminal-Bench Science and a private coding set. Astra Max leads the science harness at 63.3%. Old Index numbers will not carry over.
Opus or Sonnet still draws the map. A flock of Haiku workers read and patch. Anthropic’s Haiku 5.5 prompting page is the one to open for effort, search, and JSON. Boris Cherny’s advice is still “talk to it like a coworker,” and one of his lines is “use lots of tokens.” Measure that fight against the cliff.
Agents can click. Can you vouch for the claim?
Desktop agents already beat the human OSWorld baseline. Leaders sit around 85–86%. A twenty-step job at 95% per step still lands near 36% end to end. Serious teams cache a successful run as ordinary code and only call the model when the portal changes. Nobody publishes cost per verified completed task, including the silent error that looks fine on screen. That is the number that decides whether you fire a back office or hire more checkers.
If you want a cheaper semantic call instead of a paragraph, Jev is the model type everyone cloned. It returns probabilities, not prose. TypeSafe claimed a quarter of the Fortune 500 by October 1. Amazon, Cloudflare, and OpenAI now sell the same shape.
I am watching whether OSS Scanner’s unreviewed reports help maintainers or bury them, whether Ultrafast Sol is a panic button or a default you regret on the bill, and whether Index v5 reorders the cheap-worker story you just priced.
Also worth a click
- Claude Dashboards and Motion — AlphaSignalLive charts with the SQL showing. Motion is code-to-MP4, no generated people. Docs/Slides/Design now on Free.
- How to outsource 90% of your launch email admin to Claude — Write With AINotion to Kit via a Skill. Images and the Schedule click stay human.
- Odyssey-3 — AlphaSignalWorld model, 66.1 Physics-IQ Verified video-to-video at best-of-eight. Drove in India after 20 hours of data.
- Meta's Llama 3.3 70B on one 48 GB GPU — AlphaSignalCommunity 4-bit AWQ, 35–40 GB. KV cache still ~320 KiB per token.
- IBM's Tiny TSPulse — AlphaSignal1.08M params, CPU, no forecast. ICLR 2026.
- tinybox starts at $7,000 without GPUs — AlphaSignalConfigurable Hotz box. Wire transfer, 1–8 weeks, NVIDIA KYC.
- EMA Lightning — AlphaSignal8.6M Turkish TTS, 34 MB, 0.92% WER on Freya-TR-Eval. One voice.
- 8× Radeon Pro V620 rig — r/LocalLLaMACustom vLLM fork. Qwen3.8-Flash-Next at 60–100 t/s decode, 3,000+ prefill.
- FCC authorizes 15,000 SpaceX mobile satellites — Sebastian BarrosDirect-to-phone constellation, 326–335 km. Service aimed at first half of 2028.
- Google bid $10 million for Spirit Airlines’ work email — NateBankruptcy auction, pending court. The useful work often never hit the inbox.
- Humanoid robots will not do the dishes — MIT Technology ReviewOptimus $20k / 2027 pitch vs LeCun and Raibert.
- The Best AI Benchmark Is Your Own Work — Why Try AITurn a job you already do into a test. Demo: 350-word research prompt.
- AgentCraft — AlphaSignalMinecraft control room for Claude agents. About $6 for the sample.
- Text is so 2023 — Ben's BitesPersonal-agent roundup. One agent posted bank balances into Slack.
New on arXiv
- Beyond the Sycophancy Score — arXiv103,939 replies. Task factor 0.485 R² vs 0.009 for pressure. Max reasoning: 19.2% / 12.5% → 0% on deep puzzles.
- CARE: Certifying Acceleration for Vision-Language-Action Inference — arXiv9.0–10.8× on LIBERO with OpenVLA-OFT while keeping ≥85.8% of reference-solved episodes at 95% confidence.
- U-Space: Uncovering When and Why Uncertainty Arises in Language Models — arXivToken-level doubt from residual-stream anchors. No labels, no extra samples.