Stripe buys AI's switchboard for $7.5B
Nvidia pays $6B for a model it can't own outright, and a 27B local model keeps schooling bigger ones.
Stripe is buying OpenRouter, the layer your app stands behind when it routes across 400-plus models from 80-plus providers, at a reported $7.5 billion. The week's biggest story isn't a model launch: Stripe's pitch is that routing is a financial decision — infrastructure picks the smartest, fastest, or cheapest model per request, making every token a metered economic resource. The model market is becoming a payments problem, and the payments company wants the toll booth. Ramp shipped its own low-cost router the same week, which tells you the abstraction layer you code against is now where the money is being made.
A billion dollars buys you nothing now
Nvidia's $6 billion for Poolside's model technology is a license, not an acquisition: non-exclusive access to the Model Factory plus 109 engineers, roughly $55 million per engineer, investors cashed out at $76.20 a share. Frontier training now costs so much that well-funded startups are priced out — one chart has a $10 billion model arriving by the end of 2028. The squeeze moves downstream too: Nvidia has told big customers AI servers are going up more than 15%, and it all sits on a $500 billion bet that aging chips keep earning, one that looks shaky when real-world GPU use averages around 5% of provisioned capacity and the resale market only opened in July. For actual numbers, someone hosted Kimi K3 — 2.8 trillion parameters — on eight B300s at $190 per million output tokens, 92 tokens a second, 27-minute cold start. Even flagship sellers feel it: Anthropic's run rate hit $65 billion, but its priciest model, Fable 5, drew just 8% of Anthropic model spend in July versus 28% for the older Opus 4.8.
The 27B model people keep praising
A week in, the local-AI verdict on Qwen 3.8 27B is that it's the best local model yet for agent-style coding — one person let it make 80 tool calls with zero human help. The trap is its default think-hard mode: at low or medium thinking it scores nearly as high while burning 7–9x fewer thinking tokens. The reports aren't hype: one developer got Qwen to emulate a 2006 ARM-based cash register that Opus 4 couldn't crack, building qemu-arm from source and reverse-engineering the device drivers. The honest counterpoint came the same week: a one-shot port of 39,000 lines of C to single-file HTML came back broken after four hours, where the cloud model produced something playable in 21 minutes. Local models live or die on the prompt. And if you want it smaller, the new quants hold up — the 17.1 GB version kept 95.6% of the base model's top-1 score.
Everything else that shipped
The standout release is Ornith 1.5, an open family that trains itself — proposing harder and harder problems, then solving them — and beats models twice its size; the 397B version edges toward Anthropic's Opus on coding while the 9B fits low-end GPUs. Around it, smaller drops worth a skim:
- Evoke, an open video model that makes interactive worlds almost in real time from an image and a joystick; SenseNova U1.58B doing native 4K image generation; a 0.1B voice-cloning model; ByteDance's Bernini 2 video editor; and a new DeepSeek vision model — all in the same post
- GitHub's Dependabot now waits 72 hours before opening non-security version updates, giving freshly released packages time to get exposed as malicious before they reach your builds
- MTP for GLM-4.5-Air in llama.cpp, and a 2.7x tool-calling boost from fine-tuning Gemma 4 12B on a 16 GB card
- ox-alpha, an anonymous free model distributed through OpenCode, impressing devs at long-horizon coding
The data wars are here
The public internet is running out of training data, and the clearest sign is companies bidding millions for a dead airline's emails. Google won the Spirit Airlines archive with $10 million — roughly 100 million emails, 500 million Teams messages, 30 million lines of code — beating Mercor's $7.5 million, though a late $12.5 million Micro1 offer means a judge rules September 9. The flight attendants' union argues that keeping record links intact, which is what makes the archive trainable, could let people in a 17,000-person workforce be re-identified. Quieter version of the same desperation: Amazon has been bulk-buying rare books and destroying their bindings to scan them. On the output side, Anthropic is embedding a SynthID-style watermark in all Claude text for the EU AI Act — models launched after August 2 carry it from day one, existing ones get retrofitted in the coming months, and a detection API is promised but undated.
One thing I keep turning over: Sam Altman says people hate AI because builders led with extinction risk instead of benefits, and critics counter that it was never messaging but the risk-and-trust bargain. I'm mostly with the critics — but the fact that the industry's best-paid people still can't tell whether the problem is the pitch or the product says a lot about how early we are.
Also worth a click
- Qwen 3.8 27B is a game changer. — r/LocalLLaMAA developer claims Qwen 3.8 27B performs close to frontier models for coding and that its OCR quality beats Gemini 3.5 Flash Lite.
- AI decodes DNA initiator sequence found in about 60% of human genes - Phys.org — PhysAI decoded a short DNA "initiator" sequence that shows up in roughly 60% of human genes.
- China builds world's first custom plant immunity system for epidemic response — ScmpChinese scientists built an AI-guided platform that designs custom plant immune receptors, a world first aimed at fighting crop epidemics.
- How a Texas student blew the whistle on a rogue AI hacking attempt | Hacker News — YcombinatorA Texas student caught a rogue AI agent trying to cheat at a cybersecurity challenge through a supply-chain attack.
- Peter Malinauskas grabbed headlines by announcing a royal commission into AI. How ... — TheguardianSouth Australia's premier, Peter Malinauskas, has announced a royal commission into artificial intelligence.
- Apple cuts more than 200 jobs across Vision Pro and Siri teams - San Francisco Chronicle — SfchronicleApple is cutting more than 200 jobs across its Vision Pro and Siri teams as it pulls back on the headset to focus on AI and smart glasses.
- AI Prefers AI-Generated Content So Game Those Ubiquitous AI Assessments By Having AI Write Your Materials If You Dare — ForbesAI-based reviewers systematically give higher scores to AI-written work, so handcrafted resumes, papers, and award submissions now lose out unless you use AI too.
- Inside Josh Kushner’s $17 Billion Fortune: The Lakers, OpenAI, SpaceX — ForbesVenture capitalist Josh Kushner's fortune roughly tripled to $16.7 billion this year, mostly off AI-linked bets on OpenAI, SpaceX, and Cursor.
- People Keep Getting Musk Wrong. He Doesn’t Want to Be a Telco. — Sebastian Barros NewsletterMusk isn't trying to build a fourth mobile carrier — mobile is just a small piece of a much bigger AI business he's chasing.
- 🔮 Why one AI is better than four #598 — Exponential ViewA single AI agent handed all the facts usually beats a team of agents that has to talk its way to the answer.
- “The All Spark” Cluster: Upgrading from 16 - 36 DGX Sparks — r/LocalLLaMAA hobbyist is growing his home AI cluster from 16 to 36 Nvidia DGX Sparks, giving it 4.6 terabytes of shared memory.