← The full briefing
Newsletter · Tuesday, 15 September 2026

Gemini Live thinks while it talks

Google’s voice model took the speech bench, Factory printed a $5B ticket, and a replay test showed why “77% accurate” is not the same as “it will do that again.”

The morning pile was about who gets to sit in the lab and who is marking the text you paste. Evening added the product that actually shipped: a live voice model that can keep talking while it works a tool chain, a coding-agent round that triples a five-month-old valuation, and a reliability paper that should change how you read every agent leaderboard. If you build, the useful question is not “did they pause.” It is whether the voice you bought is doing the thinking, and whether the agent that passed once will pass the same prompt tonight.

The mouth is not the brain

Google’s Gemini 3.8 Live is two models. Live is the cheap, fast lane. Extended Thinking is the one that reasons in steps and speaks status updates while the API call is still in flight. Extended Thinking sits first on Artificial Analysis’s Speech-to-Speech Index at 82.6, with 68.6% on tau-Voice and 97.7% on Big Bench Audio. It will switch among 97 languages without a restart. It can look at a camera or a shared screen. Google has not published latency, a per-minute price, or how long a session can last. Generated audio carries SynthID. If you transcode it, test whether the mark survives.

That split is the same design OpenAI’s GPT-Live-1 already shipped. The voice layer handles interruptions. Astra or Sol in the back does the tools. Live-1 is first on the same index at 81.5, two tenths ahead of Grok Voice Think Fast 2.0 High, and 67.9% on Tau-Voice with Astra. First audio is 1.24 to 1.34 seconds. Grok is 0.70. The voice API is $0.05 a minute. If you are building a tutor or a support line, pick the backend on purpose. “The voice model” is not a complete order.

The factories got more expensive

Factory raised $200 million at $5 billion, up from $1.5 billion in April. Total funding is now over $400 million. Blackstone is investor and customer. The pitch is Droid plus a router that is said to cut token spend more than 60%, including air-gapped installs. They did not publish current revenue or how that 60% was measured. Treat the router claim as a homework assignment, not a spec.

Zoom out and the check looks small. Jessica Wachter’s accounting says hyperscaler spend through 2027 is near $1.1 trillion and that those firms need 2.7 times more productivity by 2030 to break even after capital cost, a 15% return, and chip depreciation. This year’s build is about $750 billion against $150 to $200 billion of AI revenue. Alphabet just posted a $5.9 billion free-cash deficit. Meta’s Louisiana site grew from a $10 billion, 2 GW announcement to a $50 billion, 5 GW plan wrapped in four-year leases.

OpenAI’s charity is spending on a different bottleneck. The Foundation is paying for medical datasets models still lack — $500,000 for bankrupt-biotech files, $40 million for cancer-vaccine data at UNC. Teslo’s point: about 70% of drug-development time and money is the clinical grind, and that grind is still a black box.

It passed. Ask it again.

IBM’s AppWorld replay is the eval lesson of the day. A ReAct agent on GPT-4.1 hit 77.4% average success across five repeats and succeeded on all five runs for only 53.0% of tasks. They call the difference the consistency gap. Guidelines distilled from one saved trace cut that gap to 12.0 points without dropping average accuracy. The analyzer needs no grader and no live re-run — one extra sample of five completions per decision. If your demo is a single green check, you have not measured what a user will feel.

The same reliability problem showed up in a sales costume. Salesforce in Claude is 37 skills, writeback only after the seller clicks yes, and about 7,000 sellers already on GitLab, Siemens, and Legora. Permissions inherit the CRM role. Stale stages become the agent’s personality.

The rulebook, the mark, and the invoice

Microsoft’s draft Code of Conduct still treats future MAI models as staff who must stop when a person says so. Roadmap into 2027, not a description of what is running today. Dario’s ask is the other half: put METR-style evaluators at desks under AEF-1. xAI, OpenAI, and Anthropic are named as cosigners in the title. A model that knows it is being tested may look aligned only then.

Anthropic is watermarking Claude’s text as it generates it. Light edits survive. A full rewrite, or an open-weight pass, can strip it. You cannot check. CrofAI sold Kimi K3 at $2/$10 and routed it to GLM 5.3 Flash. The sites 404’d the same morning. Name the model on the invoice.

If you make things today, Jay’s /animate refine loop and the ElevenLabs MCP inside Claude are the two how-tos that transfer. One prompt is everyone else’s reel. The skill is the feedback. What I am watching next is whether Google prints a price and a time-to-first-audio — and whether anyone reports Pass^k next to the pretty average.


Also worth a click
New on arXiv