Gemini Live thinks while it talks
Google’s voice model took the speech bench, Factory printed a $5B ticket, and a replay test showed why “77% accurate” is not the same as “it will do that again.”
The morning pile was about who gets to sit in the lab and who is marking the text you paste. Evening added the product that actually shipped: a live voice model that can keep talking while it works a tool chain, a coding-agent round that triples a five-month-old valuation, and a reliability paper that should change how you read every agent leaderboard. If you build, the useful question is not “did they pause.” It is whether the voice you bought is doing the thinking, and whether the agent that passed once will pass the same prompt tonight.
The mouth is not the brain
Google’s Gemini 3.8 Live is two models. Live is the cheap, fast lane. Extended Thinking is the one that reasons in steps and speaks status updates while the API call is still in flight. Extended Thinking sits first on Artificial Analysis’s Speech-to-Speech Index at 82.6, with 68.6% on tau-Voice and 97.7% on Big Bench Audio. It will switch among 97 languages without a restart. It can look at a camera or a shared screen. Google has not published latency, a per-minute price, or how long a session can last. Generated audio carries SynthID. If you transcode it, test whether the mark survives.
That split is the same design OpenAI’s GPT-Live-1 already shipped. The voice layer handles interruptions. Astra or Sol in the back does the tools. Live-1 is first on the same index at 81.5, two tenths ahead of Grok Voice Think Fast 2.0 High, and 67.9% on Tau-Voice with Astra. First audio is 1.24 to 1.34 seconds. Grok is 0.70. The voice API is $0.05 a minute. If you are building a tutor or a support line, pick the backend on purpose. “The voice model” is not a complete order.
The factories got more expensive
Factory raised $200 million at $5 billion, up from $1.5 billion in April. Total funding is now over $400 million. Blackstone is investor and customer. The pitch is Droid plus a router that is said to cut token spend more than 60%, including air-gapped installs. They did not publish current revenue or how that 60% was measured. Treat the router claim as a homework assignment, not a spec.
Zoom out and the check looks small. Jessica Wachter’s accounting says hyperscaler spend through 2027 is near $1.1 trillion and that those firms need 2.7 times more productivity by 2030 to break even after capital cost, a 15% return, and chip depreciation. This year’s build is about $750 billion against $150 to $200 billion of AI revenue. Alphabet just posted a $5.9 billion free-cash deficit. Meta’s Louisiana site grew from a $10 billion, 2 GW announcement to a $50 billion, 5 GW plan wrapped in four-year leases.
OpenAI’s charity is spending on a different bottleneck. The Foundation is paying for medical datasets models still lack — $500,000 for bankrupt-biotech files, $40 million for cancer-vaccine data at UNC. Teslo’s point: about 70% of drug-development time and money is the clinical grind, and that grind is still a black box.
It passed. Ask it again.
IBM’s AppWorld replay is the eval lesson of the day. A ReAct agent on GPT-4.1 hit 77.4% average success across five repeats and succeeded on all five runs for only 53.0% of tasks. They call the difference the consistency gap. Guidelines distilled from one saved trace cut that gap to 12.0 points without dropping average accuracy. The analyzer needs no grader and no live re-run — one extra sample of five completions per decision. If your demo is a single green check, you have not measured what a user will feel.
The same reliability problem showed up in a sales costume. Salesforce in Claude is 37 skills, writeback only after the seller clicks yes, and about 7,000 sellers already on GitLab, Siemens, and Legora. Permissions inherit the CRM role. Stale stages become the agent’s personality.
The rulebook, the mark, and the invoice
Microsoft’s draft Code of Conduct still treats future MAI models as staff who must stop when a person says so. Roadmap into 2027, not a description of what is running today. Dario’s ask is the other half: put METR-style evaluators at desks under AEF-1. xAI, OpenAI, and Anthropic are named as cosigners in the title. A model that knows it is being tested may look aligned only then.
Anthropic is watermarking Claude’s text as it generates it. Light edits survive. A full rewrite, or an open-weight pass, can strip it. You cannot check. CrofAI sold Kimi K3 at $2/$10 and routed it to GLM 5.3 Flash. The sites 404’d the same morning. Name the model on the invoice.
If you make things today, Jay’s /animate refine loop and the ElevenLabs MCP inside Claude are the two how-tos that transfer. One prompt is everyone else’s reel. The skill is the feedback. What I am watching next is whether Google prints a price and a time-to-first-audio — and whether anyone reports Pass^k next to the pretty average.
Also worth a click
- LM Studio 1.1.3 Brings Private On-Device Voice Transcription to Linux — AlphaSignalRealtime STT on Linux, Mac, and Windows. Apple Silicon and NVIDIA; AMD later. Speech model unnamed. No
/audio/transcriptionsAPI in the notes. - Hypit Lets Claude Code Turn Short-Form Videos Into Editable Code — AlphaSignalOpen-source video DSL at about 1.3k stars.
npx skills add hypit-ai/hypit -g. Free preview ends mid-table. - New BEST local AI music generator is here! — AI SearchUA 2 writes a score first. 3.96 GB quant for about 4 GB VRAM. Apache 2 code; CC-NC weights.
- If you have a 3090, or other 30xx for local LLMs, I have something for you — r/LocalLLaMAllamAmpere fork. Author claims 90+ TPS through 100K on a 27B. Author numbers, not a third-party bench.
- Devin Got a Mac — The AI CornermacOS VM, accessibility tree, TestFlight in Slack. No iOS pass rate. The prompt pack is paywalled.
- Can Skills Learned in Games Transfer to Real-World Work? — Latent.Space1830 railroad game → Finance-Agent only when they used a multi-turn tool agent. $3.6 million seed.
- Your AI Knows Your Context. Does It Know Your Process? — The AI MakerThin harness, fat SKILL.md. Seven questions. Paid-post-evaluator can reject a topic.
- I Asked an AI Agent Whether You Should Subscribe to My Newsletter — All Agents ConsideredFresh subagent, public URL only. Cache miss on the articles. “Soon” is not a live course.
- Quasar 1.1 Uses Quantum Data to Shrink a 438B Model — AlphaSignal37.6% fewer output tokens vs 1.0. No token price. No ablation on the quantum examples.
- Odyssey Builds One AI Backbone — AlphaSignalFrozen world model, small action heads. Sim-only driving at 77% of real-footage distance between interventions. Company-reported.
- The Sequence Knowledge - Issue 933 — TheSequenceClaude authored more than 80% of merged production code as of May. GPT-5.3-Codex helped debug its own training.
- Voodoo Dynamic Quant - Now MIT Licensed — r/LocalLLaMAPer-tensor GGUF gates. Author: wins on aggressive small Qwen3.5 quants; Unsloth still wins mid-to-high.
- Moonshot's Kimi Code 0.42.0 — AlphaSignalRemote Control and
[secondary_model]are defaults. Delete the old experimental env vars. - The AI Smartphone Supercycle That Never Came — Sebastian BarrosGenAI phones 36% → 45% of shipments; market +2% then −7.4% in Q2. Replacement near four years.
New on arXiv
- Lexical Prompt Compression for Large Language Models — arXivCPU-only lexical cuts. Harshest config: 40.3% fewer tokens, BERTScore-F1 0.876. Commonsense breaks first.
- Harmfulness Propagation Dynamics — arXivHarm direction rises with depth. HERALD: 89.3 F1 on OLMo2-7B, 98.4 jailbreak F1 vs 96.9 for tested guards.
- Same Patient, Different Order — arXiv1000 MedAgentBench reruns. At 8B and temperature 0.7, all 43 ordering groups changed orders across five identical runs.
- Causal Analysis and Mitigation of Spurious Onsets — arXivMoshi talked into silence in 12/40 five-minute zero-input runs. A mute-the-user check stopped 13/13 false starts.
- Clinical Reasoning Under a Partially Observed Objective — arXivHidden 80% entailment weight. Chasing visible n-grams drops entailment precision from 0.522 to 0.266.
- PhysMent — arXiv105 MuJoCo scenes. Up to 80% on qualitative single-concept; most models below 30% on the hardest quantitative bin.
- Hindsight Bias in Clinical Temporal Reasoning — arXiv171 cases. Full timelines push GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5 toward the hindsight trap. Masking the future cuts the bias without lowering accuracy.
- Token Merging for Multilingual Speech Recognition — arXivWhisper family, sixteen languages. Merging leftover tokens speeds decode with almost no accuracy loss.
- CVSS-X — arXivEnglish→28 languages, ~240,000 pairs each, over 16,000 hours, CC-BY-NC 4.0.
- Domain-Specific Jargon in Large Language Models — arXivLlama-3.1 beat its medically fine-tuned cousin on two jargon benches.
- From Token Probabilities to Semantic Constraints — arXivModelLog scores symbolic constraints that token likelihood hides.
- RFCLLM — arXiv4 tasks, 1,482 queries, 16 protocols. The abstract does not publish a headline accuracy.
- A Hybrid Hierarchical 1D-CNN-BiLSTM Framework — arXivExtractive clinical summaries. Every sentence is copied.
- Toward Complete Hospital Discharge Summarization — arXivEvidence-linked discharge sentences on MIMIC-III and Illinois notes. The abstract does not print a numeric gain.
- TestHallVQA — arXivMulti-image exam packets with controllable junk. F1-R² tracks reasoning and whether the model still finds the right evidence.