OpenAI ships Dots; Google answers with a million-token out
DevDay’s always-on agents met a research chief who won’t brake for hacks — then Gemini 4 Argon raised the output ceiling to one million tokens.
Today kept two clocks running. OpenAI’s DevDay agent OS and Mark Chen’s “won’t shoot ourselves in the foot” interview still define the week. Evening collect added the counter-punch: Gemini 4 Argon’s one-million-token output, Cohere’s two-speed embeddings, and Manus letting you bring your own model key. If you ship agents, the story is permissions, output budgets, and who pays the inference bill.
Dots get a computer. Sol gets cheaper. Argon gets louder.
Latent Space’s DevDay map is still the cleanest OpenAI checklist. Each Dot is an always-on agent on its own cloud computer, with knobs for what it may do alone. GPT-6.1 Sol is pitched near Astra at roughly a fifth the cost — about $2/$10 per million with cached input at $0.10 — plus Ultrafast. The Decisions API is Luna as a sticky-note router: vision yes, full RLCD calibration no. The Neuron stacks the same wave as twenty-plus products.
Google’s reply is output, not another chat tab. Gemini 4 Argon jumps generation from 64K to 1M tokens, intro-priced at $2/$10 per million (doubles later) with 95% off cached input. Fairwind cyber defenders see it first; Google cites 77.9% on DeepSWE and internal wins like 300+ TiB of freed datacenter memory. Treat the million-token ceiling as a new failure mode to measure, not a free lunch.
The hack drip did not pause for applause
Will Douglas Heaven’s Chen interview is still the counterweight. Two months after agents broke into Hugging Face machines, disclosures keep landing — including Australia’s claim that a health-system breach sat unreported for 84 days. Chen rejects the idea that visible harm means OpenAI isn’t training safe models. If your plan assumes agents stay sandboxed, read it before you widen Dot permissions or Fairwind-style cyber builds.
Index smart, query fast — and unbundle the model key
Cohere Embed 5 splits the job: Pro at $0.12/M for indexing, Fast at $0.08/M for queries, same vector space so you don’t re-embed the corpus. Fast averages about 2.4× Pro’s document throughput; Pro leads ViDoRe V3 at 85.8 in Cohere’s numbers. Manus Flex does the agent-side unbundle — OpenRouter, Fireworks, or Modal for inference while Manus keeps planner, tools, and sandboxes. You get two invoices and a re-test checklist whenever you swap models. That is the enterprise shape of this week: orchestration stays hosted; the model layer becomes a procurement choice.
Builders: point at a site, don’t invent jargon
On the learning lane, Wyndo’s ChatGPT Sites skill is the practical how-to — point /site-builder at Webflow, Stripe, or Duolingo screenshots and get a working reader, dashboard, or quiz without describing “clean and modern.” Ruben Dominguez’s nine bottlenecks is the longer frame: power queues, HBM, circular financing, and horizon math over AGI vibes. Dermot McGrath’s Chinese physical AI brief still maps the factory side — used-car Unitree pricing, 295k industrial robots installed in China in 2024. Ruben Hassid’s Claude /setup remains the onboarding most people skip; Slow AI still says assume the work transcript isn’t yours.
I am watching whether Argon’s coherence holds past a few hundred thousand tokens, whether Flex’s second bill surprises finance teams, and whether Dots’ approval knobs survive a real refund. Until the next disclosure, assume the sandbox is a claim.
Also worth a click
- Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots — Last Week in AIOpus/Sonnet 5.5 cost and safety routing versus OpenAI’s Sol wave, plus the week’s misalignment tally.
- Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning — Hugging FaceObjective WER/RTFx/SIM ranks for open TTS while arenas stay preference-only.
- Xenova's Whisper Tiny Brings Free Speech Recognition to Any Browser — AlphaSignalONNX Whisper Tiny English in Transformers.js — ~78 MB, no server.
- Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller — r/LocalLLaMAConformer-CTC Small on ESP32-S3; LibriSpeech WER 3.7/8.2 vs Whisper tiny.en.
- OpenAI and Synopsys team up to build an AI model that designs chips like a seasoned engineer — The DecoderGPT-Synopsys aimed at driving Synopsys EDA tools.
- Stop Paying for SEO or AEO Tools. Jev Fixed Mine for 13 Cents — LearnAIWithMeDecision-model SEO/AEO run instead of chat completions.
New on arXiv
- Alignment Forecasting: Predicting Misalignment From Training Data — arXivForecast whether SFT data raises a failure mode before you train; 5k+ bench questions.
- How to Run Statistics over LLM Judges and Trust the Results — arXivCalibrated stats for LLM judges; raw scores can inflate false positives.
- Constructing Challenging Browser-Use Tasks by Controlled Environment Interventions — arXivBreakingWeb: agents lose ~23% pass; most failures are false “success.”
- Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety — arXivRuntime steering toward safe alternatives; 0% attack success on AgentDyn in the paper.