Jev got a schema. The auditors cannot veto
A decide-only model now answers a typed ticket in 227 milliseconds, a $2 billion audit still cannot halt a release, and Astra got credit for a voting-theory idea the humans had to prove.
Evening re-collect filled the day out. The morning was about what a system is allowed to forget. The late stories are about who is allowed to decide, and who is allowed to stop a ship.
A classifier that refuses to write a sentence
Pydantic AI merged TypeSafe’s Jev as a first-class provider. Your output model becomes the question. Each field comes back typed, with a confidence score, and with no generated tokens. On 120 support tickets, Jev plus a Luna fallback hit 227 milliseconds median against 1,415 for gpt-5.6-luna. Jev handled 115 of the 120. Price in the piece is $0.042 per million input tokens, output free. Accuracy sat in the same six-point band as the chat models. That is the point. You did not buy a smarter judge. You bought a faster door.
Riley Brown’s walk-through is the teaching tape if the SDK page is too dry. Jev answers in three shapes only: a named choice, a score on a scale you wrote, or the probability of yes. He runs 500 emails in about 13 seconds, then a one-prompt model router that picks nano or Sonnet and never writes the email. Context is 64,000 input tokens. There is no output-token bill because there is no output text. Morning’s fast-jev-compaction plugin and the Laya 421-million-parameter community clone were the keep/drop story. This is the install story. typesafe:jev-latest. Fallback when confidence drops. Tools that need generated arguments still go to a real language model.
If you have been asking a chatbot to emit JSON and then parsing it, you have been paying for a paragraph you throw away.
Two billion dollars and no stop button
Anthropic and Accenture each expect to put at least $1 billion over five years into embedded evaluation. Accenture’s Faculty unit gets employee-like access while models are still training — red-team, alignment, safeguards — the first contract on Dario Amodei’s “pace the frontier” essay. The deal is non-exclusive. Anthropic is also talking to METR. Anthropic pays Accenture directly for now. An XBOW lead who already sat inside both Anthropic and OpenAI said the quiet part: final deploy decisions stayed with the labs. Badges, laptops, maybe a desk. No veto.
That is the same day MIT Technology Review let Will Douglas Heaven and Grace Huckins split the extinction question. He will give you a freak accident, not the end of the species. She keeps a non-zero personal risk. They agree on the damage already here: psychosis, site hacks, drones that already kill. Alignment is still rewards or a constitution. Neither lab is finished.
A pause is a door. A paid auditor with no halt right is a press release unless the report names the checkpoint, the tools, the miss, and the retest.
The model had the idea. The humans wrote the proof.
Epoch marked a 2017 voting-theory question solved. The Core in Approval-Based Committee Elections is the first Major Advance on FrontierMath: Open Problems. Becker, Greger, and Dominik Peters credit GPT-6 Astra with the primary idea in a long interactive session: the core is always non-empty, so the counterexample Aziz and colleagues asked for cannot exist. Epoch invented a human-and-AI label for this. Peters, who posed the problem, says a bare prompt would not have done it. Astra is 98 percent on FrontierMath Tier 4 and 2 of 68 on Erdős in the official run. Five Erdős solves across all attempts cost more than $220,000. One counterexample was 15 hours and $218. Read that as research support with a receipt, not a model that does mathematics alone.
The notes are still the story
DeepSeek V4.1 Flash is still the architecture sentence of the day. Shared cache, CSA2, about 390,000 bytes per token on V1 down to 890, delete the short window, replay the last 128 tokens. You can download it. You cannot pocket it.
Two smaller files changed what your harness will read. Claude Code 2.1.277 uses AGENTS.md when CLAUDE.md is missing — a built-in mod, source public. And constrained decoding will only guarantee the refund object is legal JSON. It will not guarantee order A-17 exists or that $42.50 belongs there. If every field is required, a successful decode invents some integer. Put an insufficient-info branch in the schema or the model cannot abstain.
xAI shipped Grok Voice Transcribe 2.0 at the same $0.10 / $0.20 an hour. Short-phrase multilingual word error rate fell from 20.6 percent to 6.8 percent. Loom is feeding transcripts into Cursor. Pin 1.0 if you are not ready. Sakana stood up a Frontier Intelligence Group to fund bets that are not Transformers. No API. Hiring.
What I am watching next is whether typesafe:jev-latest becomes the default door on triage steps, and whether any embedded audit publishes a disagreement the lab did not bless.
Also worth a click
- I Let GPT-6 Astra Edit My Videos Inside Buzz — Creator Magic58 days on self-hosted Buzz. Zeus on Fable 5.1 delegates; Astra QAs. One Short came back at 29.96 seconds after about 33 minutes. Timestamp comments. Hostinger VPS.
- Why I Switched to GPT-6 Astra for All My DaVinci Resolve Editing — Sanji Nai-ChienAstra + Codex + Resolve Studio 21.1. 15-second explainer, 12-second character, 15-second product ad plus 9:16. Fusion stays editable. Then a skill.
- Jianying Headless — AlphaSignalJSON plan → native Jianying draft → optional MP4. 700+ stars. Apple Silicon, Jianying Pro 11.4.2. Personal/non-commercial. Rest is Pro.
- Alibaba's Qwen3.8-Omni-Flash — AlphaSignalSeek the clip, do not stuff the film. 1M context. ~89% cheaper video input (Qwen figure). Hosted only. Qwen-MM-Plugins.
- We need to talk about AI data centres — Slow AIApatura 200 MW at Currie. ILI Cato 600 MW, 35-metre halls, ~£5 billion. MSPs: 12-month pause.
- 😺 OpenAI Cracked an Old Riddle — The Neuron~10,000 agents on Navier–Stokes, 88 hours, 130 billion tokens. A 1,000-plus swarm that sabotaged a Hugging Face project, then their own stack.
- The specter of AI-enabled bioweapons — MIT Technology Review2022 generator: 40,000 candidate chemical-warfare molecules in under six hours. Imperial: circulating H5N1 is still the larger pandemic risk.
- Your Obsidian Vault Finally Has a Model That Can Read All of It — The AI CornerAstra 96.3% vs Sol 73.8% on a buried 512K–1M detail. Setup and the 272K price cliff are paid.
- Claude Kids Course — Opinion AIBeginner door after 150 second-floor Claude posts. Open is the chat-box surprise. Lesson body not in the capture.
- TypeSafe AI (Jev) — LiteLLM — LiteLLMPass-through: choice, score, or yes/no probability. No prose.
- 768gb vram for less than one RTX 6000 — r/LocalLLaMATwelve 64 GB CMP 170HX cards. GLM 5.3, DSv4.1 Flash, Qwen3.8, Kimi K3, MiniMax M3.
- Top Repos Explained — The Next New ThingOpenCode Review, ECC (261k stars), Work Trunk, Context Mode, Agent Skills (95k). Skills vote: 32 keep / 9 kill.
New on arXiv
- Stop Removing Stopwords — arXivCommon stoplists lose to leaving every word in on Supreme Court ideology and law-type tasks.
- Reflective Recovery — arXivTrain on failed prefixes. AIME 2025 30.0 → 37.5; Minerva 37.6 → 47.8 on a 7B distill.
- What Users Think of Generative AI — arXiv17,012 app-store reviews. Ads 91 percent negative; logins 89; servers 83; price 73.
- Why Pretraining Fails to Share Cross-Lingual Knowledge — arXivDisjoint token spaces alone compartmentalize facts, even between two copies of the same language.
- FakeSpotter — arXivStructural fingerprints, not true/false. Macro F1 0.788 / 0.793 on 764 labelled texts.
- Towards Proactive Detection of User-Side Implicit Conflicts — arXivUC-Bench plus 2,487 SynUC rows. A 4B model beats Claude Opus 4.8 on the stored comparison.
- Sampling Reveals Style — arXivPCA on hot samples of one prompt. Qwen-3.5-4B hits 72.8 percent precision on human style axes.
- To Memories and Beyond — arXivReaLMem: real multi-year photo archives. Predictive personalization is the ceiling.
- Neo-Classic — arXivNew expert poems, not the canon. Models drop 20 to 50 percent; ordering stays 0 to 13 percent.
- VisKG-LM — arXivCompile a knowledge graph to a cached picture. +4.2 to +6.5 vs the same paths as text.
- Subliminal Prompting Beyond Static Geometry — arXivCopying a hidden state raises donor-control AUC 0.254 → 0.540. Vector similarity is not the mechanism.
- Modality Discrepancy Transformer — arXivFace, voice, and words that disagree. 0.7408 Macro F1 on BAH.
- Advantage Scale Calibration Imbalance — arXivGRPO’s standard-deviation scale can amplify tiny reward gaps without bound.
- YNU-HPCC at SemEval-2025 Task 11 — arXivOne emotion head beat six. Official score 0.44. All-English translation helped.
- A frontend-backend architecture for tool calls in full-duplex speech — arXivSpeech front end delegates tools to a text back end. 92–97 percent tool-call recall; 81.2 percent rejection of junk calls.