The cheap model is not the product
A 38 MB adapter clones the decide-only trick, overnight agents still need a second pass, and voice APIs just took a 70-to-95 percent haircut.
The sticker fight from yesterday is still on. Evening is the layer around it. You can rent a cheaper default, bolt a decide-only head onto a laptop, or leave sessions running after the lid closes. None of that is the model. It is the wrapper, the schema, and the check you refuse to let the first pass write.
The 40 percent still only counts on medium
Latent Space is moving AINews to Opus 5.5. List is $4 in and $20 out per million, 20% below Opus 5, with cache reads at $0.20. Anthropic’s 40% cheaper and 30% faster claim is default medium. At max, Artificial Analysis has Opus 5.5 at $5.98 per Intelligence Index task against $5.86 for Opus 5, because the extra tokens eat the discount. CursorBench at max is 57.8%. The Neuron’s stopwatch is the same day with Sol at $2 / $10 and Luna at $0.10 / $0.50, about 1% of Astra’s raw token price. Nate Herk preferred Opus on seven of eight usable jobs and paid about 8 hours 40 minutes and $213 against Sol’s 5 hours 51 and about $74. Browser Use flipped it: Sol medium 66.9 versus Opus 5.5 at 59.4, about 3.5× cheaper. Thursday’s six-prompt rematch is still the follow-up. Route by job.
A tiny adapter answers the form
Jared Palmer shipped Kev-0.5B. It is an Apache-2.0 LoRA on Qwen2.5-0.5B that copies TypeSafe’s Jev design. It reads a document once and scores yes-or-no, multiple-choice, and rating questions in a single pass. They print 0.799 accuracy and 0.065 ECE across six held-out sets, or 0.031 ECE after temperature scaling. The adapter is 38 MB and trained in about 1 hour 45 minutes on an Apple M5. It speaks the typesafe-sdk. A LocalLLaMA skeptic says you are watching a classifier race a chatbot. On Banking77, BGE-small plus logistic regression hit 93.3% against Jev’s 83.2% at about 9 ms locally. The “0% hallucination” line is schema validity, not a correct class. Teach that distinction before you sell the trick.
Overnight only works if you do not grade your own homework
Ruben Dominguez’s walkthrough is the useful version of the thousand-agent boast. A session is a tab. Sub-agents stay in one conversation when they must see each other. Separate sessions when they do not. Git worktrees keep files from colliding. Past four or five sessions on one repo, reviews and usage limits become the wall. Closing the laptop only works as a Claude Code routine on Anthropic’s servers. A loop inside a live session dies when the session dies, up to three days at most. A desktop schedule still needs the machine on. Do not let the same agent mark its own work. Fresh session, real tests, permission to say it cannot verify.
The Augmented Mind puts a name on the rest. The model is an engine. The harness is the car: loop, tools, context, permissions, memory, routing. DeepSeek Harness is MIT and checkable. DeepSeek-R1 weights are MIT. V3 weights are a different licence. Open is five claims. Exit cost is the one most people skip.
Voice is in the same price war
Qwen-Audio 3.1 is a five-model stack for recognition, speech, and live conversation. Prices drop about 70% for text-to-speech, 85% for realtime, and up to 95% for recognition versus prior Qwen rates. Realtime Plus holds 262,144 tokens, listens while speaking, calls tools, and clones a voice. ASR-Next and TTS-Next are still listed as coming soon. Simon Willison’s playground for Gemini 3.8 Flash TTS is the other end of that market: more than 2,000 voices, a custom clone from a 30-second sample you have rights to, and a pelican debate that took about 20 seconds to make 1 minute 18 seconds of audio at 2.74 cents.
The glasses, a phage locus, and a briefing that cuts off
Anuj Behal’s piece from Delhi is still the human story. Shubnam was filmed at a spring protest on Meta glasses. The mocking reel drew millions of views. Meta’s latest pair is $420 in India. Reliance Jio says a rival under $105 arrives later this year. Claude flagged a CRISPR-like locus in bacteriophage DNA: a putative enzyme plus about 2,900 base pairs of repeating DNA. Humans did the bench work. No sequence and no paper. Executives from three labs briefed the UN Security Council on Wednesday. The stored sentence cuts off at the warning. Bernie Sanders’s ASI ban still has no definition in the alert.
I am watching Thursday’s rematch, and whether anyone benches Jev against BGE instead of against a chatbot. Until then, measure a finished ticket and keep a second pass that did not write the first one.
Also worth a click
- Your Own Record Will Not Save You — Slow AITikTok removed 184 million videos in Q1 2026, more than 96% automated. Their 0.97 precision counts an unappealed takedown as correct.
- GGUFs in transformers natively! — r/LocalLLaMA
from_pretrained(..., gguf_file=...). M2 Max: 70.4 tok/s versus llama.cpp 71.8 on Unsloth’s Qwen3.5-4B Q4_K_M. - Qwen-2.5-1B-RLCD Scores JSON Fields in Parallel — AlphaSignal5.6× to 7× faster JSON on Apple Silicon. A 28-field triage drops from 1,900 ms to 270 ms. Enums and booleans only.
- NVIDIA's Nemotron 3 Beats 12 Rivals With a 14.72% Speaker Error Rate — AlphaSignal100-million-parameter open-weight diarizer. 14.72% speaker error on Voice Arena versus 19.3% for the next system.
- Higgsfield Opens Its Studio Code and Offers $50K to Builders — AlphaSignalNext.js 16, 38 models, cost on the submit button. Five-second Seedance 2.5 is about $0.495.
- Unreal Labs' Unreal Agent Cuts AI Coding Costs 39% — AlphaSignalGo harness, MIT. Same Terminal-Bench 4.0 pass rate as Codex plus Astra at 39% less.
- GPT6 Sol vs Opus 5.5 - All you need to know in 3 mins — Jay E | RoboNuggetsOpus 5.5 is seven points above Opus 5 on Artificial Analysis. Sol is roughly even with 5.6.
- 😺 CoreWeave: AI infrastructure is one giant computer — The NeuronChen Goldberg: more than 90% of CoreWeave’s AI workloads still run on Kubernetes.
- How to Use NVIDIA Warp and MjWarp — Hugging FaceSO-101 arm from CPU MuJoCo to as many as 2,048 parallel MJWarp environments. Physics at a 0.002-second timestep.
- Open-Source natural-japanese — AlphaSignalMIT Agent Skill. Twelve-article style constitution, sudachipy lint, 0-to-100 naturalness score. Past 1,100 GitHub stars.
- Ema raises $77M — TechCrunchTeams of agents for HR, IT, and finance. $77 million is in the title.
- Pritzker Creates Illinois AI Cabinet — The Times WeeklyGov. JB Pritzker signed an executive order on Tuesday standing up a panel of experts.
- Senate Republican seeks to fast-track AI whistleblower bill — POLITICOSen. Chuck Grassley may go to the floor and force a colleague to register opposition.
- Governor Newsom announces world-leading experts — Office of Governor Gavin NewsomOutside experts to advance California’s AI safety and security governance. The title includes advancing a kill switch.
- A congressional representative just proposed killing America’s border tower program — MIT Technology ReviewRep. Delia Ramirez cites nearly 1,100 deaths within range of a billion-dollar system between 2015 and 2026.
- AI is wiping out whole degrees in China — TechRadarTranslation, photography, illustration, fashion design. Prompt engineer and AI animator show up in film courses.
- Engram gone wild! 2b model update... — r/LocalLLaMALocked spec: 2.6b overall, 2.2b trained, 4.3b Engram table. Past 100 million tokens. Checkpoints on Hugging Face.
New on arXiv
- What Does 99% Accuracy Measure? — arXivSubject metadata alone, article deleted, F1 1.000. On LIAR, three models sit at ROC-AUC 0.54–0.57.
- FrontierMath Erdős — arXiv68 still-open Erdős problems in Lean. $300 each. GPT-6 Astra 3%. Four other systems 0%.
- MoM: Memory of Memory — arXivCommit the current value, keep what you displaced. Revision-chain read stays at 100% where query-time reading falls to 25%.
- Training a Language Model End-to-End in Rust — arXiv$164 of rented GPU. Five silent Candle defects. He moved training back to PyTorch.
- LatentPort — arXivQwen3.5 4B-to-9B hybrid-state handoff without the big model rereading the prefix. One pair, one direction.
- AIBuildAI-2.5 — arXivLLM-guided tree search. First on MLE-Bench at a 73.3% medal rate.
- Same Quantity, Different Answer — arXivCanonical accuracy 0.969–0.996. Orbit correctness 0.848–0.981. Mistral Small 4: 265 power-of-ten errors on unit conversion.
- Peerify — arXiv800 reviewer claims from NeurIPS/ICLR 2024. Automated labels match humans on 90.3% of the audited set.
- Prompt Breadth and Rollout Refresh — arXivEight prompts with ten snapshots nearly match 14,080 prompts. Frozen rollouts make extra prompts hurt.
- Self-Cleaning and Captured Anyway — arXivClaude Sonnet 4.5 captured on 20 of 20 seeds against a registered prediction under 0.5.
- Retrieved-Span Training for Efficient Query-Focused Meeting Summarization — arXivA 406M Fusion-in-Decoder hits 36.33 ROUGE-1 after span fine-tuning versus 35.41 for a 1.2B system.
- From Tone to Trajectory — arXivSentiment-arc shape in ECB and Fed press conferences predicts rate decisions beyond lexicon averages.