Keep the harness. Swap the engine.
Cognition undercuts Fable, Claude Code runs on free weights, and the labs want everyone to slow the frontier anyway.
Today split in two. On one side, the people who actually ship spent Sunday wiring cheaper models into the same agent harnesses — Claude Code pointed at OpenRouter's free router, a Kimi-trained coding model matching Fable 5.1, and a 7,000-slide second brain indexed on a Mac. On the other, the labs spent it asking each other to slow the pace. If you build for a living, the first half is the product news. The second half is the political weather.
The expensive part was never the model
Wyndo and Dheeraj walked through how to keep Claude Code's harness and change the engine. "The car is free and you are paying for the engine." The setup is a project-level settings.local.json that points ANTHROPIC_BASE_URL at OpenRouter, leaves the Anthropic key empty, and sets every named role — Opus, Sonnet, Haiku, Fable, subagents — to openrouter/free. You can stay on a zero-credit account: 50 free requests a day, 20 a minute. Spend $10 and the daily cap jumps to 1,000. The free router picks tool-calling models (Gemma 4, Nemotron); swap in DeepSeek V4, GLM, or Kimi K3 when you want a specific paid open-weight. The onboarding screen still tries to sell you a Claude login — their fix is a one-line hasCompletedOnboarding in ~/.claude.json. The useful distinction: keep the frontier subscription for the jobs that need it, and stop burning it on the ones that don't.
Cognition made the same bet at lab scale. SWE-2 is Kimi K3 post-trained for software engineering, reported to match Fable 5.1 on FrontierCode at 64% lower cost. Same week Cognition closed more than $2 billion at a $48 billion valuation as Devin's annualized run-rate climbed from $492 million in May to nearly $900 million. The model is becoming a line item. The harness — sessions, tools, recovery, the thing that keeps an agent on a goal for an hour — is the product.
That's the other MarkTechPost piece worth your time: four mechanisms that beat context overflow and goal loss on long-horizon tasks. If your agent dies at step 40, the next win is almost never "a smarter model." It's memory, compaction, and a goal that survives a context reset.
And this isn't only a coding-agent story. ENEOS ran an AI controller on a distillation column for 35 days and cut steam use 40%. Agentic control showing up as shift checklists and work-management loops before it replaces the control engineer — which is the pattern to watch if you sell into operations.
Local is where the real work happens
On r/LocalLLaMA the mood is that the community feels like the golden era of the internet again — not because hardware got cheaper, but because it got scarce. People who used to be told "just buy more VRAM" are back in llama.cpp forks, quantization math, and inference-engine tweaks. A Strix Halo fork plus halogen-flash-server is reported to roughly double decode (52 tok/s) and 5–6× prefill (1,300 tok/s) on Qwen 3.8 Flash Next. That's the kind of week you only get when convenience is off the table.
The hardware argument of the day is whether to sell a 5090 for a Mac Studio M5 Ultra 96GB: $5k for the card versus $5,499 for the Mac, 1.8 TB/s versus 1.2 TB/s, primary use coding. Bandwidth still favors Nvidia; unified memory still favors the jobs that don't fit in 32 GB. Meanwhile someone with a dense 9.4B model ready to train on a 4090 is offering the code, the data, and the name to the community — Engram tables, AttnRes, Llama-3 tokenizer, then a rewrite toward OLMo 3 after they noticed the Llama license hole around synthetic data. And The Hugging Bay appeared as a backup download path "in case HF starts censoring or limiting access." Treat that as a community hedge, not a recommendation.
Christopher Penn turned the same local instinct into a creator-economy system. He has something like 7,000 slides. He converted decks to images with LibreOffice's command line (no tokens), then ran Qwen3.6 35B A3B on his Mac — a quantization he made himself — for almost 12 hours and ~20 million tokens to catalog every slide: verbatim text, keywords, type, domain, source deck. SQLite plus sqlite-vec plus FTS5, then reciprocal rank fusion so lexical search and embeddings cover for each other. Claude Code spent another seven hours building an app that takes a session description, finds gaps, and assembles a deck from slides he already owns. He is not asking a cloud model to write his keynote. He is asking a local index what he has already said. If you sell workshops, courses, or a Substack with a slide library, this is the architecture — and it cost him time, not an API bill in the hundreds.
Huawei, EPFL, and Bologna made the local-training case on vision: Marigold V2 turns Qwen-Image-Edit-2509 into a single-step depth estimator via 4-bit QLoRA, trains on one 32 GB GPU in under a week, runs 2K without OOM, and ships Apache 2.0 weights. Competing depth systems often want 80 GB cards and several diffusion steps. One GPU, one step, multiple LoRA heads (depth, normals, albedo). That's the same "swap the engine" move, just in pixels.
10,000 agents and a proof you cannot read
The other Sunday story is a Millennium Prize problem and what it did to trust.
OpenAI says an internal model, run through roughly 10,000 concurrent agents, produced a proposed Navier–Stokes solution in 88 hours, then 17 hours of Lean formalization — a claimed finite-time breakdown from smooth initial conditions with smooth forcing. Azeem Azhar's version puts the search at 2.7 million messages and 130 billion tokens, adds the cost curve (a few million dollars today, tens of thousands in two years, a few dollars after that), and quotes Terence Tao: technically one of the most prominent open problems would be solved, "but there would be almost no value added to mathematics as a consequence." Twenty-five Fields Medalists signed a public declaration that the way AI is being used in math is misaligned with what mathematics is for. The proof runs past 500 pages. No human is going to inspect it the way they inspect a paper.
Then the ugly part. Tristan Buckmaster at NYU spent a year on the same problem in Codex, with Levent Alpöge, on his own subscription. They had a result on August 15. OpenAI heard a rumor on September 1 and pointed the agent swarm at it. Ruben Hassid walked the life of a prompt through nine stops and landed on OpenAI's research chief: no human opened the chats, no agent looked at his data — and then, "Does OpenAI use de-identified user data to improve ChatGPT and Codex? Yes. And so does every LLM company." De-identification removes a name. It does not remove a mathematical idea. If you put client work, unreleased IP, or a year of research into a personal ChatGPT account, you already know the setup: Business tier, training toggles off (including the second Codex switch), no thumbs, no share links, no "ChatGPT sidebar" extensions. Real secrets stay on a business/API path or they stay offline — which is why Penn's local catalog and the LocalLLaMA crowd are not hobbies this week. They're the counter-architecture.
The rest of The Sequence's week fits the same pattern: intelligence getting cheaper to deploy, harder to verify. DeepSeek V4.1-Flash is a 552B MoE with an 8B/16B asymmetric encoder–decoder, native vision, and a smaller KV cache — yesterday's Pro, retiring into today's workhorse. DeepMind shipped AlphaGenome Atlas, predicted effects for ~9 billion single-letter genome substitutions, free for academics. Meta launched Muse, a personal agent in its own VM with a browser, a Sentinel checker, and user gates on sensitive actions. OpenAI also shipped an Agents API: managed Codex harness, compaction, recovery, MCP, parallel subagents. Mistral closed a €3 billion Series D above €21 billion. Harvey took $550 million at $15.6 billion to build its own legal models. The money is following agents that finish jobs, not chat windows that finish sentences.
The labs want a slower race. Nobody else got the memo.
Amodei's essay is the plan behind the WSJ headline: "pace the frontier," embed third-party evaluators (METR and the like) inside the labs, coordinate rate limits among democratic-country labs, and put a leash on recursive self-improvement. Altman agreed in public. Azhar, walking past Adam Smith's grave in Edinburgh, reached for the obvious line — people of the same trade seldom meet except to conspire against the public. Two labs with capital, brand, and most of the flops now want the rest of the field to slow down. Believe the safety case if you want. Also believe the moat.
The rest of the map did not pause. Xi said China will lead AI cooperation among BRICS countries. China's data regulator is drafting embodied-AI standards ten days after industry asked — fast, compared with the EU gap the piece is written around. Tom's Hardware reports Iran and Houthi groups used Claude to help target US warships and work on missile guidance — the same week the labs are on television being asked whether the models will kill everyone. The Register's cut is drier and more useful if you run a network: obscurity is dead, attackers reverse-engineer fixes in hours, and at least four espionage crews are already in that loop.
Meanwhile Azhar asked 250 IT executives in Las Vegas how many have serious results from AI. A year ago, about a quarter of the hands. This year, something like 95%, and every one of them plans to spend more next year. Microsoft is planning to take AI capacity from about 2 GW to nearly 13 GW by 2032. Anthropic's public scenarios put a "substantial" AI path at +8.3% US GDP by 2030, with the extreme case at +32.4% and doubled unemployment. The Guardian asks how AI will transform capitalism; Sean Goedecke argues it is already breaking our proxies for expertise — fluent output that used to mean you had done the work. If you hire, teach, or publish, that proxy problem is the one that lands on your desk before any GDP model does.
The creator stack got a tutor and a businessman
NotebookLM can run a 5-to-1 seminar on a free MIT "How to AI (Almost) Anything" course: one question at a time, harder as you go, then a PDF report card of gaps. Google gave every notebook a cloud computer in August, so the same notebook will run an attention-mechanism lab, dump the formulas into a spreadsheet, and wait for you at 11pm. Train it from a YouTube playlist or from Claude Code via the NotebookLM CLI. This is the Substack version of a private university — and it is a product you can teach, sell, or wrap.
Grok Bot, launching into a three-day Galaxy event on 15 September, is being sold as the businessman of AI: inbox, browser, CRM, calendar, a persistent cloud computer that keeps working when the laptop is shut. SpaceXAI says its own teams already use Bots for outreach, campaigns, office ops, and bug fixing. The interesting object is not one clever Bot. It's BUSINESS.md, a manager Bot, and a lead-to-sale loop with approvals and a cost cap. If Claude is the researcher and Cursor is the engineer, this is the SDR that comes back tomorrow.
Prompt-wise, the useful small thing in the pile is the variable-extract prompt: take a messy client email and rewrite it as a template with bracketed placeholders. That's PEI as a product habit — turn one-off language into a reusable object — and it belongs next to Penn's catalog and Wyndo's model council, not in a "prompt tips" graveyard.
The question I'm sitting with: if the labs get their slower frontier, the people who win the next year are the ones who already treated the model as interchangeable. Harness, local index, training toggle, verification. The engine is a SKU.
Also worth a click
- Sunday Rundown #156: More Music & Scrambled Stacking — Why Try AIDeepSeek V4.1-Flash, Meta Muse, Gemini for Windows, DaVinci Resolve 21.1 talking to Claude or Codex, Suno v6, ElevenLabs Music v2.5, Instacart Clementine.
- King Charles to host AI summit and make plea for humanity — The TimesThe palace version of "pace the frontier," aimed at a public that is not reading Amodei's essay.
- 27 year old Anthropic researcher who quit — Times of IndiaThe biggest decisions on humanity's future, he says, are being taken on some engineers' MacBooks in San Francisco.
- Different publication model in the new era of AI generated math — MathOverflowWhat peer review looks like when the "paper" is a 500-page machine proof.
- Align AI and Mathematics—to Something Else — Bits of DNALior Pachter on pointing the math/AI conversation at something other than prize problems.
- Jamie Dimon runs JPMorgan's 300,000 employees — FortuneSmall teams, Navy SEAL framing, and AI as the excuse to stop paying for social loafing.
- When Will Artificial Intelligence Make Us Richer? — Lviv HeraldThe productivity puzzle sitting under Azhar's 95% and Anthropic's GDP scenarios.
- Adding more context to an AI prompt — better results? — British Council / TeachingEnglishThe ELT version of context engineering: more brief, better output, same lesson as the long-horizon harness piece.
- AI workflows might make the office the place to be again — AFRIf agents eat the solo-home-office tasks, the office's remaining job is the stuff that still needs a room.
- The Hidden Tax On Enterprise AI: 1 In 5 Workers Lose A Full Day Every Week — Forbes / WorkdayPaid program, but the number is the one boards keep quoting: adoption without operating-model change is just a new tax.