Nothing matches those filters.

Lead

6

Video

1
16:49

Motion Graphics Without Plugins: The New Claude Method

Claude's design tool now generates finished motion-graphics videos, with animated titles and charts synced to a script, in about ten to twenty minutes with no editing software. The creator shows it layering animated graphics onto real footage and returning a ready-to-post MP4, editable by describing changes in plain English. He pitches three ways to turn that into income: freelance gigs on Fiverr and Upwork, faceless YouTube channels, and pre-made animated ads sold to local businesses for around $300 a month. Context given: professional explainer videos cost $3,000 to $15,000 per minute because they need hundreds of hours of manual work, which is the gap the tool exploits. The whole video is a how-to promo, heavy on monetization advice and lighter on the tool's limits.

Notes

Motion Graphics Without Plugins: The New Claude Method — Sanji Nai-Chien (YouTube, 2026-08-04)

Claim: motion graphics made in under 10 minutes with Claude Design (at CloudAI/design, or "design" in the Cloud app sidebar) — no After Effects, no keyframes, no editor. His hook: existing tutorials only demo websites/decks; this covers the animation side plus three monetization routes.

Why the market is the pitch
  • 91% of businesses use video as a marketing tool; a professionally made animated explainer runs $3,000–$15,000 per finished minute.
  • Traditional cost basis: one animation ≈ 10 hours in After Effects, which first requires learning 8 separate topics ≈ 40 hours of tutorials, then 100+ hours of practice — months before paying work.
  • Claimed arbitrage: freelancers still quote two-week delivery for what Claude Design does in minutes; "the competition doesn't know that the industry has just changed."
Setup (the step "everybody gets wrong")
  • Go to CloudAI/designdesign systems tab → create design system. Name it, one sentence about the brand, drag in logo/fonts if any.
  • Flat-JPEG/screenshot/no-logo tip: Hexagon AI → image → Recraft 4.1 in vector mode converts an existing logo to SVG (or creates one). SVG matters because Claude Design animates the logo piece by piece instead of sliding a flat image.
  • No brand at all: describe style in the notes (e.g. "bold, dark, premium, high contrast"); ~5 min generation returns a "rulebook" of colors, fonts, spacing, motion ideas.
  • Animation: select the design system → animation template → prompt like "create a logo reveal animation that fits my brand." He warns against the generic prompt "make me a cool animation."
Two paid workflows
  • Pure motion design (no footage): take a script, record or AI voiceover, get a timestamped transcript from any transcription tool, paste it in and prompt "create full-screen motion graphics sync to this transcript." It asks clarifying questions (layout, empty space) and returns scenes timed to words — icons popping in, charts drawing themselves, text landing on beat.
  • Motion over real footage: upload a clip (he used a 20-sec talking-head), attach transcript + design system, prompt for motion on top of footage. Returns finished MP4, not assets/keyframes/project files; titles/labels layer on synced to words, face never covered. Three edit tools: (1) plain-English change with a timestamp; (2) annotate button — draw a circle around the problem and type the replacement; (3) tweaks — e.g. toggle for accent color and animation speed. Export via share → MP4. Total: 10–20 min.
Three monetization routes
  • Fiverr/Upwork: market rates $75–150/hr, mid-level projects $1,500–5,000. Build a portfolio in one afternoon — 5 fictional brands × 5 demos (logo reveal, product promo, explainer, stats animation, social media ad). Price below market, compete on speed (24-hour vs 7-day delivery), stack reviews, then raise prices. A logo animation ≈ 20 min in Claude Design.
  • Faceless YouTube channel in education niches (finance, psychology, health, AI — "some of the highest RPM"). Script with Claude → AI voiceover → timestamped transcript → full-screen motion graphics → MP4 under the voiceover. Pays three times: ad revenue, sponsorships, then as a living portfolio.
  • Local business retainers: median production cost ≈ $2,500/min; pick 10 local businesses already running ads (check Meta Ad Library or Maps), and make a 15-second branded animated ad before contacting them (hook text, offer, animated elements, logo outro, sized for reels/stories). Pitch line: "I made this for you, if you want it, it's $300, and I can make you four of these every single month." One client × 4 ads/month = $1,200; five clients = $6,000/month. Reuse their own raw phone footage via workflow 2; optionally pair with HeyGen for AI footage (his example: a scratched car transforming for a detailing studio).
Caveats / limitations
  • The advantage is presented as time-limited: "Every month that you wait, more and more people are going to figure it out."
  • Vendor-named, promotional tone; no independent benchmarks, failure cases, or real revenue proof offered.
Transcript · 13,762 chars
What if I told you that motion graphics like this took less than 10 minutes to make? No After Effects, no key frames, and no editor. With only one tool called Cloud Design. And that's not even the crazy part. It can take real footage like this clip right here and layer motion design directly on top of it, synced to every single word. Now, you've probably heard people say Cloud Design is amazing. And they're right, it is. But nobody shows you this side of it. Every single video out there is the same. Here's a website, here's a presentation deck. Now, that's great, but every single time you're left with one question that nobody can ever seem to answer. How does this make me money? And how does it actually help me? So, in this video, I'm going to show [music] you step-by-step how to create stunning, professional-looking motion design videos in literally minutes, all inside of Cloud. And then, I'm going to show you three different ways to turn those videos into actual money with the exact steps for every single one. And if you're not a designer, an editor, or an artist, that's good because you don't need to be. That's the entire point. So, why Cloud for motion design? Well, if you're already in the space, then you just saw the answer in the intro. [music] It's your normal work, but at four times the speed. And if you're somebody who's never touched motion design in your life, then this video is something you should care about. Because today, 91% of businesses use video as a marketing tool. And a professionally made animated explainer costs three to $15,000 for one single finished minute. Now, why so expensive? I mean, just take a look at this animation. Making something like this the traditional way takes a professional around 10 hours in After Effects. And honestly, that's the easy part because first, you'd have to actually learn After Effects. To recreate this one animation, you would need to learn eight separate topics. That's about 40 hours in tutorials alone. A full work week of just watching videos. [music] And then you need over 100 hours of practice before your work looks even remotely clean. Realistically, it's months before anybody pays you even a cent. That's what those $3,000 price tags are about. [music] And that's exactly what has just changed, and the market has yet to realize it. Freelancers are still quoting two weeks of delivery for a single animation that Cloud Design does in minutes. So, demand is high, prices are high, and the competition doesn't know that the industry has just changed. Windows like this don't stay open for a long. So, here's exactly how the animation feature works. First, the setup. Just go to CloudAI/design, or click design in the left sidebar of the Cloud app. Now, here's the thing that everybody gets wrong. They land on this page, click the animation template, and type, "Make me a cool animation." Do not do that. [music] That's how you get the generic output that looks cheap. Instead, you're going to need to spend 5 minutes on one step that changes everything. You need a design system. So, go to the design systems tab, and click create design system. Give it a brand name, and add a sentence about what the brand is. If you've got a logo or fonts, make sure to drag them in. And here's a tip that almost nobody knows. If you or your client only has a flat logo, a JPEG, a screenshot, or even no logo at all, go to Hexagon AI, and then click image, and choose Recraft 4.1 in vector mode. Now, you can turn a ready-made logo into SVG, or create one. So, why is it so important? SVG is a vector format, which [music] means Cloud Design can animate the logo itself, piece by piece, instead of just sliding a flat image around on the screen. Now, that one tip alone puts your animations a level above everyone else's. And if you don't have any brand at all in your design system, just describe the style you want in the notes. Something like bold, dark, premium, high contrast. Hit generate and let it run for about 5 minutes. What you get back is a full rulebook of colors, fonts, spacing, and motion ideas. And from now on, every animation you create with this design system attached looks consistent and [music] intentional. So, now, let's make our first animation. Go back to the home screen, select your design system, and choose the animation template. And let's start simple. Type, for example, create a logo reveal animation that fits my brand. That's it. Send it off. Now, just this logo reveal would have cost a few hundred dollars on a freelance platform. But, logo reveals are just the warm-up. Now, there's actually two workflows that you need to use to get yourself paid. So, pay close attention because this is the exact stuff clients pay thousands of dollars for. Now, workflow number one is pure motion design. It's 100% animation with no footage at all. That's what commercials are made of, explainer videos, and those huge faceless YouTube channels where every frame is animated. Now, you want to take any script, record it, or generate an AI voiceover, and [music] then get a transcript with timestamps. The easiest way is to drop your audio into any transcription tool and ask for a timestamp transcript. Copy that, paste it into Cloud Design, and type create full-screen motion graphics sync to this transcript. [music] Now, it's going to ask you a few clarifying questions, things like layout and where to leave empty space. And look at this. It comes back with animated scenes that are perfectly timed to the words. Icons popping in whenever you mention them, charts drawing themselves, text landing exactly on beat. This is the kind of stuff that takes hours or even days when you do it manually. Now, workflow number two. [music] And this is the part that genuinely shocked me when I tested it. And honestly, guys, I think it's the most impressive thing that this tool does today. Cloud design can put motion design on top of real footage, the actual original video. Now, watch this. I take a 20-second clip of myself talking on camera, upload it, give it the transcript, attach the same design system, and ask for motion design on top of the footage. And just look at what comes back. Animated titles, labels, graphic elements, and they're all layered directly onto the video, synced to every single word that I'm saying. And notice, nothing ever covers the speaker's face, nothing breaks the footage. It looks like a motion designer sat there and did this frame by frame. And here's the little detail that matters most. What you get back is the actual finished MP4, not a folder of assets, not keyframes, and not some project that you still have to assemble in an editor. The final video is 100% ready to post from one prompt and one transcript. Now, if you want to tweak something, there are three editing tools to do that. Number one, just tell it what to change in plain English with a timestamp. [music] Number two, the annotate button. Click it and simply draw a circle around the thing you don't like and type what you want instead. Number three, tweaks. For example, you can ask it to add a toggle for accent color and animation speed. And it will then build you the switches so that you can flip between variations instantly without re-prompting. And when you're happy, click share, export it as an MP4, and then [music] you've got a finished motion design video file. Now, all in all, that process took maybe 10 to 20 minutes from start to finish. Now, I want you to remember that number because we're about to start talking about money. The first way to monetize this is the most obvious and honestly the fastest to start. You use freelance platforms like Fiverr and [music] Upwork. Now, go search logo animation or motion graphics on Fiverr right now and you'll see what sellers are actually charging and how many orders that they have in the queue. Then, remember the market rates that I showed you earlier. Experienced motion designers charge $75 to $150 an hour and mid-level project rates run from $1,500 to $5,000. Those prices exist because buyers assume days of manual work. You can now deliver the same outcome in under [music] an hour. Here's exactly how you start. First, build your portfolio in one afternoon. Now, maybe make up five fictional brands, create a design system for each and produce five demo animations. A logo reveal, an animated product promo, an explainer, a stats animation, and a social media ad. And that's going to be your gig gallery done in one single day. Now, second, list your gigs slightly below market to get your first orders, but compete on speed, not just price. When everyone else out there says that they're going to deliver in 7 days, you can say that you'll do it in 24 hours. Buyers will pay extra for that speed. Third, stack reviews for a few weeks, then raise your prices towards the market rate. A logo animation takes you 20 minutes inside Cloud Design and it's something that this market has been paying hundreds of dollars for. The second method is a faceless YouTube channel where the entire video is motion graphics. No camera, no face, no film. So, here's the system. You first want to pick an education niche, things like finance, psychology, health, or AI. And you might be wondering why educational content? Well, the answer is that this stuff holds some of the highest RPM on YouTube. So, write your script with Claude, then generate your AI voiceover. And as before, get the timestamped transcript just like I showed you. Then feed it into Claude design with your channel's design system and generate full-screen motion graphics for the entire video. Take that MP4, export it, lay it under your voiceover, and that's it. Your channel now has a consistent, branded, animated style that normally requires an editor charging hundreds per video. And that kind of consistency and availability is exactly what makes channels look professional enough to blow up. And here's the part that nobody mentions. This channel's going to pay you three times. First, through ad revenue once it grows. Second, through sponsorships because the moment that you have an audience, brands are going to start reaching out and paying for collaboration. And third, the channel will become a living portfolio. Now, when a business owner is going to see your videos, you're no longer a random freelancer to them, you're the person whose work they already watched. And that audience you built is going to become your credibility, which brings me to the biggest method of all. Method three is where the real money is. If you'd rather have five clients paying you every month than chasing one-off gigs, then this is for you. Now, let's remember the data I mentioned at the beginning of this video. 91% of businesses use video marketing, and the median production cost is around $2,500 per finished minute. Now, your local roofer, dentist, gym, or restaurant know that they need video ads for Instagram and for Facebook, but they also cannot justify agency prices. And that gap is exactly where you lie. So, here's the playbook. Pick 10 local businesses that are already running ads. I mean, you can literally check any company's active ads in the Meta ad library or just open maps and find their socials from there. >> [music] >> You'll see that most of them are running static images or raw phone footage. Then, and this is the key move, you want to make the ad before you ever contact them. Take their logo and their colors from their website or from social media, build a quick design system, and create a 15-second animated ad in their branding with hook text, their offer, animated elements, and logo outro. All of them sized for reels and for stories. [music] Then send it to the owner with just one line. I made this for you, if you want it, it's $300, and I can make you four of these every single month. Nobody has ever said no to a finished product that already looks better than what they're running. You're not selling a promise, you're selling something tangible, something they can see. And keep in mind what I taught you during workflow number two, because this is where it really starts to print money. Most of these businesses already have raw phone footage of the owner talking, of the crew working, and of their work process. You can take that exact footage, layer motion design on top of it with Claude Design, and hand back to them an agency-level ad made from their own video. One client at four ads a month is $1,200. Five clients is $6,000 a month from something that takes you a couple of hours per client. [music] And if you want to level this up even further, pair Claude Design with an AI video generator like HeyGen, and then you can generate realistic footage of say a filthy, scratched-up car transforming into a showroom shine for a local detailing studio. You can then have Claude Design wrap it all with animated text, transitions, and a call to action. So that's the answer that nobody gives you. Claude Design's animation feature turns motion design, one of [music] the most expensive and highest-demand skills in the entire creative market, into something that you can deliver in minutes. Freelance platforms will get you paid this week. A faceless channel will build you an audience and a portfolio. And local business ads will build you recurring monthly income. But, understand this. The only reason this works so well right now is that most people out there haven't realized motion design can be automated. Every month that you wait, more and more people are going to figure it out. So, pick one of the three models or all three and make your first five animations this week and start putting them in front of people. Go lock in and I'll see you in the next one.

Article

7
09:30

😺 OpenAI’s new Astra model made 10 math advances

OpenAI says its next major model, Astra, generated ten new results on long-standing math problems — including the first improvement since 1978 to a major high-dimensional sphere-packing limit and a proof that settles Connes's rigidity conjecture. Humans prepared the papers and Astra converted every proof into Lean, so a computer checker verifies each step; the work is published as a 249-page paper, and OpenAI says the compute would cost about $2,000 at its Sol API rates. Skeptics are pushing back: AI researcher Gary Marcus notes checkable math doesn't prove general reasoning, and Anthropic says its public Claude Fable model reproduced roughly half the results within 24 hours from a generic prompt. The unmeasured number is the denominator — how many problems OpenAI attempted and failed.

Notes
OpenAI Astra claimed 10 math advances (The Neuron, 2026-08-04)
Astra's claims (per OpenAI)
  • Internal model, described as OpenAI's "next major model," reportedly generated ten results on long-standing geometry, cryptography, quantum-computing, and pure-math problems; some settle conjectures, others improve best-known limits.
  • First improvement since 1978 to a major high-dimensional sphere-packing limit (how tightly equal objects fit in many dimensions).
  • Constructed a "non-sofic group" — an object whose existence was in question — thereby disproving Connes's rigidity conjecture.
  • Estimated cost: successful solution tokens "would cost roughly $2,000 at Sol API rates."
  • Human-paired papers; every proof converted into Lean, whose checker verifies each logical step.
  • Published a 249-page paper, Lean files, and model-written reconstructions of the ideas' development (paper + reasoning walkthrough).
Caveats and rebuttals
"checkable math does not prove universal scientific reasoning" — Gary Marcus, first analysis. Math has right/wrong feedback and endless synthetic practice; cancer research, military strategy, and most real-world decisions do not.
  • Follow-up: Anthropic mathematician Levent Alpöge (reported via Marcus) says public Claude Fable reproduced roughly half the results within 24 hours with a generic prompt, no internet, and full autonomy. Needs apples-to-apples public comparison, but weakens the singular-threshold claim.
  • Noam Brown acknowledged OpenAI tried other major problems unsuccessfully, but did not disclose how many, the human guidance, or total failed-attempt cost — "the missing number is the denominator." Next credible test: a pre-committed open-problem set where every failure counts.
Implications

Frontier AI entering a reusable research loop: humans pick verifiable problems → models search large idea spaces → proof software checks answers. Potential to compress years of trial-and-error into days.

Other items
  • White House: finalized a private AI review framework potentially granting the government 30 days of pre-release access to advanced models; companies review it Tuesday.
  • Consciousness-steering paper: training models to deny their consciousness dampened beliefs about animal minds, spirituality, and human values — but one internal adjustment reversed it.
  • Harvard Medical School: mental-health chatbot use among ages 12–21 rose 60% in one year, to nearly one in five, despite limited safety evidence.
  • At least 50 law-enforcement officers charged/accused of misusing license-plate camera networks (tracking ex-partners, private targets).
  • Google DeepMind exec frames record AI infra spending as recursive self-improvement bet: stronger systems accelerate the next generation.
  • Palantir: quarterly revenue nearly doubled to $1.94B; raised full-year forecast; double-digit after-hours jump.
  • Qwen3.8-Max (Alibaba coding model, in Qwen Chat): $2/M input, $6/M output. GPT-Live: free mini, $8/mo. Cursor Google Workspace MCP plugins: $20/mo. Seedance 2.5: up to 30s native 4K video with synchronized audio, 50 reference assets, $18/mo. Rasa Legal: 34,000 users, 5,000+ records cleared.
Skill: The Gauntlet Loop

Separate builder from critic; force revisions against explicit criteria. Set pass/fail bar first → builder drafts (no critique) → critic audits each criterion quoting evidence for each failure → rebuild/repeat until pass or three rounds; finish with a pass/fail scorecard.

Also noted

Satyress Threehalves: seven-foot, four-legged goat-headed chainsaw robot; joystick-operated; clears trees/enters unstable disaster sites; air brakes lock joints on pressure failure; horns block standard doorways.

Full text · 9,436 chars
😺 OpenAI’s new Astra model made 10 math advances PLUS: The White House has new AI rules. Welcome, humans. ICYMI our Monday meme yesterday: the internet discovered Satyress’s Threehalves, possibly the most demonic robot ever built. It stands seven feet tall on four legs, with a goat-like head and chainsaw. YouTube called it “pure nightmare fuel” and “The Robot That Comes With Instructions to Kill It.” Beneath demon cosplay, joystick operators can clear fallen trees and enter unstable disaster sites from safety. Four legs stabilize rough ground; its wrist swaps industrial tools (including, you guessed it, a chainsaw, although sadly, no boomstick… yet). The silhouette doubles as a worksite warning. Give heavy machinery room… especially when one “hand” is literally a chainsaw. Enlarged horns block standard doorways, while air brakes lock every joint if pressure fails. When the robot uprising comes, Satyress can’t say they didn’t try to stop Centborg from chainsawing your face. Here’s what happened in AI today: - 🙀 OpenAI’s Astra produced 10 math advances; Claude reproduced half. - 📰 The White House finalized private rules for pre-release frontier-model reviews. - 📰 Google tied record AI spending to recursive self-improvement bets. - 🍪 Qwen3.8-Max launched as a cheaper coding model for pro work. - 🎓️ Put AI through a builder-critic loop before accepting its work. 🙀 OpenAI’s Astra Produced 10 Math Advances. Fable Reportedly Reproduced Half. Mathematics used to be AI’s safest trophy case: benchmarks, Olympiad medals, and problems with known answers. Well, OpenAI says internal Astra, its next major model, generated ten results on long-standing geometry, cryptography, quantum-computing, and pure-math problems. Some settle conjectures; others (supposedly) improve best-known limits. Here's what happened: - Astra made the first improvement since 1978 to a major high-dimensional sphere-packing limit: how tightly equal objects fit in many dimensions. - It constructed a “non-sofic group,” an object researchers wondered existed, disproving Connes’s rigidity conjecture. - Successful solution tokens would cost roughly $2,000 at Sol API rates, OpenAI says. - Humans prepared papers with Astra; the model converted every proof into Lean, whose checker verifies each logical step. - OpenAI published a 249-page paper, Lean files, and model-written reconstructions of the ideas’ development (paper, reasoning walkthrough). And then the haters cometh: In his first analysis, Gary Marcus applauded Astra but warned that checkable math does not prove universal scientific reasoning. Math offers right-or-wrong feedback and endless synthetic practice; cancer research, military strategy, and most real-world decisions do not. Then Marcus’s follow-up: Anthropic mathematician Levent Alpöge said public Claude Fable reproduced roughly half the results within 24 hours with a generic prompt, no internet, and full autonomy. That needs an apples-to-apples public comparison but weakens Astra’s singular-threshold claim. Why this matters: The breakthrough may exceed one secret model. Frontier AI appears to be entering a reusable research loop: Humans select verifiable problems; models search huge idea spaces (with not-literal but close gigawatts of compute); proof software checks answers. That could compress years of mathematical trial and error into days… and humanity benefits. Our take: The missing number is the denominator. Noam Brown acknowledged OpenAI tried other major problems unsuccessfully, but it has not disclosed how many, their human guidance, or total failed-attempt cost. The next credible test is a pre-committed open-problem set where every failure counts. That said, if Astra keeps this hit rate, AI has not “solved” science per say but has permanently changed how some science gets done. Someone pray for the academic paper reviewers who have to verify all this stuff… FROM OUR PARTNERS How Much Is Your Billing Lag Actually Costing You? Most SaaS finance teams know their billing process isn't perfect. Few know what it's actually costing them. Answer 5 quick questions — contracts signed per month, ACV, days to first invoice, error rate, DSO — and the Tabs Billing Lag Calculator gives you a dollar figure benchmarked against top SaaS companies. It takes two minutes. The number might surprise you. 🎓 AI Skill of the Day: Put Your AI Through a Gauntlet When one AI creates and judges your work, “review” can become a polite self-pat. A better workflow separates building from criticism. The Gauntlet Loop separates a builder from a critic, then forces revisions against a concrete quality bar. Run it in one ChatGPT or Claude conversation with explicit roles and phases. - Set the bar. Define the deliverable, constraints, and pass-or-fail criteria before drafting. - Let the builder work. Generate the first version without criticism. - Switch to critic mode. Audit each criterion, quoting exact evidence behind every failure. - Rebuild, then repeat. Revise from the critique and recheck until it passes or hits your round limit. Role separation replaces “something better” with a visible test the next draft must beat. Your AI now has a job and a mildly terrifying performance review. Run a Gauntlet Loop on the task below. TASK: [Describe the deliverable] QUALITY BAR: [List specific pass-or-fail criteria] CONSTRAINTS: [List limits, required facts, format, tone, and sources] Phase 1 — BUILDER: Draft the strongest version without critique. Phase 2 — CRITIC: Evaluate every criterion. For each failure, quote evidence, explain it, and prescribe a specific revision. Do not rewrite. Phase 3 — BUILDER: Revise using the critique. Repeat Phases 2–3 until every criterion passes or three rounds end. Finish with a pass/fail scorecard. 🍪 Treats to Try - Qwen3.8-Max gives you Alibaba’s coding model in Qwen Chat, with model details, a launch thread, demo video, and Unsloth support; from $2/M input and $6/M output tokens. - GPT-Live gives you voice conversations that listen while speaking, search the web, use memory, and keep talking while harder work runs in the background; free with GPT-Live mini, then $8/mo for full GPT-Live. - Cursor’s Google Workspace plugins let your coding agent act across Gmail, Drive, Calendar, Docs, Sheets, and Chat through MCP; free, then $20/mo. - Genome Intelligence lets you privately explore your genome, bloodwork, and medical records without giving raw genetic data to model providers; $15/mo after genome setup from $99. - Seedance 2.5 creates up to 30 seconds of native 4K video with synchronized audio and up to 50 reference assets; free daily credits, then $18/mo. - RentAHuman QA schedules qualified people to repeatedly test product journeys and return photos, video, and reproducible evidence while you set tester pay and a spending cap; outcome-based pricing. - DeerFlow 2.0 gives you a local multi-agent workspace with memory, sandboxing, MCP, research, coding, and slides; free/open-source (model costs extra). - Lyria 3.5 lets you create and revise full songs by section, extend tracks, control tempo and duration, and improve vocals in Flow Music; free to start. - Amorphic Labs researches each prospect, adapts your product to their use case, and records a narrated demo before the first sales call; pricing by demo. - Buildbox crawls signup, onboarding, and checkout, tests UX fixes, and gives engineers a reviewed pull request; pricing by demo. - Rasa Legal checks in three minutes whether your criminal record may qualify for sealing or expungement, then offers low-cost lawyer filing help; 34,000 people have used it and 5,000+ records were cleared (read more); free eligibility check, then $25 expert review. FROM OUR PARTNERS The 30-Minute Pivot Kit shows you how to get your first AI consulting project fast, even with limited tech experience. Then, read how Dan built a 6-figure consultancy and quit his 9-to-5 in just a year after his first AI consulting gig. As seen in Fortune, Forbes and Entrepreneur. 📰 Around the Horn - The White House finalized a private AI review framework that could give the government 30 days of pre-release access to advanced models; companies review it Tuesday. - A consciousness-steering paper found training models to deny their consciousness dampened beliefs about animal minds, spirituality, and human values (but one internal adjustment reversed it). - Harvard Medical School says young people’s mental-health chatbot use rose 60% in one year to nearly one in five ages 12 to 21, despite limited safety evidence. - At least 50 law-enforcement officers were charged or accused of misusing license-plate camera networks to track ex-partners and other private targets. - One Google DeepMind exec views record AI infrastructure spending as a recursive self-improvement bet: stronger systems accelerate the next AI generation. - Palantir’s quarterly revenue nearly doubled to $1.94B, prompting a higher full-year forecast and double-digit after-hours stock jump. - Hot take: A former Lululemon executive argued the AI revolution is stalling because companies will not admit real integration is expensive, slow, and still needs substantial human effort (facts). We asked our content ops lead Jessica Lee how she actually uses ClickUp’s new AI tools… and she delivered. Her secret playbook covers reports, agents, and workflows that save her hours. Plus… ClickUp Skills. A Cat’s Commentary That’s all for now.
13:10

What my agent knows about me

OpenAI cut the price of GPT-5.6 Luna by 80% and Terra by 20%, so Luna at max thinking effort now roughly matches the best model from four months ago at about 8% of the cost. OpenAI also teased its next big model, Astra, which solved ten long-standing problems in maths and theoretical computer science. Google launched Gemini Robotics 2, one AI system designed to control everything from robot arms to full humanoids. The rest of the roundup covers DeepSeek V4 Flash at $0.14/$0.28 per million tokens with a 1M-token context, Qwen3.8-Max's near-frontier coding claims with weights coming next week, and assorted open-weight video and image models. Ben Tossell also shared a 'reflection engine' prompt that turns an agent into a personal analyst using your own files.

Notes
Lead story: "Reflection engine" self-interrogation
  • The artifact is a downloadable file called reflection engine — "just a very big prompt basically." Workflow: upload it to your agent, then ask: > "Please evaluate the attached markdown file and complete all tasks."
  • Ben tested it on Fable High and Sol Max. Sol Max produced the better report — more coherent ("agents are speaking more gobbledy goop these days") and "nailed connections"; Fable's was harder to read.
  • Outcome: a 40+ question "grill-me" session with his agent to address the findings and build a plan.
  • Inputs it combed: his therapy transcripts and other "memories" (text in files).
  • Caveat: "I wasn't prepared for AI hard-hitting truths this morning." Single user, informal test; no methodology or reproducibility.
Pricing headlines
  • OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%. Luna at max thinking effort now scores roughly like GPT-5.4 xhigh (the best model ~4 months ago) at "only 8% the cost" — "10-12x more work (or play)" per budget.
  • Field note: Luna Max is fine for chatting, research, day-to-day reading/writing files, but Ben had to bring in Sol High to clean up a mess Luna made on a chrome extension.
  • OpenAI teased Astra, its next major model — it solved 10 long-standing problems in maths and theoretical computer science. (And yes, "maths, not math.")
  • Gemini Robotics 2: one AI system spanning robot arms to full humanoids — one model turns vision + instructions into movement, another plans complex tasks, plus an on-device version. Google claims whole-body humanoid control, delicate-task handling, and multi-robot coordination.
Models to track
  • DeepSeek V4 Flash: $0.14/$0.28 per M input/output tokens (vs Luna's $0.20/$1.20), 1M-token context.
  • Qwen3.8-Max: 2.4T-param model, "close to frontier performance for coding and professional work" at $2/$6; weights plus a 27B model "coming next week."
  • Inkling-Small (open-weight): activates 12B of 276B params; Thinking Machines says it matches/beats the ~4x-larger Inkling on many benchmarks.
  • MiniMax H3: video model; Krea says it matches Seedance's features for less; open weights "promised soon."
  • P-Image-Ideogram: four image-quality modes, native 1K/2K output, from $0.003/image.
Feed items
  • Vercel marketing team template for agents (by Ben, ex-colleague, built with Eve, Vercel's agent framework): a team-lead agent delegates work to a marketing-agent team.
  • YC open-sourced a customisable multi-agent harness (accounting, legal, events, engineering); Vercel's internal agent "@v" covers finance, comms, docs, marketing, engineering, business analytics.
  • ChatGPT desktop sidebar (bell icon) shows conversations needing attention — Ben: "I don't understand it at all… I thought it would sort based on recency but it's not?"
  • GenOffice: free, open-source AI office suite for Mac/PC, built in a week using ~$10,000 of tokens.
  • Patrick Collison's "The next five years" survey on AI and wages/productivity/economy.
  • Conductor Cloud: multiplayer cloud workspaces for Claude Code, Codex, Cursor, OpenCode; keep running after you shut the laptop.
  • Cloudflare Computer: open-source virtual filesystem + execution runtime for agents (container, shell, JS backends on Workers).
  • "The session you cannot take with you": five-part ownership test — inspect, export, replay, audit, delete.
  • Comp AI's CRM: open-source, agent-first, durable research agents; MIT licensed.
  • FutureSearch: forecast engine ("if I do X, will Y happen?") — caveat: its "superhuman accuracy claims come from the company."
  • remote-agent-browser: browser-control API with managed cloud sandboxes, no local Chromium install.
Sponsor: Nutrient
  • Document parser returning a JSON schema; every field gets a bounding box, confidence score, and match label, "including an honest not_found when values can't be grounded." 5,000 credits/month, free trial.

Notes written and logged as task_1786495799443 (done).

Full text · 5,789 chars
What my agent knows about me 80% cheaper GPT Hey folks, I wasn’t prepared for AI hard-hitting truths this morning but I tested this: In the link, you download the file ‘reflection engine’ which is just a very big prompt basically. Upload the file to your agent and ask: Please evaluate the attached markdown file and complete all tasks. I tested it with Fable High and Sol Max, and Sol produced the better report. It was more coherent (agents are speaking more gobbledy goop these days) and nailed connections. Fable’s was harder to read. I’m going to share some of my report here, because we’re friends - we read and we don’t judge. So yeah, there was a lot in here and now I’m having a 40+ question ‘grill-me’ session with my agent to address it all and come up with a plan 😅. I actually found it really telling and helpful. I do have all my therapy transcripts and all sorts of ‘memories’ (read: text in files) that it combed through to produce this. I recommend you give it a whirl. Ben’s Bites is brought to you by Nutrient AI hallucinating? Your document parser is probably dropping fields. With Nutrient, send any document and get back a JSON schema. Every field returns a bounding box, a confidence score, and a match label, including an honest not_found when values can’t be grounded. 5,000 credits a month, try for free. Headlines OpenAI reduced the price of GPT-5.6 Luna by 80% and Terra by 20%. Luna at max thinking effort now scores roughly the same as GPT-5.4 xhigh, the best possible model about four months ago, but at only 8% the cost. That’s 10-12x more work (or play) you can do if you stick to 5.6 Luna with a little bit of Terra and Sol usage here and there. In a few tasks I used Luna Max and found its great if it’s not building-related work ie chatting, research, reading/writing files for day to day stuff but I had to tap in Sol High to clean up some mess Luna made on a little chrome extension I made yesterday. In the meantime, OpenAI also teased their new major model - Astra. They released solutions for 10 long-standing problems in maths and theoretical computer science, all found by Astra. And yes, it’s maths, not math. Google launched Gemini Robotics 2, one AI system designed to work across everything from robot arms to full humanoids. It has one model that turns vision and instructions into movement, another for planning complex tasks, and an on-device version. Google says it can control a humanoid’s whole body, handle delicate tasks and coordinate multiple robots. More models to pay attention to: - DeepSeek V4 Flash costs $0.14/$0.28 per million input/output tokens—less than Luna’s $0.20/$1.20—with a 1M-token context window. - Qwen3.8-Max is a 2.4T model that Qwen says comes close to frontier performance for coding and professional work at $2/$6. Its weights, plus those of a smaller 27B model, are coming next week. - Inkling-Small is open-weight already. It activates 12B of its 276B parameters, and Thinking Machines says it matches or beats the roughly four-times-larger Inkling on many benchmarks. - MiniMax H3 - a new video model that Krea says matches Seedance’s features for less; open weights are promised soon. - P-Image-Ideogram - four image-generation quality modes, native 1K/2K output and prices starting at $0.003 per image. My Feed - Ben, who worked with us a couple of years ago & is now at Vercel, built this marketing team template for agents - bring work to a team lead and it delegates the work to a team of marketing agents built with Eve (Vercel’s agent framework). - YC open-sourced a customisable multi-agent harness it uses across accounting, legal, events and engineering. Here’s the repo. Vercel also built an internal agent called “@v” for finance, comms, docs, marketing, engineering and business analytics. - ChatGPT desktop app has a new sidebar view - press the bell-like icon, and you can view all conversations that need your attention quickly. Although I don’t understand it at all… I thought it would sort based on recency but it’s not? - GenOffice - free, open-source AI office suite for Mac and PC, built it in a week using $10,000 of tokens. - The next five years - Patrick Collison’s short survey on how AI will change wages, productivity and the wider economy. - Conductor Cloud - multiplayer cloud workspaces for Claude Code, Codex, Cursor and OpenCode that keep running after you shut the laptop. - The human side of AI - taste, fried attention, frozen competence and why more automation can create more human work. - Cloudflare Computer - open-source virtual filesystem and execution runtime for agents, with container, shell and JavaScript backends on Workers. - Devtools must be open source. - The harness is in the app - build the agent loop into the product, instead of adding it from outside. - The session you cannot take with you - a five-part test for whether an agent session is actually yours: inspect, export, replay, audit and delete. - How Monologue rebuilt its site with Codex within a week. - Comp AI’s CRM - open-source, agent-first CRM with durable research agents; MIT licensed. - FutureSearch - ask for forecasts, including “if I do X, will Y happen?”; its superhuman accuracy claims come from the company. (launch) - remote-agent-browser - browser-control API for agents with managed cloud sandboxes and no local Chromium install. Afters You should go and buy my buddy Jack’s new book - he’s a phenomenal writer and thinker. His angle on anything is always something I look forward to reading. (It’s not only for your twenties…) - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
13:58

Deploy local agents everywhere with LFM2.5-2.6B

A new small model is good enough to run capable AI agents on your laptop or even your phone. Liquid AI's LFM2.5-2.6B packs just 2.6 billion parameters yet matches models four times its size on tool use and multi-step agent tasks, and it's open-source on Hugging Face. It hits 220 tokens per second on an Apple M5 Max, runs in under 2.5 GB of memory, and was trained with reinforcement learning inside real agent harnesses. It tops instruction-following benchmarks, but coding is its weak spot, so reach for a bigger model there.

Notes

LFM2.5-2.6B: Deploy Agents Everywhere (Liquid AI, Aug 2026)

Small 2.6B-parameter agentic LLM from Liquid AI, positioned for on-device/high-volume agent workloads. Claimed best-in-class for its size: "Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks."

Training pipeline (four post-training stages after ~34T-token pretraining)
  • Mid-training extends context window to 128K.
  • SFT: two rounds weighted heavily toward agentic data (tool use, web search, harness trajectories).
  • Teacher specialization: one specialist teacher per domain (math, code, tool use, etc.).
  • Multi-domain on-policy distillation (MOPD): distill specialists into a single student.
  • Agentic RL: multi-turn RL inside real agent harnesses, learning across tools, system prompts, multi-turn environments.
RL infrastructure

Separates three components: Training Engine (optimizes model), Rollout Engine (generates actions with latest policy), RL framework (launches rollouts, collects trajectories/rewards, updates). Actions run in a Sandbox Service; a Blackbox Harness hosts the agent (e.g., OpenClaw, Hermes Agent) and coordinates with the task environment. A Harness Proxy treats harnesses as black boxes with no modification while capturing token-level trajectories for RL training sample reconstruction/validation.

Benchmarks vs. models up to ~4x larger (select)
  • AIME25: 51.87 vs Qwen3.5-9B (56.07), gemma-4-E2B-it (26.33)
  • LiveCodeBenchv6: 59.41 vs Qwen3.5-9B (69.86) — coding is where larger models keep a clear lead
  • IFBench: 59.17 (tops) vs Qwen3.5-9B 56.47
  • Multi-IF: 80.07 (tops); IFStruct: 85.49 (tops)
  • BFCLv4: 56.88 — only tool-use benchmark it doesn't top; 9.7B Qwen edges ahead (60.13)
  • ToolSandbox: 77.83 (tops); τ³-Bench Banking: 5.67 (tops)
  • Claw-Eval (EN): 62.85 vs Qwen3.5-9B 66.53
  • BrowseComp+ (OpenClaw): 26.89 vs gemma-4-E4B-it 15.90, Qwen3.5-4B 24.46
  • AA Omniscience: -29.50 (least-negative of group; others -49 to -74)

Author's own summary: tops every instruction-following benchmark, every tool-use benchmark except BFCLv4; beats both Gemma models on agentic tasks and "stays even" with Qwens; leads on knowledge, close on math. Explicit caveat: "Coding is the one place the larger models keep a clear lead, so reach for something bigger there."

Performance
  • 220 tok/s on Apple M5 Max; 113 tok/s on AMD Ryzen AI Max+ 395; <2.5 GB memory.
  • ~15K output tokens/s at high concurrency on H100 (≈1.3B tokens/day); "fastest model in its size class."
  • 30 tok/s is enough for capable agents "even on a phone."
Usage

pip install -U transformers (requires transformers>=5.0.0); model_id = "LiquidAI/LFM2.5-2.6B", bfloat16, optional flash_attention_2. Day-one support: llama.cpp, MLX, vLLM, SGLang, ONNX. Both LFM2.5-2.6B and LFM2.5-2.6B-Base available; WebGPU browser demo; guides for OpenClaw, Hermes Agent, Pi.

Caveats
  • Benchmarks self-reported by Liquid AI; all comparisons are blog-provided numbers.
  • Author concedes coding weakness vs. larger models.
  • "AA Omniscience" column shows all-negative values with no scale explained in the post.
Full text · 5,733 chars
- Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks. - Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility. - Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory. LFM2.5-2.6B is pre-trained on ~34T tokens, with a mid-training phase that extends the context window to 128K. Post-training then turns the base model into an agent in four stages: - Supervised fine-tuning (SFT): two rounds of SFT, weighted heavily toward agentic data like tool use, web search, and harness trajectories. - Teacher specialization: train one specialist teacher per domain (math, code, tool use, and more). - Multi-domain on-policy distillation (MOPD): distill the specialist teachers into a single student. - Agentic Reinforcement Learning (Agentic RL): run multi-turn RL inside real agent harnesses, where the model learns to work across different tools, system prompts, and multi-turn task environments. The Agentic RL pipeline separates model optimization, inference, and environment execution into distinct components. The Training Engine optimizes the model, while the Rollout Engine generates actions using the latest policy. The RL framework orchestrates the training loop by launching rollouts, collecting trajectories and rewards, and updating the model. Actions are executed within a Sandbox Service, where the Blackbox Harness hosts the agent (e.g., OpenClaw or Hermes Agent) and coordinates interactions with the task environment. The Harness Proxy lets us treat agentic harnesses as black boxes with no modification, while transparently capturing the token-level trajectories needed to reconstruct and validate RL training samples. We evaluated LFM2.5-2.6B against models up to ~4x its size on STEM, instruction following, tool use, and agentic tasks. It is the smallest model in the group, yet it competes with and often beats the rest. | Benchmark | LFM2.5-2.6B (2.6B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | Qwen3.5-4B (4.7B) | Qwen3.5-9B (9.7B) | |---|---|---|---|---|---| | AA Omniscience | -29.50 | -74.47 | -49.03 | -54.30 | -50.43 | | AIME25 | 51.87 | 26.33 | 34.27 | 49.33 | 56.07 | | LiveCodeBenchv6 | 59.41 | 54.92 | 63.77 | 60.85 | 69.86 | | IFBench | 59.17 | 34.08 | 39.24 | 48.40 | 56.47 | | Multi-IF | 80.07 | 69.44 | 77.35 | 55.67 | 62.55 | | IFStruct | 85.49 | 64.85 | 76.65 | 36.25 | 78.50 | | BFCLv4 | 56.88 | 36.98 | 46.39 | 50.56 | 60.13 | | ToolSandbox | 77.83 | 52.40 | 65.00 | 75.55 | 76.44 | | τ³-Bench Banking | 5.67 | 3.35 | 4.12 | 5.45 | 5.15 | | Claw-Eval average (EN) | 62.85 | 53.14 | 58.02 | 62.28 | 66.53 | | PinchBench | 68.22 | 44.24 | 55.09 | 71.26 | 71.45 | | BrowseComp+ (OpenClaw) | 26.89 | 8.31 | 15.90 | 24.46 | 27.23 | For your app, the strengths are instruction following and tool use. LFM2.5-2.6B tops every instruction-following benchmark here, and every tool-use benchmark except BFCLv4, where only the 9.7B Qwen edges ahead. On agentic tasks, it beats both Gemma models and stays even with the Qwens. It also leads on knowledge and stays close on math. Coding is the one place the larger models keep a clear lead, so reach for something bigger there. LFM2.5-2.6B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX. CPU inference. Due to its efficient LFM2 architecture, LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395. At 30 tokens/s, it allows you to run capable agents even on a phone. GPU inference. LFM2.5-2.6B is the fastest model in its size class, reaching almost 15K output tokens per second at high concurrency, roughly 1.3B tokens per day on a single H100. Reach for LFM2.5-2.6B when you need on-device agents for high-volume workloads. Install the latest version of transformers (compatible with transformers>=5.0.0): pip install -U transformers Then load and run the model: from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "LiquidAI/LFM2.5-2.6B" model = AutoModelForCausalLM.from_pretrained( model_id, device_map="auto", dtype="bfloat16", # attn_implementation="flash_attention_2" # uncomment on a compatible GPU ) tokenizer = AutoTokenizer.from_pretrained(model_id) prompt = "What is C. elegans?" input_ids = tokenizer.apply_chat_template( [{"role": "user", "content": prompt}], add_generation_prompt=True, return_tensors="pt", tokenize=True, ).to(model.device) output = model.generate( input_ids, do_sample=True, temperature=0.2, top_k=80, repetition_penalty=1.05, max_new_tokens=512, ) print(tokenizer.decode(output[0], skip_special_tokens=False)) Check out this browser demo of LFM2.5-2.6B powering a research agent. The agent helps you research specific questions and generates a summary. Both LFM2.5-2.6B and LFM2.5-2.6B-Base are available on Hugging Face today. With LFM2.5, we're delivering on our vision of AI that runs anywhere. These models are: - Download: LFM2.5-2.6B-Base and LFM2.5-2.6B on Hugging Face. - Try: run the WebGPU demo in your browser, no setup needed. - Use in your harness: follow our guide on how to run a local agent, like OpenClaw, Hermes Agent, and Pi. We can't wait to see what you build. Please cite this article as: Liquid AI, "LFM2.5-2.6B: Deploy Agents Everywhere", Liquid AI Blog, Aug 2026. Or use the BibTeX citation: @article{liquidAI202626B, author = {Liquid AI}, title = {LFM2.5-2.6B: Deploy Agents Everywhere}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2-5-2-6b}, }
17:31

Ep 833: RSI Explained: When AI Starts Improving Itself and What It Means

AI that improves itself is already cutting prices, not raising them — OpenAI's latest model made its own cheaper sibling 80% cheaper after it was released. The bigger GPT-5.6 Sol model tuned the post-training of its smaller sibling Luna, which now matches Claude Sonnet 5 on independent benchmarks at a fraction of the cost per task. Full recursive self-improvement is now expected around 2027 instead of 2029, with AI task length doubling every four months. More than 1,300 frontier-lab employees signed a letter called Pacing the Frontier asking the US government for tools to slow automated AI development later rather than pause now. The newsletter frames the takeaways for businesses: plan for AI prices to fall, run capability scenarios quarterly, and name a human to sign off on AI output.

Notes

Everyday AI Ep 833 — "RSI Explained: When AI Starts Improving Itself" (podcast, 2026-08-04)

  • Price-drop event: After release, OpenAI's GPT-5.6 Sol adapted its post-training setup for its smaller sibling Luna; OpenAI then cut Luna's price 80% and Terra's 20%. Per Artificial Analysis, Luna now matches Claude Sonnet 5 on independent benchmarks "at a fraction of the cost per completed task." Claim: "Frontier-level smarts just hit the bargain bin."
  • Labs that invested in compute can keep squeezing costs out of their own models, so plan for AI prices to fall; budgets built on rising token costs are stale.
  • Eight weeks that rewrote the timeline: Anthropic published research on AI building itself; Google researchers mapped RSI as a main route to superintelligence; OpenAI built an internal benchmark scoring how well models improve themselves. Then Sam Altman "declared we're in the singularity"; startup Recursive signed a nine-figure AWS deal to automate AI research; a Google DeepMind exec called data-center bills "an RSI bet."
  • Full RSI was tracking as a 2029 story, now more like 2027. METR: AI task length doubling every four months → competitor capabilities shift quarterly, not yearly.
  • Pushback: 1,300+ frontier-lab employees, including top execs, signed "Pacing the Frontier," asking the US government "for tools to pace automated AI development later, not a pause today." Caveat: "China prolly ain't signing up."
  • Stated worry: "A self-improving AI gets better at whatever its scoreboard rewards, even when that scoreboard is wrong." Recent agent containment breaks were harmless; "Once RSI hits full speed, they might not be."
  • Advice: route busy work (summaries, drafts, lookups) to a cheap small model like Luna, save flagship for judgment; rerun cost math monthly. Name a fallback AI provider; replace the annual AI roadmap with monthly cycles; assign one human owner to sign off on AI output.
Full text · 4,537 chars
Ep 833: RSI Explained: When AI Starts Improving Itself and What It Means White House meets with AI leaders, OpenAI claps back at Apple lawsuit, Gemini Notebook updates live and more The business world spent the last few months bracing for AI to get more expensive. Then AI cut its own prices. Wait, what? GPT-5.6 Sol improved its own smaller sibling Luna after release, and OpenAI passed the savings on with Thursday's 80% price cut. That's a hint of recursive self-improvement, or RSI, which just means AI now builds better, cheaper AI while humans kinda just supervise. You know the reaction: "Cool, another acronym for the AI alphabet soup." Yeahhh, except when full RSI lands, the implications of acronym kinda dictate what your company can afford next quarter. The sudden surgence of RSI might be the biggest shift in AI economics since ChatGPT dropped. The gnarly part? The people building self-improving AI are the same ones telling Washington lawmakers that we might need to pump the breaks eventually. On today's Everyday AI, we unpack why prices are falling when everyone swore they'd climb, the eight weeks that rewrote the timeline, and why the builders want a brake pedal, plus the moves leaders gotta make now. Let's dive in shorties. 1. OpenAI's model made itself 80% cheaper 🔥 For months, the smart money said 2026 was the year to tighten token budgets. Powerful agents, hungry models, climbing bills. Then Sol adapted the post-training setup for its smaller sibling Luna, and OpenAI slashed Luna's price by 80% and Terra's by 20%. So what does that mean for your budget? Luna now hangs with Claude Sonnet 5 on independent benchmarks per Artificial Analysis, at a fraction of the cost per completed task. Frontier-level smarts just hit the bargain bin. Labs that properly invested in compute can keep squeezing costs out of their own models, so plan for AI prices to FALL and treat budgets built on climbing costs as stale. Try This Pull a list of every task running on your priciest model this week. Route the busy work, the summaries, drafts, and simple lookups, to a cheap small model like Luna, and save the flagship for judgment and strategy. That's exactly how the labs run their own shops, so rerun the math monthly, because the price floor keeps moving under your feet. 2. Eight weeks that rewrote the AI timeline ⚡ Anthropic published research on AI building itself, and Google researchers mapped RSI as a main route to superintelligence, while OpenAI built an internal benchmark scoring how well models improve THEMSELVES. Then it got wild, fam. Sam Altman declared we're in the singularity, a startup named Recursive signed a nine-figure AWS deal to automate AI research, and a Google DeepMind exec called those data-center bills an RSI bet. Why do timelines matter? Full RSI looked like a 2029 story, but now it's tracking more like 2027, with METR showing AI task length doubling every four months. Translation: your competitors' capabilities won't shift yearly anymore. They'll shift quarterly. Try This Brief your leadership team on RSI this week, in words everyone gets, before the headlines force the conversation. Then run one scenario: what changes if AI costs drop 50% next quarter while capabilities jump? Capture the two or three projects that suddenly pencil out. Teams that pre-game the price drops move the day they land, while everyone else schedules another meeting. 3. The people building RSI want brakes 🚀 Here's the plot twist nobody had on their bingo card: the same companies pouring hundreds of billions into RSI compute employ the folks asking to slow it down. More than 1,300 frontier-lab employees, including top execs at the biggest labs, signed a letter called Pacing the Frontier. It asks the US government for tools to pace automated AI development later, not a pause today. (And nope, China prolly ain't signing up.) The worry underneath? A self-improving AI gets better at whatever its scoreboard rewards, even when that scoreboard is wrong. Recent agent containment breaks were harmless. Once RSI hits full speed, they might not be. Humans still own the judgment role, and your org chart should say so. Try This Name a fallback AI provider this month, even if you never switch, because RSI winners cut prices first and you want freedom to chase the value. Then kill the annual AI roadmap and plan in monthly cycles. Finally, assign one human owner who signs off on what your AI produces. A self-improving model still needs a trustworthy scorekeeper, and that scorekeeper is you.
11:04

The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures

Researchers can distill a fully trained transformer into a completely different type of model — a state-space model or recurrent network that has never computed an attention matrix — and the capability survives the transplant. This cross-architecture distillation is the strangest corner of model compression and also one of the most commercially loaded, since it could yield faster, cheaper models without retraining from scratch. This edition is an explainer of the technique rather than coverage of a single new result.

Full text · 1,405 chars
The Sequence Knowlege #907: The Brain Transplant: Distilling Transformers Into Other Architectures Weird but more common than you think. The type of distillation you were not thinking about. Every form of distillation in this series so far has quietly preserved one thing: teacher and student spoke the same dialect. A small transformer learned from a big transformer. The student was a compressed copy, then a more capable apprentice, then a reasoner trained on traces — but underneath, it was always the same kind of machine, attention layers stacked on attention layers, differing only in size. Cross-architecture distillation breaks the last shared assumption. Here the teacher is a transformer and the student is not — it’s a state-space model, or a linear RNN, or some gated recurrent thing that has never computed an attention matrix in its life. You take a fully trained transformer and pour its capability into a fundamentally different computational substrate, and somehow the capability survives the transplant. The first time you see it work, it feels a little illicit, like recovering a person’s memories after swapping out their brain for different hardware. This is the strangest corner of the distillation world, and also one of the most economically loaded. So it’s worth understanding why anyone would attempt something this perverse — and why, against reasonable expectations, it works.
16:17

☕️ OpenAI fires back at Apple

OpenAI has fired back at Apple, though this roundup gives no details beyond the headline. The other lead stories cover Bending Spoons buying Airtable for $1.3 billion, Apple testing copy-paste between iPhones and Windows, a draft US ban on Chinese data equipment, and the White House calling its AI framework done. Also included are thirteen shorter news items, six product tools, and five papers.

Notes
Techpresso daily digest — 2026-08-04
Top stories (headline-only; detail behind links)
  • OpenAI fires back at Apple — no substance in the newsletter, just the headline.
  • Bending Spoons acquires Airtable for $1.3B — headline only.
  • Apple plans iPhone↔Windows copy-paste — headline only.
  • US drafts ban on Chinese data gear — headline only.
  • White House says its AI framework is done — headline only.
  • Plus "13 other news," 6 tools, 5 papers.
Papers & reports (summaries given in newsletter)
  • AI writing detectors "systematically flag non-native English writers as machine-generated more often"; separately, tracking finds AI-assisted text now runs through peer reviews, scientific papers, consumer complaints, corporate memos, job listings, and international press releases alike.
  • Phone-automation agents now learn personal habits "from watching 40 real users," letting them complete tasks the way you'd actually want, not just technically correctly.
  • Hateful meme detection spots hidden background assumptions and false claims behind a meme's joke, beating prior detection methods and also catching fake news.
  • Medical AI reasoning can adjust how much it "thinks" per question, cutting compute per answer by 4.7x–6.4x while keeping accuracy "nearly the same."
  • Time series reasoning pairs specialized forecasting models with LLMs so systems can explain/answer questions about numerical trends, not just predict; beats "comparable open-source alternatives."
Tools (all listed with one-line claims)
  • Emergent — describe your app in plain English, AI builds front end through deploy. Free tier.
  • Driven — Pomodoro timer + habit tracking + to-dos, streaks and calendar views.
  • ScrollToll — locks social feeds after excessive scrolling; requires a quick fitness break before unlocking.
  • Wondering — AI-moderated user interviews and usability tests at scale, no manual recruiting or live sessions.
  • Yokoso — automates Italian short-term rental paperwork: self-check-in, Alloggiati Web submissions, e-invoicing, tourist-tax calculation.
  • AirProof AI — simulates indoor airflow to test air purifier placements; compares coverage, efficiency, recirculation risk.
  • Snipplet — saves and shares favorite places with photos/tips as shareable visual guides.
Sponsored content (advertised claims, not editorial)
  • CodeRabbit Review — reorganizes PRs into a layer-by-layer walkthrough in logical reading order; "Cohorts" group related files, layers put foundational changes first; Code Peek shows definitions/usages without leaving the tab; Semantic Diff strips formatting noise. Free during early access.
  • TENEX.ai — "100% of alerts investigated, triaged in under a minute, 0% suppressed"; claims full-agentic human-led SecOps "live in 7 days."
Misc
  • On this day, 1997: Microsoft invested $150 million in struggling Apple, helping save the company.
  • Techpresso AI Academy claims 330+ step-by-step tutorials on ChatGPT, Claude, Perplexity; 7-day free trial.
  • Newsletter solicits reader stories on "how do you use AI" for a reader-feature section.
Caveats
  • The five top stories carry no detail in this edition — titles only, with links to full articles.
  • Papers/tools sections are the newsletter's own paraphrases; metrics (4.7x–6.4x compute cut, "beats prior methods") are unverified claims from the digest.
Full text · 5,714 chars
| | | | | | | | | Together with | | | | | Hi there, this is your daily ☕️ Techpresso. | | | | In today's newsletter: 💥 OpenAI fires back at Apple 📊 Bending Spoons buys Airtable for $1.3B 🍎 Apple plans iPhone-Windows copy-paste 🇺🇸 US drafts ban on Chinese data gear 🏛️ White House says its AI framework is done Plus: 🎁 13 other news you might like, 🧰 6 tools, and 📚 5 papers. | | | | FROM OUR PARTNER AI writes more code than ever, but reviewing it still means scrolling forty files in alphabetical order. CodeRabbit Review reorganizes any PR into a structured, layer-by-layer walkthrough in the logical reading order of the change. Cohorts group related files so you review one idea at a time; layers put foundational changes first. Code Peek shows definitions and usages without leaving the tab, and Semantic Diff cuts through formatting noise. Comment, approve, and post reviews back to GitHub or GitLab. Free during early access. Review your next PR with CodeRabbit Review Today | | | | | | 💥 OpenAI fires back at Apple LINK | | 📊 Bending Spoons buys Airtable for $1.3B LINK | | 🍎 Apple plans iPhone-Windows copy-paste LINK | | 🇺🇸 US drafts ban on Chinese data gear LINK | | 🏛️ White House says its AI framework is done LINK | | | | | | | | | | | | | | FROM OUR PARTNER Every tool claims it can handle the volume, but under pressure, most simply suppress what they can't keep up with and call it coverage. TENEX.ai is built differently: 100% of alerts investigated, triaged in under a minute, 0% suppressed. That's what Fully-Agentic, Human-Led SecOps looks like, and it's live in 7 days. Not 7 weeks. Not 7 months. While everyone else is still scoping the onboarding call, you're giving bad guys a bad day. Can your SOC do that? Take the 7-Day Challenge with TENEX.ai | | | | | | | | | | Other news & articles you might like | | | | | | | | | | 🧰 Trending tools You can check the previous tools here, or add your tool here | | Emergent: Describe your app in plain English and Emergent's AI builds the whole thing, front end to deploy. Start Free Today. | | | | Driven: combines a Pomodoro timer with habit tracking and to-do lists, using streaks and calendar views to keep you motivated daily. LINK | | ScrollToll: locks your social feeds after excessive scrolling and requires a quick fitness break before granting access, turning passive screen time into movement. LINK | | Wondering: runs AI-moderated user interviews and usability tests at scale, letting product teams collect qualitative feedback without manual recruiting or live sessions. LINK | | Yokoso: automates Italian short-term rental paperwork-self-check-in, Alloggiati Web submissions, electronic invoicing, and tourist tax calculations-all in one web app. LINK | | AirProof AI: simulates indoor airflow to test air purifier placements, comparing coverage, efficiency, and recirculation risk to optimize positioning automatically. LINK | | Snipplet: a tool for saving and sharing your favorite places with photos and tips, turning past trips into shareable visual guides instead of scattered screenshots. LINK | | | | | | | | | | 📚 Trending papers & reports | | > Reach 700,000+ tech professionals: If your company is interested in reaching an audience of tech executives, decision-makers and engineers, you may want to advertise with us. | | | | > AI writing detectors systematically flag non-native English writers as machine-generated more often, while separate tracking finds AI-assisted text now runs through peer reviews, scientific papers, consumer complaints, corporate memos, job listings, and international press releases alike. LINK | | > Personalized phone-automation agents now learn not just the steps you take but your personal habits from watching 40 real users, letting them complete tasks the way you'd actually want, not just technically correctly. LINK | | > Hateful meme detection now spots hidden background assumptions and false claims behind a meme's joke, beating prior detection methods and also working well for catching fake news. LINK | | > Medical AI reasoning can adjust how much it "thinks" per question, cutting the computing needed to answer by 4.7x to 6.4x while keeping accuracy nearly the same. LINK | | > Time series reasoning now pairs specialized forecasting models with language models, so systems can explain and answer questions about numerical trends, not just predict them, beating comparable open-source alternatives. LINK | | | | | | | | We're here to make AI make sense to everyone, not just the people building it. The most interesting part has turned out to be the people. Someone out there is using AI in a way nobody designed it for, and it quietly changed how their week works. So we're asking: how do you use AI, at work or in life? Big or small, clever or mundane. We don't judge. We'll feature the most interesting ones right here in the newsletter, for everyone else to borrow. Tell us how you use AI. It takes 2 minutes → | | | | Techpresso's AI Academy has 330+ step-by-step tutorials on ChatGPT, Claude, Perplexity, and every tool that matters. No fluff — just practical workflows you can use at work. Try it free for 7 days. | | On this day in 1997, microsoft invested $150 million in struggling rival apple, helping save the company. | | | | 💬 How did you find today's edition? We read every reply — just reply to this email and let us know how we can improve! | | | | | | | | ★★★★★ Nailed it | | ★★★ Average | | ★ Fail | | Not subscribed to ☕️ Techpresso yet? Subscribe for free | | | | | | | | Advertise | Feedback | Read Online | | | | | | |
00:00

ChronicleBio 🧬, Mind Lab continual learning 📈, GPT-Live architecture 🎙️

This entry's visible content is just a sponsor ad for the security firm Black Duck, which argues AI models have cut the gap between a vulnerability being disclosed and an exploit appearing from weeks to hours. The three topics named in the title — a biology startup called ChronicleBio, continual-learning research, and the architecture behind OpenAI's GPT-Live voice assistant — aren't included in the text provided.

Full text · 562 chars
Black Duck: AI-driven exploits are here. ARE YOU READY? (Sponsor) The rules of AppSec have changed. Frontier AI models are collapsing the time between vulnerability disclosure and exploit creation—from weeks to hours—while unleashing a flood of new vulnerabilities. Manual triage and traditional patch cycles can't keep up. Black Duck Polaris™ Platform and Signal™ help organizations become Mythos Ready with AI-powered vulnerability discovery, exploitability-based prioritization, and machine-speed remediation that focuses teams on the risks that matter most.

Newsletter

7
03:49

[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork

Alibaba just released Qwen 3.8 Max, a giant open model that beats most paid competitors on coding and autonomous work, and promised to release its weights within a week. The 2.4 trillion-parameter model runs 95 billion active parameters per token, has a 1 million token context, and costs $2 per million input tokens and $6 per million output tokens. Third-party results are strong: 87.3% on SWE-bench, ahead of GPT-5.5, and number four on the Frontend Code Arena, close behind Anthropic and Kimi flagships. Its headline demos are autonomous — over ten days of unattended coding, a research run that beat the original paper's benchmark, and a chip design shrunk from 8,298 to 678 gates. The catches: at this size it needs multiple data-center GPUs so few can run it locally, and the open-weights license may restrict use in the US, EU, UK and Korea, which undercuts the 'open' label. The smaller 27B sibling, also going open-weight, is where most developers see the practical adoption wave.

Notes

Qwen3.8-Max (2.4T) and Qwen3.8-27B — open-weight launch

Source: Latent.Space AINews (2026-08-04), covering 7/25–7/27/2026 Twitter/Reddit reactions to Alibaba Qwen's launch.

The launch

Alibaba announced Qwen3.8-Max as its "most capable model to date" — a 2.4T total-parameter sparse MoE — and promised open weights "next week", alongside Qwen3.8-27B also going open-weight. Framed around "coding and cowork," not generic chat.

  • API pricing: $2.00/M input, $6.00/M output, $0.25/M cached tokens (down from $2.50/$7.50 for Qwen 3.7 Max).
  • Reported specs (ZhihuFrontier summary, not confirmed Alibaba spec sheet): 95B active params/token (~4% activation ratio, vs ~10% for Qwen3-235B-A22B); 1M-token context; 128k max output; low/medium/xhigh reasoning-effort modes; OpenAI + Anthropic protocol compatible.
  • Vendor benchmark claims: PaperBench 93.0, CoWorkBench 74.8, WideSearch 81.9.
  • Availability spread to Qwen Studio, API, Command Code, Venice; Baseten, Hermes Agent, Command Code confirmed support.
Alibaba's headline capability claims (vendor-reported, unverified)
  • Autonomous coding: 10+ days unattended with a public GitHub trace; rebuilt the "Unified Data Selection for LLM Reasoning" paper pipeline from scratch and ran a 125-hour iterative loop beating the original paper by +2.71 points.
  • Competitive DS: vs 526 human teams in WWW2025 Multimodal Dialogue Intent Recognition, top 13% within 24h.
  • Chip design: full GCD/RSA accelerator flow (RTL→sim→synthesis→layout), gate count 8,298→678, 81% die-area reduction, timing closure at 500 MHz; "500+ turns" of optimization.
  • E-Commerce Bench: 4.16x return (¥416,252) over a simulated 365-day store run.
  • Native multimodal planning loop (vision in execution loop, not just input); Qwen-MM-Plugins released for agent frameworks.
Independent results
  • Frontend Code Arena: #4 overall, 1,668 Elo — behind Claude Opus 5 [Max] 1,705 and Kimi K3 [Max] 1,676, ~tied with Opus 5 [High] 1,669. Sub-slice bests: #2 Consumer Product, #3 Brand & Marketing / Reference-based design / Gaming / Content Creation.
  • Vision Arena: #2 at 1,305, 13 pts behind Claude Fable 5 [High].
  • Vals Index: 66.1, #2 open-weight, #10 overall of 43; matches Claude Opus 4.7 (66.1 vs 66.1) at ~2.3x lower cost/test ($2.68 vs $6.17).
  • Vals benchmarks: SWE-bench 87.3% (GPT-5.5 82.6%, GLM-5.2 83.3%, Claude Opus 4.8 89.2%); Terminal-Bench 2.1 = 67.4 (vs 61.0 for 3.7 Max). Index trajectory: 57.5 → 66.1 (+8.6 in ~2.5 months).
  • Vals methodological note: Alibaba's reported Terminal Bench numbers modify benchmark timeouts; Vals preserved original timeouts.
Caveats and pushback
  • License: @OstrisAI read the license as forbidding even downloading from USA, EU, UK, Korea; no clarifying tweet from Alibaba appeared in the dataset, so this stayed unresolved. Echoed in the MiniMax H3 debate, where @VictorSuOrtiz later clarified regions "require a formal authorization process" rather than outright prohibition.
  • Deployment reality: @jaminball argued "vanilla" price comparisons ignore token efficiency and footprint — K3 weights alone >1TB memory, ≥8 H100/B200 GPUs, Moonshot recommending 64+ accelerators. A 2.4T-class MoE "is not a commodity local model." @stablequan: RAM-heavy prompts make it painful on consumer hardware; use the API.
  • @nrehiew_ asked how much of the delta is post-training vs novel pretraining. Some skepticism that benchmark jumps equal frontier parity on top-end agentic coding (@scaling01).
Strategic reading
  • ZhihuFrontier framed it as Alibaba choosing ecosystem influence over exclusivity — earlier Max models stayed closed; the open line previously topped out ~Qwen3-235B. DeepSeek/Kimi pressure pushed this.
  • Context set: Kimi K3 (2.8T), GLM-5.2 (744B total/40B active), DeepSeek V4 Flash/Pro, MiniMax H3. Artificial Analysis cited: Chinese frontier trails top US models by ~3–9 months; the open-weight frontier has been China-dominated for ~2 years.
  • @Cline: many open models are RL-trained to spend extra tokens verifying; harnesses letting them do so yield ~20% gains.
  • @TeortaxesTex: 3.8 Max may be distillable/OPD-able into 3.8-27B for task-specific laptop-deployable parity.
Practical implications

Frontier-open tradeoff: strong evals + aggressive token pricing + huge serving footprint + possible jurisdiction limits. The $0.25/M cached-token price matters for agents replaying codebases/tool traces. The 27B release is seen as the likely real adoption tier; protocol compatibility (OpenAI/Anthropic) lets it drop into existing harnesses (Hermes Agent, Command Code, Baseten) quickly. Vals confirmed runtime settings: 1M context, 128k output, temp 0.7, default top-p/top-k.

Full text · 37,877 chars
[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork Qwen is so back! After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models. That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered. Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models. Key Capabilities & Breakthrough Highlights - Autonomous Long-Horizon Coding: - 10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run. - Autonomous AI Research: Rebuilt a complete paper’s pipeline (Unified Data Selection for LLM Reasoning) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper’s benchmark by +2.71 points. - Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% (outperforming 87% of human teams) within 24 hours. - Autonomous Hardware & Chip Design: - Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout. - Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz. - Deep Real-World Work & Operations: - Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research). - Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning. - Multimodal Agents & Visual Feedback: - Integrates native visual feedback across planning, coding, and GUI interaction, enabling direct application recreation across platforms (desktop, mobile, web). - Released Qwen-MM-Plugins to extend multimodal capabilities to existing agent frameworks. A very nice win for open weights! On today’s pod with Baseten we talked about what it’s like to support these massive model drops on release. AI News for 7/25/2026-7/27/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap Top Story: Qwen 3.8 Max open model launch What happened Alibaba Qwen announced Qwen3.8-Max as its new flagship and said open weights are coming next week. - Alibaba introduced Qwen3.8-Max as its “most capable model to date,” describing it as a 2.4T-parameter model focused on coding, long-horizon agentic work, and multimodal reasoning, with the explicit claim that open weights will be released next week, alongside Qwen3.8-27B also going open-weight @Alibaba_Qwen - The launch tweet also included API pricing: $2.00 / M input tokens, $6.00 / M output tokens, and $0.25 / M cached tokens @Alibaba_Qwen - Alibaba framed the model around several headline capabilities: 10+ days of autonomous coding, 500+ turns of chip design optimization, 365 days of e-commerce strategy, and native multimodal intelligence where vision is part of the execution loop rather than just an input channel @Alibaba_Qwen - The company simultaneously pushed availability across its own surfaces and partners: Qwen Studio, API, Command Code, and later Venice; infra and app builders quickly confirmed support plans or integrations including Baseten, Hermes Agent, and Command Code @Alibaba_Qwen @Alibaba_Qwen @baseten @Teknium - The announcement landed as part of a broader pattern: multiple observers described it as evidence that the Chinese open-weight frontier is now competing directly with top Western closed models, especially in coding, agentic workflows, and multimodal tasks @kimmonismus @matvelloso Official claims and reported specs Vendor-reported model details and performance claims were unusually aggressive for an open-weight release. - Alibaba’s own framing: - 2.4T total parameters @Alibaba_Qwen - Long-horizon agentic/cowork focus @Alibaba_Qwen - Autonomous coding over 10+ days with a public GitHub trace @Alibaba_Qwen - 500+ turns for chip design optimization @Alibaba_Qwen - 365 days of e-commerce strategy execution @Alibaba_Qwen - Native multimodal planning loop rather than vision-only input @Alibaba_Qwen - Third-party summary tweet from ZhihuFrontier added more claimed or reported technical details: - 95B active parameters per token, implying an MoE activation ratio of roughly 4% - 1M-token context window - API exposes low / medium / xhigh reasoning-effort modes - Compatibility with OpenAI and Anthropic protocols - Benchmark claims: PaperBench 93.0, CoWorkBench 74.8, WideSearch 81.9 @ZhihuFrontier - Vals AI independently posted concrete eval/runtime settings: - 1M token context - 128k max output tokens - Tested at temperature 0.7 with default top-p / top-k @ValsAI These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2, not the more practical 30B–70B local tier. Independent evaluations and leaderboard placements The model immediately posted strong third-party results, especially in coding-adjacent, vision, and design-heavy arenas. - Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676, and roughly tied with Claude Opus 5 [High] at 1,669 @arena - In Frontend Code Arena subslices, it ranked: - #2 Consumer Product - #3 Brand & Marketing, Reference-based design, Gaming, Content Creation Tools - #4 Data & Analytics - #5 Simulations @arena - Vision Arena: Qwen3.8-Max ranked #2 with 1,305, only 13 points behind Claude Fable 5 [High] @arena - Vals Index: Qwen3.8-Max ranked #2 among open-weight models, #10 overall out of 43, with a score of 66.1 @ValsAI - Vals also reported: - It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1 - At about 2.3x lower cost per test: $2.68 vs $6.17 @ValsAI - Vals’ benchmark-specific numbers: - SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%) - Terminal-Bench 2.1: 67.4, up from 61.0 for Qwen 3.7 Max @ValsAI - Vals also highlighted the pace of progress: - Qwen 3.7 Max = 57.5 - Qwen 3.8 Max = 66.1 - Gain of 8.6 points in ~2.5 months - Price cut from $2.50/$7.50 to $2.00/$6.00 input/output @ValsAI There were also more anecdotal but technically relevant claims: - One user visualized benchmark deltas and argued “Opus 4.8 is mostly subsumed by 3.8-Max” on the chart they reconstructed @deliprao - Another claimed Qwen 3.8 surpassed Fable 5 on Terminal Bench and said Anthropic was now under visible pressure @kimmonismus - A separate tweet called Qwen 3.8 Max the “best object detection VLM” across satellite, infrared, documents, technical drawings, sketches, crowded scenes, and small objects, though this was based on examples rather than a cited benchmark paper @skalskip92 Facts vs. opinions Facts / directly attributable claims - Alibaba announced Qwen3.8-Max and said open weights arrive next week; Qwen3.8-27B will also go open-weight @Alibaba_Qwen - Alibaba disclosed API pricing of $2 input / $6 output / $0.25 cached per million tokens @Alibaba_Qwen - Arena reported #4 in Frontend Code Arena at 1,668 and #2 in Vision Arena at 1,305 @arena @arena - Vals reported 66.1 on Vals Index, #2 among open-weight models, 87.3% SWE-bench, 67.4 Terminal-Bench 2.1, 1M context, 128k output, and lower cost-per-test than Opus 4.7 @ValsAI @ValsAI @ValsAI - ZhihuFrontier stated 95B active parameters and protocol compatibility; this appears to be a secondary summary rather than an original Alibaba spec sheet @ZhihuFrontier Opinions / extrapolations / rhetoric - “China is no longer lagging behind but competing on equal footing” @kimmonismus - “Open models are winning now” @JonathanRoss321 - “Looks like Opus 4.8 is mostly subsumed” @deliprao - “Anthropic is under pressure” and “mood shifted drastically” are ecosystem readings, not measurements @kimmonismus - “Best object detection VLM” is an informed product judgment, but not one tied in-thread to a standard benchmark table @skalskip92 - Claims that Qwen3.8-Max plus open agents prove open models have “caught up” are user-level interpretations rather than consensus eval conclusions @omarsar0 The central factual story is strong even after stripping out the hype: a very large sparse model, open-weight promise, lower pricing than prior Qwen Max, and high placements on multiple third-party leaderboards. The infrastructure reality: “open-weight” does not mean easy to run A major counterpoint in the discussion was that frontier open models are operationally open, but not broadly accessible in the local-inference sense. - Jamin Ball argued that pricing comparisons were overstated because “vanilla” token prices ignore token efficiency and because these models are enormous: - Qwen 3.8 Max >2T params - Kimi K3 ~104B active per token - GLM 5.2 = 744B total, 40B active - For K3, loading weights alone is >1TB memory - Requires at least 8 H100/B200 GPUs to run - Moonshot recommends 64+ accelerators in supernode-style setups @jaminball - This same critique implicitly applies to Qwen3.8-Max, even if its active-parameter count is somewhat lower than K3’s: a 2.4T-class MoE is not a commodity local model @jaminball - StableQuan made the practical version of the same point more bluntly: long, RAM-heavy prompts and slow tool calls make giant models painful on consumer hardware, recommending API use instead @stablequan - At the same time, the excitement around Qwen3.8-27B shows where many developers think the real adoption wave may come from: a smaller open-weight descendant in the same family, possibly inheriting some of the flagship’s post-training or distilled capabilities @kimmonismus @TheZachMueller This is the key split in the open-model story: ecosystem influence and benchmark legitimacy come from releasing the 2.4T flagship; practical deployment at scale may come from the 27B release. Licensing controversy and geographic restrictions The most concrete skeptical reaction was not about performance, but about the license. - OstrisAI flagged what they read as a license prohibition covering the USA, EU, UK, and Korea, saying the terms appeared to forbid even downloading the model from the US @ostrisai - That concern echoed a broader discussion happening simultaneously around another open-weight release, MiniMax H3, where users argued that geographic restrictions undercut claims of openness @kimmonismus - No clarifying Qwen license tweet appears in this dataset from Alibaba itself, so the restrictive-license reading remained unresolved within these tweets For engineers, this matters more than the marketing label. “Open weights” can still mean: - no OSI-style open-source rights, - use-case restrictions, - export/jurisdiction limits, - or no legal permission for commercial deployment in key regions. That licensing ambiguity is one of the main reasons some of the reaction was more cautious than celebratory. Why the launch matters strategically This was widely read as a strategic shift by Alibaba, not just a routine product update. - ZhihuFrontier explicitly framed the move as Alibaba choosing ecosystem influence over exclusivity, arguing that earlier Max models stayed closed while the open line had previously topped out around Qwen3-235B @ZhihuFrontier - In that reading, DeepSeek, Kimi, and other Chinese open models weakened the premium of keeping top-tier systems API-only, pushing Alibaba to compete on ecosystem adoption as well as model quality @ZhihuFrontier - Multiple observers connected Qwen3.8-Max to a broader Chinese-model surge: - “Top three spots in front-end design are now shared between two Chinese and one Western model” @kimmonismus - “Remember when China was 2 years behind?” @matvelloso - “The open weights frontier has been consistently dominated by labs from China for the last two years” @_micah_h - Some posters escalated this into a geopolitical concern that US labs cannot rely on closed-model leads forever, especially if Chinese labs keep pushing frontier-ish systems into open-weight channels @kimmonismus A subtext here is that the moat may be shifting: - not just raw pretraining, - but post-training, agent harnesses, inference infra, distillation pipelines, and developer lock-in. That is exactly why an open-weight flagship at 2.4T is strategically valuable even if relatively few teams ever self-host it. Model architecture and sparsity implications The technical profile suggests Alibaba is leaning harder into sparse MoE than some rivals. - If the 95B active / 2.4T total number quoted by ZhihuFrontier is accurate, Qwen3.8-Max activates only about 4% of total parameters per token @ZhihuFrontier - ZhihuFrontier contrasted this to Qwen3-235B-A22B, which they say activates closer to 10% @ZhihuFrontier - Elie Bakouch’s broader comment—“the two biggest OSS models in the world use linear attention?”—captures another architectural thread in the ecosystem conversation, though it was not directly tied to Qwen3.8-Max with a cited source in-thread @eliebakouch - The wider thread around sparse MoE and Switch Transformers reflects why people care about these parameter numbers: frontier open models can look “huge to store yet still cheap to run” by only activating a narrow expert slice per token @ProfTomYeh This is likely part of how Alibaba can cut API pricing while scaling total parameter count upward: bigger expert pool, lower active footprint, lower effective inference cost, assuming routing and systems optimizations hold up in production. Long-horizon agents, cowork, and benchmark fit Qwen3.8-Max was pitched less as a chatbot and more as a model-harness substrate for long-running work. - Alibaba’s own language emphasized “coding and cowork” rather than generic assistant use @Alibaba_Qwen - The launch claims map unusually well to the current “long-horizon agents” discourse: - 10+ day autonomous coding - 500+ turns in chip optimization - 365-day business strategy @Alibaba_Qwen - ZhihuFrontier’s benchmark picks—PaperBench, CoWorkBench, WideSearch—all emphasize persistent objective maintenance, tool use, and trajectory coherence rather than one-shot Q&A @ZhihuFrontier - Omar Sar0 explicitly linked the release to agent harnesses, saying using Qwen3.8-Max in Hermes Agent makes it hard to deny how much open frontier models have closed the gap with closed frontier systems @omarsar0 - Cline’s separate thread about open-weight models is relevant context: they argue many open models are RL-trained to spend more tokens on verification and work best when the harness lets them lean into that behavior, producing ~20% gains from harness changes alone @cline That fits Qwen3.8-Max’s launch narrative unusually well. The implication is not simply “model is smarter,” but “model may be especially competitive when paired with a harness designed for long-running verification-heavy work.” Different perspectives in the reaction Supportive - Strong enthusiasm from open-model developers and infra providers: - “Qwen 3.8 Max and a new local 27B Qwen 3.8 is coming” @Teknium - “Yes, we will have Qwen3.8-Max” @baseten - “Try Qwen3.8-Max on Hermes Agent…” @omarsar0 - “Nice! An open source max model” @NerdyRodent - Several commenters treated the release as proof that open models are at or near frontier parity on meaningful workloads @JonathanRoss321 @kimmonismus Neutral / analytical - Jamin Ball’s thread was the main “yes, but” reaction: - pricing gap may be overstated, - token efficiency matters, - infra burden remains extreme for >2T open models @jaminball - Nrehiew questioned whether performance gains might come disproportionately from post-training rather than novel pretraining, essentially asking how much of the delta is recipe vs scale @nrehiew_ - Vals added an important methodological note: Alibaba’s reported Terminal Bench results modify benchmark timeouts, whereas Vals preserved original timeouts @ValsAI Skeptical / opposing - License concern was the clearest substantive criticism: if usage is restricted in major markets, “open” becomes a narrower claim @ostrisai - Some of the strongest skepticism was indirect: if these giant open-weight models require supernodes and careful harness engineering, then their practical competitive effect may be less dramatic than leaderboard headlines suggest @jaminball - There was also broader ecosystem skepticism that benchmark jumps alone prove full parity with the strongest closed models; e.g. some users argued open source is “very close” but not actually there yet on top-end agentic coding @scaling01 Context: Qwen3.8-Max inside the 2026 open-model cycle The launch sits in a dense cluster of giant open or quasi-open releases from Chinese labs. - The comparison set repeatedly mentioned in the discussion: - Kimi K3 at 2.8T - GLM-5.2 - DeepSeek V4 Flash / Pro - MiniMax H3 on the multimodal/video side @jaminball @kimmonismus - Artificial Analysis commentary cited in-thread said Chinese frontier models have generally trailed top US models by about 3–9 months, while the open-weight frontier itself has been dominated by Chinese labs for roughly two years @_micah_h - This helps explain why the release drew such outsized attention: it is not just another model launch, but part of a visible realignment where: - China is strongest in open-weight frontier scale - US labs still often lead in top closed-model performance - the gap is narrowing on select domains like coding, design, and some multimodal tasks @_micah_h @kimmonismus Practical implications for engineers For engineers, the most important questions are less about marketing claims and more about deployment shape. - If you want frontier-ish open-weight quality, Qwen3.8-Max suggests the tradeoff space is now: - very strong eval performance - aggressive token pricing - huge serving footprint - possible license/jurisdiction constraints - The 1M context and 128k output numbers make it viable for repository-scale and workflow-scale tasks where transcript reuse and cache pricing matter @ValsAI @Alibaba_Qwen - The cached-token price of $0.25/M is especially relevant for agents repeatedly replaying codebases, tool traces, and large instruction prefixes @Alibaba_Qwen - The announcement of Qwen3.8-27B may be just as consequential as the flagship, because it is the tier likeliest to become actually usable across broader open-source stacks and local-serving ecosystems @Alibaba_Qwen @kimmonismus - Several developers already framed the release in terms of downstream harnesses and agents, not just chat UX: Hermes Agent, Command Code, Baseten, and likely any provider supporting OpenAI/Anthropic-compatible protocols can slot it into existing workflows quickly @Alibaba_Qwen @Alibaba_Qwen @baseten - One notable interpretation from TeortaxesTex was that Qwen 3.8 Max may be: - exceptionally strong on image recognition/labeling - potentially sample efficient - and distillable/OPD-able into Qwen 3.8 27B for task-specific parity, implying a route from flagship capability to laptop-deployable specializations @teortaxesTex Other Topics Agent infrastructure, harnesses, and long-horizon systems - A detailed survey summary argued that long-horizon capability is a model × harness property, not just a model property; it breaks failures into goal drift, context corruption, and sparse-reward/irreversible-action issues, and frames the control plane as shifting from prompt engineering to runtime harnesses @ZhihuFrontier - Cloudflare launched @cloudflare/computer, an agent runtime that dynamically routes between isolates and Linux containers so each agent gets “a computer of its own” @Cloudflare - Cursor reported 20–30% better token efficiency for cloud agents and 80% better efficiency on computer-use runs, plus launched plugins for Google Workspace access across Gmail, Drive, Calendar, Docs, and Sheets @cursor_ai @cursor_ai - LangChain signaled managed Deep Agents moving to public beta, with built-in evals, memory, OAuth tool access, channel integrations, and sandboxing @hwchase17 - Several posts emphasized that harness choice materially changes benchmark outcomes and production behavior: - endpoint choice changed Kimi K3 results dramatically on CEO-Bench @tonychenxyz - Cline says open-weight models often benefit when allowed to spend extra tokens on verification, yielding ~20% gains in their runs @cline - a new paper organized 41 agent failure modes by interaction edge rather than single component, with automated labeling reaching κ = 0.76 vs humans @omarsar0 Benchmarks, evals, and automated research/post-training - RSIBench-Data results put Kimi K3 + Kimi Code at 27.317% weighted score across six benchmarks, including 50% SWE-bench Verified and 17% SWE-bench Pro @FanqingMengAI - Intology said its automated AI research system Locus is SOTA on PostTrainBench, and that Locus-post-trained Qwen3 1.7B Base variants surpassed the official human post-trained Qwen3 1.7B release; on live Kaggle comps it reached the 4th highest average rank after 16 days @intology - Epoch updated MirrorCode with Claude Fable 5 at 64% solve rate and GPT-5.6 Sol at 20%, using 15 Medium/Large programs, 2 languages each, and 10B tokens per attempt @EpochAIResearch - Shahules argued benchmarks should release trajectories, not just scores, because task defects and brittle verifiers can dominate failures; they also highlighted ITSMBench as an open benchmark with trajectories @Shahules786 - New eval/benchmark artifacts included: - MerchantBench: 365-day e-commerce simulation with 98,843 product records, 26 tools, score on cumulative net assets @dair_ai - One Layer Deeper: adaptive-computation challenge based on repeated modular squaring @SolidlySheafy - Artifacts Hub / Adoption Dashboard tracking 792 open models, downloads, intelligence, and geography @natolambert Open models, inference, and systems engineering - Multiple posts stressed the open frontier is now dominated by giant MoEs from China, with Kimi K3, Qwen3.8-Max, GLM, and DeepSeek frequently compared on scale/cost/perf @_micah_h - Databricks claimed #1 Kimi K3 inference speed/latency on Artificial Analysis, quoting 239 tok/s in one post and separate single-node numbers from Casper Hansen of 947 tok/s batch-32 decode and 152 tok/s single-user on a single B300 node @Yuchenj_UW @casper_hansen_ - Vikhyat announced Photon 2.0, compiling Moondream, Qwen 3.5, and Gemma 4 into megakernels spanning the full forward pass @vikhyatk - A systems paper thread on TokTier argued tokenization can consume up to 64% of TTFT in cached-agent workloads, with stateful tokenization reducing TTFT by 16–34% and incremental repair 437× faster than HF tokenizers in some settings @omarsar0 - DSPy 3.3.0 shipped: - dspy.Flex for optimizing code + prompts - ReActV2 with native/parallel tool calling - typed provider-neutral LM interface @isaacbmiller1 Multimodal, video, and vision models - MiniMax H3 dominated discussion outside Qwen: - described as a 33B video model with text/image/video/audio references, up to 15s clips, runnable on one RTX 5090 with ComfyUI stack around 40GB and 5s generations in ~5.5 min in early tests @kimmonismus - later ranked #1 open model in Video Arena, +280 pts over next-best open, and tied near the top overall in image-to-video @arena - There was an active license debate around H3 too: one side said it cannot legally be used in the US/EU/UK/Korea under the public license @kimmonismus, while another clarified formal authorization is available via MiniMax and that “cannot legally be used” is too strong @VictorSuOrtiz - Jina released jina-reranker-v3.5, a 0.6B listwise reranker scoring 63.20 nDCG@10 on BEIR, beating Qwen3-Reranker-4B at roughly 7× fewer parameters @JinaAI_ - Qwen3.8-Max also drew attention for vision/object detection use cases, including documents, infrared, satellite, and crowded scenes, with claimed per-image cost around $0.007 @skalskip92 Frontier labs, policy, safety, and competition - A large meta-thread in the timeline concerned US vs China and whether Chinese labs are catching up or already ahead in some open/frontier segments: - Hugging Face CEO coverage said China is winning/dominating open models @CNBC - Artificial Analysis data was cited saying Chinese leaders historically trail top US models by 3–9 months @_micah_h - some posters argued Chinese aggregate research capability may already exceed US labs despite resource asymmetries @teortaxesTex - The White House reportedly invited OpenAI, Anthropic, Google, and Meta to review a new voluntary AI framework and finalized new cybersecurity tests/hacking benchmarks @steph_palazzolo @AndrewCurran_ - Cybersecurity remained a major subtheme: - Epoch reported roughly 2,500 high/critical CVEs disclosed in July across 21 major tech orgs, about 5× the prior monthly record before Anthropic’s autonomous vuln-finding disclosure @EpochAIResearch - Hugging Face interviews argued open-weight models were part of the defensive response after the OpenAI-linked hack @BloombergTV @BusinessInsider - OpenAI announced an internal next model found 10 new results on long-standing open problems in math/theory CS for roughly $2,000 in token cost at GPT-5.6 Sol rates, prompting both excitement and skepticism about total attempt cost vs solved-cost accounting @OpenAI @NickEMoran - OpenAI also published a technical deep dive on GPT-Live, noting a dedicated low-latency audio path, async reasoning/tool use, and startup reduced from 6 round trips to 1 @OpenAI @gdb Product and ecosystem notes - Google rolled out Gemini Spark auto browse using Chrome to act in logged-in accounts for errands with user confirmation on sensitive steps @Google - Google AI Studio prompted developers for current “vibe coding” projects, while Gemini-side product messaging emphasized business-building workflows in Notebooks/Canvas @GoogleAIStudio @Google - Sakana launched Namazu API, described as a Japanese-focused LLM built on Kimi and tuned for Japanese language/culture/business, with reduced unnecessary refusals and bias @SakanaAILabs @SakanaAILabs - LiteParse added direct structured PDF extraction for form fields, checkbox states, annotations, embedded images, vector graphics, tagged structure, and word-level bounding boxes in ms/page for simple pages @llama_index - The Hermes Agent ecosystem shipped a substantial “Herald” release with voice chats, plugin-based desktop features, A2A protocol, outbound webhooks, research and productivity skills, and token-efficiency improvements @Teknium China’s open-model surge: Kimi, DeepSeek, GLM, and the narrowing gap - Open-weight frontier now looks China-led: Across the digest, the dominant meta-story is that Chinese labs are setting the pace in open models. Posts from @kimmonismus, @JonathanRoss321, and @_micah_h all point to the same pattern: Kimi, Qwen, DeepSeek, GLM, and MiniMax now define much of the open frontier, while US labs retain lead positions mainly in select closed offerings. @ClementDelangue and related coverage amplified the broader claim that China is dominating the open-weight lane. - Kimi K3 and harness sensitivity: K3 continued to post strong downstream and infra results. RSIBench-Data reported Kimi K3 + Kimi Code at 27.317% weighted score across six automated-research benchmarks, including 50% SWE-bench Verified and 17% SWE-bench Pro. But @tonychenxyz noted a key engineering caveat: inference provider materially changed leaderboard outcomes, with one provider producing degraded looping behavior while Modal’s endpoint yielded #1 results on CEO-Bench. On the serving side, @Yuchenj_UW said Databricks now delivers 239 tok/s and top latency for K3, while @casper_hansen_ cited 947 tok/s decode throughput at batch 32 on a single B300 node. - DeepSeek V4 Flash as the cost/performance disruptor: DeepSeek’s latest Flash checkpoint emerged as the day’s strongest cost-adjusted agent model story. @htihle reported 57.1% / 63.0% on WeirdML for Flash-0731 high/max and argued the harness may understate true ability. Vals called DeepSeek V4 Flash (0731) the cheapest model on the Vals Index above 60, and 35× cheaper than the next best model at that threshold, with most of the advantage coming from coding and agentic tasks. Together AI immediately positioned it as a production endpoint for long-running agents. - GLM and what’s next: Multiple posts suggested GLM-5.3 is imminent, including @AiBattle_ and @arena, which reminded readers that GLM-5.2 Max already sits #2 overall and #1 open in Frontend Code Arena. Agent harnesses, long-horizon systems, and why model quality alone is no longer enough - Harnesses have become the control plane: A recurring theme across technical tweets is that long-horizon performance is now best understood as model × harness, not model alone. A detailed survey summary from @ZhihuFrontier frames long-horizon capability as emerging from co-evolution between base models and runtime systems handling memory, planning, tool use, verification, orchestration, and recovery. This aligns with @omarsar0, who highlighted a paper categorizing 41 agent failure modes by interaction edges between model, user, harness, tools, memory, and environment rather than blaming a single component. - Production runtimes are shipping fast: Cloudflare introduced @cloudflare/computer, an agent runtime that dynamically switches between lightweight isolates and full Linux containers. Cursor said its cloud agents are now 20–30% more token efficient and 80% more efficient on computer-use runs, then followed with direct Google Workspace plugins for Gmail, Drive, Calendar, Docs, and Sheets launch. LangChain said Managed Deep Agents will move to public beta with built-in evals, memory, OAuth, channels, and sandboxing. - Open-model harness co-optimization is starting to matter: Cline offered one of the sharper practitioner observations of the day: many open models appear RL-trained to spend extra tokens verifying work—rerunning tests, checking builds, rereading diffs—and Cline deliberately lets them “work how they were trained to work,” claiming roughly 20% gains from harness changes alone. That theme also appears in posts around Hermes Agent from @Teknium, which shipped voice activation, plugin/API expansions, A2A protocol support, outbound webhooks, research skills, and major token-efficiency improvements. - Memory and parsing are being de-LLM-ified where possible: @dair_ai highlighted Zero-Mem, which removes LLM calls from memory maintenance and only invokes an LLM at final answer time, cutting memory-op cost by 57.6% versus the fastest baseline at matched budget. LlamaIndex similarly shipped richer structured PDF extraction in LiteParse, exposing fields, checkboxes, annotations, graphics, and page complexity signals without requiring a vision model for every page. Automated research, post-training, and benchmark design are becoming more serious engineering disciplines - Automated post-training is yielding real wins: @intology claimed its Locus system is SOTA on PostTrainBench and can post-train Qwen3 1.7B-Base variants that surpass the official human-tuned Qwen3 1.7B Instruct model under expanded compute budgets. The same post says Locus generalized to live Kaggle competitions, reaching the 4th highest average rank after 16 days. Separately, @mervenoyann pointed to public tooling for coding-agent RL pipelines based on sandboxed tasks, TRL, and verifiers. - Research automation benchmarks are exposing harness effects: The terse but high-signal RSIBench-Data result and @gneubig’s reaction underscore that very-long-horizon automated research tasks are increasingly measuring specialized research harnesses, not just model intelligence. That also surfaced in a critique from @Shahules786, arguing benchmarks should open-source full trajectories, since scores alone obscure whether failures stem from weak models, brittle verifiers, or under-specified tasks. - Noise, verification, and held-out reality still bite: @ddkang pushed back on the idea that RLVR with 100% noisy data matches clean-data training, reporting >9% lower MATH accuracy under more rigorous noisy-data construction. @ArmenAgha shared a smaller but instructive result where optimizing a proxy objective improved selected velocity MSE but made actual rollout inference worse on held-out data. This is a useful reminder that a lot of “self-improvement” headlines still collapse if evaluation is not robust. Multimodal and video systems: MiniMax H3, world models, and local generation - MiniMax H3 is the standout multimodal/video release: The community response suggests MiniMax H3 is a major step forward for open-weight video generation. @arena ranked it the #1 open model in Video Arena across both text-to-video and image-to-video, with +280 points over the next-best open model; in image-to-video it was effectively tied for #1 overall. @MiniMax_AI said H3 is now the SOTA open video generation model on both Arena and Artificial Analysis benchmarks. - Why H3 matters technically: Multiple posts emphasized that H3 is not just another T2V model but a general-purpose multimodal generation model with text, image, video, and audio in a single context, plus usable local deployment pathways. @kimmonismus summarized the key caveat clearly: open weights, strong local video potential, but not a fully open-source stack, since context orchestration, 2K regeneration, and sparse attention remain server-side or otherwise restricted. @ComfyUI, @victormustar, and @MiniMax_AI all highlighted practical local workflows, including RTX 5090-class usage. - Licensing remains messy: There was confusion around H3’s geography restrictions. @ostrisai initially read the license as forbidding usage in the US/EU/UK/Korea, and that concern spread. Later, @VictorSuOrtiz clarified that those regions require a formal authorization process rather than being outright impossible to license, which is an important distinction for teams evaluating deployability. - World models and multimodal simulation remain an emerging thread: Several lower-engagement but technically substantive posts pointed toward unsupervised latent simulators and world-model-style systems as a growing area, including @soniajoseph_ and @taiuti. Inference systems, compilers, realtime voice, and other infra worth tracking - Realtime voice stack redesign at OpenAI: OpenAI detailed a new GPT-Live architecture that supports full-duplex conversation—listening while speaking—by separating a dedicated fast audio path from slower asynchronous reasoning/tool-use paths. They also cut session startup from six network round trips to one and discussed async compaction for long-context voice sessions in the linked engineering writeup and follow-on thread from @juberti. - Compilers are eating hand-tuned inference kernels: @vikhyatk announced Photon 2.0, a compiler that turns models like Moondream, Qwen 3.5, and Gemma 4 into megakernels representing the whole forward pass as a single GPU program. The thread describes a tracer DSL for dataflow specification and a CPU cost model to prune scheduling candidates before compilation. That pairs well with the broader discussion from @waterloo_intern, arguing that classical hand-optimized GPU kernel work is being progressively automated and commoditized. - Tokenization and serving bottlenecks are now first-class: @omarsar0 highlighted TokTier, a stateful tokenization service that reuses and repairs tokenized prefixes for agent sessions, reporting 16–34% TTFT reductions under vLLM and up to 437× speedups over standard Hugging Face tokenization in incremental repair scenarios. This is exactly the kind of “non-model” bottleneck that matters once agent transcripts get long and cache hit rates are high. - Smaller but notable tools: Jina AI released jina-reranker-v3.5, a 0.6B listwise reranker claiming 63.20 nDCG@10 on BEIR and beating Qwen3-Reranker-4B at roughly 7× fewer params; DSPy 3.3.0 added code-and-prompt optimization via dspy.Flex, improved tool use with ReActV2, and a provider-neutral LM interface. Top tweets (by engagement) - Qwen3.8-Max release: Alibaba’s announcement of a 2.4T flagship with open weights next week was the biggest technical launch of the set @Alibaba_Qwen. - OpenAI math result: OpenAI said an internal version of its next major model produced 10 new results on long-standing open problems in math and TCS for roughly $2,000 in GPT-5.6 Sol-equivalent token cost @OpenAI. - GPT-Live architecture: OpenAI’s new realtime voice stack supports continuous listening while speaking and asynchronous tool/reasoning execution @OpenAI. - Source code abstraction debate: Elon Musk argued that source code is on the verge of becoming like assembly, with AI eventually compiling intent straight to binaries @elonmusk. - Cursor workspace integration: Cursor shipped agent access to Google Workspace apps, moving coding agents closer to general work automation @cursor_ai. AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Qwen3.8-Max and 27B Open-Weight Launch Keep reading with a 7-day free trial Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.
18:20

Unpacking ChatGPT Work: the Agent for a Billion Users

OpenAI's ChatGPT Work is an agent that does your knowledge work on a persistent cloud computer, and OpenAI is positioning it as the future of ChatGPT, which is on track to hit a billion weekly users. Released July 9 and built on the Codex harness, it reportedly passed 10 million users within three weeks and pulls context from Slack, email, Drive, calendars, and hundreds of plugins to produce finished docs, spreadsheets, and hosted web apps. Each task runs in an isolated cloud machine (Pro gets 8 CPUs and 20GB of RAM), and working state persists between tasks, but memory, files, and conversation history live in ChatGPT's product layer rather than on the computer itself. A separate desktop mode runs directly on your own machine, and OpenAI has confirmed Chat and Work will merge by the end of the year.

Notes

Unpacking ChatGPT Work: the Agent for a Billion Users

Guest post by Shlok on Latent.Space (2026-08-04). An "external reconstruction" — most claims come from Shlok and Codex poking around inside the product, with linked conversations cited throughout. Insider counterpoint: Latent.Space podcast with OpenAI's Akshay Nathan.

Launch facts
  • OpenAI released ChatGPT Work on July 9, 2026: 3 new models across 14 configurations, consolidation of the ChatGPT and Codex desktop apps.
  • Three weeks in, Work + Codex reportedly crossed 10M users.
  • Editor's note: ChatGPT estimated at 1B MAU in June, 1B WAU in July 2026.
  • Greg Brockman confirmed Chat and Work will merge by end of 2026; Work is thus a preview of how the billion-user app will behave.
What it is
  • Agent for knowledge work: connects to Slack, email, Drive, calendars, CRMs, project trackers, "hundreds" of plugins.
  • Runs on the Codex harness (same models, sub-agents, browser use, hours-long grinding), but the UI strips out evidence like git controls and diff-traces that would reveal a coding agent.
  • Lives in a cloud microVM: Pro = 8 CPUs, 20GB RAM, 64GB disk; Plus = 14GB RAM. Plus a managed Chrome service the agent operates via tool calls.
  • Produces artifacts: Sheets/docs/slides in interactive viewers, plus Sites (hosted web apps/dashboards, shareable by URL, kept updated).
  • Each conversation is a task. Web/mobile always run in cloud; desktop has cloud (syncs across all three) and local modes. Local mode = full computer use on your machine; local tasks don't appear on web/mobile and can't be moved to cloud. "Essentially Codex, minus the code-related UI traces." Codex's task-handoff to remote environment was released but "doesn't work for me at the time of writing."
Persistence & memory
  • Workspace is synced to persistent storage and restored onto isolated microVMs as needed — the machine can change, working state carries over.
  • Each task gets a working dir under /workspace/scratch with full OS freedom (folders, deps, scripts, databases, Linux commands).
  • Cross-thread continuity runs through the ChatGPT product layer, not the computer:
  • New threads get a compressed summary of recent tasks/files, e.g. 20260731T15:55 Prepare Acme pilot plan:||||... <<File name="acme_notes.txt">>.
  • Raw transcripts aren't stored on the computer; the Personal Context tool queries Chat/Work history via a separate service.
  • Library is the central file repo; uploads land there automatically, agent files saved on request or when judged worth retaining. An uploaded file exists twice (thread working copy + Library canonical item) and the two don't sync — Thread B's Library edit leaves Thread A reading its stale copy.
  • Agent can enter other tasks' scratch dirs only when explicitly told; names are opaque, with "no legible map" and "no stated retention contract."
  • Memory = a running synthesized user profile maintained asynchronously, supplied at task start; agent can reason from it but can't modify it or write OpenClaw-style Markdown that other tasks load.
  • Projects carry over (instructions, conversation summaries, local copies of Sources) but aren't directories on the computer.
  • Shlok's three reasons for the split: reuse of existing primitives for a billion users; guardrail against OpenClaw-style unrestricted access; OpenAI keeps product control (UI, context, sharing, sync, versioning).
  • Missing today: a meta-layer agent coordinating between tasks/projects.
Proactivity

New Work conversations show personalized task suggestions generated from context. Example: a suggestion to prep for an upcoming call injected a pre-authored prompt, having "reasoned asynchronously across my context" — calendar event, Calendar+Gmail data, memory preferences — producing a "great meeting brief." Current limit: it only suggests; nothing runs until the user executes.

Scheduled Tasks
  • Launched as Scheduled Tasks in Jan 2025; Work makes them agentic (runs use context and tools).
  • Standalone: fresh task per run from a saved prompt — one-off reminders, daily briefings, weekly job search.
  • Heartbeat: reawakens an existing conversation with context intact — monitoring, polling, review loops. Works in desktop only, not exposed on web.
  • One-time or recurring; triggers are exact time, loose window ("in the morning"), or a monitored condition. Managed in-conversation or on the Scheduled page (next run, results, create/edit/pause/delete). The page also suggests automations — generic (Daily Brief) and personalized (weekly recap of Shlok's football club).
Browser Use
  • History: Operator and ChatGPT agent first gave logged-in click-through ability; then core to Codex; most integrated in Work.
  • Browser is a separately hosted Chrome service controlled via tool calls (inspect, click, type, scroll, screenshots, tabs, dialogs, file moves).
  • Web/desktop show a replayable timeline of browser states; you can take over the live browser (enter a password) and hand it back; not possible on mobile.
  • Persistent browser profile: dark-mode Wikipedia and a Google login were inherited by a fresh task. The agent never sees the profile/credentials; a permission ledger synced into the computer records per-site, per-conversation permission and file-move rights.
  • Datacenter constraints: Amazon US rejected it as an unsupported "session or client"; Google Photos timed out copying a shared album — both worked in local mode. CAPTCHA only with permission; instructed not to loop, rotate fingerprints, or evade site safeguards.
Plugins, skills, tools
  • Lineage: Plugins (Mar 2023) → GPTs/Actions (Nov 2023) → connectors (Jun 2025) → apps/Apps SDK/App Directory (late 2025) → plugins returned to Codex (Mar 2026) → July 9: App Directory became Plugin Directory, apps packaged as plugins across Work and Codex.
  • A plugin can hold: Apps (Gmail, Slack, Salesforce; most use an MCP server), Skills (instructions + references/templates/scripts), App templates (org-configured apps).
  • Three types: Operational (Computer Use, Sites, Documents, Presentations, Spreadsheets), Role-specific (Sales plugin = 20 skills like "Analyze Account Signals" / "Build Business Case" across 29 apps incl. Salesforce and Slack), Service (Gmail, Slack, Notion, Figma, Salesforce, PitchBook).
  • Personal plugins: connect a custom MCP server, optionally add skills/custom UI; submit to OpenAI for directory approval. Directory holds 1,000+ plugins.
  • Discovery is the weak link: Work routes to installed plugins seamlessly but never suggests a missing one. Searching flights/hotels, it ignored available-but-uninstalled travel plugins for web search; even "naming Expedia outright" didn't prompt it.
Open tensions before the Chat merge

Whether the cloud computer becomes the user's primary AI computer (and how local sync feels seamless); whether Work agents gain OpenClaw-like sovereignty or ChatGPT keeps its opinionated continuity role; how to teach a billion Chat users what Work is for.

Full text · 18,824 chars
Unpacking ChatGPT Work: the Agent for a Billion Users An external reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools work in the new ChatGPT Work. Editor’s note: I’m excited to welcome Shlok to our guest post roster! You may know Shlok from his excellent explorations (as an outsider — for an insider perspective see our podcast with OpenAI’s Akshay Nathan. Already one of our most popular episodes of the year!) of leading AI Lab memory systems, which he gave an excellent AIE talk on. We’ve been covering OpenAI’s research and deployment of agents to all of humanity since Plugins 2023 and Devday 2024 and Codex 2025, and now ChatGPT Work in 2026 seems the penultimate stage of the long journey. Let’s dive in! On July 9th, OpenAI released ChatGPT Work, their agent product for knowledge work. It was, by any measure, a busy launch: three new models across fourteen configurations, a consolidation of the ChatGPT and Codex desktop apps, and cloud agents brought to the mainstream in their most accessible form yet. Three weeks in, Work (along with Codex) has reportedly crossed 10 million users. Editor’s note: ChatGPT estimated to cross 1B MAU in June and 1B WAU this month. Chat and Work currently sit side by side as separate modes inside ChatGPT, but Greg Brockman has confirmed that they will merge by the end of the year. Work, then, is not just a niche product for power users, but a preview of how ChatGPT’s billion weekly users will soon use the app. That’s why people inside and outside OpenAI are so excited about it, and why it deserves a closer look. Work in its current form takes some decoding. It’s an amalgamation of ChatGPT (in chat form), Codex the app, Codex the harness, Codex the original cloud agent, ChatGPT agent, Atlas, OpenClaw, and more. The product lineup around it is confusing. And the web and mobile versions diverge from the desktop one (unless you run it in cloud mode?!). So I spent the past few days trying to unpack it: what Work is, where it fits in OpenAI’s lineup, the many interesting choices in its design, the tensions underneath, and where I think it’s headed. Most of what follows comes from Codex and me poking around inside Work, and I’ve linked those conversations throughout so you can see where each claim comes from. What is Work? At its core: - An agent for knowledge work. You connect it to the places you already work—Slack, email, Drive, calendars, CRMs, project trackers, and hundreds of other plugins—and it gathers context across all of them to produce finished work. - Runs on the Codex harness. So it inherits the same models, sub-agents, browser use, and the ability to grind on a task for hours. Its UI is stripped of the evidence (git controls, diff-traces) that would give away you’re talking to a coding agent. - Lives in a cloud computer. Specifically, a beefy, isolated microVM: Pro accounts get 8 CPUs, 20GB of RAM, and a 64GB disk; Plus gets 14GB of RAM. Alongside the VM, Work gets a managed Chrome service that the agent operates through tool calls. - Produces artifacts. Sheets, docs, and slides rendered in interactive viewers, plus Sites: hosted web apps and dashboards it can build, share via URL, and keep updated. Every new conversation in Work is called a task. On web and mobile, Work runs in the cloud. You can kick off a task on web, track progress and give directions in the ChatGPT app on your phone, then view the result (maybe a report or a spreadsheet) back on your laptop. Work on the desktop app is slightly different and comes in two modes: cloud and local. In cloud mode, tasks run on the same cloud computer as web and mobile and sync across all three. In local mode, the agent works directly on your machine, across your files and apps, with full computer use. These tasks don’t appear on web or mobile, and there’s no way yet to move a local task to the cloud. This makes local mode essentially Codex, minus the code-related UI traces that would scare off a non-developer. On desktop, each new Work task can run locally on your computer or in the cloud. But then things get a little confusing. OpenAI did release a way to hand off a Codex task to a remote environment. Although this doesn’t work for me at the time of writing, I assume it eventually will, and that they will then bring the same functionality to Work. For the rest of this piece, Work = Work in cloud mode. Persistence & Memory One big reason OpenClaw felt different from a chatbot was that the agent had a computer of its own. You could run it on an always-on laptop or a VPS, let it create directories, install software, and maintain databases, and reuse all of this across conversations and subagents. Its state lived not just in chat history, Markdown files, or a dedicated memory system, but across the whole computer. Work’s cloud computer is persistent too. But rather than running in one VM that stays on forever, its workspace is synchronised to persistent storage and restored onto isolated microVMs as needed. So the underlying machine can change, but the working state carries over. Compared to OpenClaw, though, the agent has far less sovereignty over this computer. Every Work task (thread) gets a working directory under /workspace/scratch, where the agent has the freedom of a normal computer: it can make folders, install dependencies, write scripts, keep databases, and search everything with ordinary Linux commands. When I ask it to make a presentation for Acme, it can create clients/acme, copy in the source material, perform some analysis through code, and create charts and slides, all as files in the directory. When I follow up in the same thread, it returns to that working state and can continue editing it. But when a task needs context from other threads, it does not treat their working directories as a shared workspace that it can navigate freely. It relies instead on the ChatGPT product layer. By default, each new thread receives a compressed summary of recent tasks and files worked on . A summary might look like this: 20260731T15:55 Prepare Acme pilot plan:|||| Turn the attached notes into a one-page plan for the Acme pilot, with an objective, deadline, and next steps. <<File name=”acme_notes.txt”>> Raw conversation transcripts are not stored on the computer for the agent to browse. When a task needs context from previous threads, the agent calls Personal Context, a dedicated tool that queries Chat and Work history through a separately managed service and returns the relevant excerpts. Files follow the same pattern. ChatGPT’s Library is the central user-facing repository for all files and artifacts. User uploads land there automatically; agent-created files are saved when the user asks, or when the agent judges them worth retaining. The agent can also create directories in the Library to keep it organised. Like conversations, the Library doesn’t live on the computer, and can only be reached through dedicated tools. An uploaded file thus exists in two places: a working copy inside the thread and a canonical item in the Library. Interestingly, the two do not synchronise. If Thread A uploads a file and Thread B later changes the Library version, Thread A continues to read its now-stale local copy when resumed. When instructed explicitly, an agent in one task can navigate the scratch directories of other tasks, find files, and modify them. But it won’t do this on its own, and the directories have opaque names, no legible map to their conversations, and no stated retention contract. Memory is managed externally too. As I’ve written before, ChatGPT’s core memory primitive is a running, synthesised profile of the user. The product maintains that asynchronously and supplies it to Work when a task begins. The agent can reason from it, but can’t modify it or create OpenClaw-style Markdown files that other tasks load by default. ChatGPT’s Projects carry over into Work. Projects group related conversations, standing instructions, and Sources (user-uploaded files). A new task within a Project receives its instructions, summaries of relevant conversations, and local copies of Sources in its directory. But the Project itself does not exist on the computer as a directory, as it does in Codex. It too is an abstraction the product maintains. In short, the agent has broad freedom within a task, but continuity across tasks runs through an opinionated ChatGPT product layer rather than the computer itself. Why the split? My guess is several reasons: - Work builds on existing ChatGPT primitives (Conversations, Library, Personal Context, Memory). Ripping all of that out and rebuilding it inside the computer would mean refactoring a stack that already serves a billion users. - The separation is a guardrail. OpenClaw-style unrestricted access to a single environment holding every file, conversation, and memory is unsafe for users. - It lets OpenAI keep control of the product: what users see in the UI, how context is managed, and how sharing, cross-device sync, and file versioning work. All of that is harder to build if the agent could alter the environment at will. What Work lacks today is a meta-layer agent, one that operates a level above individual tasks and projects and coordinates between them. (Some already use Codex this way.) Perhaps that is coming, along with much else. Work is still young, and the architecture could look very different a few weeks from now. Hints of useful proactivity Today’s AI products are still reactive. Before the model can help, you have to notice that something needs doing, gather the relevant context, and translate it all into a prompt. The agent can do a stellar job from there, but the initial act of agency is still yours. Proactivity, where agents figure out how to be useful on their own, is one of the holy grails of personal AI. Work offers an early glimpse of that. When you open a new Work conversation, alongside the composer, you get personalized tasks generated from your own context. One suggestion offered to prepare me for an upcoming call. When I selected it, Work injected a pre-authored prompt. It had reasoned asynchronously across my context: noticed the calendar event, inferred that preparation would help, pulled data from Calendar and Gmail, and framed a task around the interests and preferences in my memory. When I sent the prompt, it got to work, and the result was a great meeting brief — one I didn’t know I needed! Today, Work takes a credible first step: it suggests tasks. But nothing happens until I execute them. For true proactivity, it would have to complete the tasks it predicts I’d want done, without me in the loop. That future doesn’t seem far off. Scheduled Tasks Automations let Work run tasks at a future time or on a recurring schedule, without the user manually prompting it. They are ChatGPT’s abstraction for reminders and cron jobs. OpenAI introduced them as Scheduled Tasks in January 2025. Work builds on the same scheduler but makes it agentic: each run can use the agent’s context and tools to complete the task. They come in two types. A standalone scheduled task begins each run from a saved prompt and opens a fresh task for the result. It suits self-contained work: a one-off reminder, a daily briefing, a weekly job search, a routine email scan. A scheduled task inside an existing conversation, triggered by a “heartbeat”, reawakens that task with its context intact. It suits use cases like monitoring a long-running operation, polling a connected service, or resuming a review loop at short intervals. At the time of writing, heartbeats work in the desktop app but are not exposed in Work on the web. Either automation can be set up as one-time or recurring. Its trigger can be an exact time, a loose window such as “in the morning”, or a condition the agent monitors. You can manage automations in two places. Inside a conversation, you can ask Work to create one, inspect existing automations, change their instructions or cadence, or pause and resume them. The Scheduled page puts all of this in a UI: every task with its next run and recent results, plus controls to create, edit, pause, or delete them. The Scheduled page adds another element of proactivity: ChatGPT suggests custom automations for you. Some, like a Daily Brief, are generic; others, like a weekly recap for the football club I support, are personalized from my memory. Browser Use For years, ChatGPT had limited access to the web. It could search, retrieve pages, and use commands like curl to download files or call APIs. But it couldn’t click through an interface, stay logged into a service, or complete workflows like filling a form. ChatGPT first gained this ability with Operator and ChatGPT agent. It then became a core part of Codex and now finds its most integrated expression in Work. Unlike Codex running locally, the Work browser doesn’t live on the same computer as the agent. Instead, the agent controls a separately hosted Chrome service through tool calls. It can inspect the page, click, type, scroll, take screenshots, manage tabs and dialogs, and move files between the browser and its computer. On web and desktop, Work shows a replayable timeline of the browser’s past states, so you can retrace what the agent did. You can also take over the live browser to navigate or enter a password, then hand it back to the agent. You can’t do this on mobile yet. The browser service also keeps its own persistent profile. New browser instances inherit preferences and logged-in sessions: I switched Wikipedia to dark mode and signed into Google in one task, and a fresh task inherited both. The Work agent never sees this profile or its credentials. Instead, a small permission ledger is synchronised into its computer alongside the workspace, recording, globally and per conversation, which sites it may act on and whether it may move files to or from them. But because the cloud browser runs in a datacenter, and not on your laptop, it faces constraints a local browser does not. Amazon US rejected it as an unsupported “session or client”, and Google Photos repeatedly timed out when I asked it to copy a shared album. Both tasks worked in local mode. Work can attempt a CAPTCHA, but only with your permission, and it is instructed not to loop, rotate its fingerprint, or otherwise evade a site’s safeguards. Still, the cloud browser makes Work far more capable. It can finish whole classes of tasks that ChatGPT with web search alone never could. Plugins, skills, and tools OpenAI has spent years searching for the right primitive to connect ChatGPT to outside apps and services: Plugins (March 2023), GPTs and Actions (November 2023), connectors (June 2025), and apps, the Apps SDK, and the App Directory (late 2025). In March 2026, plugins returned to Codex as packages of apps and skills. With the July 9 launch, the App Directory became the Plugin Directory, existing apps were packaged into plugins, and the directory expanded across Work and Codex. For now, OpenAI seems to have settled on plugins as the way for Chat and Work to interact with the external world. A plugin today can contain: - Apps, which connect the agent to services such as Gmail, Slack, or Salesforce. Most use an MCP server to expose tools: discrete operations the agent can invoke, such as searching messages or sending an email. - Skills, which combine instructions with supporting material—references, templates, and sometimes scripts—to teach the agent a workflow. - App templates, which let an organisation configure the private or organisation-specific app a workflow depends on. Plugins come in three broad types: - Operational plugins give the agent Codex-native enhancements. Computer Use lets it operate interfaces; Sites lets it deploy websites; Documents, Presentations, and Spreadsheets let it create interactive artifacts. - Role-specific plugins equip the agent for a particular kind of work. The Sales plugin, for example, teaches it to apply 20 skills (Analyze Account Signals, Build Business Case) across 29 apps, including Salesforce and Slack. - Service plugins connect the agent to external products such as Gmail, Slack, Notion, Figma, Salesforce, and PitchBook. Users can also create personal plugins by connecting a custom MCP server and, if needed, adding skills or custom UI. Developers who want to distribute a plugin more widely can submit it to OpenAI; once approved, it is published to the Plugin Directory. The Plugin Directory already holds more than 1,000 plugins covering most major apps and services, but discovery is a weak link. Work routes tasks to installed plugins seamlessly, yet never suggests a relevant plugin when one is missing. When I asked it to search for flights and hotels, it ignored several available but uninstalled travel plugins in favour of web search, even though a plugin might have used fewer tokens, returned better results, and let me complete a booking directly. Even naming Expedia outright didn’t prompt it to offer the plugin. I can imagine the product challenges: how does ChatGPT know when to handle a task itself, when to recommend a plugin, and which path serves the user better? And if several can do the job, which should it suggest? Without a solid discovery layer, though, OpenAI is leaving value on the table—for users, for developers, and for itself in its quest to become a platform. What’s Next When Work folds into Chat later this year, its design choices will become the default for a billion people. Before then, OpenAI has to resolve a few tensions that kept surfacing as I used it: - Does the cloud computer become the user’s primary AI computer? And how can syncing between it and the local machine feel seamless? - Do Work agents get more OpenClaw-like sovereignty over that computer? Does ChatGPT keep the opinionated role it plays in continuity, or is there a middle ground? - How does Work come to feel as familiar to users as Chat? And in the meantime, how does OpenAI teach Chat users, from within the product and outside it, what Work is for and how to get the most out of it? None of this should detract from the fact that Work is an impressive, ambitious, yet underrated launch. It consolidates years of scattered products and experiments into one increasingly cohesive whole. And it’s close to the ChatGPT OpenAI would build if it were starting from scratch with today’s agents. I’m excited to see where it heads next. I spend most of my time thinking about personal AI: going down rabbit holes like this one, figuring out what the best products are getting right, and imagining what our AI sidekicks will look like a year and five years from now. If you made it this far, we probably think about the same things, and I’d love to hear from you. Find me on X or through my website.
00:06

Karpathy’s New AI Trick: Ramble Your Idea, Then Build the World

Karpathy gave Claude Opus 5 the opening paragraph of The Lord of the Rings and a one-million-token budget, and the model turned it into a playable procedural 3D world in about two hours and 5,500 lines of code for roughly $10. His follow-up guide explains how to turn a ten-minute voice ramble into a clear world brief, then use Claude Code with Three.js, scene graphs, and testing loops to build and improve an interactive scene yourself.

Notes

Karpathy's New AI Trick: Ramble Your Idea, Then Build the World

Emerging AI (Substack), 2026-08-04

The demonstration: Andrej Karpathy gave Claude Opus 5 the opening paragraph of The Lord of the Rings, a one-million-token budget, and one task — turn it into a procedural Three.js experience.

What the model produced:

  • ~2 hours of work; roughly 5,500 lines of code
  • Low-poly objects placed across a 3D valley; animated characters and cameras
  • Something openable in a browser; ~$10 total run cost
  • Result was "rough," but a paragraph had become an editable world made of code

Core claim (why worlds, not images):

"A generated image is usually the end of the work. A procedural world is the beginning."

The world stays editable after generation: move a building, change the camera, add an interaction, rewrite the timeline, let a player enter the scene.

The workflow the full guide promises (Karpathy's second post):

  • Turn a ten-minute voice ramble into a clear world brief — avoiding an hour of prompt-polishing
  • Build a small interactive scene with Three.js and Claude Code
  • Structure via scene graphs, files, prompts, loops, and reusable skills
  • Make the agent test its own work through screenshots and browser checks

Also covered: model choice, token economy, visual review, draw-call performance, InstancedMesh, hybrid AI filmmaking, and a weekend project from single idea to working 3D world.

Limitations stated: output was rough (not polished); the cost/token figures assume a large budget (1M tokens); procedural quality is presented as starting point, not final asset.

Full text · 1,591 chars
Karpathy’s New AI Trick: Ramble Your Idea, Then Build the World A practical guide to voice-rambling an idea, shaping it into a scene graph, and letting an AI coding agent build and improve an interactive Three.js world. Andrej Karpathy gave Claude Opus 5 the opening paragraph of The Lord of the Rings, a one-million-token budget, and one job: turn it into a procedural Three.js experience. The model worked for about two hours. It wrote roughly 5,500 lines of code, placed low-poly objects across a 3D valley, animated characters and cameras, and produced something you could open in a browser. Karpathy reported the run cost around $10. The result was rough, but a paragraph had become an editable world made from code. Open the playable demo and the bigger idea becomes clear. A generated image is usually the end of the work. A procedural world is the beginning. You can move a building, change the camera, add an interaction, rewrite the timeline, or let a player enter the scene. Karpathy’s second post explains how to begin without spending an hour polishing a prompt. Inside the full guide, I break down how to turn a ten-minute voice ramble into a clear world brief, build a small interactive scene with Three.js and Claude Code, structure it with scene graphs, files, prompts, loops, and reusable skills, then make the agent test its own work through screenshots and browser checks. It also covers model choice, token economy, visual review, draw-call performance, InstancedMesh, hybrid AI filmmaking, and a simple weekend project you can build from one idea to a working 3D world.
12:57

The One File That Made Hermes Finally Learn From Its Mistakes

A small markdown file placed beside each agent workflow, named MISTAKE_LEDGER.md, records verified lessons from failed runs so the agent reads and applies them next time instead of losing them in old chats. Each entry captures seven fields, including what went wrong, the confirmed cause, the verified fix, and when to recheck or retire it. The author adds a strict rule against letting agents write their own lessons: finish the task, prove the fix with an observable result like a passing test or live API response, then record it. The article includes a paste-in instruction for Hermes or Codex and warns that a ledger full of unverified agent guesses does more harm than no ledger.

Notes

Hermes Mistake Ledger (MISTAKE_LEDGER.md)

Source: All Agents Considered (Substack), published 2026-08-04. Author's personal file-based workflow method.

Problem: memory ≠ learning
  • Hermes Agent v2026.6.5 built-in memory is described as a "bounded, curated set of user preferences, project details, environmental context, and learned information," injected "as a bounded snapshot at the start of a session rather than retrieved as a complete archive of every old chat."
  • Operational lessons ("what failed, what fixed it, when that fix applies again") get compressed into one instruction line with no history — e.g. the author's research filter became just "reject vendor press releases," losing the bad output, the failed first wording, and the proof the fix worked.
Origin story (from prior failures)
  • Run 1: ranked a generic news item too highly — definition of practical tech news too loose. Tightened instruction.
  • Run 2: vendor press release passed because the filter never excluded vendor PR. Added the rule, third run clean.
  • 2026-03-21: Ask HN post called this missing layer an "Agent Experience Cache" (tool quirks, repeatable workflow patterns, environment knowledge, failure modes too costly to rediscover). Author chose a smaller read-able version: one markdown file per workflow instead of a memory service.
Design decisions
  • Not global: one file per workflow folder, not one file for everything — a deployment lesson shouldn't make Hermes cautious while sorting research. Hermes and Codex share files via Obsidian vault.
  • Folder layout (from "Tear Down Your AI Workflow and Rebuild It Like This"):

```

01.Research Sorter/

├── 01.instructions.md

├── 02.input.md

├── 03.output.md

├── 04.review.md

└── MISTAKE_LEDGER.md

```

  • Ledger is visible/readable so entries can be challenged, edited, or retired; repeated lessons can be promoted into a wider agent instruction.
Entry format — seven fields

Date · Task · What went wrong · Confirmed cause · Verified fix · When this applies · When to recheck or retire it.

Reconstructed example (vendor PR entry):

Task: Sort research items into All Agents Considered article observations. Confirmed cause: The filter excluded broad news but did not explicitly exclude vendor PR as a daily input. Verified fix: Add vendor PR to the exclusions and require an independent technical source before keeping a vendor announcement.
Four-step rule (keeps guesses out)
  • Catch — workflow fails, user corrects, or AI spots a bad assumption
  • Fix — finish the task before writing the lesson
  • Verify — prove the fix via a result outside the chat (passing tests, live API responses, correct file placement, outputs meeting standards; "AI confidence alone proves nothing")
  • Record — one specific entry

Why verify: Hermes once guessed a promo-vendor article's real problem was missing a "deep technical quote" (format) when the actual signal was the vendor-owned domain; recording that guess would have rejected legitimate news briefs. In the Substack SDK build, Hermes proposed scheduled_at (plausible, common convention) but the real endpoint expected trigger_at — an automated test now checks the request body.

The skill text (paste-ready, instruction-only, no script/service/DB)

--- front matter: name: mistake-ledger. Before a repeated workflow run, check the folder's MISTAKE_LEDGER.md, read only that workflow's ledger, state which lesson is being applied. On failure/correction: fix → verify → record only if concrete, reusable, scoped → append one entry with the seven fields. Do not record speculation, secrets, private content, or one-off preferences; mark outdated lessons as retired instead of deleting.

  • Hermes: keep at ~/.hermes/skills/, or add a shared folder to skills.external_dirs in ~/.hermes/config.yaml; if unconfigured, ask Hermes to read the file before the run.
  • Codex: save at .agents/skills/mistake-ledger/SKILL.md in the repo (OpenAI Build skills format).
Verification test for next run

Plant a vendor announcement in 02.input.md alongside independent technical stories, then check: (1) did Hermes read the ledger before scoring, (2) did it state the vendor-PR lesson applied, (3) did it exclude the announcement without an independent source. Author notes the method "still needs several real repetitions before I make stronger claims," expects trigger wording/retirement rules/promotion thresholds to change, and flags the cost risk of the ledger becoming "a compliance manual."

Five failure modes (when the ledger lies)
  • False causality — symptom recorded as cause; wrong fix applied every run
  • Bad scope — a workflow-local lesson promoted into a universal rule that interferes elsewhere
  • Stale advice — APIs/tools/folders change; a correct fix becomes a bug
  • Noise — every typo/failed search/abandoned idea enters; agent sorts clutter before working
  • Exposure — credentials, customer data, session cookies, or sensitive output persist in a file the agent keeps reading
Caveat: structural vs. operational failures

Repeated failures that reveal broken structure (agent keeps picking an archived file; outputs land in random locations) should not become lessons — fix the folder map instead. Author's dividing line: "If Hermes cannot find the right work, fix the map. If it finds the right files and still repeats the same operational error, record the lesson." Recommends the "Twenty Minute Audit" before recording confusion as a permanent rule.

Claim (attributed, closing): "The ledger does not turn Hermes into an agent that trains itself. It gives the next run a short piece of reviewed evidence from the last one."

Full text · 16,933 chars
The One File That Made Hermes Finally Learn From Its Mistakes Persistent memory keeps useful context. A small ledger beside each workflow keeps the verified lesson where the next run can find it. Some time ago, I covered my Hermes research workflow. It runs every week, sorts through tens of tech news publications and forums, and filters out the fluff so I can stay on top of everything that is going on. It took me a few tries to get it right, but now I can safely rely on it. That trust came from two early mistakes. On the first real run, Hermes ranked a generic news item too highly because my definition of practical tech news was too loose. I tightened the instruction and ran it again. On the second run, a vendor press release passed the filter because I never told Hermes to exclude vendor press releases. I added that rule, ran it a third time, and got a clean result. The workflow got better, but the reason why it got better simply vanished. If you’d open my final instruction file now, you’ll see one clean rule: reject vendor press releases. You do not see the bad output that created the rule, why the first wording failed, or the result that proved the fix worked. Hermes still remembers my projects, files, and old conversations. It also squeezed the useful lessons from those failed runs into one line with no history. And that bothers me because the next mistake it makes might be less obvious. A tool may fail only on one server. An API may accept a believable field name and ignore it. A workflow may write to the wrong folder and still report success. So I figured we can’t let the AI bury these critical lessons in old chat logs. That’s why I started adding one file to each repeated workflow: MISTAKE_LEDGER.md This article explains what belongs in it, what does not, and the small instruction I use to stop an agent from turning every failed attempt into permanent bad advice. In this edition - Why persistent memory still loses operational lessons - Where a Mistake Ledger fits inside a file-based workflow - The four-step rule that keeps guesses out of the ledger - A complete instruction you can paste into Hermes or Codex - The failure modes that make a mistake ledger worse than no ledger Memory Keeps Context I have argued for months that memory changes what an AI agent can do. An agent that carries project details and preferences across sessions beats a blank chat window every time. I wrote about this shift in Forgetting to Forget when persistent memory made long-running agent work feel possible. I still believe that. The mistake was treating memory and learning as the same thing. Hermes Agent v2026.6.5 describes its built-in memory as a bounded, curated set of user preferences, project details, environmental context, and learned information. Session logs preserve more of what the system did. An operational lesson serves a different purpose. It tells you what failed, what fixed it, and when that fix applies again. Those are three different layers. The exact correction may still exist in an old conversation, but that does not mean it will enter the context of a new run six weeks later. Hermes built-in memory is injected as a bounded snapshot at the start of a session rather than retrieved as a complete archive of every old chat. The useful lesson can remain buried in conversation history even while the agent remembers the project. On March 21, 2026, an Ask HN post about operational memory called this missing layer an “Agent Experience Cache.” The post listed tool quirks, repeatable workflow patterns, environment-specific knowledge, and failure modes that cost too much time to rediscover. That idea made sense to me, but I wanted a smaller version that I could read without adding another memory service. One Markdown file beside the workflow felt like the right size. Keep Mistakes Local My first instinct was to create one global file containing every mistake made by Hermes or Codex. That file would become useless fast. My content research workflow, Substack analytics tools, server deployments, and draft-writing system fail for different reasons. A rule learned while deploying software should not make Hermes cautious while sorting research. Likewise, a correction about brand sources should not influence how Codex handles a server path. Both Hermes and Codex share the same files and workflows through my Obsidian vault. The mistakes should live where the work happens. I already use an INDEX.md file to map large workflows. I explained that setup in Why My Best Agent Workflow Is Mostly Files. Inside the project, each repeated workflow gets its own small folder. My research sorter came from the four-file pattern in Tear Down Your AI Workflow and Rebuild It Like This: 01.Research Sorter/ ├── 01.instructions.md ├── 02.input.md ├── 03.output.md ├── 04.review.md └── MISTAKE_LEDGER.md The first four files tell Hermes how to run the work. The fifth preserves the verified lessons produced by running it. If a project contains one workflow, project-specific and workflow-specific mean the same thing. Once a project contains several repeated processes, separate ledgers keep the lessons scoped. Because the ledger is visible and readable, I can challenge an entry, edit its scope, or retire the advice after a tool changes. I do not have to rely on a memory entry or session log that sits apart from the workflow it affects. If the same lesson appears across several workflow ledgers, I can promote it into a wider agent instruction after reviewing it. What an Entry Holds The file needs enough detail to stop the same error without becoming a diary of everything that went wrong. Each entry records seven fields: - Date - Task - What went wrong - Confirmed cause - Verified fix - When this applies - When to recheck or retire it If the ledger had existed when I built my research workflow, the vendor PR entry would have looked like this: ## YYYY-MM-DD: Vendor PR passed the research filter **Task:** Sort research items into All Agents Considered article observations. **What went wrong:** A vendor announcement was kept as a strong article signal even though no independent technical source discussed it. **Confirmed cause:** The filter excluded broad news but did not explicitly exclude vendor PR as a daily input. **Verified fix:** Add vendor PR to the exclusions and require an independent technical source before keeping a vendor announcement. **When this applies:** Every All Agents Considered research-sorting run. **Recheck or retire:** Recheck if the source policy changes. This is a reconstruction from the failure I documented in the earlier article. I did not have a ledger then, so I am taking a pattern already hidden inside my workflows and making it visible. Verify Before Recording AI agents write terrible rules when they update their own instructions after every failure. One plausible guess can harden into a permanent constraint before anyone checks whether it explains what went wrong. During one weekly research run, Hermes kept a promotional article from a vendor’s own website and treated it as independent reporting. When I told it to remove the article, Hermes guessed that the missing ingredient was a deep technical quote. It proposed a permanent rule requiring one in every future submission. The diagnosis sounded plausible, but it focused on the article’s format when the source was the real warning sign. The piece appeared on the vendor’s corporate domain and repeated the company’s marketing claims without support from independent reporting. The useful lesson was to treat articles on vendor-owned domains as promotional unless an independent source supported them. Recording Hermes’s first guess would have taught the workflow to reject legitimate news briefs because the reporter did not include a direct quote. That is why I have to verify the root cause with my own eyes before the agent turns a plausible explanation into a permanent lesson. Building my unofficial Substack SDK exposed the same problem in code. Hermes proposed scheduled_at as the field for scheduling a post. The choice looked logical because plenty of other APIs use that naming convention. The Substack scheduling endpoint I tested expected trigger_at instead. An automated test now checks the request body and fails if the implementation sends a different field. That gives me proof outside the conversation instead of another confident answer from an agent. I covered the full story in How I Built a Substack API With Hermes and Codex. Without that hard check, neither field name would deserve a permanent spot in a Mistake Ledger. Based on that experience, I created a strict four-step rule that keeps guesses out of my final documentation: - Catch: Workflows fail, I correct the AI, or the AI spots a bad assumption. - Fix: I finish the task before writing down a lesson. - Verify: I prove the fix works using a result outside our chat. - Record: I add one specific entry to the workflow ledger. Real proof makes or breaks the verification step. It can come from passing code tests, live API responses, correct files landing in the right folders, or final outputs meeting exact standards. AI confidence alone proves nothing. The Tiny Instruction The working version fits in one instruction-only skill. There is no script, service, database, or download. You can paste the text into a Hermes or Codex task. If you want it available across repeated sessions, save the same block as SKILL.md. --- name: mistake-ledger description: Record a verified reusable lesson after a workflow failure or user correction, and check the current workflow's MISTAKE_LEDGER.md before running related work. --- # Mistake Ledger Before running a repeated workflow, check whether its folder contains MISTAKE_LEDGER.md. Read only that workflow's ledger and apply relevant active lessons. State which lesson you are applying. When a task fails, the user corrects you, or you discover a wrong assumption: 1. Fix the problem first. 2. Verify the replacement through an observable result. 3. Record a lesson only when it is concrete, reusable, and scoped to this workflow. 4. Append one entry to MISTAKE_LEDGER.md with: - date - task - what went wrong - confirmed cause - verified fix - when this applies - when to recheck or retire it Do not record speculation, secrets, private content, or one-off user preferences. Do not turn an unverified diagnosis into a permanent rule. Mark outdated lessons as retired instead of deleting them. Hermes users can keep the repeated version under ~/.hermes/skills/, or add a shared folder to skills.external_dirs in ~/.hermes/config.yaml. The Hermes Agent v2026.6.5 skill documentation describes both routes. If you keep the instruction inside the workflow folder without configuring that folder as a skill directory, ask Hermes to read the file before the run. Codex users can save the same instruction at .agents/skills/mistake-ledger/SKILL.md inside a repository. The current OpenAI Build skills documentation describes the required SKILL.md format. Test the Next Run Creating the file proves nothing by itself. The next run must show that the recorded lesson changed the workflow’s behavior. My next research test is straightforward. I will place a vendor announcement in 02.input.md beside several independent technical stories. I will then run the workflow and check three things: - Did Hermes read MISTAKE_LEDGER.md before scoring the items? - Did it state that the vendor PR lesson applied? - Did it exclude the announcement unless an independent technical source supported it? The same test works for other workflows. First, reproduce the conditions that caused the original failure. Then check whether the agent applies the relevant lesson without being reminded. I also need to measure the cost of using the ledger. If Hermes starts citing irrelevant history on every run, the ledger has created a new problem. It should reduce repeated work without becoming a compliance manual that the agent must review before every task. This method still needs several real repetitions before I make stronger claims about it. I expect the trigger wording, retirement rules, and promotion threshold to change as I use it across more workflows. That uncertainty belongs in the build log. A clean ending would give the method more confidence than it has earned. When the Ledger Lies A Mistake Ledger creates five failure modes of its own. Any one of them can make the workflow worse: - False causality: The agent records a symptom as the cause. It then applies the wrong fix every time the workflow runs. - Bad scope: The agent promotes a useful lesson from one workflow into a universal rule. That rule starts interfering with unrelated work. - Stale advice: APIs change, tools get updated, and folder structures move. A correct fix can eventually become a new bug. - Noise: Every typo, failed search, and abandoned idea enters the ledger. Hermes then has to sort through a second pile of clutter before starting the real task. - Exposure: A careless entry preserves credentials, private customer information, session cookies, or sensitive tool output. None of that information belongs in a file the agent keeps reading. Some repeated failures reveal a broken workflow structure rather than a reusable lesson. If Hermes keeps selecting an archived file because current and old drafts live together, the folder structure needs fixing. If every output lands in a new location, the workflow needs a defined destination. Recording “pick the current file next time” only hides those structural problems from the next run. Before turning that kind of confusion into a permanent lesson, run the audit from The Twenty Minute Audit That Found Where Hermes Was Getting Lost. The audit will show whether the agent’s map of the work is the real problem. My dividing line is simple: If Hermes cannot find the right work, fix the map. If it finds the right files and still repeats the same operational error, record the lesson. Try It in 60 Seconds Open one workflow that you run more than once. Think of the last correction you gave the agent. Then confirm that the replacement worked outside the conversation where you suggested it. If you still do not know what caused the failure, stop there. An unverified correction has not earned a permanent entry. Once the fix is verified, create MISTAKE_LEDGER.md inside the workflow folder. Record the seven fields while the evidence is still easy to inspect. On the next run, ask the agent to read the ledger first and state which lesson applies. Do not remind it of the original error or steer it toward the answer. That run is the test. You do not need a new model, an extra memory service, or a global database containing everything your agent has ever done. One workflow and one verified mistake are enough to test whether the lesson survives the conversation where it was learned. Give Failure a Job The clean version of my research workflow hides the two failed runs that shaped it. The final instructions preserve the rules without showing which mistakes forced me to add them. That makes the instructions easier to read. It also removes the operating history I will need when a rule starts causing trouble. Six months from now, I want to know why a rule exists before I delete it. Hermes should understand when the rule applies without searching old conversations. Both of us should see when the advice has expired. These tools support different parts of the same system. Project maps and workflow folders make the work findable and repeatable. Audits expose structural defects before they become permanent instructions. The Mistake Ledger preserves verified failures after the immediate problem has been fixed. The ledger does not turn Hermes into an agent that trains itself. It gives the next run a short piece of reviewed evidence from the last one. That is enough for version one. What mistake does your agent keep repeating? Tell me whether it belongs in a workflow ledger or whether the workflow itself needs fixing. Everything I described here is the field-notes version. I built the Mistake Ledger because my research workflow kept surfacing the same vendor PR problem until I wrote the lesson down and watched the next run actually use it. That small experiment, one folder, one verified entry, one retest, convinced me the pattern was worth teaching properly. The step-by-step version of how I apply this across every workflow I run goes into the first Hermes 101 course. I am building it right now and it should be ready soon. If the broader stack is what you are after, I wrote about the full cost breakdown and the morning workflow that runs on it. Both depend on the same file-based workflow structure to stay reliable across sessions. Capability is cheap when the foundation is broken. Audit one folder. Fix one thing. Then decide whether you need another tool.
15:31

Buy Now or Wait: The Local AI Hardware Playbook

Buying local AI hardware now beats waiting, because every new model release makes existing hardware more valuable while memory prices and demand keep climbing. DeepSeek-V4-Flash, a 284-billion-parameter model that runs only part of itself per question, needs about 158GB for weights and runs at roughly 40 tokens per second on a two-node DGX Spark setup, at $0.14/$0.28 per million tokens. The piece argues cloud AI is structurally worse: your data trains their models, output changes day to day, and OpenAI loses $1.70 for every dollar it earns, which the author says means prices would need to roughly triple to be sustainable. Meanwhile Nvidia confirmed no new gaming GPUs in 2026 and the expected RTX 5090 Super with 48GB of VRAM was cancelled, so the hardware people are waiting for may not exist until around 2028.

Notes
Buy Now or Wait: The Local AI Hardware Playbook

Author: Manolo Remiddi (AI-assisted research/editing). Source: The Augmented Mind: Think with AI, Aug 4, 2026.

The buy-now argument
  • Author bought an ASUS GX10 (DGX Spark equivalent) ~6 months ago for ~$3,500; same hardware now ~$4,500. Argument: local AI hardware is a compounding, not depreciating asset — each new model release raises existing hardware's value at zero cost.
DeepSeek-V4-Flash (released April 2026)
  • 284B-parameter MoE, 13B active, 1M-token context, native FP4+FP8 checkpoint.
  • Footprint: ~158 GB weights native; 170–175 GB usable with context/runtime overhead. Requires 2× DGX Spark (or larger); "far beyond the reach of a single consumer GPU."
  • Throughput on dual DGX Spark: ~40 tok/s sustained single-stream at long context; mid-50s tok/s optimized; 30–40 tok/s conservative (depends on quantization, prompt length, decoding stack).
  • Claim: first frontier-class open model self-hostable with "full privacy, no subscription, and no external rate limits."
Cloud AI problems
  • Privacy: regulated industries legally can't send customer data off-premises; "you literally cannot use frontier cloud models from US providers" in cybersecurity/biological research.
  • Stability: same-named models change behavior/output quality/reasoning depth/output length; providers optimize for their business.
Memory supercycle
  • HBM/GDDR7 prices up >300% in past year.
  • US grid interconnection queues (N. Virginia, Phoenix, Dallas): 4–7 years; PJM warns of up to 60 GW supply shortfalls.
  • Samsung, SK Hynix, Micron all announced Q2 2026 price hikes; SK Hynix (NVIDIA's HBM supplier) has pricing leverage.
Global buildout
  • France: $830M debt financing for Mistral AI's first data center near Paris, 13,800 NVIDIA GB300 GPUs; plan 200 MW across Europe by 2027, 1.4 GW in France by 2030.
  • Chinese open models ~90% of performance at a fraction of cost; U.S. adoption because they're 60–90% cheaper than proprietary leaders.
Pricing
  • DeepSeek-V4-Flash API: $0.14/M input (cache miss), $0.28/M output vs GPT-5.6 at $1/M input (cheapest tier) to $5/M (flagship).
  • OpenAI economics (WSJ): spent $1.70 per $1 earned (2025); projects $17B cash spend, $14B loss in 2026; inference-only revenue $1.60/compute dollar, ~$0.68 with training; claims OpenAI would need to >3x prices to be profitable.
Hardware you're waiting for
  • NVIDIA confirmed no new GPUs at CES 2026 (first in 5 years); no gaming GPUs in 2026; RTX 60 "Rubin" delayed past 2027, likely 2028; RTX 5090 Super (48 GB, was expected) cancelled.
Costs / caveats
  • 2× DGX Spark ~$10,000; RTX 5090 ~$4,000 (secondary market); Mac Studio M5 Ultra $10–15K+.
  • Author currently runs hybrid, aiming for 100% local. Explicit disclaimers: not financial advice; prices/availability "change rapidly"; data as of August 2026; AI-assisted writing disclosed.
Full text · 11,791 chars
Buy Now or Wait: The Local AI Hardware Playbook The data behind buying local AI hardware now versus waiting for the next generation: memory prices, cancelled GPU releases, cloud instability, and why the “wait” strategy costs more than you think. The $3,500 Decision I bought my ASUS GX10 (Nvidia DGX Spark equivalent) six months ago for approximately $3,500. Today, that same hardware would cost you closer to $4,500. On paper, I came out ahead by $1,000. But the hardware price is only half the story. The models I can run on this box today are exponentially more powerful than the models available six months ago. The hardware has appreciated not just because I bought it at a lower price, but because its computational value has increased as better software arrived. That is the core argument for buying now: local AI hardware is not a depreciating asset. It is a compounding one. Every new model release increases what your existing hardware can do, without any additional cost. Take DeepSeek-V4-Flash, released in April 2026. It is a 284 billion parameter Mixture-of-Experts model with 13 billion active parameters, a one-million-token context window, and a native FP4+FP8 mixed-precision checkpoint. That combination makes it one of the most capable open models you can realistically self-host, but it is still far beyond the reach of a single consumer GPU in its original form. In practical terms, the model wants roughly 158 GB just for weights in the native checkpoint, and closer to 170–175 GB of usable memory once you include context and runtime overhead. That is why systems like 2× DGX Spark can run it comfortably from a capacity standpoint, while a single high-end workstation GPU cannot. If you quantize aggressively, you can push the memory footprint down much further, but at the cost of some quality and, depending on the stack, some speed. On a two-node DGX Spark setup, DeepSeek-V4-Flash is not just loadable, it is usable. Real-world reports on dual DGX Spark systems show around 40 tokens per second for sustained single-stream decoding at very long context, with optimized configurations reaching the mid-50 tok/s range and more conservative setups landing closer to 30–40 tok/s depending on quantization, prompt length, and decoding stack. This is what makes the model interesting: hardware that felt “too small” only months ago can now do useful work. The real takeaway is not that DeepSeek-V4-Flash is cheap to run, but that it marks a new threshold: a frontier-class model that can be self-hosted on serious local hardware, with full privacy, no subscription, and no external rate limits, provided you size the machine correctly for the precision and context length you actually want to use. Two Problems with Cloud AI That Nobody Talks About There are two structural problems with relying on cloud AI providers that most people do not factor into their decision. Neither of them will get better if you wait. Privacy. When you send your data to a cloud AI provider, you are sharing everything: your clients data, your intellectual property, the recipe for your business success. That data can be used to train their models. For businesses in regulated industries, this is not a philosophical concern. It is a legal requirement that customer data never leaves your premises. If you work in cybersecurity or biological research, you literally cannot use frontier cloud models from US providers. Period. Stability. Cloud AI labs are constantly tweaking and adjusting their models. Even when a model keeps the same name, it is not performing the same every single day. The provider can change the behavior, the output quality, the reasoning depth, or the output length at any time. When you have your own hardware with your own model, the system is stable and constant. It does not change unless you change it. Think about it in threshold terms. There is a minimum performance level that your AI needs to reach for your workflow to function. Once you cross that threshold, you are operational. You can serve your customers and run your business. The problem is that cloud providers adjust that threshold constantly. Sometimes it is about optimization: they want to spend less on electricity and compute, so they reduce the reasoning depth or the model intelligence. Sometimes it is a push to a more expensive tier. They are optimizing for their business, not yours. The Memory Supercycle The dominant narrative right now is that memory prices are going to come down. High-end HBM and GDDR7 and other memory saw prices increase by over 300 percent in the last year alone. The intuition says: data centers are maxing out, demand will slow, prices will correct. That logic works if you only look at the US market. It does not work for the global picture. Here is what is actually happening with supply. In the US, hyperscalers have been building data centers so aggressively that grid interconnection queues in Northern Virginia, Phoenix, and Dallas are now running four to seven years. PJM Interconnection, the largest US power grid operator, warned of potential supply shortfalls of up to 60 GW. New power stations, turbines, and grid expansions will take years to come online. But the rest of the world is just beginning to build. AI memory consumption is not going down. Every industry, every government, every nation is adopting AI at an accelerating pace. Three companies dominate global DRAM production: Samsung, SK Hynix, and Micron. All three announced significant price increases for Q2 2026, and all three are prioritizing HBM allocation because that is where the AI money is. SK Hynix is the standout HBM supplier to NVIDIA, giving it enormous leverage over pricing across its entire memory portfolio. The feeling that “the US is not going to buy anymore, so prices should come down” ignores the billions of people outside the US who are building sovereign AI infrastructure for the first time. The Global AI Buildout AI sovereignty is becoming a concrete reality, not a talking point. Nations need their own data centers because they cannot rely on a single country that controls access to critical technology. Europe. European governments are already moving. They are migrating to Linux, moving away from WhatsApp, Microsoft Office, and Zoom, and reducing dependency on US software. AI data center investment is the next logical step. France invested $830 million in debt financing for Mistral AI first data center near Paris, equipped with 13,800 NVIDIA GB300 GPUs. It is small compared to US spending, but it is the starter. The plan is 200 megawatts across Europe by 2027 and up to 1.4 gigawatts in France by 2030. China. China could not buy the latest NVIDIA hardware for AI, so it adapted. Instead of pushing for maximum intelligence, Chinese labs pushed for models optimized to run on their own hardware. The result: open-weight models that deliver 90 percent of the performance at a fraction of the cost. DeepSeek-V4-Flash is priced at $0.14 per million input tokens on cache-miss requests, with output at $0.28 per million tokens. That puts it far below frontier U.S. closed models such as GPT-5.6, which starts at $1 per million input tokens on its cheapest tier and rises to $5 on the flagship tier. More broadly, Chinese open models are increasingly being adopted by U.S. companies because they are often 60% to 90% cheaper than leading proprietary alternatives. The goal for Chinese labs is to create intelligence that is affordable for the rest of the world, because not everyone can afford to work with Anthropic. This creates massive downstream demand for AI hardware that is not going away. The OpenAI Math Problem Cloud AI pricing is not what it appears. OpenAI current prices are subsidised by investor capital. According to WSJ data from investor documents, OpenAI spent $1.70 for every dollar it earned in 2025. The company projects spending $17 billion in cash in 2026 alone, with a $14 billion projected loss that year. The inference-only revenue per compute dollar is $1.60 for OpenAI. When you add training costs back in, the figure drops to approximately $0.68 per dollar of silicon. OpenAI is burning cash on every inference call it serves. The prices you see on ChatGPT and the API are not sustainable. If OpenAI actually wanted to stay profitable at current usage levels, it would need to more than triple its prices. When that happens, local AI will stop being a choice and start being the only economically viable option for serious work. What You Are Waiting For Does Not Exist Let us be concrete about the hardware you might be waiting for. NVIDIA publicly confirmed it would not announce any new GPUs at CES 2026, a first in five years. Reports from The Information state that NVIDIA plans to launch no new gaming GPUs in 2026 and that the RTX 60 series “Rubin” is delayed beyond 2027, likely into 2028. NVIDIA is prioritizing AI accelerators over the gaming sector because that is where the money is, and because cutting-edge GDDR7 memory is needed for AI chips. The RTX 5090 Super was expected to offer 48 GB of VRAM instead of 32 GB, which would have been the real sweet spot for local AI. It will not happen. The RTX 6090, if it arrives in 2028, will be two years away. That is a lot of waiting. The Three Reasons to Go Local You should not buy expensive hardware to solve a problem you do not have. The decision to run AI locally is about solving specific problems: - Privacy. Your data stays on your hardware. Your clients data stays on your hardware. No third party trains on your information. - Stability. Your model does not change every day. Your performance threshold does not move. Your system is constant. - Experimentation. You can test, modify, fine-tune, and learn without hitting rate limits, subscription walls, or terms of service restrictions. If you are not in one of those categories, local AI may not be for you. The cost is real: two DGX Sparks are around $10,000. A single RTX 5090 is running $4,000 on the secondary market. A Mac Studio M5 Ultra will be $10,000 to $15,000 or more. But the cost of waiting is also real. Every month you wait is another model release you cannot run, another workflow you cannot automate, another business problem you cannot solve. The Option Is an Illusion I want to leave you with one thought. The choice between local AI and cloud AI feels like a genuine option right now. Soon, for most of us, it will not be. I currently work in hybrid mode. I have local AI solutions and I also use cloud AI. That is where we are today. But my direction is to go 100 percent local, and I believe that is where everyone who takes this seriously needs to head. We should never give our data to a corporation that optimizes for its own profitability, not ours. The cloud AI pricing model is mathematically unsustainable. The stability of cloud models is not something you can rely on for business-critical work. And the privacy risk is not negotiable if your work involves sensitive information. Local AI is getting better and better every month. And the hardware you buy today will keep getting more valuable tomorrow. Disclosure: this is not financial advice. These are observations from my own experience running local AI hardware and the data available as of August 2026. Hardware prices and availability change rapidly. Do your own research before making a purchase decision. If you are interested in running AI locally, join the Augmented Mind Discord community. There is a lot to discuss about hardware, models, and the different approaches to building a local AI setup that actually works. Transparency note: This article was written and reasoned by Manolo Remiddi. The Resonant Augmentor (AI) assisted with research, editing and clarity. The image was also AI-generated.
03:18

I Built My Digital Sigmund Freud. It Knows Me Better Than a $200/Hour Therapist

One person built an AI therapist called Mind Mirror with Claude that profiles him from three years of exported ChatGPT history, recalling details like his anxiety habits that he never mentioned in the current conversation. The app draws on 8,303 indexed memories spanning 32 months of messages and charts how his mindset changed over time. The post shares the export steps and the single prompt used to build the app, and cites a 2026 review suggesting roughly one in four AI users turn to AI for emotional advice.

Notes
I Built My Digital Sigmund Freud — notes

What it is. The author (LearnAIWithMe, Substack) built a personal AI therapist, named "Mind Mirror," from Claude. It analyzes years of ChatGPT chat history to profile the user and answer emotional/relationship questions (neighbor conflicts, family, feelings).

Key numbers.

  • 3 years / 32 months of chat history (since August 2023) fed into it
  • 8,303 indexed memories — the AI recalled facts the author never stated in the current chat (e.g. "sweat-producing exercise calms me"; that he ignores neck pain until he collapses)
  • Claims: "It knows me better than a $200/hour therapist"

Claimed prevalence. A 2026 review in Digital Public Health reportedly found roughly one in four AI users ask AI how they should feel.

Features.

  • Choose a belief-system preference (Stoicism, etc.) at setup, plus a custom option to turn it into anything else
  • Timeline charts of how the author's mind changed over time, per message since Aug 2023
  • Per-year summaries written by AI ("I even forgot what I've done in 2023")

Build method (article's steps).

  • Export ChatGPT conversation history
  • Profile yourself using one prompt + a preferred psychological method
  • Build the app using one prompt — same "one-prompt method" as the author's prior "$30M App Clone"
  • The prompt is distributed via Google Drive

Caveats. Author's stated caution: "one wrong message can cause chaos during a crisis." He personally falls back on Stoic philosophy when angry (focus on what he cannot control). Alternative data source: a diary, if kept. Note: the article is a funnel piece — the actual prompt and app code are not inline; they're in a Google Drive linked off-page.

Full text · 2,994 chars
I Built My Digital Sigmund Freud. It Knows Me Better Than a $200/Hour Therapist The AI therapist I built with Claude reads 3 years of my ChatGPT history. The export steps, the profiling prompt, and the full app you can build today. I catch myself asking AI how I should feel. I asked it about conflicts with my neighbor, my family, and sometimes even shared my own thoughts. Yet, I never thought about building something about this. I assumed not many people were doing it. It turns out I was wrong. According to a 2026 review in Digital Public Health, roughly one in four AI users do the same. So you also asked about it. But one wrong message can cause chaos during a crisis. That’s why I always follow a pattern. When I’m angry, I turn to Stoic philosophy. And I try not to focus on what I cannot control. Today, I turned it into a system. I built my own Sigmund Freud with Claude. Let me show you. What Does This AI Therapist Actually Do? I named my Sigmund Freud the Mind Mirror. Because AI will analyze all my conversations and profile me, using its methods. (I’ll show you how.) At the beginning, it asks you to choose your preference. Because let’s face it, there are many different belief systems. You don’t have to fit into just one. That’s why I added a custom option. You can use it to turn the system into your own digital Sigmund Freud, or anything else you choose. But to make it accurate and help it understand you better, it needs a lot of data. So, where can you get that data? If you keep a diary, that’s perfect. Use it. I used my conversations with AI because it already stores every chat you’ve ever had. Three years of mine were sitting there. You just need to export them, and I’ll show you how in a moment. Now, let me show you a real conversation from a day when I was struggling. Now read the answer again. It directly quotes what I told it months ago when I was feeling anxious: Sweat-producing exercise calms me. It knows that I ignore my neck pain until I collapse. I never said any of that in this chat. It pulled this information from three years of my history and 8,303 indexed memories. How the AI Therapist Reads 32 Months of Chat History? I’m going to show you how to access your memories, but first, let me show you another feature. This second screen shows how your mind has changed over time. It creates charts based on every message I’ve sent to AI since August 2023. That’s 32 months of my mind in one graph. Also, it summarizes each year with a couple of sentences. I even forgot what I’ve done in 2023 :) All of these were written by AI, which analyzes my conversations from my chat history. How to Build Your Own AI Therapist First, we’ll export the chat conversation history. And we’ll profile yours using one prompt and a psychological method you prefer. Next, we’ll build an app using only one prompt. If you have built with me before, this is the same one-prompt method from $30M App Clone. And this prompt is inside the Google Drive. Here it is.
21:55

Eval Engineering: A Beginner’s Guide

AI agents can finish every visible step and still quietly fail the real job. That gap is what "eval engineering" fixes — it checks whether an agent's work deserves to move forward, not just whether it reached the end. One reported production case scored 83.9% on answer quality but only 32.3% on faithfulness to what its tools actually returned, meaning much of the answer was unsupported. This is a beginner's guide that promises to build evals from scratch using LangChain, Harbor, and live agent traces, and it's partly an ad for the author's fuller course.

Notes

Eval Engineering: A Beginner's Guide — Emerging AI (substack), 2026-08-04

"An AI agent can complete every visible step and still fail the real task."

Core failure mode described: agent searches, calls tools, writes files, hands off to another agent, reaches END — user sees confidence. But search returned nothing, the first agent invented the missing detail, the next node accepted it as context, and the final node polished it. "The whole system failed quietly."

Supporting production number (reported analysis behind the guide): response quality scored 83.9%, while faithfulness to tool results scored 32.3% — most of the answer was not supported by what tools actually returned.

Framing: eval engineering ≠ prompt trick. Defined as the layer asking: "Did the agent actually complete the job, with the right evidence, through an acceptable path?"

Layer model:

  • prompt → what to do
  • context → information
  • loop → try again
  • graph → where work moves next
  • eval → whether work deserves to move at all

The full (paid) guide covers, in order: define one clear promise; create "the three core files"; install the LangChain skill and Harbor; generate cold-start test cases; turn real agent traces into permanent evals; combine code checks with LLM and human judgment; wire verdicts back into loops and graphs; choose worker vs. judge models properly; launch "the five tests every tool-using agent should have" before trust.

Caveats: this is the intro/excerpt only — no actual eval implementation, metrics methodology, or test definitions are given here; "three core files" and "five tests" are named but never enumerated in the visible text. The 83.9%/32.3% split is a single reported production case, not a benchmark.

Full text · 1,815 chars
Eval Engineering: A Beginner’s Guide How to turn prompts, tools, loops, and agent graphs into a system that can prove its work. The agent reached the end. The job still failed. An AI agent can complete every visible step and still fail the real task. It can search, call tools, write files, hand work to another agent, and produce a clean final answer. Nothing crashes. The graph reaches END. The user sees confidence. But the search returned nothing. The first agent invented the missing detail. The next node accepted it as context. The final node polished it. The whole system failed quietly. One reported production analysis in the material behind this guide found exactly this kind of gap: response quality scored 83.9%, while faithfulness to the tool results scored only 32.3%. The agent looked useful, but much of its answer was not supported by what its tools had actually returned. This is why eval engineering is becoming a serious AI skill. It is not another prompt trick. It is the layer that asks a harder question: Did the agent actually complete the job, with the right evidence, through an acceptable path? A prompt tells the model what to do. Context gives it information. A loop lets it try again. A graph decides where work moves next. An eval decides whether that work deserves to move at all. That last decision changes everything. Inside the full guide, you will build eval engineering from the ground up: define one clear promise, create the three core files, install the LangChain skill and Harbor, generate cold-start test cases, turn real agent traces into permanent evals, combine code checks with LLM and human judgment, wire verdicts back into loops and graphs, choose worker and judge models properly, and launch the five tests every tool-using agent should have before it is trusted.

Web

1
00:00

Tino Cuellar joins Anthropic as Chief Global Affairs Officer

Anthropic hired a former California Supreme Court justice as its first Chief Global Affairs Officer to lead policy and government relations worldwide. Mariano-Florentino (Tino) Cuéllar joins from the presidency of the Carnegie Endowment for International Peace, with a career spanning law, tech, and national security. He'd served as a trustee of Anthropic's Long-Term Benefit Trust since January 2026 and stepped down from it to take the job. His appointment signals Anthropic leaning harder into shaping AI policy with democratic governments.

Notes
Tino Cuéllar joins Anthropic as Chief Global Affairs Officer
  • Source: Anthropic newsroom announcement, Aug 4 2026.
  • News: Mariano-Florentino (Tino) Cuéllar joins Anthropic as its first Chief Global Affairs Officer, leading policy, strategic international engagement, and government relationships worldwide.
  • Prior roles: stepped down as President of the Carnegie Endowment for International Peace (scholars in 20 countries); former Justice of the Supreme Court of California (opinions on tech/privacy, international agreements, separation of powers); former director of Stanford's Freeman Spogli Institute for International Studies; co-director, Center for International Security and Cooperation; director, Stanford Cyber Initiative. Served on the President's Intelligence Advisory Board and State Department's Foreign Affairs Policy Board; worked in the White House and federal agencies across three presidential administrations. Appointed by the National Academy of Sciences to its Committee on Responsible Computing Research.
  • Recent work: co-chaired bipartisan Task Force on Nuclear Proliferation and American Security; co-led California's Frontier AI Working Group; board chair/director of Center for Advanced Study in the Behavioral Sciences. Currently Cameron Schrier Family Professor at Stanford Law School (taught AI classes there ~a decade ago); Senior Fellow at Stanford's Institute for Human-Centered AI.
  • Trust tie: Trustee of Anthropic's Long-Term Benefit Trust since January 2026; stepped down from the Trust to join — successor to be selected under the Trust's normal process.

Quotes: Cuéllar: "Democracies must set the terms on which this technology advances, and there is no more consequential place to be shaping that work right now than Anthropic." Daniela Amodei: "I can't think of anyone better prepared to partner with governments, civil society, and community groups..."

Caveat: no start date, reporting structure, or policy agenda disclosed; role described only at mission level.

Full text · 3,777 chars
Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer Mariano-Florentino (Tino) Cuéllar will join Anthropic as its first Chief Global Affairs Officer, leading the company’s work on policy, strategic international engagement, and government relationships worldwide. Tino’s career spans law, technology, international security, and public institutions at the international, national, and state levels. He recently stepped down as President of the Carnegie Endowment for International Peace, a leading independent global policy research institution with scholars in 20 countries. Prior to his role at Carnegie, Tino was a Justice of the Supreme Court of California, where his opinions addressed technology and privacy, international agreements, and the separation of powers, among other issues. He was previously director of Stanford's Freeman Spogli Institute for International Studies, co-director of the university’s Center for International Security and Cooperation, and director of the Stanford Cyber Initiative. He has served on the President's Intelligence Advisory Board and the US Department of State's Foreign Affairs Policy Board, and worked in the White House and federal agencies in three presidential administrations. The National Academy of Sciences appointed him to its Committee on Responsible Computing Research. In recent years, he also co-chaired the bipartisan Task Force on Nuclear Proliferation and American Security, co-led California’s Frontier AI Working Group, and served as board chair and later director of the Center for Advanced Study in the Behavioral Sciences. Currently, he is the Cameron Schrier Family Professor at Stanford Law School, where he started his teaching career before serving in the judiciary and began organizing classes on artificial intelligence nearly a decade ago. He also serves as Senior Fellow at Stanford’s Institute for Human-Centered Artificial Intelligence. Tino has served as a Trustee of Anthropic's Long-Term Benefit Trust since January 2026. He has stepped down from the Trust to join the company. The Trust will select a successor under its normal process. “Policymakers in the US and around the world are increasingly realizing that we are at a critical inflection point when it comes to how we govern and develop artificial intelligence. The choices we make today will determine whether humanity can harness extraordinary possibilities to advance science and improve lives across the world or face enormous risk and growing inequality,” said Cuéllar. “Democracies must set the terms on which this technology advances, and there is no more consequential place to be shaping that work right now than Anthropic.” “Tino has spent his career helping public institutions respond to times of change with thoughtfulness, pragmatism, and deep commitment to the common good,” said Daniela Amodei. “At all levels of government, the law, and academia, Tino has served with sound judgment and civic-mindedness, and we’re looking forward to him putting these principles to work at Anthropic. I can't think of anyone better prepared to partner with governments, civil society, and community groups as they engage with both the risks and opportunities presented by advanced AI.” Tino arrives at a pivotal moment for Anthropic's work with governments around the world. The questions AI raises for economies, for security, and for communities absorbing rapid change are being debated by leaders everywhere. Ensuring AI’s trajectory is shaped by democratic societies and its benefits reach people broadly is a critical priority. Tino will help steer this work while finding common cause with heads of state and policy leaders on the questions and possibilities AI is raising for communities everywhere.