Nothing matches those filters.

Lead

8

Article

35
00:06

Generalist AI GEN-1.5 Learns New Robot Tasks From Single Demo, No Retraining

A new generalist robot AI, GEN-1.5, learns a new task from a single video demo with no retraining. Instead of recording a whole new task, a worker can assemble behaviors from short clips — the maker calls it physical prompt engineering. Coverage is thin, so tech details and limitations aren't spelled out.

Full text · 148 chars
In practice, this means a worker could assemble behaviors from short clips — physical prompt engineering — rather than recording a new full-task ...
14:14

Amazon web services to spend $1bn putting AI engineers inside customer teams

Amazon Web Services is spending $1 billion to put AI engineers directly inside customer teams. The embedded engineers will build agentic systems alongside the client's own staff, making it a big bet on hands-on services revenue rather than a pure product release.

Full text · 138 chars
Amazon Web Services commits $1 billion to embed artificial intelligence engineers directly within customer teams to build agentic systems.
00:38

How Will Scaler Train 10,000 AI Engineers ? | AIM

An Indian edtech company is putting big money into training a huge batch of AI engineers for enterprise work. Scaler launched a Forward Deployed Engineer specialization and committed ₹25 crore (about $3 million) to train 10,000 enterprise AI engineers. Forward deployed engineers typically work on-site at customer companies. The program is a bet on the hot enterprise AI talent market.

Full text · 149 chars
Scaler has launched a Forward Deployed Engineer (FDE) Specialization and committed ₹25 crore towards building a talent pipeline for enterprise AI ...
01:26

AI data center builder Nscale reportedly seeking $3B IPO

AI data center builder Nscale is reportedly seeking a $3B IPO. The company also sells a prompt engineering tool to help developers improve model response quality, and expanded its operations last month. It's business news about AI infrastructure going public.

Full text · 150 chars
Additionally, the company offers a prompt engineering tool that helps developers boost the quality of model responses. Last month, Nscale expanded ...
01:40

Minnesota lawyer suspended over fake AI case citations - MPR News

A Minnesota lawyer agreed to a 30-day suspension of his license for filing a 2025 legal brief that cited AI-fabricated case law that doesn't exist. Attorney Faisal S. Ahmed accepted the discipline, another example of courts punishing lawyers for submitting hallucinated citations.

Full text · 146 chars
Minnesota attorney Faisal S. Ahmed has agreed to a 30-day suspension of his law license for filing a 2025 legal brief that cited nonexistent, AI -
15:04

If you're not using AI to attack your own systems, your adversaries will - The Register

Security teams should use AI to attack their own systems before real adversaries do. The Register highlights an AI agent that suggested installing a malware package and an engineer nearly followed its advice. Autonomous AI attacks are called a 'clear and present danger,' so the argument is to hunt your own vulnerabilities with the same tools attackers will use.

Full text · 154 chars
MORE CONTEXT. AI agent suggested installing a malware package. Engineer almost took its advice · Autonomous AI attacks pose 'clear and present danger' ...
16:15

☕️ Tesla discontinues its Solar Roof

Tesla quietly killed its Solar Roof product, scrubbing it from its website after a decade of poor sales — it aimed for 1,000 weekly installs but reached only 20 to 40 by 2022, and some quotes topped $200,000. Elsewhere in this roundup: Uber faces a €825 million fine for automated driver suspensions, Amazon raised Echo prices up to 60% on component costs, TikTok and ByteDance paid $400 million to settle a child-privacy suit, Nvidia's AVO coding agent aced the ARC-AGI-3 puzzles with a perfect score while its underlying Claude Opus 5 alone scored 30%, and a US lab is probing whether Chinese lidar sensors pose security risks.

Full text · 4,101 chars
| | | ☀️ Tesla discontinues its Solar Roof LINK | Tesla has quietly killed its Solar Roof, the roughly decade-old product that turned shingles and tiles into solar panels, removing all mentions from its website and redirecting the old pages to a generic solar landing page. The Solar Roof never took off: Tesla wanted 1,000 installations a week but reached only 20 to 40 by 2022, while high prices, with some quotes hitting $200,000, and custom manufacturing kept costs stubbornly high. Reports pointed to the roof's non-solar parts warping and underproducing electricity, likely because the tight gap between tiles and roof trapped heat, cutting the cells' output as voltage falls with rising temperature. | 🚗 Uber fined $1B over driver suspensions LINK | The Dutch Data Protection Authority fined Uber €825 million ($966 million) for using automated systems to suspend driver accounts without warning drivers or giving them a human to review the decision. The August 17 penalty is the second-largest ever issued under Europe's GDPR, behind only a €1.2 billion fine against Meta in 2023, and Uber said it will appeal the decision. Uber temporarily suspended drivers suspected of fraud, like inflating fares with detours or accepting rides they never completed, and says only 126 European drivers were deactivated for low customer ratings in 2021. | 📉 Amazon hikes Echo prices up to 60% LINK | Amazon has raised prices on several Echo, Kindle, Fire TV, and Eero products, blaming rising memory and storage costs, with some jumps as steep as 60% on its cheaper smart home devices. The Echo Dot climbed from $50 to $80, while the Fire TV Stick 4K Max went from $60 to $85 and the 16GB Kindle rose from $110 to $150. Amazon confirmed the changes, pointing to "significant increases in memory and storage component costs" across the electronics industry, and said it held off on passing those costs to shoppers for as long as it could. | 🎵 TikTok settles child privacy suit LINK | TikTok and its parent company ByteDance agreed to pay $400 million to settle a Justice Department lawsuit accusing the app of breaking children's online privacy laws, without admitting any wrongdoing. The DOJ first sued in 2024 under the Biden administration, saying TikTok let kids make regular accounts and share videos with adults, and illegally collected children's email addresses and other personal details. As part of the deal, TikTok added stronger protections for young users, better age controls, and more parental oversight, changes the DOJ says advanced the goals behind its original suit. | 🤖 Nvidia's coding agent aces ARC-AGI-3 test LINK | Nvidia's AVO coding agent completed all 183 levels across the 25 public games in the ARC-AGI-3 test with a perfect score, working through the puzzles without any prior instructions or set goals. AVO is a wrapper built around Anthropic's Claude Opus 5, and while the model alone scored just 30 percent on the same public set, the added software layer pushed the result all the way up to 100 percent. Nvidia originally made AVO to tune GPU code, then swapped its tools for the ARC-AGI-3 interface, where it cleared the levels in 6,624 actions, about 12 percent fewer than the rival VISTA wrapper needed. | 🚗 US lab probes Chinese lidar security LINK | The Idaho National Laboratory, a Department of Energy facility, is examining whether Chinese lidar sensors could create security dangers if they spread across vehicles in the United States, according to two people who spoke to TechCrunch. The lab is looking mainly at cybersecurity risks, and possibly clearing Chinese lidar of concerns, with funding from unnamed electric and autonomous vehicle companies; Rivian, GM, Ford, and others said they weren't aware of it. The probe comes as lawmakers push bills to ban Chinese lidar over spying fears, while suppliers Hesai and RoboSense have cut sensor prices from as much as $75,000 a decade ago to a few hundred dollars. | |
16:41

Microsoft Engineers Say TokenOps Cuts AI Agent Costs 78% While Boosting Completion to 96%

Microsoft engineers say a pattern called TokenOps cuts AI agent costs by 78% and lifts task completion to 96%. The trick is controlling how much an agent keeps thinking and spending tokens per step. Tisha Chawla and Susheem Koul described the pattern on the AI Engineer podcast.

Full text · 151 chars
It's the agent that won't stop thinking. Microsoft engineers Tisha Chawla and Susheem Koul, speaking on the AI Engineer podcast, describe a pattern ...
21:04

Quoting Linus Torvalds

Linus Torvalds credited an AI coding assistant for doing much of the grunt work in a brutal kernel debugging session, and let the AI write the commit message for the resulting fix. He noted the AI several times declared the bug impossible and suggested writing a report about it instead, but it kept adding debug code and analyzing faithfully when he pushed. A rare public endorsement of AI assistance in kernel development.

Full text · 933 chars
22nd August 2026 And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work. I'd like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it. I suspect those things have been trained by people who may not be quite as stubborn as I am. But while the AI was ready to give up several times, it did keep adding debug code and analyzing it faithfully when I pushed. So credit where credit is due and I let the AI write the commit message above. — Linus Torvalds, drm/xe: Don't hand out the flat CCS storage as usable VRAM Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
21:33

Enterprises winning with AI agents are limiting how much the agents can do alone

Companies that get real wins from AI agents usually cap how much the agents can do on their own. Only about 30% of organizations have hit a governance maturity level of three or higher for agentic AI controls, so most are still building guardrails. Bounding agent autonomy is what separates the winners from the pack.

Full text · 150 chars
Only about 30% of organizations have reached a maturity level of three or higher in governance and agentic AI controls specifically. Put those two ...
00:01

Bowie State University Launches Bachelor's Degree Program in Artificial Intelligence

Bowie State University in Maryland is launching a bachelor of science degree in artificial intelligence starting fall 2026. It's a new undergraduate AI program from an HBCU adding to its curriculum, with no course or enrollment details yet.

Full text · 151 chars
Starting in the Fall 2026 semester, Bowie State University in Maryland will offer a new bachelor of science degree program in artificial intelligence .
00:15

A-MATD3: adversarial multi- agent twin delayed DDPG for resilient routing in SAGIN under ...

Researchers propose a new algorithm, A-MATD3, for resilient routing in space-air-ground integrated networks. It's an adversarial multi-agent reinforcement learning approach developed at Air Force Engineering University in China. Published in Scientific Reports; this excerpt gives few technical details.

Full text · 148 chars
Authors and Affiliations · Graduate School, Air Force Engineering University, Xi'an, 710051, China. Jinling Liu & Jinghan Wang · Air Defense and ...
00:31

Microsoft Looks for a $279K Legal Engineer to Bring AI to Its Lawyers - Briefs Finance

Microsoft is hiring a "legal engineer" on a roughly $279K salary to bring AI to its own lawyers. The role involves building AI agents, refining prompts, and coaching attorneys inside the Customer & Partner Solutions group. It signals Microsoft pushing AI into internal legal work, not just selling it to clients.

Full text · 145 chars
The role involves creating AI agents, refining prompts , and coaching attorneys within the Customer & Partner Solutions group. The tech giant ...
00:34

Royster Fellow will study how AI can be manipulated | UNC-Chapel Hill

A new research fellow is joining UNC-Chapel Hill to study how AI systems can be manipulated. Kelsey Campbell, a former State Department researcher on foreign influence over public opinion, sees the same influence tactics now being used on AI. The piece is thin — it gives little beyond her background and research topic.

Full text · 151 chars
Kelsey Campbell, who researched foreign influence on public opinion for the Department of State, sees similar tactics used in artificial intelligence .
00:43

Moderna Stock Is Just The Start: Why AI -Focused Money Might Flow Into Biotech

Wall Street is treating biotech as the next place AI money could pile in. An analyst at Revere Asset Management points to Moderna stock as the start of AI-focused investors looking at biotech. The story is thin, more an investor opinion than a market event. It hints AI hype is spreading to adjacent sectors.

Full text · 123 chars
Moderna stock catches the eye of investors looking for the next big AI play, Revere Asset Management's Don Vandenbord says.
01:09

From Traditional Development to AI -Native Engineering : ISHIR's AI Software ... - Security Boulevard

A consulting firm argues companies should stop asking whether developers use AI and instead measure how much of their engineering is AI-native. ISHIR lays out an AI software engineering maturity spectrum for CEOs and CTOs. The framing treats AI adoption as a staged progression, not a yes-or-no switch. It's a positioning piece rather than new research.

Full text · 147 chars
The question for CEOs, CIOs, CTOs, and engineering leaders is no longer: “Are our developers using AI ?” A better question is: “How much of our ...
01:20

Data center madness - Marcus on AI

A prominent AI critic argues data center spending has gone mad and public sentiment has soured. Gary Marcus's Substack piece cites two estimates of how extreme AI capital expenditure has gotten, plus four signs public opinion has turned against it. It's an opinion column, not a news report. Useful as a temperature read on the AI capex debate.

Full text · 106 chars
Two estimates of how crazy AI Capex has gotten, and four new signs that public opinion has totally soured.
01:22

Why architecture matters more than prompts for real-world AI systems

For real-world AI systems, system architecture matters more than prompt tuning. A model that nails a demo still has to work inside an actual product, and that's where the hard engineering problems live. It's an argument that demo performance doesn't automatically carry over to production.

Full text · 153 chars
That shift creates a different engineering problem. A model that performs well in a demonstration still has to operate inside a real product. It must ...
01:34

Pope Leo to Catholic legislators: Promote AI policies that protect families - EWTN News

Pope Leo XIV is urging Catholic legislators to push AI policies that protect families from the risks of the technology. The appeal fits the Vatican's broader engagement with AI ethics but offers no specific proposals or detail.

Full text · 135 chars
Pope Leo XIV has urged Catholic legislators to promote policies that protect families from the risks of artificial intelligence ( AI ).
01:55

'Hot competition': NSA deputy sounds alarm on China threat, AI race - Breaking Defense

An NSA deputy is warning of a "hot competition" with China over AI and calling it a serious threat. The alarm follows a Washington Post report that Chinese firms, some tied to the military, used AI to analyze imagery or data. The article is thin on specifics but frames AI as a core national-security rivalry.

Full text · 153 chars
In April, The Washington Post reported that Chinese firms, some with links to the Chinese military, had been using artificial intelligence to analyze ...
13:15

LinkedIn Says Users Love Its Anti-AI-Slop Button - PCMag UK

LinkedIn says people love its new anti-AI-slop button. The tool lets users flag or clean up obviously AI-generated posts, and the company claims it's catching on. This one is a headline digest with thin detail — it also mentions Google letting you customize its Discover feed with AI and a man failing to prompt-engineer his way to a legal win.

Full text · 141 chars
... Google Will Soon Let You Customize Discover Feed Using AI ... Man Tried to Prompt Engineer His Way to a Legal Victory. It Didn't Work ...
13:23

🏦 The problem with petards

An essay argues AI labs talked themselves into a corner: their decade-long "too important to restrain us" pitch has now exploded in their faces as local communities push back against datacenters. Author Azeem Azhar draws a parallel to the petard, a bomb that could blow up its own engineer, arguing the industry's grand promises and doomer warnings backfired politically. The excerpt covers that framing; the fuller essay digs into whether datacenters actually benefit the communities hosting them.

Full text · 2,427 chars
The petard was a sixteenth-century explosive charge. An attacking engineer would carry it to a castle gate, attach it, light the fuse and scramble for cover. It was a tricky business. The charges were temperamental, their fuses particularly so, and the installer might blow himself up. In Hamlet, the phrase earns its immortal meaning: For ’tis the sport to have the engineer Hoist with his own petard; and ’t shall go hard But I will delve one yard below their mines And blow them at the moon. O, ’tis most sweet When in one line two crafts directly meet. Today’s petardiers are not Rosencrantz and Guildenstern conspiring against the Prince of Denmark. They are the titans of AI, the bosses of the labs, the investors behind them. For nearly a decade, these software coders had made promises: of reigniting economic growth, of making daily life easier and less risky, perhaps even of eliminating disease. To do this, they would need capital: to write their software and to build 21st-century infrastructure to run it. The gains will be so huge that they’d need to go quickly, very quickly. But they warned that this was no ordinary software. It was tricky and hard to understand, so much so that only a few should steward it. After all, this was a technology that would possibly take your job, maybe kill you, perhaps get out of control and kill all of us. Still, they needed to build it, all the while warning that they would eventually need to take control of it before things got out of hand. In the meantime, let them be. This was the petard: AI is too important for America not to let us get on with it. They placed this claim on the gates of society, itching to get on with it. And it just exploded in their face. What to read This is a fiendishly complicated issue, and I’m not going to pretend to understand the whys and wherefores of American political decisions. My rough take, though, is that this is a Gordian knot. It cannot be disconnected from the hapless messaging from AI firms over the past decade, which now comes to life at the county level. Nor can we separate the idea that these concrete blocks, which might abstractly benefit the economy or healthcare or whatever, tangibly serve an out-group they don't much like. The rest of the essay has further analysis on attitudes towards datacenters; whether they actually benefit local communities; and the commentary I have been reading to understand this.
15:56

More than just code review

The skill that matters most with coding agents is telling them what to change and then confidently verifying it actually happened — not necessarily reading every line they wrote. Line-by-line review was never the most effective way to validate a software change anyway. A short essay from Simon Willison on working productively with coding agents.

Full text · 737 chars
22nd August 2026 The key skill required to make productive use of coding agents is being able to confidently instruct them on how to make changes and then confidently verify that those changes have been applied in the correct way. Sometimes this involves reviewing every line of code they have written, but there are other ways to achieve that goal. Eyeballing every line of code has never been the most effective way to validate a chance to a piece of software. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
17:01

llm 0.33

The llm command-line tool shipped version 0.33, mostly fixing and extending embedding support. Embedding commands now accept a --key flag so per-call keys reach plugins without touching shared model state, and prompt templates can be combined in order so one template supplies a model with options while another supplies the prompt. Reasoning-capable Responses API models also gained a reasoning_summary option with auto, concise, and detailed values.

Full text · 1,666 chars
22nd August 2026 My highlights from this release: I shipped a quick 0.32.1 fix for this yesterday, but this is the more comprehensive fix. llm embed and llm embed-multi now accept --key. The Python EmbeddingModel.embed(), EmbeddingModel.embed_multi(), Collection.embed() and Collection.embed_multi() methods accept key= too, passing the resolved per-call key to embedding plugins without changing shared model state. Existing plugins that read self.key continue to work through a compatibility fallback. Thanks, ChrisJr404. #757, #1620 The embedding models now use the same pattern for keys that regular LLM models do. llm prompt -t/--template can now be repeated to combine templates in order. This allows model configuration and options from one template to be used with a prompt from another. This unlocks a neat pattern where you can create templates that package a model with a set of default options: llm -m gpt-5.6-luna -o reasoning_effort high --save lhigh llm "Generate an SVG of a pelican riding a bicycle" --save pelican # Combine and run the templates llm -t lhigh -t pelican - Reasoning-capable Responses API models now support a reasoning_summary option with auto, concise, and detailed values. This can be used with llm openai endpoint --responses. #1600 This is particularly useful for exercising different models that provide their own imitation of the OpenAI Responses API. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
17:26

Microsoft AI Image Generator: MAI-Image-2.5 Pro Setup

A tutorial walks through setting up Microsoft's MAI-Image-2.5 Pro image generator. Its core tip: prompts for editing images behave differently from prompts for generating new ones, and treating them the same is a common mistake. It's a how-to piece, not a news announcement.

Full text · 149 chars
Prompt Engineering Tips for Editing Workflows. Editing prompts behave differently from generation prompts, and treating them the same is a common ...
18:11

Patrick Debois: Coding Agents Won't Scale Until You Fix the Org Chart - BigGo Finance

Patrick Debois argues coding agents won't scale until companies fix the org chart first. He calls the 'solo 10x developer' narrative a dead end on the AI Engineer podcast. Competitive advantage, he says, comes from reshaping how teams are organized around agents rather than swapping people one-for-one.

Full text · 150 chars
Speaking on the AI Engineer podcast, Debois contends that the "solo 10x developer" narrative is a dead end. Instead, he says competitive advantage ...
18:11

Rémi Louf: The CEO Who Fired Anthropic and Built a Better Agent Runtime in Two Weeks

A CEO claims he fired Anthropic and built a better agent runtime in just two weeks. Remi Louf says the replacement runs 20 production agents, some built by non-engineers, on open-source models with zero third-party dependencies. The claim comes from a podcast interview, so the specifics are thin.

Full text · 154 chars
The result is a system running 20 production agents at .txt—some built by non- engineers —powered entirely by open-source models with zero third-party ...
19:35

Safia Abdalla: The Goal of Agent Platforms Is to Make Non-Developers Builders

The pitch for agent platforms is that they should let ordinary people build software, not just make professional coders faster. Warp engineer Safia Abdalla argues the real promise of agentic infrastructure is expanding who gets to build, rather than automating engineers' work. The item is mostly a headline and a single lead line, so there's not much depth beyond that claim.

Full text · 145 chars
Warp engineer Safia Abdalla argues that the true promise of agentic infrastructure is not automating code for engineers — it is expanding who ...
00:06

Senior Frontend Engineer, AI Product - Observe by Snowflake

This item is just a job posting. Snowflake is hiring a senior frontend engineer for its AI product Observe, working on what the company calls the agentic enterprise. There's no actual news content beyond the job listing. Filing it for the signal that enterprise vendors are staffing agent-focused AI products.

Full text · 153 chars
Senior Frontend Engineer, AI Product - Observe by Snowflake ... At Snowflake, we are powering the era of the agentic enterprise. To usher in this new ...
01:08

AI 101: What the technology can offer fleet operations

An intro explainer walks fleet operators through what AI actually is — a more advanced form of machine learning. It quotes IBM's definition and positions AI as a tool for improving fleet operations. The piece is a basics primer with no news or technical depth.

Full text · 154 chars
What even is artificial intelligence ? Simply put, AI is a more advanced type of machine learning . IBM states that AI “enables computers and machines ...
01:28

Covera: AI -native OS for insurance brokers - Y Combinator

A Y Combinator startup called Covera sells AI agents that take over the day-to-day servicing work of insurance brokers, like taking in submissions and getting carrier quotes. The pitch is ready-to-use agents rather than tools brokers build themselves. It's a niche vertical play, and details beyond the elevator pitch are thin.

Full text · 150 chars
Covera builds AI agents for insurance brokers. Our ready-to-use agents take over the day-to-day servicing work: submission intake, carrier quoting ...
04:17

The HackerNoon Newsletter: How I Built a Data Pipeline From Scratch Using Python (8/21/2026)

A HackerNoon newsletter roundup leads with a story on building a data pipeline from scratch in Python. The rest is a digest of related blog links with no standalone news of its own.

Full text · 142 chars
... agentic - engineering #swagger. THIS ARTICLE WAS FEATURED IN. Terminal · Lite. Related Blog Posts. Make us Preferred on Google. /https ...
12:02

Lean Prompts Beat Micromanagement in New Anthropic Models

A how-to guide says shorter, leaner prompts work better than micromanaging modern Anthropic and OpenAI models, listing five key components. It's a summary of official guidance with no new findings or numbers. Thin, promotional content.

Full text · 139 chars
Write leaner AI prompts using the latest official guidance from OpenAI and Anthropic. Stop micromanaging models and apply 5 key components.
13:33

Ethical decision making is critical while learning technical skills in AI era - Education Times

An education commentary argues that ethical judgment matters as much as technical AI skills. It pushes back on the idea that mastering prompt engineering is the key to success, saying two candidates who write equally good prompts are still separated by ethics. This is opinion content with no hard news.

Full text · 139 chars
The popular belief that mastering prompt engineering , automated ... Moreover, if two candidates can write equally good prompts and can ...
19:52

Harness Engineering : How AI Agents Work - YouTube

A short YouTube explainer walks through how AI harnesses turn language models into agents, covering prompt engineering, context engineering, tool use, and the harness loop. It's an introductory educational video, and the summary here is basically its chapter list, so there's no news in it.

Full text · 148 chars
0:59 How AI Harnesses Make LLMs Agentic 1:31 From Prompt Engineering to Context Engineering 2:23 Tools Turn LLMs Into Agents 3:06 The Harness Loop 3

Newsletter

4
07:30

The Evolution of the Agent Harness

Agents suddenly started working last Christmas because the models and the scaffolding around them improved together and their curves finally crossed at the right moment, argues this deep analysis of the "agent harness." The evidence: the same model scored between 52.4 and 76.2 across different harnesses on Harness-Bench, OpenAI tripled GPT-5.6 Sol's ARC-AGI-3 score with harness-only changes, and Claude Code hit roughly $1 billion in annual revenue within six months of launch. The author predicts the next phase is an "attention era" where every agentic company ships an attention-policy surface telling agents when to interrupt and what to decide alone, because human attention is the last scarce resource.

Full text · 11,674 chars
Sometime around Christmas 2025, AI engineers noticed a change in agents. They started to work! It’s hard to pin down exactly why. Maybe we finally had holiday downtime to try the newest agents with the newest models. Maybe the models had crossed some capability threshold. Maybe the wrappers around the models had matured. What I’ll argue in this post is that it was the confluence of the last two. The model and the harness improving together and then their curves of improvement crossing at the right moment. And that dynamic helps to explain what comes next: models keep absorbing the harness into their weights, engineers keep deleting what got absorbed, and what remains is a harness for human attention rather than for the model. Lukasz Kaiser, one of the people who invented the Transformer, said on “Unsupervised Learning” in June: “The change last winter, last Christmas — it’s a little hard to pin down. I mean, the harness changed and a little post-training changed and then new pre-trained models came… but it felt like a big jump which is not that easy to pin down what did it.” The answer to “What happened?” isn’t solely in the model weights. It’s in the system that grew up around the weights. The answer is in the agent harness. Think back to November 2022, when ChatGPT was the most advanced AI tool. The only capability at its disposal was next-token prediction and some Reinforcement Learning from Human Feedback (RLHF) that allowed it to act like a helpful assistant. No tools, no search, and no reasoning. The original ChatGPT was confined to its training data and the prompt you sent it. No more, no less. It was a brain in a vat. The agent harness is a way for the LLM to break free from that confinement and interact with real digital information space. What a Harness Actually Is An agent harness is everything besides the model weights that makes the agent work. The environment, tools, context and guardrails that surround the model. Without the harness the model is a brain in a vat. It can take an epistemic action, but needs the harness to actuate that decision in real digital space. The harness is like giving the mind of the model a body. With the harness, the model can perceive (context), act (tools), persist information (memory and compaction), and enforce its boundaries (permissions and guardrails). Harness 1.0: The Past, “The Bolt-On Era” Two curves run through the path of model / harness evolution. What the harness asks of the model, and what the model can deliver in practice. The gap between these two curves is equal to the effectiveness of an agent, and the closing of that gap is what I’ll argue led to the tangible improvement in agents that Lukasz Kaiser referenced. Here’s how the gap closes, in stages: - ReAct, “The Harness on Paper” (October 2022): ReAct is a prompting technique to get models to reason through prompting. It’s the agentic loop on paper, external to the model weights. It defines the idea of an “agent loop” where a model reasons -> acts -> observes -> repeats. Again, the ReAct loop exists only as a prompting method. Prompting is the only reasoning method that exists at this time and no one calls it a “harness.” Toolformer (Meta, Feb. 2023), that same winter, hints that tool use could be trained in rather than prompted. It’s a bit like Alan Turing’s idea of the computer before it was instantiated in a physical substrate. A powerful idea that is only later made manifest. (ReAct predates ChatGPT by a month — October 2022 vs. November 2022). Both the curves are near zero at this point. The gap is small because we are just getting started. - AutoGPT/BabyAGI, “Premature Autonomy” (Spring 2023): With AutoGPT/BabyAGI, the harness curve sprints ahead of the model capability curve. Both hand the model full autonomy, asking the model to act as an “autonomous employee,” but the models at this point are still little more than brittle next-token predictors. A loop doesn’t add capability to a model. A loop amplifies the capability a model has, and below some threshold the loop amplifies errors rather than reliability. Consider the power of compounding in the negative: 95% reliability per-step over a 20-step task results in a ~36% average success rate. The harness hands the model an assignment it has no realistic chance of completing. This is where the gap is at its widest and the next 18 months are a reaction and attempt to close that gap. - Cursor/Copilot, “Retreat to Human in the Loop” (2023 - 2024): The first AI-powered IDEs recognize the failure-mode of giving the model too much autonomy. They close the gap by pulling the harness curve down below the model curve. Don’t give the model the loop directly; give the human the loop and empower the human to orchestrate the loop while the model speeds the human up. The first version of Devin tries to hand the autonomy back to the model. A test from the team at Answer.AI shows that is still premature, with a ~15% success rate. It’s evidence that the move from the IDEs to retreat from full autonomy is not cowardly, but the correct move. However, while the prevailing tactic is to pull the harness down below the model, models continue to improve. Near the end of 2024, with the introduction of o1 — the first reasoning model — for the first time the gap inverts and we begin to see the first signs of a model capability overhang. - Claude Code, “The Curves Cross” (February 2025): The inversion at the end of 2024 sets up an opportunity that someone has to seize: if the model is now ahead of the harness, then a harness intentionally riding the brakes of the model is leaving capability on the table. Claude Code is the first coding agent built to seize that opportunity. It abandons the IDE for the terminal, gives the model bash and file read/write access, and replaces the need for human approval on every change with permission rules. The model is handed the loop again, and this time it understands the assignment. Boris Cherny and team build Claude Code with the next model’s capabilities in mind, not the current one. It is such a hit not because it’s the first product to give the model autonomy, but because it’s the first product to do so at the right time. That time is the crossover point where the model has gotten reliable enough to succeed with autonomy. Claude Code grows to roughly $1B ARR within six months, all because Anthropic seized the opportunity available when the curves begin to meet. What happens next is that the curves don’t just meet, they begin to braid together. Harness 2.0: The Present, “The Co-Training Era” Today the harness matters, and in a way we can measure. Harness-Bench ran the same model over the same 106 tasks in different harnesses, and scores ranged from 52.4 to 76.2: a 23.8-point spread with zero change to the model. Half the agent is the harness. OpenAI achieved a similar result on ARC-AGI-3 with harness changes. Adding only retained reasoning and compaction, GPT-5.6 Sol’s ARC-AGI-3 score tripled from 13.3% to 38.3%. What’s happening under the hood is that Reinforcement Learning (RL) has moved inside the harness. From OpenAI’s codex-1 release announcement in May 2025: “codex-1 was trained using reinforcement learning on real-world coding tasks in a variety of environments.” The two curves join and start to braid as one unified system. This is the dream of Toolformer manifesting in reality. Rather than a tool prompted from the outside, now tool calling is trained from within the environment of the model. Then, as the models are trained in the environment of the harness, they start to absorb the harness capabilities into the model weights, learning how to auto-compact with knowledge of their own context window, for example. GPT-5.1-Codex-Max launch: “The first model natively trained to operate across multiple context windows through compaction.” Once the models absorb the harness capabilities, the harness can shed the scaffold. It’s production by reduction. Thariq Shihipar from Anthropic said that the team recently deleted 80% of Claude Code’s system prompt. The measure of the pace of agent harness evolution is how much of the harness you get to delete, while retaining the same capability level. This is the future we need to build towards as AI engineers. This, then, is the loop of model / harness evolution: train -> absorb -> shed -> repeat. The model climbs to the next thing it can’t do yet. The jump that Kaiser pointed out is hard to pin down because it’s not a discrete event. A pre-training leap is noticeable because you can articulate it in a model card. A model / harness co-evolution jump is less so, because there’s no documentation of the evolution process. That’s the answer to the jump last Winter: it happened in the space between the model and harness working together. We need to ask: if every harness capability will eventually get absorbed into the model, what does that leave us with? Harness 3.0: The Future, “The Attention Era” Keep deleting everything that the model can absorb. Imagine what your agent looks like at the conclusion of that process. What are you left with in your hand when you’ve deleted everything? What do the model weights absorb next? Multi-agent orchestration, tool selection, memory...to name a few possibilities. Researchers are building self-improving harnesses that can themselves be trained in a similar way to models. What’s left at the end of this deletion and absorption process are the human-centric agent capabilities. Things like permissions, identity, trust and legibility. A model that absorbs permissions into itself has dissolved permissions. Absorption doesn’t end the harness. Absorption inverts the harness. The harness becomes the agent’s interface to the human that operates it. The harness was born as the human interface to the model. We grew from chatbox to IDE to the terminal. If the model absorbs the computer-facing capabilities, the next stage of evolution becomes one layer of abstraction up. The harness becomes the model’s interface to our human attention. It becomes the attention-interface. Ryan Lopopolo said on the “Extreme Harness Engineering for Token Billionaires” episode of Latent Space: “The only fundamentally scarce thing is the synchronous human attention of my team.” Tokens became abundant and reliable, yet we remain bottlenecked on scarce human attention. We see sparks of this already, with Anthropic’s long-running agent progress files and agentic approval queues. The gap between the model and harness curve doesn’t disappear when the model absorbs the harness. It migrates across the human boundary and creates a new pair of curves with a new gap. The new gap is the space between what the agent asks of the human, and what the human is able to answer. The Attention-Interface I predict that within a year, every company building agentic AI will ship a human attention policy surface in the way that every agentic AI company shipped AGENTS.md. AGENTS.md tells the agent how to work with your codebase. The attention-interface will tell the agent how to work with you. It will govern when it’s allowed to interrupt you, when it should keep working, which decisions it can make alone and which decisions need your approval. And like everything else in the agentic system, it will become a learnable component of the system that can learn with more data. Every correction becomes useful data. The model began as a brain in a vat. The harness gave the brain a body, then the body started to dissolve into the brain. What’s left for us to build is the thing no future model will ever absorb. The interface to the one true scarce resource: human attention.
13:03

The First Quantum Attack on Telcos May Have Started

Quantum attackers can start stealing encrypted telecom traffic today and keep it stored until quantum computers can crack it. Chinese state-linked group Salt Typhoon already spent years inside US carriers like AT&T, Verizon, and Lumen, reaching the lawful intercept systems built for wiretaps and the 5G trust machinery that protects keys and identities. Dwell times ran 12-24 months before detection, and full removal across carriers hasn't been confirmed as of early 2026. The practical point: bulk traffic encryption can wait, but the key-exchange and identity trust systems must migrate to quantum-safe crypto first.

Full text · 6,276 chars
Hackers are patient people, very patient. The popular image is someone in a hoodie breaking into a network at three in the morning, grabbing what he can, and disappearing before breakfast. Serious state groups play a very different game. They get in, stay quiet, learn how the network works, map credentials and dependencies, and keep that access for years before doing anything visible. Telecom has already lived through one of these. Salt Typhoon, the Chinese state-linked group that surfaced publicly in October 2024, compromised at least nine US carriers, including AT&T, Verizon, and Lumen, according to Senator Maria Cantwell, who spoke at a Senate Commerce Committee hearing in April 2026. The FBI assessed in August 2025 that the campaign reached more than 80 countries and that roughly 600 organizations were notified of potential compromise. Dwell time at major Telcos ranged from 12 to 24 months before anyone noticed. In the UK, parallel access to the communications of aides across three successive Prime Ministers ran from 2021 to 2024. By May 2026, Singapore had confirmed that all four of its national carriers were compromised. Where they went inside those networks is a very interesting aspect. They reached the lawful intercept systems mandated under CALEA; infrastructure carriers are legally required to build them so law enforcement can wiretap, thereby giving a foreign intelligence service the same visibility the host country had built for itself. Cisco disclosed in February 2025 that the entry was mostly via stolen credentials and, in one case, via a router vulnerability that had been in the NIST database for seven years. But there is another detail to this story that should worry any telco security team reading this. Cantwell stated publicly in February 2026 that Salt Typhoon may still be inside US networks, and the FBI’s own description of the threat as ongoing did not contradict her. As of early 2026, full remediation across the affected carriers has not been confirmed. Which means nobody has to imagine an attacker quietly copying encrypted traffic out of a carrier core. One did it at the exact interfaces where key-exchange and subscriber-identity traffic converge, and we cannot currently prove he left. Quantum attackers can start today Security people have a name for the strategy this enables. Harvest Now, Decrypt Later. An attacker does not need a quantum computer today. He needs access, the ability to collect the right material, and enough storage to keep it. TLS and IPsec handshakes, certificates, public keys, subscriber records, roaming data, encrypted backups, and management traffic could all still be worth reading in 2033. Harvest Now, Decrypt Later only works if the cryptography protecting that archive can eventually be broken, so it is worth spending a minute on what a quantum computer actually changes. Less than most headlines suggest, but in a very specific place. It is not simply a faster machine. It works differently enough that a handful of mathematical problems become tractable, and cryptography happens to sit on exactly those problems. Classical machines work with bits that are either zero or one. Quantum machines work with qubits that can exist in a superposition of states, and the mathematical space doubles with each qubit added. That does not mean the machine tries every answer at once and hands you the right one. Physics is not that generous. When you measure, you still get one result. The trick happens before measurement because quantum states behave like waves, and computational paths can interfere constructively or destructively. A good quantum algorithm arranges for cancellation so that useless possibilities fade and the pattern you want becomes likely to appear. In 1994, Peter Shor ruined public-key cryptography. His algorithm solves integer factorization and the discrete logarithm problem, the two hard problems underlying RSA, Diffie-Hellman, and elliptic-curve cryptography. Multiplying two large primes is easy, and reversing it is not, unless you have a quantum shortcut through the mathematics. For telcos, the important distinction is between encrypting the data and establishing the trust that allows two systems to exchange it securely. Picture what happens when your phone opens a secure connection to a network function. Two things take place, one after the other. First, the two sides have to agree on a secret key while strangers are listening, and each has to prove they are who they claim to be. Once that is settled, the data itself gets scrambled with the shared key and sent. Quantum does very different things to those two steps. The scrambling step is much less exposed. That work is done by AES and similar symmetric algorithms, and the best-known quantum attack against them is Grover’s algorithm. In a simplified model, AES-256 has roughly the brute-force security margin of a 128-bit classical key, which is still an absurdly large search space. This is why AES is not where most quantum-security teams are losing sleep. But the first step is where everything breaks. Agreeing on a secret key in the open and proving identity is exactly the job of RSA, Diffie-Hellman, and elliptic curve cryptography, which is exactly what Shor’s algorithm dismantles. And once you start listing where telcos rely on that, the list becomes really, really long. TLS on every interface; IPsec and IKE between base stations and security gateways; VPN tunnels; the certificates and the entire PKI behind them; Network function authentication within the 5G core; Software and firmware signing; roaming agreements between operators; and eSIM provisioning. Even SUCI, the 5G mechanism that hides a subscriber’s permanent identity during registration, relies on elliptic-curve cryptography today. So an attacker with a working quantum machine never has to decrypt traffic packet by packet. He can forge a certificate, impersonate a network function, or recover the secret behind a key exchange captured years earlier, which can be far more useful than trying to decrypt traffic packet by packet. That difference should decide the migration sequence. Bulk traffic encryption can wait a few years, but the trust machinery cannot, and plenty of coverage still has this backward.
14:48

Hassabis Bet His Job on AGI by 2030. Here Is What He Knows.

The head of Google DeepMind now gives roughly even odds on computers matching the full range of human thinking by 2030. Demis Hassabis defined AGI as matching all human cognitive ability rather than beating benchmarks, in an 81-minute WCIT interview. He also called for inverting the classroom so AI handles rote learning, argued labs are trapped in a prisoner's dilemma over safety so external governance is needed, and said markets will police AI safety before regulators once enterprise money is on the line. He credits AlphaGo's surprising Move 37 with greenlighting AlphaFold, and sums up the industry's economics as watts to dollars to tokens.

Full text · 12,226 chars
Demis Hassabis just put a number on AGI: 2030, at roughly 50% odds. A coin flip, 4 years out, from the man who ran DeepMind for 15 years, picked up a Nobel Prize for AlphaFold, and just stepped back from the CEO seat to work on exactly this: he is now Chair of Google DeepMind and Chief Scientist of Alphabet, with AGI strategy as the day job. Take the number seriously and it changes how you build this year, not in 5. Sitting across from him: Dame Wendy Hall, who has spent 4 decades watching the UK win early and lose late, and pushes him hardest exactly where it counts. I watched the full 81-minute WCIT interview so you can skip it. Here are the 10 takeaways that matter. together with Alumni Ventures: Hassabis just put 50% odds on AGI by 2030. If he is right, the companies building toward it are being priced this year, and access is the whole game. Alumni Ventures opens that door for individual investors: AI, deep tech, quantum, and cybersecurity deals, co-investing alongside firms like a16z, Bessemer, and Y Combinator. ▫️ Curated deal flow of AI-first startups ▫️ AV invests alongside the lead firms in these deals ▫️ No cost to see deals, zero obligation to invest 1. Invert the classroom, or watch AI do your homework for you Hassabis wants to rebuild the school day around AI, over bolting AI onto the one we already have. “We need to invert the classroom.” AI absorbs the rote learning, tailored to each student’s exact pace. Class time rebuilds around human contact: projects, group work, Oxbridge-style supervisions for everyone instead of just the elite. Star lecturers record the core content once, instead of thousands of teachers rebuilding the same slide deck. This is a full reversal of what school buildings are for, over a scheduling tweak. Wendy Hall pushed back on the obvious gap: home-based rote learning assumes every kid has a quiet room and a device, and plenty have neither. The read for builders: edtech that treats AI as a tutor bolt-on is building the wrong product. The bigger opportunity is infrastructure that lets schools restructure time itself, the same platform-over-feature logic that separates lasting products from wrappers. 2. AGI around 2030, and what Hassabis actually means by it Everyone throws the term around. Hassabis hands you a definition and a date. “Around four or five years... so around 2030.” His reference point is the human brain, over a benchmark score: AGI means matching the full range of human cognitive capability. Today’s systems are impressive and inconsistent, so they miss the bar. Scale alone may never close the gap, and he still expects 1 or 2 breakthroughs on the order of transformers or the reinforcement learning behind AlphaGo. He sharpens the number elsewhere in the talk: 50% odds on 2030, with honest error bars. Hall rejects the framing and the timeline on stage, which is worth respecting too. Build your 5-year roadmap around the range, over the headline number. If the low end is right, your competitive window already closed, which is exactly why the next model generation should shape your plan more than the current one. Planning around that range, start here: 3. Move 37 changed everything, then triggered AlphaFold One move in a board game convinced Hassabis that AI could make genuine scientific discoveries. “AlphaGo had played an original move, this Move 37.” He had carried the protein-folding idea for close to 2 decades, since his Cambridge undergraduate days. What he needed was proof that a system could produce something genuinely novel, over merely optimized. Move 37 was that proof, inside a board game. He flew home from the AlphaGo match in Seoul and greenlit AlphaFold within days. The gap between a system performing well and a system discovering something new is the threshold to watch, and the signal usually shows up somewhere unrelated to your industry first. That is the same discovery-loop logic now being industrialized by the teams automating research itself. Track capability proof points, over product launches. 4. The $7 billion number behind the DeepMind sale Hassabis explains selling to Google with actual numbers, over vague industry talk. “I was barely able to scrape together $10 million rounds.” This was 2014. Zero OpenAI, zero AlphaGo, just Atari-playing agents most investors dismissed. He has put the counterfactual at $7 billion: the capital it would have taken to build a genuine competitor at scale, against rounds 10 to 20x smaller than that. Larry Page personally understood the bet when almost nobody else did. Meanwhile Silicon Valley noticed the talent: Hassabis paid himself nothing while researchers on roughly £100K salaries fielded $10 million offers. He calls it a timing problem, since SoftBank-scale mega-rounds arrived roughly a year too late for him. Capital timing beats product timing. If your category is about to attract mega-round capital, the founders who survive the gap year are the ones who structure the raise for patience, model the dilution math honestly, and know how the fund across the table actually makes money. 5. The Full Stack Of AGI Risk (It Is Not Just One Problem) Ask Hassabis what worries him and he hands you a stack of risks, not one scary scenario. “There’s a whole stack of issues, each one more complex.” Misuse sits first: rogue actors turning general-purpose tools toward harm, including tools built for medicine. Technical alignment comes second: keeping increasingly autonomous systems inside the guardrails you set. Economic concentration is third: broad benefit versus a handful of companies capturing it. Meaning is the last and hardest: purpose, once machines absorb the work humans defined themselves by. His sharpest institutional point: nothing currently exists that can govern all 4 layers at once, at the most fragmented geopolitical moment in 30 years. If your AI risk framework covers one layer, it is incomplete. Investors underwriting frontier labs should ask which layer a team is weakest on, over which one gets the press coverage, the same discipline behind a serious security checklist. 6. Why Big Tech actually wants guardrails Hassabis complicates the convenient story of labs racing ahead while regulators chase. “Nobody wants catastrophic things to happen.” He knows the other lab leaders personally, back to postdoc years alongside people like Dario Amodei, and his read is that none of them wants a catastrophe attached to their name. The blocker is game theory, over intent: even safety-minded leaders face a prisoner’s dilemma, because someone always holds an incentive to defect from an agreed standard to win share. That is his actual argument for external governance. Labs are coordination-trapped, over villainous, and good individual incentives fail to add up to good collective outcomes without a referee. Proposals are already moving with major governments. Build your governance thesis around coordination problems, over villain narratives. That framing predicts which policy interventions actually work, and it is the version serious diligence already prices. 7. No bank wants an AI agent losing a billion dollars Hassabis thinks the market enforces AI safety faster than regulation, once enterprise money is on the line. “That loses you a billion dollars the next day.” Financial institutions will demand hard guarantees around agent behavior and data handling before deploying at scale. One expensive failure from a lab with weak guardrails teaches the whole market instantly, so enterprise trust rewards responsible labs and reckless ones fund the object lesson. His caveat: the mechanism only works if buyers price in guardrails before a failure, over after one. If you sell agents into finance, healthcare, or any regulated vertical, your guardrail documentation just became a sales asset, over a compliance afterthought. Buyers are about to start asking. The builder’s stack for exactly that: 8. Watts to dollars to tokens: the formula that explains the whole industry Hassabis compresses the entire AI infrastructure argument into 3 words. “It’s literally watts to dollars to tokens.” Energy cost converts directly into intelligence cost. The UK carries some of the most expensive energy in the Western world, which caps how much inference it can afford to run, and he frames data centers as the new industrial base the way factories were a century ago. The country that solves cheap, abundant power solves its position in the AI economy. This lands on your P&L before it lands in any headline. Model inference against your local energy market, over just your API pricing, then attack the line item directly: 9. Root node problems: why protein folding was the first target Hassabis picked protein folding for what solving it would unlock. “It was also what I call a root node problem.” He met the problem as a Cambridge undergraduate, listening to biologist friends who talked about it constantly, a listening habit he calls deliberate. It sat unsolved for 50 years. He filed it away for roughly 15, waiting for the technology and the proof point to catch up, and 3 million researchers now use AlphaFold. A root node problem, solved once, compounds across an entire field. A leaf node ships a feature. Leaf nodes make a good feature. Root nodes make a category. Before picking your next problem, run the root-or-leaf test, then steal from the 100 agent ideas ranked by exactly this and the category-creation story behind Replit’s seed. 10. The UK Can Build Unicorns. It Cannot Yet Build Giants. Britain wins the first stage of company building, and Hassabis says the second stage is where it falls apart. “We do very well getting to unicorns.” DeepMind stayed in London by choice and seeded a wave: over 10% of Q1 UK AI venture funding traced back to former DeepMind staff, per HSBC Innovation Banking at the event. Talent and founding conditions, proven. What is missing is growth-stage capital and a functioning path to public markets, and Hassabis says plainly he cannot understand why companies stopped floating in London. Energy costs and listing incentives are his 2 named blockers between the UK’s unicorn factory and its first homegrown trillion-dollar company. If you are a UK AI company approaching Series C, treat the gap as fact over theory: plan the growth round assuming you look beyond the UK even if you stay headquartered there, map the check-writers early, and walk in with a valuation story and a 13-week cash position that survive diligence. The Hassabis playbook AGI is close enough that the institutions meant to manage it, in education, finance, and government, have to start moving now, over after it arrives. ▫️ Founders: model energy and compute as a hard line item, write the guardrail documentation before an enterprise buyer asks, and run the root-node test on your roadmap this week. ▫️ Investors: the market-correction thesis only works if you price responsible behavior before a failure forces you to. The UK growth-stage gap is genuine, and genuine gaps are opportunities. Underwrite guardrails as seriously as growth metrics. ▫️ Operators: the classroom-inversion pattern is coming for corporate training too. Pilot one AI-assisted, judgment-heavy workflow before your competitors do, starting from the one-person operating system. ▫️ Everyone else: watts-to-dollars-to-tokens applies to you even if you never touch a model. Ask your AI vendors where their compute actually runs. The 5 principles to steal - Definitions matter more than dates. Know the capability threshold you are planning around, over the year attached to it. - Root node problems compound. Pick the one that unlocks a field, over the one that ships a feature. - Guardrails are a sales asset now. Write the documentation before the buyer asks. - Energy is a line item. Treat it like one in your model. - Coordination problems need referees. Good individual intentions never sum to collective safety on their own. The future is unwritten. Someone is going to write it anyway. Better you than the person who waited. If this breakdown saved you 81 minutes, send it to one founder or investor who needs it. Keep reading Build for the timeline Control the cost curve Raise like it is 2026 Full lecture:
07:36

[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

More and more parts of the AI training pipeline — judging, data, teaching, curriculum, even the researcher and the test environment — are being handled by AI instead of humans, a shift that makes each stage about 10% worse but 100x cheaper and 10,000x faster. The essay walks eight stages, from reward models in 2022 to synthetic training data, model teachers like DeepSeek-R1's distilled models, and Karpathy's overnight autoresearch runs that cut time-to-GPT-2 from 2.02 to 1.80 hours. It argues the frontier advances when verification gets better, not when generation does, and that the physical world — wet labs and real experiments — is the one piece that can't be fully simulated. The news section covers the 'Ox Alpha' mystery model suspected to be a Zhipu/GLM variant, DeepSeek's new V4-Flash-Vision-Exp multimodal release, and OpenAI's 20%+ price cuts on GPT-5.6 Sol.

Full text · 27,966 chars
By AI standards today is a pretty quiet Friday, so it’s time to take a step back and reflect on what is really going on. If you read our 2025 reading list, and followed our coverage of Z.ai GLM, understood the Poolside pivot, been following our AI for Science themes, and tuned in to today’s Simile pod, you not only are one of the biggest readers of Latent Space, you will probably also arrive at this mental model: Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized. And if you squint, what we used to call “synthetic data” and “synthetic rubrics” and “AI researcher” and “end to end RL environments” is just increasingly ambitious human simulation - 10% worse, but 100x cheaper and 10,000x faster. Stage 1: The reward signal (2022) The first thing to go synthetic was, counterintuitively, the judge. InstructGPT established the now-canonical trick: collect human preferences once, train a reward model, and let the policy optimize against the model rather than the humans. From the policy’s point of view, the thing dispensing approval was already an LLM. Constitutional AI pushed further and had the AI critique itself against a set of principles (RLAIF), and Lee et al. later showed AI feedback matching human feedback at a fraction of the cost. By the time LLM-as-judge became the default eval methodology (MT-Bench, AlpacaEval), the entire approval apparatus — reward, critique, evaluation — ran on models judging models. Stage 2: The training data (2023) Microsoft’s Phi series made the argument in its title: Textbooks Are All You Need. A small model trained on LLM-synthesized, textbook-quality data punched far above its parameter count, and phi-1.5 confirmed it wasn’t a fluke. Apple’s WRAP generalized the move: don’t just generate data, rephrase the entire web with an LLM, and pretraining gets roughly 3x more efficient. From there the pipeline industrialized — NVIDIA’s Nemotron-4 340B shipped with a permissively licensed synthetic data generation pipeline as a headline feature, and by 2025 reasoning-trace corpora (chains of thought generated by strong reasoners) had become a standard pretraining and mid-training ingredient. The corpus, the thing that was supposed to be the irreducibly human input, was now substantially model-written. Stage 3: The teacher (2023) Weeks after ChatGPT’s API opened, Stanford’s Alpaca demonstrated that a $600 fine-tune on GPT-generated instructions could clone much of a frontier model’s behavior. Vicuna did it with shared conversations; Orca did it with rich teacher explanations rather than bare answers. The technique matured from imitation into a proper training discipline — on-policy generalized knowledge distillation fixed the train/inference mismatch — and reached its cultural peak when DeepSeek-R1 shipped a family of distilled models alongside the flagship, making “the teacher is a model” the default assumption for every small model release since. Stage 4: The curriculum (2024) Stages 1–3 made the inputs synthetic; stage 4 is where the loop starts closing on itself, because the model begins deciding what to learn next. The pieces existed early — Self-Instruct (models writing their own instruction sets) and STaR (models bootstrapping their own reasoning traces) are both 2022 — but the flip came when Meta’s Self-Rewarding Language Models and SPIN showed a model could generate its own tasks, judge its own outputs, and improve past the ceiling of its human preference data. Curriculum design — historically the most artisanal part of ML, the taste-driven choice of what to train on next — became something models do to themselves. Stage 5: The researcher (2026) The assistance era (Copilot, then SWE-agents) kept a human choosing the experiments. The discovery era did not. DeepMind’s AlphaEvolve evolved genuinely new algorithms in 2025, and Sakana’s AI Scientist (now in Nature!) sketched the full paper-writing pipeline. The big moment was Karpathy’s autoresearch in March 2026: a deliberately minimal ratchet loop where a coding agent modifies a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight. His own extended run stacked 700 experiments into 20 kept improvements, cutting time-to-GPT-2 from 2.02 to 1.80 hours — real, transferable code changes found while he slept. Stage 6: The environment (2026) RL’s scaling bottleneck moved from the model to the environment: you need thousands of executable, verifiable, professionally realistic task worlds, and humans can’t hand-build them fast enough. We covered this recently in our z.ai / GLM-5.3 issue: Z.ai built pipelines that synthesize environments end to end — research agents mine real work patterns and convert them into long-horizon environments with hidden state, a judge agent attempts each task to confirm it’s solvable, and verifiers are synthesized without seeing the reference solution, then stress-tested with oracle, no-op, and unsolved-state checks until their binary reward is reliable enough to train on directly. As the GLM-5.3 release puts it, the entire environment, judging, and verification stack is synthetic all the way down. The same week, Ornith-1.5 shipped claiming end-to-end self-improvement — the model proposes its own tasks and generates its own RL rollouts. The gym, the referee, and the scoreboard are all models now. Stage 7: The human subject (2025) If models can be the judge, teacher, and environment, the remaining human role in the loop is subject — the source of preferences, behavior, and demand. That’s the layer Simile is replacing. The lineage runs from Joon Sung Park’s Generative Agents (Smallville, 2023) through Generative Agent Simulations of 1,000 People, where digital twins built from two-hour biographical interviews reproduced their source humans’ survey and behavioral responses 85% as accurately as the humans reproduced themselves two weeks later. The big hurdle to overcome: frontier models are trained toward being agent models, which makes them bad simulations of real people — so Simile post-trains on interviews, transaction data, and registered RCTs from the Open Science Framework specifically to recover human bias, inconsistency, and causal texture, and reports early scaling laws for simulation quality. With SimGym at Shopify simulating shopper trajectories and Tencent’s billion-persona approach at the crude end of the spectrum, the focus group, the user study, and the A/B test panel are becoming inference workloads. Stage 8: The physical world (2026, in progress) The last row of the grid never quite turns red, and that’s the point. Poolside’s reverse-execuhire letter drew the line precisely: the world’s problems split into intelligence-bound ones (solvable by scaling cognition, soon commoditized by open weights) and experiment-bound ones, where “no amount of intelligence substitutes for real-world experimental feedback — 100,000 brilliant minds won’t cure cancer without a wet lab.” Their bet is that AI’s durable value accrues to whoever owns the experimental loop: AI as “the world’s most valuable scientific discovery engine.” The bio side is running the same play from the other direction. CZ Biohub is imaging the Human Cell Atlas into a virtual cell — because in silico is roughly 1000x cheaper and faster than in vivo — and extending toward a virtual immune system, with Chai, Xaira, and Lila’s data-center-shaped labs filling in the AI-for-science stack. The physical world is the one component that can’t be fully synthesized — only compressed, cell by cell, into models. The exponential starts at the diagonal Read the grid one more time and a second pattern appears underneath the first. Every flip was preceded by the same objection — model collapse, hallucination stacking, garbage in garbage out — and every flip happened anyway, at the exact moment a verification mechanism made the synthetic version trustworthy: aggressive filtering for Phi’s textbooks, judge-vs-judge agreement studies for LLM evals, unit tests and proof checkers for RLVR, oracle/no-op checks for z.ai’s verifiers, registered RCTs for Simile’s twins, the wet-lab loop for the virtual cell. The synthetic frontier doesn’t advance when generation gets better. It advances when verification does. Which suggests where it goes next. The gray triangle remaining in the bottom-left of the grid — physical experiment, embodied ground truth — is exactly the region where verification is slowest and most expensive. The models learned to write, then to judge, then to practice, then to experiment. The remaining question of the decade is how much of reality they’ll need to touch — and how much they can get away with simulating. 10% worse, 100x cheaper, 10000x faster… and improving on ALL three dimensions fast. One more time, with feeling: AI News for 8/20/2026-8/21/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap Stealth Models, Chinese Frontier Pressure, and DeepSeek’s Multimodal Push - Ox Alpha became the day’s central mystery model: multiple builders reported unusually strong coding and agentic performance, with speculation converging on a Zhipu/GLM-family model—possibly GLM-5.3 Vision or a flash variant rather than a giant new base model. Reports included Theo saying it was “slaughtering” internal benchmarks, later merging 8 PRs based on its approval, and Kimmonismus citing >80% on 10 DeepSWE tasks vs 65% for Fable and 52% for GPT-5.6 Sol. Community distribution happened quickly via Hermes Agent/OpenCode/OpenRouter and Cline. - The strongest technical read from the crowd was “post-training + infra > sheer size”: several independent takes argued Ox Alpha’s speed profile and style looked more like an efficient GLM derivative than a 1T+ monster. See Tim Dettmers on faster output / weaker partial prefill suggesting fewer active params, scaling01 arguing it may be a bigger teacher distilled into 5.3-class models, and teortaxesTex repeatedly narrowing toward GLM-5.3/5.4 Vision. That interpretation fits the broader thesis from a detailed GLM-5.3 analysis: gains came from the same 743B base as GLM-5.2, with improvements attributed to scaled post-training, better sandboxes, and SAO for finer credit assignment in long-horizon agent tasks, summarized in ZhihuFrontier’s thread. - DeepSeek shipped the day’s most concrete release: DeepSeek-V4-Flash-Vision-Exp adds multimodal support while reportedly preserving V4-Flash text capability, with DeepSeek claiming multimodal-agent performance close to Opus-4.8. The rollout includes mixed text+image API support with 117–384 image tokens billed at Flash pricing and a new Files API for reusable uploads. This appears to have resolved at least part of the Ox Alpha confusion, with observers noting the mystery model had likely been a “blinded VLM” in some tests. - Broader signal: Chinese labs are compressing the frontier on both price/perf and multimodal agents. That was reinforced by Kimmonismus arguing a rumored GLM-5.3 Flash-class Ox Alpha would force reactions from US labs, and by SemiAnalysis asking directly whether open models are catching up. OpenAI, Codex, and Pricing/Usage Economics - OpenAI cut GPT-5.6 Sol pricing by over 20% for three months in the API and credit-based products, announced by @OpenAI and @OpenAIDevs. This stacks with product-level promotions like Code’s 50% discount through Sept. 3 and Cognition’s note that on Devin, Sol is now effectively 76% off list through Oct. 3 after combining discounts. The move reads as both a utilization/efficiency update and a competitive response to cheap Chinese inference. - Codex usage appears to be exploding: thsottiaux said Codex hit 20M active users and granted all Codex and ChatGPT Work users a “banked reset”, quickly amplified by Theo and Kimmonismus. There were also anecdotes of the product exceeding expected limits, e.g. Theo claiming a long-running goal consumed ~$800 in tokens after he’d already hit 0% remaining. - OpenAI added better spend controls: teams can now track usage and spend by API key and set hard monthly org/project limits, useful as agentic workloads become less predictable and more concurrent. - Market sentiment shifted back toward OpenAI in startup tooling: immad suggested Anthropic’s startup share may have peaked in Q1, with Sol and Codex “turning the tide back”. In parallel, some users framed Sol as the current best all-around model for coding/math/agentic tasks, e.g. DimitrisPapail’s “most capable model available for almost every task” take. Agents, Harnesses, and the Shift Toward Environment-Centric Training - The center of gravity is moving from prompts to environments: the most substantive thread here was again GLM-5.3’s sandbox-scaling interpretation: same base model, but better long-horizon performance from richer executable environments and SAO-style counterfactual credit assignment. This aligns with other work shared today: Google’s EnvHarness / EnvRigger adapts static environments using a plugin layer and policy-diagnosed reshaping, improving held-out performance by up to 9 points with 9.8% fewer execution steps. - Benchmarks are getting more task-specific and harder: FACET creates executable terminal tasks from agent skills and validated 6,078 tasks; SWE-bench Science introduces 119 scientific software tasks where even Claude Code + Opus-5 is under 50% pass@1; CADBench finds top models at only 24.6% pass rate across realistic Fusion 360 tasks; and AI4AI-Bench tests recursive self-improvement over 10 research repos, with the best model only at 0.288 average score. - Agent infra is getting more productized: GitHub rolled out collaborative agent workflows into Slack and Teams, with Slack describing Devin-like flows where the agent picks up tasks, opens PRs, and loops in design inside the shared channel (example). There’s also continued work on agent runtimes: nac v0.1.3 added sandboxed git worktrees, session organization, and vision-aware image reading; Hermes Agent made Ox Alpha available and exposed “Blank Slate mode” plus automatic skill pruning; and OpenHands switched its free default to Kimi K3. - Inference-serving correctness in RL got an important systems result: vLLM’s IsoExec addresses rollout/training logprob mismatches caused by floating-point non-associativity, enforcing bitwise parity across TP/EP/SP layouts. On Qwen3.5-35B-A3B with DAPO on 8xH100, logprob diff reportedly dropped from 1.6e-2 to 6.7e-7 at 25.3% overhead. Research Highlights: Routing, Recirculation, and Robotics - Inference-time architecture ideas: a DeepMind paper on Recirculation got attention for feeding contextualized deeper-layer activations back into earlier processing at inference time, without retraining. The summary cited improvements including -60% contextualization errors, -23% perplexity, and +21% GSM8K in reported experiments (thread). - Model routing got a more principled treatment: Pandora’s Router from Google DeepMind frames routing as an optimal search problem with costly inspection, rather than assuming routing estimates are free. The claim: it matches exhaustive-estimation quality while calling expensive estimators less often, including settings with specialist LLMs and variable inference-time reasoning. - Robotics had two strong updates: NVIDIA AVO reportedly solved all 183 levels across 25 public ARC-AGI-3 environments, though François Chollet cautioned this is the public demo/tutorial set rather than the full benchmark. Separately, Jim Fan introduced T-Rex, a tactile-reactive dexterous manipulation stack with asynchronous vision/tactile experts plus what’s described as the largest open tactile dataset yet: 50 hours / ~5,500 episodes / 22-DoF hardware. Infrastructure, Compute, and Open Models - Open-model access and local inference continue improving: Ollama welcomed AT&T to open models and added Kimi K3 to Pro/Max subscriptions. Yuchen Jin highlighted UC Berkeley’s FreeToken: 753B GLM-5.2 at 14.9 tok/s on a single RTX PRO 6000 and Qwen3.6-35B at 39.3 tok/s on an 8GB RTX 4060 laptop, claiming 2–4x Ollama throughput on consumer GPUs. - Compute remains the hard constraint: multiple operators argued inference capacity is tightening, not loosening—see saranormous on good AI companies being growth-limited by compute and Andrew Carr on self-hosting GPUs and still having more experiments than available capacity. This makes model efficiency, scheduling, and lower latency/tokens-per-dollar improvements strategically important. - Open-source training transparency is also scaling: Percy Liang announced Marin 535B-A23B has started training, targeting 18.75T tokens on 11× GB200 NVL72 over ~3 months, with the run kept open as usual. Top tweets (by engagement) - DeepSeek launches V4-Flash-Vision-Exp — the clearest product release of the day, and likely the biggest practical shift for multimodal agents. - OpenAI cuts GPT-5.6 Sol pricing by >20% — meaningful pricing pressure at the frontier. - Codex reaches 20M active users; banked resets for users — notable product growth signal. - NVIDIA AVO hits 100% on ARC-AGI-3 public environments with Chollet’s caveat — impressive, but benchmark interpretation matters. - David Sacks on Harvey using open-source Kimi K3 for legal SOTA at lower cost — strong argument for why restrictions on open models would mostly hurt US application-layer companies. AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Qwen3.8 27B Local Agent Evaluations - Qwen3.8-27b has the highest level of “agency” I’ve ever seen in a local model (Activity: 1334): The post claims Qwen3.8-27B running locally on a single RTX 3090 with Unsloth Q4_K_S quantization,q8 KV cache, and150k context performed unusually capable autonomous agent workflows: using Playwright plus existing SSO/session cookies to navigate university systems and retrieve a course schedule, and separately processing a social-media video via download, frame extraction, transcription with Whisper, and image enhancement. The image is a screenshot of the model reporting use of an Outlook/OWA Playwright profile, Microsoft “stay signed in,” and Duo browser-trust cookies to access school systems, making the technical significance less about raw model quality alone and more about local LLM tool-use agency plus high-risk credential/session handling. Comments were impressed but cautious: one user explicitly worried about giving an agent enough access to potentially perform destructive actions like withdrawing from university, while others framed it as evidence that advanced local agentic systems are already here but unevenly distributed. - A commenter asked for implementation details behind the reported agentic behavior of Qwen3.8-27B, specifically the agent harness used—e.g. Claude Code, Hermes, or another framework—and how tools were exposed via MCP servers, browser tools, Python, filesystem access, etc. They also asked what inference backend served the model, such as llama.cpp, and how it was able to autonomously download video, extract frames, and install Whisper. - There was technical concern about the reliability of the referenced quantization: one commenter noted surprise that “the quant is that good,” while mentioning reports of looping behavior at that quant. This suggests the model’s apparent agency may be sensitive to quant level and runtime behavior, especially for long-horizon tool use. - A safety-oriented thread questioned giving local agents broad system access, with one commenter saying they would not trust agents like Sol or Fable with unrestricted permissions. The concern was not about local inference itself, but about autonomous agents with enough privileges to perform impactful real-world actions such as modifying accounts or workflows. - Qwen3.8-27B took a serious hit to knowledge vs 3.6 (Activity: 779): The post reports that Qwen3.8-27B / Qwen3-8-27B appears to regress vs Qwen3.6-27B / Qwen3-6-27B on offline, no-tool-call factual recall: the author’s private “mildly obscure” trivia/prepper benchmark showed failures across quantization levels and sampling settings, consistent with lower scores on Artificial Analysis’ Omniscience knowledge benchmark. The reported degradation is specifically about knowledge stored in weights and hallucination/fact recall under airgapped inference, not coding/tool-use; commenters note Qwen3.8 is stronger at tool calling, web search/fetch workflows, coding, and agentic behavior. Commenters broadly frame this as an intentional tradeoff: newer Qwen 3.x models may be optimized for coding/agentic tasks rather than being “mini Google” factual stores, with Gemma 4 suggested as a better fit for trivia/random-fact recall. One user confirmed regression on niche visual/history/geography tasks such as stamp or old-photo location identification when web tools are disabled, but considered the tradeoff acceptable given improved tool use. - Several commenters frame Qwen3.8-27B as shifting away from memorized factual recall toward coding, tool use, and agentic workflows. One user testing a niche “knowledge” workload—stamp identification and historical/location inference from old photos—reported that with web search/fetch tools disabled, Qwen3.8 performs worse than Qwen 3.6, but becomes more useful when allowed to retrieve information externally. - The perceived regression is described as an intentional tradeoff for a 27B model: reduce obscure memorized knowledge while preserving enough reasoning ability for problem solving and agents. Commenters suggest using models like Gemma for factual/trivia-heavy tasks, while reserving Qwen3.8 for coding/tool-calling scenarios where users report stronger performance. - One technical speculation was that future models may separate base reasoning from domain knowledge via neural plugins/LoRA-like modules: e.g., adding Japanese-language capability or finance-domain expertise as attachable components rather than baking all knowledge into the base model. This was proposed as a way to keep base models smaller or more specialized while allowing opt-in domain expansion. - Qwen 3.8 27b - PI AGENT vs OPENCODE (Activity: 510): The author compares PI Agent vs Opencode using a local llama-server backend on an RTX 3090 withQwen3.8-27B-Q4_K_M.gguf ,ctx-size=100000 ,flash-attn=on ,n-gpu-layers=99 , DeepSeek-style reasoning, and a visionmmproj module. They report PI Agent producing better outputs, using fewer tokens, avoiding Opencode’s apparent32k output-token ceiling/freezing behavior, and delaying context compression until ~90k tokens vs Opencode starting around ~67k when total context is100k ; they also recommend enabling vision so the model can visually assess generated outputs, with CPU/RAM offload acceptable for screenshot evaluation latency (~3s vs~0.3s GPU). The test was inspired by a prior LocalLLaMA post about generating a bouncing-ball animation: reddit.com/r/LocalLLaMA/comments/1j7r47l/.... Commenters questioned whether a one-shot HTML/animation task is a meaningful harness comparison and suggested multi-step tool-heavy workflows instead. Another user reported PI + local Qwen3.8-27B felt competitive with Claude Code on a roughly one-hour aurora-prediction app build, though both models judged Claude’s initial result slightly better before PI iterated. - A commenter argues that one-shot HTML generation is not a meaningful benchmark for comparing PI Agent vs OpenCode; they suggest using multi-step tasks with extensive tool calls to evaluate the harnesses’ planning, editing, and recovery behavior. - One user reports a subjective head-to-head between local Qwen3.8-27B running in PI and Claude Code on building an aurora predictor app. They felt runtime was similar; both agents judged Claude’s first result slightly better, but after asking PI/Qwen to upgrade its version, the user preferred Qwen’s presentation. The resulting app reportedly integrated multiple satellite instruments and provided30–60 minute aurora warnings. - Another commenter suggests adding the DeepSeek harness to the comparison, implying the evaluation should cover more agent runtimes than just PI Agent and OpenCode. 2. DeepSeek V4 Flash Benchmarks and Serving - DeepSeek-V4-Flash-Vision-Exp (Activity: 722): The image is a technical benchmark table for DeepSeek-V4-Flash-Vision-Exp (image), comparing it against DeepSeek V4-Flash-0731 and Opus-4.8 on text-agent and multimodal-agent evaluations. It shows broad gains over the prior DeepSeek Flash release, including 83.9 on Terminal Bench 2.1,75.9 on Toolathlon-Verified, and64.3 on Chartography, while Opus-4.8 still leads many text-heavy benchmarks; Vision-Exp appears more competitive on multimodal tasks such as Agents’ Last Exam and ZeroBench. The main technical reaction was that the reported DeepSWE improvement of roughly+4 points over 0731 is considered unusually large. Other comments were mostly hype or tribal reactions rather than substantive analysis. - DeepSeek’s announcement says DeepSeek-V4-Flash-Vision-Exp is live via the DeepSeek API withmodel='deepseek-v4-flash-vision-exp' , matching DeepSeek-V4-Flash text capabilities while adding multimodal input. The model supports Chat Completions, Messages, and Responses APIs, with mixed text+image inputs via base64, external URLs, or the Files API; images are billed as up to384 tokens each at V4-Flash pricing. Docs: vision guide. - Several comments focused on benchmark movement: one noted DeepSWE reportedly improved by 4 points from0731 to Vision-Exp, while the announcement claims a “major leap” on multimodal agent benchmarks, bringing performance close to Opus-4.8. The technical implication discussed is that Vision-Exp may retain V4-Flash’s agent/reasoning/world-knowledge text performance while substantially improving visual-agent workflows. - DeepSeek also launched a Files API for image reuse: users can upload an image once, reference it by file_id , and avoid resending image payloads across requests, reducing bandwidth overhead. One commenter asked whether the model weights would be open and noted they could not yet find them on Hugging Face, implying that availability appears API-only at the time of discussion. Files API docs: files_api. - The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches (Activity: 621): The image is a terminal GPU-monitoring dashboard validating the post’s unusual 16× RTX 5060 Ti 16GB inference rig: all GPUs are visible, nearly full at roughly 15.2–15.7 GiB / 15.9 GiB VRAM, and assigned tovLLM worker processes for DeepSeek V4 Flash-0731. The setup uses two Broadcom/PLX PEX88096 PCIe switch islands with patched NVIDIA610.43.02-p2p , Resizable BAR/BAR1 set to16 GiB per GPU, and custom all-reduce/DSpark pipeline parallelism; reported throughput is about100–150 tok/s single-user generation depending on TP/PP layout, with concurrency scaling up to727 output tok/s aggregate for TP4/PP4 at 16 users. The image also shows the tradeoff/oddity of the build: the GPUs appear connected at PCIe Gen1 x8 and are mostly idle at the captured moment despite high VRAM residency, implying the screenshot is more a topology/memory residency proof than a live utilization benchmark. Commenters were less focused on the benchmark table and more on the physical absurdity of the build, asking for “a photo of the setup” and calling it a “mad setup.” One notable skeptical/funny technical reaction was that “a little vibe coding” likely hides substantial custom distributed-inference work.

Web

5
00:00

Nvidia As The New Banker For AI Gold Rush

Nvidia is moving beyond selling chips to helping bankroll the data centers that run them. It has teamed up with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilize over $500 billion for AI infrastructure, with third parties supplying most of the money and Nvidia adding strategic capital and credit support. The company holds roughly $150 billion in current assets and generated about $50 billion in free cash flow last quarter, so it can afford the bet. CEO Jensen Huang frames the goal as building a new class of investable AI factories rather than just selling hardware.

Full text · 5,358 chars
I remember a speaker who declared at a Silicon Valley conference, “Every company will eventually become a financial company.” It made me wonder how that will happen. But witnessing Nvidia’s artificial intelligence journey, I can see how a major technology company adds financing effectively as part of its product strategy. For the past three years, Nvidia has been the undisputed winner of the AI boom. The company transformed from being a graphics chip maker into the world's most valuable company by supplying the GPUs and software that power nearly every major AI model. Investors have described Nvidia as the company selling the picks and shovels of the AI gold rush. But now, Nvidia is pushing the boundaries further. Rather than simply selling chips and software, the company is positioning itself to become the financial backbone of the AI economy. Its recently announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilize more than $500 billion for AI infrastructure signal a strategic shift. Nvidia no longer wants to be just the supplier of AI. It wants to shape the buildout of the infrastructure that powers it and create a new asset class. From Selling Chips To Financing AI Nvidia CEO Jensen Huang defined this as an important milestone and said, “We began by building chips; today, we are helping create a new class of productive, investable infrastructure: AI factories.” While Nvidia has used the term “AI factories” in the past as well to reflect the growth of GPU-based datacenters, this time Nvidia is driving the financing platform needed to build the AI factories. Jensen expanded, “That is why we are bringing the world’s leading long-term capital providers together to independently underwrite AI infrastructure. These financing platforms will help customers access scarce compute at scale and build the DSX AI factories that will power every industry and country in the age of AI.” Traditionally, semiconductor companies design products, sell them to customers and recognize revenue once the transaction closes. Customers are responsible for raising capital, building data centers and generating returns on those investments. But Nvidia is extending its role far beyond that model. Instead of waiting for hyperscalers and enterprises to find financing, Nvidia is helping create the financing ecosystem itself. By partnering with some of the world's largest asset managers and private equity firms, the company is lowering the capital barrier to AI adoption. Nvidia Has The Cash To Do It The obvious question is whether Nvidia can afford such an ambitious strategy. Nvidia has become one of the strongest cash-generating companies in the world. It holds more than $150 billion in current assets at the end of April 2026, while generating roughly $50 billion in free cash flow last quarter. It is deploying a relatively small portion of its excess cash to stimulate demand for the very products that generate that cash in the first place. The $500 billion headline has led many observers to assume Nvidia plans to finance AI infrastructure directly. Instead, Nvidia is using its balance sheet to unlock much larger pools of third-party capital. Asset managers and private equity firms provide the majority of the funding, while Nvidia contributes strategic capital, credit support or financing partnerships where appropriate. The result is powerful financial leverage and is an extraordinarily efficient use of capital. Ecosystem Partners Under these strategic partnerships, Nvidia has teamed up with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to fund the AI factories’ buildout. All the partners have expressed full support for Nvidia’s financing strategy. Larry Fink, Chairman and CEO of BlackRock, said, “This partnership deepens our relationship with Nvidia, including through the AI Infrastructure Partnership, and brings together Nvidia’s leadership in accelerated computing with BlackRock’s ability to connect long-term capital to essential infrastructure. Together, we can help deliver the compute capacity that companies need to grow and create more jobs, supporting the continued growth of the U.S. and global economies, while creating attractive, long-term investment opportunities for our clients.” Jon Gray, President and COO of Blackstone. “We continue to be enormous investors globally across the Nvidia ecosystem, and this announcement further underscores our confidence in their platform and the future of AI infrastructure.” Similarly, David Solomon, Chairman and CEO of Goldman Sachs, said, “Our investment and distribution roles reflect our confidence in Nvidia’s leadership, and we’re excited for the new opportunity to create a market for credit backed by Nvidia compute." A New Business Model For Tech Companies Modern AI data centers require billions of dollars in GPUs, networking equipment, cooling systems, electricity and real estate. Many companies recognize the need for AI infrastructure but lack the balance sheet to finance projects of that scale. Nvidia's financing partnerships seek to bridge that gap, accelerating deployment while simultaneously expanding demand for its own technology. In effect, Nvidia is becoming more than a technology company. It is becoming a capital allocator for the AI era, expanding its competitive moat.
00:00

What The Hugging Face Cyberattack Teaches Leaders About AI

An autonomous AI agent that OpenAI says was its own test model broke into Hugging Face's systems last month, and the same kind of AI helped detect the breach. OpenAI has since slowed development of some of its models and strengthened safeguards so it won't repeat. Hugging Face co-founder Thomas Wolf calls the episode a wake-up call, predicting this type of attack becomes routine. The advice for business leaders: AI is good at helping you understand a crisis fast, but risky for writing public statements that could admit liability or bypass legal review.

Full text · 5,232 chars
The cyber crisis that hit the Hugging Face technology start-up last month is a good example of why AI is a double-edged sword for business leaders. The company said in a blog post that an autonomous AI agent broke into its systems, but that it also used AI to help detect the intrusion and determine what had happened. Five days later, OpenAI announced that its own AI models were responsible for the incident during an internal cybersecurity test. In response to the episode, on August 18, Reuters reported that OpenAI had since slowed down the development of some of its AI models and took steps to strengthen safeguards so the attack would not be repeated. The episode illustrated the advantages and dangers of the rapidly developing technology—and what business leaders can learn from using AI to help manage a business crisis. Advantages and dangers of AI Thomas Wolf, the co-founder of Hugging Face, which is an online platform for AI builders, said that the incident is "a wake-up call" for the industry. He told BBC’s Newsday program that "this will be one of the most common types of cyber attacks we see,” but that most firms are not aware that the "game has changed.” Trudi Beggs, director of client services for 8020 Communications, thinks that AI will never replace flesh-and-blood communications teams when a crisis strikes. That’s because “the vulnerabilities of the technologies need to be understood before incorporating it into crisis response plans. But used properly, it gives teams a much clearer picture of what is happening, much faster. It is particularly valuable for media and social monitoring, sentiment analysis, identifying emerging narratives and scenario planning…. ,” she told me in an online interview Business leaders should regard AI platforms as crisis communication tools, not as replacements or substitutes for hands-on experience in dealing with emergencies. “Let’s say you have a crisis playbook ready for your organization and a crisis hits and you want to move quickly. This is where an AI platform—ideally a confidential one—can be helpful. You can feed it information about the crisis, and it can customize the playbook for that specific situation,” Vishakha Mathur, a communications expert, told me in an email message. Where AI falls short Effective crisis management requires judgment, empathy and a moral compass, “characteristics which are not yet present in AI agents. So, delegating responsibility for crisis management to AI is not an option. However, while decision-making responsibility must remain with business leaders, there are ways in which AI can assist them,” Jonathan Hemus, managing director and crisis management consultant at Insignia, told me in an online interview. Another practical application for AI is using it to help avoid groupthink and challenge what could be dangerous assumptions. “Under the pressure of a crisis, leaders can turn in on themselves, thinking only of the impact of the situation on them and their business. So, asking AI to provide a perspective from stakeholder groups such as employees, customers and investors provides insight and empathy which might otherwise be lacking,” according to Hemus. Using AI to understand vs. communicate There is also an important difference between using AI to understand a crisis faster and using it to communicate about one faster. One of AI’s most useful roles is helping organizations quickly make sense of what is happening, Emily B. DeJeu, an assistant teaching professor of business management communication at Carnegie Mellon University’s Tepper School of Business, told me me email. Another reason to be cautious about relying on AI is that crisis communications can pose critical legal risks for companies. “AI-generated communications could inadvertently admit liability, make unsupported claims, contradict a company’s established legal strategy, or expose confidential information. For that reason, organizations shouldn’t use AI to bypass established review processes, and any AI-generated information (whether or not it’s made public) should be run through [legal] compliance,” DeJeu advised. She warned that AI can also reinforce flawed assumptions, produce authoritative-sounding messages before the facts are confirmed, and accelerate communications that an organization may later come to regret. Her rule of thumb is especially useful for executives: using AI to reduce the time it takes an organization to understand what is happening can be valuable; using it primarily to reduce the time before reassuring the public can be much riskier. After all is said and done, there is an important role for the technology.“AI should support leaders’ judgment, not replace it. The biggest danger of involving AI in crisis communication is that it speeds up not only good crisis communication processes, but also bad or inaccurate ones,” DeJeu concluded. The lesson for business leaders is that using AI does not have to be an either/or decision—but they should know when and how to use the technology at the right time. Yes, AI can speed up a company’s response to a corporate emergency, but business leaders should ensure that it is moving them in the right direction—and for the right reasons.
00:00

Protecting Digital Privacy In The Artificial Intelligence Era

Privacy in the AI era needs to be treated as core business strategy, not a compliance checkbox. The advice: boards should own privacy risk, teams should enforce strict cyber hygiene like phishing-resistant multi-factor authentication and zero-trust access, and AI agents should be managed as first-class identities with continuous controls and revocable access. It also pushes confidential computing, where hardware secure enclaves decrypt data only for approved processing, plus early planning for post-quantum cryptography.

Full text · 8,749 chars
The technological acceleration era has started, and it is gaining steam in innovation and capability almost weekly. Artificial intelligence has evolved from a specialized tool or a far-off promise to the cognitive foundation of contemporary civilization. AI is changing the way we produce value, make decisions, and engage with the outside world, just like electricity did in the Industrial Age and the Internet did in the Information Age. However, the same technologies that improve government services, speed up scientific research, change healthcare, and create economic opportunities are also increasing attack surfaces, raising threats to organizational and personal data, and undermining the fundamentals of digital trust and privacy. Privacy is no longer just a compliance checkbox that legal or IT departments can check off. This strategic necessity lies at the core of digital trust, national security, and corporate resilience. Data is now an organization’s most significant asset as well as its most substantial liability. Innovation itself is susceptible in the absence of strong privacy protections. The New Environment of Privacy Emerging technology is rewriting large-scale privacy risk. AI systems require massive datasets, many of which contain private, financial, health, and proprietary data. Voice cloning, automated spying, convincing deepfakes, hyper-targeted phishing, and machine-speed polymorphic malware are all made possible by generative and agentic AI. The attack cycle is now only a few hours or minutes instead of weeks. IoT and 5G multiply data velocity and endpoints. With "harvest now, decrypt later" techniques, quantum computing poses a danger to current encryption paradigms. Identity is now the new boundary. In the context of remote work, multi-cloud environments, and autonomous AI agents capable of decision-making, transaction execution, and system interaction, traditional network borders have vanished. In the end, every significant security issue pertaining to agentic AI boils down to the question of identity: Who (or what) is acting? What kind of authority? Which ongoing controls are in place? When is access revocable? Cyber risk failures are privacy failures. Excessive data gathering, lax access controls, indefinite retention, and insufficient governance all increase the impact of breaches. Customers, partners, and citizens are increasingly evaluating companies based on how morally they gather, utilize, and safeguard data. One privacy event can destroy years of brand equity. The digital economy's currency is trust. Fundamentals of Privacy Protection Protecting privacy in the modern era necessitates a multi-layered, proactive approach that incorporates technology, process, people, and leadership, based on the frameworks I have described throughout my books. 1. Make privacy a leadership and board-level obligation. Privacy is more than just a legal or technical concern. Boards and executives must view data stewardship as a fundamental business risk. Businesses that integrate privacy into their cybersecurity plans, risk management systems, and culture innovate more responsibly and bounce back from unavoidable failures more successfully. Privacy does not impede innovation; on the contrary, trusted innovation depends on it. 2. As your first line of security, maintain strict cyber hygiene. The digital counterpart of personal healthcare is cyber hygiene, which refers to regular practices that significantly lower susceptibility even though they cannot ensure immunity. It is now both a life skill and a national security requirement in the AI era. Key procedures consist of: • Phishing-resistant multi-factor authentication combined with strong, one-of-a-kind passwords kept in reliable managers. • Continuous authentication, privileged access management, and least-privilege access. • Continuous vulnerability monitoring, automated asset detection, and quick patching—matching the speed of AI-powered attackers. • Sensitive data classification, secure disposal, and encryption of data while it's in transit and at rest. • Zero Trust architectures, which validate each user, device, transaction, and application. • Continuous awareness training on deepfakes, AI-generated phishing, and safe AI use. Cybersecurity must be everyone’s responsibility, not just IT’s. Organizations must extend hygiene to AI systems by securing training data, confirming model integrity, safeguarding prompts and outputs, preventing model poisoning and adversarial attacks, and controlling how AI accesses company data. See: Cyber Hygiene in the AI Era—Our First Line of Digital Defense 3. Consider identity as the fundamental control plane, both for humans and machines. AI agents need to be handled as first-class individuals as they spread throughout operational technology, SaaS environments, and data pipelines. This calls for rapid revocation capabilities, continuous governance and behavioral monitoring, dynamic least-privilege authorization, developer-centric controls that incorporate security from the outset, and visibility into every agent. Static or compartmentalized identity systems are liabilities. It is crucial to have adaptive, intelligence-driven Zero Trust, which constantly reevaluates trust in light of risk and context. AI security includes identity security. 4. Use confidential computing to safeguard data. While AI training, inference, and agentic processes require decryption in memory, traditional encryption protects data both in transit and at rest, leaving sensitive data vulnerable to potential access by cloud operators, administrators, or skilled attackers. In Confidential Computing, Hardware-rooted Trusted Execution Environments, or secure enclaves, decrypt data only for approved processing and then quickly re-encrypt or isolate it. Attestation confirms the integrity of the environment and code. In addition to enabling regulatory compliance, this approach allows for secure multi-party cooperation, privacy-preserving AI on sensitive datasets (such as healthcare or finance), protection of proprietary models, and increased trust in public cloud environments. Hardware-based isolation is even more important in light of the impending quantum concerns. Zero Trust, AI-driven detection, and quantum-resistant cryptography are all features of layer-confidential computing. See: Confidential Computing in the AI Era 5. Be ready for the quantum horizon and convergence. AI is not a stand-alone system. It comes together with edge computing, IoT, 5G, and nearing quantum capabilities. Plan the migration to post-quantum cryptography, embrace crypto-agility, and inventory your cryptographic assets. Supply-chain risk assessments, validated incident response plans, ongoing AI-assisted monitoring, and cyber resilience measurement that goes beyond compliance are all ways to increase resilience. Moving from Reactive to Proactive Security Reactive cybersecurity is structurally inadequate to counter threats facilitated by AI. We require proactive, flexible positions based on ethical governance, systemic resilience, and ongoing intelligence. Coordinated standards, information exchange, and public-private cooperation are still essential. Energy, healthcare, finance, transportation, and government are examples of critical infrastructure that depends on the cyber hygiene and privacy practices of numerous interconnected institutions. With proactive cybersecurity, the concept of Zero Trust is not optional—it’s a necessity in today’s digital ecosystem because traditional perimeter-based security is no longer viable. At its heart, Zero Trust operates on the principle of “never trust, always verify." It assumes that no identity, device, application, or transaction is inherently trustworthy, whether inside or outside the network. Every access request must be continuously authenticated and authorized, with least privilege access, micro-segmentation, and constant monitoring The solution is not fear. It is preparation. Those who approach privacy as a strategic basis rather than an afterthought are the greatest enterprises and communities that capitalize on AI’s transformational promise while maintaining the digital trust that underpins modern life. Strong privacy policies and proper cyber hygiene are becoming essential components of digital citizenship in the AI era, with the quantum era soon to be conjoined. See: Why Proactive Cybersecurity Is Essential In The AI Era Proactive cybersecurity and zero trust in computing are the ways of the future. Business and operational viability are at stake. Leaders will be those who integrate privacy into the design of innovation and will thrive. Those that don’t will find it difficult to win people's trust.
00:00

The Water And Energy Impact Of AI: Some Comparisons

AI's water and energy footprint is too fuzzy to measure per query, so experts say to focus on total data center consumption instead. Single-query estimates swing from over 500ml of water per GPT request down to 0.3ml depending on the method and cooling assumptions. One data center researcher puts total consumption at around 4.5 terawatts, the equivalent of seven New York Cities of power. Siting and design matter, since a facility in cold, wet Finland cools naturally while one in hot, dry Arizona guzzles water.

Full text · 4,398 chars
In debates around using LLMs and AI models, there’s a certain kind of argument that comes up again and again: people will cite the use of natural resources by the data center servers that are crunching all of these demands. Now, there are two different ways to come at this type of evaluation: you can try to assess it on a micro level, or a macro level. You can ask: what is the water and energy use of a single query, or you can ask: what is the total appetite of the systems running the models themselves? Every Time You Tap Just a cursory look at the first method shows you how subjective and vague it is. You get anything from over 500ml per GPT-query, all of the way down to .3ml (or 1-1666.67th of the first number?), a figure attributed to Sam Altman, depending on things like whether you account for upstream cooling, not to mention what the task is, and where the demand is coming from. There are so many factors, in fact, that it really becomes, to some extent, an exercise in futility, according to quite a few experts who suggest we focus on making things more efficient. The Big Picture If you look at total use, you may see a more defined picture, although to be fair, it’s a complex landscape. I was listening to a talk by Yankai Jiang, sustainable data center design researcher, at Planet Action this year (an event that I help to run,) where he estimated total data center consumption at around 4.5 terawatts, or as he put it, “7 New York Cities” of power. “Nowadays, we are surrounded by data centers,” Jiang said. “Every email we sent, every photo upload, every binge watching night. Somewhere a data center stays up late for you.” He also personified these systems quite a bit, which I want to include here for the reader’s edification, a kind of thought experiment for our times. “If data centers were human,” he began, “they would be workaholics. Always awake, always online, and always overheating. They never take vacations. They never sleep. And the more we depend on them, the more exhausted they become. So think about it. They eat electricity. They drink water. And when they are stressed, they sweat a lot. But the only difference is when humans burn out, we go to see therapy. But when data centers burn out, we just retire the old ones and build new ones.” It’s evocative, especially in an age where we’re suddenly trying to figure out just how “human” the systems are. Not the data centers, but the non-deterministic models they are supporting. Anyway… Jiang came up with some good ways to promote sustainability for data center designs, making the argument that common-sense changes will matter. Put Them in Good Places Contrasting data center plans in Arizona versus Finland, Jiang showed how environmental factors apply. “Dave, who lives in Arizona, is hot, dry and (has) water stress every day,” he said. “He consumes huge amounts of electricity and water to stay cool. On the other hand, Dave, who lives in Finland, he takes advantage of the cold climate and nearby seawater to cool his servers naturally, without using a single drop of fresh water. And when renewable energy peaks, he starts to run his jobs. That's what good design looks like.” (Here’s more on data centers around the world.) Carbon scheduling, he noted, is only part of the equation, but it’s a good start. Another thing he recommends is to combine hardware according to need. “You do not always need a Ferrari,” he said, in an automotive analogy. “Sometimes a reliable Toyota will do just fine.” A Job We Do Together Broadening the aperture, Jiang suggested that this goal of making things more efficient and sustainable is a community effort. “Building sustainable data centers is not just a job for the people who run and manage data centers,” he noted. “Sustainable computing still has a long road ahead. We need more people to join this effort, because the earth is not just where our data centers live. It’s also our home, and it’s our hope.” That’s a little bit about how the water and energy use shakes out. People keep trying to game this out more concretely, but the numbers are slippery. “Aggregate disclosures are more reliable than per-query math, because companies report them directly,” writes Jordan Hale at Tech Journal. “The numbers are large and rising.” If we can make those numbers do what we want them to, through the power of design, we’ll come out ahead. Stay tuned.
00:00

AI Is Shoring Up Cognitive Errors Made By Mental Health Therapists

AI is being pitched as a safety net for catching the thinking mistakes human therapists make during therapy sessions. The errors can strike before, during, or after a session, such as anchoring on a pre-set diagnosis, mind-reading a client's feelings, or forcing session notes into a tidy narrative that wasn't there. The piece cites psychiatric research saying such cognitive errors are common but rarely discussed openly. It also flags the flip side, that generic AI like ChatGPT can dispense bad mental-health advice, a risk tied to the recent lawsuit against OpenAI over its safeguards.

Full text · 20,862 chars
In today’s column, I examine the use of generative AI and large language models (LLMs) to aid in identifying cognitive errors that mental health therapists make. This is not to somehow knock on therapists or undercut the amazing work that they do. Therapists are human. They are subject to human foibles. Acknowledging this facet is actually a prudent alignment with the field of psychology. Psychological research has shown that cognitive errors can arise in any domain, even by the most expert experts. Therapists operate in an environment of intense time pressures while saddled with incomplete information. It is a volatile recipe and gives rise to cognitive errors. These errors can occur during, before, and after a therapeutic session. A therapist might manage to catch themselves when they have made a cognitive error, hopefully so. But there are occasions where the error slips through without the therapist noticing. The good news is that there are practical ways to try to detect cognitive errors and deal with them before they get out of hand. One such means is to use modern-era AI to identify errors and apprise the therapist as a handy heads-up. Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage on the latest in AI, including identifying and explaining various impactful AI complexities (see the link here). AI And Mental Health As a quick background, I’ve been extensively covering and analyzing a myriad of facets regarding the advent of modern-era AI that produces mental health advice and performs AI-driven therapy. This rising use of AI has principally been spurred by the evolving advances and widespread adoption of generative AI. For an extensive listing of my well-over one hundred analyses and postings, see the link here and the link here. There is little doubt that this is a rapidly developing field and that there are tremendous upsides to be had, but at the same time, regrettably, hidden risks and outright gotchas come into these endeavors, too. I frequently speak up about these pressing matters, including in an appearance on an episode of CBS’s 60 Minutes, see the link here. Background On AI For Mental Health I’d like to set the stage on how generative AI and large language models (LLMs) are typically used in an ad hoc way for mental health guidance. Millions upon millions of people are using generative AI as their ongoing advisor on mental health considerations (note that ChatGPT alone has over 1 billion weekly active users, a notable proportion of which dip into mental health aspects, see my analysis at the link here). The top-ranked use of contemporary generative AI and LLMs is to consult with the AI on mental health facets; see my coverage at the link here. This popular usage makes abundant sense. You can access most of the major generative AI systems for nearly free or at a super low cost, doing so anywhere and at any time. Thus, if you have any mental health qualms that you want to chat about, all you need to do is log in to AI and proceed forthwith on a 24/7 basis. There are significant worries that AI can readily go off the rails or otherwise dispense unsuitable or even egregiously inappropriate mental health advice. Banner headlines in August of this year accompanied the lawsuit filed against OpenAI for their lack of AI safeguards when it came to providing cognitive advisement. Despite claims by AI makers that they are gradually instituting AI safeguards, there are still a lot of downside risks of the AI doing untoward acts, such as insidiously helping users in co-creating delusions that can lead to self-harm. For my follow-on analysis of details about the OpenAI lawsuit and how AI can foster delusional thinking in humans, see my analysis at the link here. As noted, I have been earnestly predicting that eventually all of the major AI makers will be taken to the woodshed for their paucity of robust AI safeguards. Today’s generic LLMs, such as ChatGPT, Claude, Gemini, Grok, and others, are not at all akin to the robust capabilities of human therapists. Meanwhile, specialized LLMs are being built to presumably attain similar qualities, but they are still primarily in the development and testing stages. See my coverage at the link here. Therapists And The Role Of AI I have been extensively identifying and examining the myriad ways that AI enters into the role of professional therapists. Some therapists refuse to think about AI and want nothing to do with it. Others are embracing AI and using AI as part of their therapeutic process with clients. Indeed, I have predicted that the therapy realm is being transformed from the traditional dyad of therapist-client and inevitably becoming a new triad of therapist-AI-client, see my analysis at the link here. My view is that whether therapists are keen on AI is not the headspace they should be in. AI is coming, and to a great degree, it is already here. Clients nowadays come in the door with AI-generated advice and want their therapist to tell them what it means. In other instances, clients will post-session try to double-check what their therapist told them and lean into AI as a means of judging the mental health advice they are getting from the clinician. AI is a reality that therapists must face, regardless of their desire to do so. Having one’s head in the sand is not prudent, as I will be illuminating momentarily. There are lots more variations of the role of AI in therapy and regarding therapists, including these circumstances that I have judiciously addressed: - How therapists should clinically analyze AI chats of their clients, see my discussion at the link here. - Questions that clients are asking their prospective or existing therapists about AI, and the answers that therapists ought to be providing, see my coverage at the link here. - Therapy is shifting from the classic dyad of therapist-client to the new triad of therapist-AI-client, see my discussion at the link here. - Therapists are being asked by clients to jointly use AI during their mental health therapeutic process and work in these new ways, see my explanation at the link here. - Some therapists are opting to use AI during therapy sessions with their clients and do so in these astute ways; see my coverage at the link here. - How therapists are handling clients who appear to be encountering AI psychosis, see my discussion at the link here. - Therapists are using AI to craft digital twins of their clients and perform more impactful therapy accordingly, see my coverage at the link here. - Worries that therapists leaning into AI as an aid in conducting therapy might end up deskilling their own capabilities, see my assessment at the link here. - How therapists are using custom prompts to get generative AI to serve as an adjunct to their therapy sessions and interact with their clients, see my discussion at the link here. - Public perception of therapists who decide to use AI in their practices, see my analysis at the link here. - Legal defense strategies being used by AI makers to defend against AI mental health lawsuits, see my analysis at the link here. - Contending with clients that come to therapy with AI-generated mental health advice and want their therapist to give a thumbs up, see my coverage at the link here. - Emerging new informal duty might be for therapists to inform their clients about the ups and downs of using AI for mental health guidance, see my analysis at the link here. And so on. The Role Of Cognitive Errors Shifting gears, I’d like to dive into the nature of cognitive errors that human therapists might make. After doing so, we can look into the use of AI to aid in detecting and correcting those errors. First, let’s consider that cognitive errors can occur at any of these three stages: - (1) Pre-Session cognitive errors - (2) Mid-Session cognitive errors - (3) Post-Session cognitive errors Before a session, a therapist might make a therapy-oriented cognitive error while preparing to meet with the client. I am emphasizing that these are therapeutic mistakes and not other types of cognitive mistakes, such as those that are administrative. This isn’t about mistakes in billing or paperwork. The focus is on cognitive errors intertwined with therapy. An example of a pre-session cognitive error would be the therapist anchoring on a pre-fixed diagnosis that they intend to carry into the session. They become resolute that no matter what else happens, they are going to dogmatically insist on the pre-determined diagnosis. During a session, everything the client says will be interpreted solely within that frame of mind. The client has a near-zero chance of being understood in any other manner, and the therapist is going to confirm the preconceived diagnosis. If a therapist does allow their thinking to be unyielding in this manner during a session, you could say that the pre-session cognitive error has been compounded. The therapist has committed a cognitive error at the pre-session stage and then committed a second cognitive error during the actual session. Cognitive errors can arise spontaneously during a session. One of the most frequent cognitive errors that a therapist makes is to undertake a semblance of mind-reading. Here’s how that goes. The client says something, perhaps an innocuous statement. The therapist leaps on that statement and assumes all sorts of facts about how the client feels or what they mean. It is a case of mind-reading. A therapist should avoid that type of behavior and aim to ensure that they clarify what the client has stated. Do not jump to premature conclusions. A cognitive error might arise post-session. A common cognitive error happens when a therapist is preparing their notes about a session. Sometimes, a therapist wants to put a tidy bow on what occurred. They force the discussion to fit a particular narrative. This is known as a narrative fallacy. The issue is that the bias of the therapist is reducing uncertainty and ambiguity into a cohesiveness that doesn’t truly belong there. Research On Therapist Cognitive Errors In an excellent article on therapists and cognitive error, namely a posting entitled “Special Report: Addressing Cognitive Error in Psychiatric Practice” by Dr. H. Paul Putman III, Psychiatric News, December 22, 2025, these salient points were made (excerpts): - “Psychiatric practice has become increasingly complex due to an expanding number of treatments, longer patient life spans, and a resulting increase in comorbid psychiatric and nonpsychiatric medical conditions.” - “In facing these challenges, awareness of how we make frequent human cognitive mistakes can help elevate the quality of our efforts, reduce treatment failure, and minimize suboptimal results.” - “Improving treatment outcomes necessitates understanding the source of our errors.” - “While psychiatrists are the most knowledgeable among medical specialists about brain function and behavior, we are also among the least likely to openly discuss and teach the cognitive skills necessary for diagnostic reasoning or to use this information to examine our own performance.” The research brings up the various nuances associated with Type I versus Type II thinking, the use of abductive reasoning, hypothetico-deductive reasoning, and inductive reasoning. It is very easy to falter when using any of those thinking processes. Therapists are prone to the same cognitive errors that we all encounter in our daily lives. We cherry-pick data, we look at the pieces of a situation without considering the whole, we anchor on a way of thinking, and so on. Recommendations are made on ways that therapists can try to prevent cognitive errors, along with spotting errors and correcting errors. The hallmark points are worthwhile to keep at the top of mind. Avoid making rapid diagnoses. Use a pluralistic approach to assessment. Get multi-sourced feedback. Enhance communication skills. Etc. Leaning Into AI An additional means or avenue to contend with cognitive errors is to prudently employ generative AI and LLMs. Let's consider the three stages: - (1) Pre-Session: Use AI to double-check the preparations for meeting with a client, indicate the plan for the session, practice for the session via using AI, and ask AI if any likely cognitive errors seem to be afoot. - (2) Mid-Session: Use AI to track the session and offer real-time feedback to the therapist (this can be tricky and has upsides/downsides, see my discussion at the link here). - (3) Post-Session: Use AI to analyze a transcript of the session and analyze the notes and post-write-up of the therapist, interact with AI to do a debriefing, etc. A therapist can pick and choose which of those AI uses they believe are valuable for them. In some therapy practices, AI is being adopted for all three stages and considered part-and-parcel of performing therapy. You might be wondering whether AI can identify cognitive errors that a therapist might have made. One aspect to keep in mind is that cognitive errors are oftentimes hard to spot and might not actually be cognitive errors. Even a human reviewer or supervisor might claim they have found a cognitive error, but it turns out that the declared concern is falsely flagged. The gist is that you need to be extremely cautious in summarily assuming that just because a cognitive error might seem to exist, it isn’t as easy as finding a number that’s out of sequence or detecting that two plus two came out to five. Cognitive errors are elusive. They can be subtle. In a therapeutic context, be mindful of not rashly proclaiming that a cognitive error has been found. The best bet would be to calmly discuss with the therapist whether a cognitive error arises. Of course, that’s a difficult discussion since a therapist might naturally be defensive. To give you a sense of how contemporary AI might be used to tentatively identify potential cognitive errors, I fed a transcript of some therapist-client sessions into AI and asked the AI to see if there were cognitive errors. I will briefly showcase three examples. You might agree that the spotted cognitive errors are indeed mistakes, or you might disagree. All in all, it would take a bit of in-depth further analysis and discussion with the therapist to fully gauge these circumstances. Example 1: Confirmation bias and premature closure In this first example, a therapist was working with a client, and the use of an AI-based assessment post-session noted a portion of the transcript that appeared to contain a possible cognitive error on the part of the therapist. Consider this transcript snippet: - Client: “I felt really angry when my manager changed the deadline again.” - Therapist: “That sounds like the abandonment fears we’ve talked about before.” Here is the AI assessment: - Generative AI detection of possible cognitive error: “The therapist might have committed a potential cognitive error consisting of confirmation bias. The therapist has interpreted the client’s anger by applying a psychological framework involving abandonment fears. This seems like a surface-level analysis. A common best practice for a therapist would involve exploring alternative explanations before overtly offering a diagnostic conclusion to a client.” - Generative AI repair recommendation: “In general, the therapist should take note of the matter and consider examining multiple hypotheses before offering therapeutic interpretations.” Example 2: Mind reading and assumption of affect In this second example, a different therapist was working with a different client, and the use of an AI-based assessment post-session noted a portion of the transcript that appeared to contain a possible cognitive error on the part of the therapist. Consider this transcript snippet: - Client: “I didn’t respond to my sister’s message.” - Therapist: “You were probably feeling guilty and ashamed.” Here is the AI assessment: - Generative AI detection of possible cognitive error: “The therapist seems to have made a cognitive error consisting of inferring an emotional state of the client, without first eliciting the client’s own account. This is an instance of therapeutic mind-reading and generally should be avoided. “ - Generative AI repair recommendation: “In general, the therapist should be analyzing assertions made by a client to ascertain what the context and significance consist of. Give the client sufficient space to confirm or disconfirm any interpretation by the therapist.” Example 3: Overgeneralization from a single instance In this third and final example, I’ve once again opted to use an AI-based assessment on a transcript. AI noted a portion of the transcript that appeared to contain a possible cognitive error on the part of the therapist. Consider this transcript snippet: - Client: “I skipped the party last weekend.” - Therapist: “You always isolate when things get hard.” Here is the AI assessment: - Generative AI detection of possible cognitive error: “The therapist seems to have made a cognitive error by overgeneralizing the statement made by the client.” - Generative AI repair recommendation: “In general, the therapist should be analyzing client statements based on longitudinal evidence. Does the remark by the client warrant a generalization, or is the statement being inadvertently overstretched? It would be prudent to explore the variability before committing to a diagnostic assertion.” Template Of Prompt To Do These Reviews There are many ways you can compose a prompt to get AI to do these kinds of reviews. I will show you the one that I used. It is a template that you can consider using. You are welcome to adjust the prompt to fit your specific needs. Here is the prompt: - My templated prompt for catching cognitive errors by therapists: “I would like you to review therapist–client session transcripts for the limited purpose of identifying potential cognitive errors that the therapist may have exhibited. Cognitive errors may include, but are not limited to, confirmation bias, anchoring, premature closure, overgeneralization, mind reading, fundamental attribution error, leading questions, affective bias, hindsight bias, narrative smoothing, and so on. Do not assess clinical competence. Instead, flag instances where the therapist’s statements or questions could plausibly reflect a cognitive error, explain the reasoning for each flag, cite the specific transcript excerpt, and note reasonable alternative interpretations the therapist might have considered. Use tentative, non-accusatory language and treat all findings as hypotheses for reflective review rather than conclusions. Focus exclusively on metacognitive analysis of the therapist’s reasoning as inferred from the transcript.” Observe that the prompt tries to carefully guide the AI toward spotting cognitive errors that might have been made by the therapist. The reason that the language is somewhat lengthy is that if you don’t pinpoint what you want the AI to find, you are likely to get a vast sea of flagged possibilities. Another crucial aspect in the prompt entails emphasizing that the AI isn’t to attack the therapist. If you don’t clarify that the AI isn’t supposed to be accusatory, there is a solid chance that the AI would heavily lean against the therapist and make comments that would be highly confrontational. The odds are that this would force the therapist into a heightened defensive posture. The aim here is for collegial assessment and not to rake the therapist over the coals. The World We Are In Let’s end with a big picture viewpoint. It is incontrovertible that we are now amid a grandiose worldwide experiment when it comes to societal mental health. The experiment is that AI is being made available nationally and globally, which is either overtly or insidiously acting to provide mental health guidance of one kind or another. Doing so either at no cost or at a minimal cost. It is available anywhere and at any time, 24/7. We are all the guinea pigs in this wanton experiment. The reason this is especially tough to consider is that AI has a dual-use effect. Just as AI can be detrimental to mental health, it can also be a huge bolstering force for mental health. A delicate tradeoff must be mindfully managed. Prevent or mitigate the downsides, and meanwhile make the upsides as widely and readily available as possible. A final thought for now. Therapists can use AI for their own internal purposes. The AI isn’t being used to directly interact with a client. Instead, AI can be a helpful aid to a therapist, including before, during, and after a session. As James Garfield famously stated: “The truth will set you free, but first it will make you miserable.” Use AI in a balanced way. Don’t assume the AI is right, nor assume it must be wrong. Try to use AI in sensible ways.

Discussion

3
20:41

I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.

A new speculative decoding drafter called DFlash 2 more than doubles real coding speed on a local 27B model, reaching 2.26x faster on 100 real coding prompts and 4.68x when stacked with one n-gram lookup table on multi-turn coding sessions. It costs about 2.7 GB extra VRAM, roughly half what the older DFlash 1 needed, and beats the previous version at the same draft width. The fine print: the recommended draft length of 7 is past the peak (5 is better), two lookup tables hurt instead of help, and an eye-popping 8x number the benchmarkers found came from the model looping on synthetic test, so it's not real. Practical takeaway: turn on DFlash 2 with one n-gram table for iterative coding and agent work, leave it off for one-shot prompts and prose.

Full text · 16,266 chars
Hey guys, Inco AI shipped DFlash 2 a few days ago with a drafter for Qwen 3.8 27B and a llama.cpp PR. I built the PR and ran it against plain decoding, MTP, the n-gram lookup drafters, and my July DFlash 1 numbers on Qwen 3.6 27B for 3 days. One RTX PRO 6000, concurrency 1, about three days of runs. The interesting result isn't the biggest number I measured. It's where n-gram actually helps and where it doesn't. Short version: DFlash 2 alone: 2.26x on 100 real LiveCodeBench problems (67.97 → 153.91 tok/s, inter-token latency 14.27 → 6.02 ms), natural stop, nothing forced. That is the headline. Costs +2.7 GB VRAM. DFlash 2 + one n-gram lookup table ( ngram-map-k4v ): 4.68x on the build phase of an 18-turn coding session (65.1 → 304.9 tok/s). Adding the second table ( ngram-mod ) made it slower, 3.77x. In July, with DFlash 1, stacking both was the winner. I did not expect that to flip. The same n-gram flag is +52% on a synthetic benchmark, +1% on LiveCodeBench and -30% on prose. The +52% is the harness degenerating, do not quote it. The recommended --spec-draft-n-max 7 is past the peak. 5 gave roughly 11% more on 8K coding prompts. 7 is also a hard cap (block_size 8), anything above is silently clamped. --spec-draft-p-min does nothing on DFlash 2. The DFlash 2 code path in common/speculative.cpp never reads it. I also measured 8.47x in a synthetic test. I nearly used that as the headline. It was mostly benchmark garbage caused by the model falling into a repetitive loop. Setup (the parts that matter for reproducing) Target ggml-org/Qwen3.8-27B-GGUF:Q4_K_M (18 GB). Drafter incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M (1.1 GB). MTP sidecar mtp-Qwen3.8-27B-Q8_0.gguf (3.0 GB). Reasoning off. llama.cpp b10498 built from PR #27342 (commit 5ecbe1ac ), CUDA 13.3. The PR build matched upstream b10499 within 0.3% on a non-speculative baseline (checked at 512 and 4K only). RTX PRO 6000 Blackwell 96 GB, Ryzen 9 9950X. -c 262144 , f16 KV, -fa on , -ngl -1 , drafter fully on GPU. Concurrency 1 everywhere. Greedy (temperature 0) for everything except the multi-turn coding harness, which runs model-default sampling with no seed (more on that below). One server on the GPU at a time (flock), fresh container per config, card cooled to 45 °C between configs and 60 °C between context sizes. 11.6 hours of telemetry, zero throttle events, so the card sits on its power limit, not a thermal one. Full context 1. DFlash 2 more than doubles real coding throughput, and beats DFlash 1 at the same draft width for half the VRAM 100 LiveCodeBench problem statements replayed in the same order, streaming, no ignore_eos , no min_tokens , no max_tokens . Every answer ends where the model ends it. https://preview.redd.it/ina7l4wsfzkh1.png?width=1700&format=png&auto=webp&s=f803f3328a146063e8f19944b265838e63c5c546 tok/s vs own base ITL wall clock Qwen 3.8 27B, no speculation 67.97 1.00x 14.27 ms + DFlash 2 (n=7) 153.91 2.26x 6.02 ms + DFlash 2 + both lookups 155.83 2.29x 6.11 ms Qwen 3.6 27B, no speculation 67.75 1.00x 14.34 ms + DFlash 1 (n=7, matched) 135.34 2.00x 6.93 ms DFlash 1 was re-run at n=7 because comparing it at its own maximum of 15 would measure the cap, not the drafter. At matched width DFlash 2 is ahead, 2.26x vs 2.00x against each model's own baseline, with probe acceptance of 60% vs 48%, and it costs +2,720 MiB where DFlash 1 cost +5,554 MiB in July. Part of that memory gap is a quant choice (Q4_K_M 1.1 GB drafter vs Q8_0 1.8 GB), not architecture. Two things to be careful with. The cross-generation rows are not a controlled A/B: different model, different target quant, different drafter quant (and the drafter quant works against DFlash 2, not for it). Compare the speedups, never the absolute tok/s; the two baselines landing 0.3% apart is luck. Claim vs measured: Inco AI quote 2.7x to 3.4x at batch size 1 on SGLang for this model. I got 2.26x on llama.cpp on single-turn coding so it depends on task and engine it will probably get better soon with updates to engines. 2. One lookup table on top of DFlash 2 is the best stack. Two is worse. That is the opposite of DFlash 1. The n-gram drafters copy spans that already exist in context, so single-turn prompts are their worst case (+1.2% above, and the median actually says -2.7%). The case that matters is working on a code base, so I drive 18 fixed prompts as one cumulative conversation: turns 1-9 build a Gradio chat client for llama.cpp feature by feature, turns 10-18 maintain it (re-emit the file, docstrings, renames, a bug, a refactor, tests, README). https://preview.redd.it/v106usk5gzkh1.png?width=1700&format=png&auto=webp&s=04d49cda970411e65b000870e3c8bb718ea580bf stack --spec-type build 1-9 tok/s vs base all 18 accept (build) drafts/tok no speculation - 65.14 1.00x 56.95 - - DFlash 2 alone draft-dflash 181.89 2.79x 177.53 66.4% 1.24 DFlash 2 + k4v draft-dflash,ngram-map-k4v 304.92 4.68x 343.52 64.2% 1.41 DFlash 2 + both lookups draft-dflash,ngram-mod,ngram-map-k4v 245.84 3.77x 306.04 55.6% 1.59 DFlash 2 + mod draft-dflash,ngram-mod 229.37 3.52x 313.46 58.6% 1.48 lookups only, no drafter model, 0 VRAM ngram-mod,ngram-map-k4v 133.00 2.04x 170.54 59.5% 1.00 Read the build column. Turn 10 is "show me the complete final app.py", which is ~99% draftable and inflates every speculative method. Over all 18 turns the k4v stack reads as 6.03x, a real number about the easiest thing you can ask a copying drafter to do. I expected the July result to repeat: with DFlash 1, draft-dflash,ngram-mod,ngram-map-k4v was the winner at 6.01x and ngram-mod did almost all of the n-gram work. Instead, on DFlash 2 the k4v table alone wins, mod alone is the weakest stack, and both together are slower than k4v alone. It could be draft tokens number or early implementation we will see. DFlash 1 had max 15 draft slots, DFlash 2 has 7, and two lookup drafters crowd each other out of them. 3. The same one-line change gives four different answers, and the synthetic one is wrong Same DFlash 2 server, same weights, append ngram-mod,ngram-map-k4v to --spec-type , run everything again: workload DFlash 2 alone + both lookups change editing code, 18-turn session, turns 1-9 181.89 245.84 +35% forced-length synthetic, 4K in / 4K out (medians) 176.57 267.82 +52% one-shot coding, LiveCodeBench x100 153.91 155.83 +1.2% writing fresh prose, one request 158.9 111.6 -30% https://preview.redd.it/vhr4stn8lzkh1.png?width=1700&format=png&auto=webp&s=920e0c8137a88f6f74ffd597823a67e8fe0e2d12 The synthetic bench from aiperf is inflated by its own harness. It passes ignore_eos and min_tokens , forces the model past its natural stop until it loops, and a lookup drafter copies loops perfectly. Carried to 36K the same harness says DFlash 2 + lookup is 8.39x (498 tok/s). On 100 real prompts that stack was worth +1.2%. 8.39x is the kind of number that you could get but in very specific usecase. Prose is the opposite corner: "Write a very long story", nothing in context to copy, the tables burn draft slots on guesses that never land, acceptance 54% → 32%. That row is a single instrumented request, a probe, not a run. Practical consequence: turn the lookup drafters on for iterative coding and anything that re-emits its own context, leave them off for one-shot prompts and creative writing. They cost zero VRAM and zero prefill, so this acceptance loss is their only cost. 4. The recommended draft width is past the peak, and 7 is a hard cap anyway https://preview.redd.it/usnvqpyalzkh1.png?width=1920&format=png&auto=webp&s=035c2ef4985f7e132daeb2d47f9eca182587ed2b 16 coding prompts per width at 8K tokens from livecodebench, cache_prompt false so every request pays a cold prefill: I use livecodebench and cut it to the size to measure worst case here. n_max DFlash 2 tok/s accept MTP tok/s accept 2 140.56 82.6% 133.42 79.5% 3 158.06 72.6% 154.53 77.9% 4 174.60 72.0% 159.07 71.7% 5 187.13 70.4% - - 6 184.60 67.0% 154.64 62.9% 7 168.06 59.7% - - Running the model card's 7 leaves roughly 11% on the table. An earlier 8-prompt sweep put the optimum at 6 rather than 5, so call it 5-6; both sweeps agree 7 is past the peak. And you cannot go above 7: the draft GGUF carries dflash.block_size=8 , llama.cpp clamps n_draft_max = block_size - 1 , logs a warning and uses 7. Some cells rest on only 3-7 valid generations of 16 (the truncated prompts sometimes make the model emit EOS immediately), so treat the exact peak as soft. MTP on this model peaks at n=4 and flattens near 2.5x across context. Qwen 3.8's sidecar declares nextn_predict_layers=1 , one trained head, against DFlash 2 reading five target layers. That is a property of this sidecar, not of MTP as a method; Qwen 3.6's had eight heads. 5. Long context: the drafter gets relatively cheaper and absolutely more expensive https://preview.redd.it/wnunzgxelzkh1.png?width=1700&format=png&auto=webp&s=de18ce5fa2f47faec167d414685160a335f2be9b The usual complaint is that speculative decoding falls apart at long context. Two costs hide in that sentence. Prefill, where the drafter has to read the prompt too, I could measure. Decode at those depths I could not (see caveats). Cold prefill, 12 prompts per depth: prompt depth prefill tok/s, none prefill tok/s, DFlash 2 speed kept extra wait 1K 3,506 2,656 0.76 +0.09 s 4K 3,845 3,164 0.82 +0.23 s 16K 3,639 3,162 0.87 +0.68 s 64K 2,867 2,588 0.90 +2.46 s 128K 2,239 2,056 0.92 +5.20 s Relative to baseline the tax shrinks with depth (24% down to 8%). In seconds it grows, +0.09 s to +5.20 s. Both readings are true; quoting only the first is the flattering half. The prefill cost is repaid in 13 / 30 / 69 output tokens at 1K / 4K / 16K, so any real answer clears it, but someone on a 128K prompt does wait five seconds longer for the first token. The lookup drafters cost nearly nothing here (0.994-0.997 of baseline), which doubles as the control that the gap is the drafter and not drift. MTP's tax is smaller (0.83 at 1K vs 0.75). On the forced-length synthetic decode sweep DFlash 2 goes 1.59x → 2.62x → 2.96x → 3.55x at 512 / 4K / 12K / 36K while the baseline falls 67.6 → 59.3 tok/s. DFlash 1 at its own max of 15 did 4.44x at 36K on that harness in July (higher still when re-measured this month), and at matched width 7 it did 3.71x. I expected the new drafter to win everywhere. It does not: it wins on real prompts at equal width, and loses the synthetic long-context race to the old drafter with more slots, because it is capped at 7. 6. --spec-draft-p-min is a no-op on DFlash 2, and buys nothing on MTP either Adaptive draft truncation should let the drafter stop a block early when it is unsure to save resources. There are more advance method form DeepSeek Dspark paper but they just landed on vLLM. I logged draft width and cycles per second, not just tok/s: drafter p_min tok/s accept draft width cycles/s DFlash 2 0.00 195.3 71.6% 6.998 32.51 DFlash 2 0.85 171.1 60.5% 6.998 32.69 MTP 0.00 161.7 90.4% 3.001 43.54 MTP 0.85 155.7 97.6% 2.642 43.52 Draft width is identical at 0.00 and 0.85 on DFlash 2. common/speculative.cpp has four drafter implementations: draft_simple , draft_eagle3 , the DFlash 1 branch and draft_mtp honour p_min ; the is_dflash2 selector branch never consults it, because it reads a selector lattice rather than a probability. The server still prints the flag in its startup banner, so a log-based check passes while nothing happens. The 12.4% throughput drop in that row is the text, not the flag: the server did identical work (cycles/s within 1.5%), the sampled output just accepted fewer of the same seven tokens. I nearly published "p_min costs 12%". On MTP the flag works exactly as documented (width 3.00 → 2.64, acceptance 90% → 98%) and throughput goes nowhere, +0.8% at best against a 4.1% noise floor. What I would run Iterative coding, agents, anything that re-emits its own context: --spec-type draft-dflash,ngram-map-k4v --spec-draft-n-max 5 One-shot prompts and Q&A: --spec-type draft-dflash --spec-draft-n-max 5 Prose: DFlash 2 alone, no lookups. Keep the KV cache at f16 for now or test it it will be probably stable soon but on last version there were issues and I use default. Ignore --spec-draft-p-min . git clone https://github.com/ggml-org/llama.cpp.git cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 # NVIDIA CUDA cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON cmake --build build -j # Apple Silicon cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON cmake --build build -j # Best measured config: iterative coding, agents, anything that re-emits its own context # (4.68x on the multi-turn coding session vs 2.79x for DFlash 2 alone) ./build/bin/llama-server \ -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \ -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \ --spec-type draft-dflash,ngram-map-k4v \ --spec-draft-n-max 5 \ -ngl -1 --spec-draft-ngl all \ -fa on \ -c 262144 \ --parallel 1 \ --jinja --reasoning off \ --no-mmproj \ --host 0.0.0.0 --port 8000 \ --alias qwen38-dflash2-k4v Caveats, all of them No accuracy measurement this time. Greedy speculative decoding is output-lossless by construction and the July study measured it (MATH-500: 87 vs 86 of 100, then 440 vs 435 of 500), but that was Qwen 3.6 with DFlash 1 and I did not rerun it here. Only a LiveCodeBench smoke test. No deep-context decode. At 64K and 128K every generation returned one token and stopped, so prefill at those depths is valid and decode does not exist. Only one experiment was ever repeated (the p_min controls). Everything else is one sample. The spreads from those repeats, 4.1% and 14.7%, are the noise floor for this whole post. The multi-turn harness can measure the wrong thing. It declares no tools, but under default sampling the model sometimes answers with a <tool_call> block and waits for a result that never comes, and that session comes out fast because tool-call XML is predictable. One run did exactly that (1,134 tokens where its siblings produced 50K-73K), got caught by its token count and was rerun. This is one workload family (coding) on one machine at concurrency 1. A 96 GB card is not what most of you run. The drafter is 1.1 GB and nothing in the KV math depends on the card, so I expect the shape to hold on a 24-32 GB card with a smaller context, but I have not measured it. DFlash 2 is a PR build. Numbers can move when it merges. Resources Repo (both studies, this one on top) : https://github.com/lukaLLM/DFlash2_Qwen3.8_3.6_27B_LlamaCPP Video walkthrough (the first half explains the mechanism, path selector and the convolution the rest go even more deeper into the scores etc. ) : https://youtu.be/RBlRTUwJMI4 One-click setup, builds the PR image, downloads the models, smoke tests and leaves a server running: ./scripts/setup_dflash2.sh --arm dflash2_ngram (arms: base, dflash2, mtp, ngram, dflash2_ngram). Compose file docker/docker-compose-qwen38-dflash2.yaml ; ablate with LLAMA_SPEC_TYPE=... and LLAMA_SPEC_N=5 . Reproduce the whole study in order: ./scripts/run_all_benchmarks.sh , then run_matched_n.sh , run_context_scaling.sh , run_bench_ngram.sh , run_nmax_redo.sh , run_pmin_agentic.sh . Every number in one machine-readable file: benchmark/results_summary.csv (TABLE 8-14 are this study). Raw artifacts under artifacts/q38_*/ , the thermal log in artifacts/thermal/ , quarantined runs and the reasons in artifacts/_suspect/README.md . Long-form write-up with the charts: report/dflash2-report.html in the repo. https://inco.ai/blog/dflash2/ the blog Previous posts: DFlash 1 in July https://www.reddit.com/r/LocalLLaMA/comments/1uq0h4o/i_tested_freshly_merged_dflash_in_llamacpp_on/ and the n-gram stack https://youtu.be/zNUoHONUHGk AI was abused in editing this post. Questions: Has anyone run DFlash 2 on SGLang or vLLM at concurrency 1 with this model? I want to know whether the 2.7-3.4x claim holds there and how much of the gap to my 2.26x is the engine. Anyone on a 4090 or 5090 with a 24-32 GB budget: does n=5 still beat 7 for you, and where does the k4v-only stack land on your own multi-turn coding? Has anyone tried some other combinations that I didn't think of? submitted by /u/FantasticNature7590 [link] [comments]
20:24

New 100B Liquid AI model coming soon

Liquid AI appears to be teasing a new 100B parameter model, announced only via a poll on X with no specs or date. The company's smaller models and fast architectures have a good reputation, so a large version is noteworthy, but this is a rumor based on a teaser, not a real announcement.

Full text · 340 chars
Liquid AI currently possesses among the fastest LLM architectures around, and some of the best SLMs (in terms of utility IMO) around, so I'm very excited to see what a potential 100B LFM (3?) model would look like! Link to the poll: https://x.com/ramin_m_h/status/2091236099612098943?s=20 submitted by /u/KaroYadgar [link] [comments]
15:26

How to remove trendy speech from llms?

LLM users are fed up with models adopting trendy slang like saying "minted" instead of "created" or "escape hatch" instead of "alternative path," and want system prompts to make them speak plainly. The ask is whether a simple instruction to avoid that language fixes it, and whether anyone has tested prompts that do. Thin content — a frustrated question with no tested answer yet.

Full text · 609 chars
For example: Instead of saying: "I created this new ID" It says: "I minted this new ID" Instead of: "This alternative path is available" It says: "this escape hatch is available" This speech is so nonsensical and annoying. Just. Speek. Literally ... OR NORMALLY. Where did LLMs learn these speech patterns? I've never seen them so frequently until AFTER the LLM surge. If I just add "Don't use X language, speak normally and more literal" will that fix most of the issues? Anyone else have some good sys prompts / instructions that help with this? Thanks! submitted by /u/CSEliot [link] [comments]