Nothing matches those filters.

Lead

21
😺 The ACTUAL ChatGPT 3 moment for robotics (one-shot learning)The NeuronChatGPT search now uses the site:operator at scaleSimon WillisonOpenAI chases Anthropic's biz customers with zero data retention pledge - The RegisterTheregister1,357 AI medical devices cleared, 3 actually tested on patient outcomes - Research journalsPlosCoSnitch: Researchers Asked Copilot How to Hack It, Then Built | Machine BriefMachinebrief☕️ Ex-Meta engineer testifies against ZuckerbergTechpressoThe Download: polycrisis support networks and a hydrogen gold rushMIT Technology ReviewLatent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without RetrainingarXivAbliteration Mitigation via Refusal AliasesarXivMuse Video leaks 📹, Ramp Router launch 🔀, why Stripe bought OpenRouter 💰TLDR AITencent begins testing its new flagship model Hunyuan Hy4r/LocalLLaMAQwen3.8-27b has the highest level of "agency" I've ever seen in a local modelr/LocalLLaMAStripe Dates The Singularity To Jan. 1 And Prices It At $8 BillionForbesReddit Nearly Vanishes From ChatGPT Citations After OpenAI Search Change, Report SuggestsForbesModular Delivers Openness And Accelerator Portability At ModCon 2026ForbesMicrosoft's GitHub Under Siege As SpaceX's Cursor Takes The AI StackForbesServal Wants To Replace ServiceNow With AI That Builds Enterprise AutomationForbesOde To A Whole New Kind Of ConsultingForbesHow to 8x Your Code Output Using Context EngineeringThe AI Corner[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling LawLatent.SpaceVisions of AI: Automating repetitive grunt Coding tasksAI Supremacy

Video

2
19:46

Free Grok Bot and all of this week’s news

A free, self-hosted version of the expensive Grok agent bot is now out from Inbox Zero, letting people run it on their own computer with whichever model they choose instead of paying hundreds of dollars a month. The weekly roundup also covers specialized named bots as a rising interface trend, a desktop tool called Meridian that recaps your workday, a local "second brain" app called Clipto that indexes everything on your machine for other agents, plus a vibecoded time-tracking app, an open-source interview-prep repo, an iPhone remote for Codex, and a directory of community-built bots.

Notes
  • Inbox Zero (free Grok bot): Free, self-hosted clone of Grok-bot (normally ~$200/mo) you run on your own computer using any model — no lock-in to Gro models. Identical look/function. Includes "your own remote computer." Monetization: eventual hosted version, still cheaper than Grok. Get from GitHub.
  • Bots-as-interface trend: Cory/Andrew see "bots" (named agents with specific functions) as the mass-adoption path — average people can't grasp an all-in-one finance/marketing bot but understand a specialized one ("a bot that does my grocery ordering"). Even Hermes added bot mode. Andrew hopes Codex adds it soon.
  • Meridian: Desktop tool (shown on GitHub) that watches what you do all day, then at day's end rewinds and builds a calendar of what you did. Pairs with daily standups: check off items (e.g., "shipped the activity feed") and copy/paste into your Slack standup report. Beyond team sharing, lets you ask your agent "where am I wasting my time?" Andrew notes the limitation of prior approaches — his OpenClaw end-of-day report only knew what they worked on together, not what he did elsewhere. Net positive: more visibility into your actual work.
  • Clipto: "Second brain" that indexes all your desktop data (videos, meetings, voice memos, screenshots, docs) locally for privacy, making it searchable not just in-app but within tools like Codex, ChatGPT, Cursor. Example: search a YouTube snippet you watched ("I've got the exact snippet of that video... send you the link"). Inspirational anecdote: someone vibecoded a solution for a family member's medical-practice problem, got no clients for 2 months, went all-in on SEO, then landed 2 paying customers in one day (one via ChatGPT search, one via Google search).
  • Simple clock app: "Know who's working, pay them by the hour" — clock-in/out app, another vibecoded-from-a-problem story.
  • Sunkit: Open-source GitHub repo for developers — store notes, projects, proof of work, up-to-date tech info to stay "interview ready." Clean layout; e.g., an Operating Systems section with 11 topics and "Start learning."
  • iPhone Codex remote: A developer built his own version of the Codex "remote control" hardware (a TV-remote-like device that sold out a month or two ago) as an iPhone app instead. From your phone: switch to voice mode, change chats, fork a working session, custom speed-dial buttons.
  • Bot Directory AI: Site to list bots built on Grok-bot. Copy an agent's prompt (e.g., an "accounting expert") into a new bot and it self-configures. Applicable beyond Grok — works in Codex/any agent. Shows people's agents, plugins (e.g., AI video generator using "AI tuber"). Andrew's addition: a dedicated bot whose system prompt is just "go find trending news topics for our show" — cleaner than an automated task.
Transcript · 13,179 chars
I've got a free version of the bot that everyone's talking about that costs 200 bucks a month. You ever look back at your day and say, "What did I get done?" I've got an agent that will tell you exactly what you got done every single day. And someone built a second brain that remembers everything that's on your computer and makes it accessible to every agent you're using. All that and so much more coming up. Presented by Zapier, the AI automation company. >> All right, Corey. First thing this week, everyone's been talking about Grockbot and how amazing it is, but they're also talking about how it's 200 bucks. In comes a company called Inbox Zero, and they created a free version of it that you can just run on your computer and use whatever model you want. You don't have to use the Gro model. What do you think of this? I think this is going to be appealing to a lot of people because I think people don't like being locked in to the Gro models even though they are really good and have gotten a lot better >> and they don't like having to pay, you know, $200 or $300 a month for the cursor or the Gro heavy subscription to be able to use it. And this looks identical. So, I think this is going to take off for sure. >> I even like how they have your own remote computer in here, which I think is is beautiful. They're going to make money from this by having an eventual uh hosted version. And so if you don't want to run it on your computer, you'll be able to do it there and it'll still be cheaper than Grock. And if you want, you can just get it directly from um from GitHub. >> Yeah. And so as you're watching this episode here, leave a comment below with what your favorite news story or tool is from this week. It helps us decide uh which ones we should include next week, and we always like to see which ones resonate with you guys the most. Speaking of, everyone is now realizing that bot mode, meaning like having these agents with names and specific functions, is the way to go. And even Hermes agent added it. What do you think of this whole idea of in addition to messages which you can see on the left side here uh they're doing bots as an interface? >> Yeah, so it seems like they're every it seems like this is where everything is going. Like everybody's kind of following this trend because I think people really like the idea of like oh I have a a bot that just does my grocery ordering or it just does my you know XYZ task >> cuz I think for the I think for agents to really get mass adoption it's got to be that way. Like I think the average person can't think of a a bot that does like all their finance or all their marketing or whatever. So by making them super specialized, that's what I think makes it more easy for people to kind of grasp in their minds as far as how they would use them, if that makes sense. >> I I hope Codeex adds this in very soon. >> Oh, that'd be great. >> Meridian, here's what this is. You ever look back at the end of the day and you say, "Man, I'm so tired, but did I even do anything? What did I do? And should I be doing all of it? Or maybe some of it I should be sending over to other people?"ID. Meridian is a tool that will sit on your desktop. In fact, here I'm going to show it on GitHub because I've got this beautiful video here. And throughout the day, it's just kind of watching what you're doing and then at the end of the day, you say, "Okay, tell me." It will rewind your day and it will put together a calendar of what you did. And if you are using uh like a daily standup, if you're using a channel on Slack to report what you did, it will make it easy for you to pick the things that you want, copy and paste them into your daily standup. Like here on the left, you might check off that you shipped the activity feed and then it goes directly into the daily standup. I think beyond just team sharing, it's good to look back on your day and see what you did. And I like then being able to give this to my agent and saying, "What should I not spend time on or what am I where am I wasting my time? Where's my time going?" Yeah, I think this is a really good way to like one hold yourself accountable so that you at the end of the day it's like you know did I do the things that move the needle or did I just do a bunch of busy work >> and to your point like when especially when openclaw was first getting big that's one of the things I had my open claw do is send me a report at the end of the day of everything that we worked on >> the the problem with that is it it only knew the stuff that we worked on together it obviously wasn't watching me do other things so all in all I think this is good I mean anytime you can get more visibility into what it is you're actually doing. I think that's definitely a net positive. >> Really interesting idea. I like seeing it. We'll have a link to that and everything else of course below. Clipto, you know, I don't know if you've done this, Corey, but I will say I know that I saw a goat or I know that I saw um uh what's it just I don't I'm trying to think of a of a thing that I've seen or a place that I've been. With Apple Photos, I can just go in and say, "Show me that. Show me goats." And I could find a goat. I know what it was. I was actually going to rehome our goats. And so I said, "Show me pictures of my goats." And I got to see pictures of my goats in time. I was going to list my lawn mower cuz we're moving to Mexico and I can't have the driving lawn mower with me in Mexico. I said, "Show me pictures of my driving lawn mower." And it showed me the whole thing. Well, I would like to be able to do that for aspects of video. Show me something within a video. Show me something in a meeting. Show me something in my docs in screenshots on my computer. Well, that's what Clipto is about. It will actually go through all your data on your own desktop so you have privacy, organize it, tell uh record what's in the videos, what's in the meetings, what happened in the voice memos that you wrote and have it be accessible not just within its app, but within the tools like codecs, like chat, GPT, cursor, etc. that you use. So they all have the ability to search. So this is this is almost like a this is one of those solutions that we're seeing around like a second brain is kind of what it looks like. They're trying to just give you the ability to just like go about your daily life and work and then your second brain kind of builds itself and compounds on itself. I mean when we were first discussing this before we hit record the um you know the thing that came to mind for me is like the other day I was looking for a YouTube video that I watched like I couldn't find it anywhere. or I spent like 10 minutes searching and it sounds like Clipto will be like, "Oh yeah, like I've got the exact snippet of that video if you just describe it and then it can just like send you the link." So that was the kind of the basic use case that I had in mind, but now that we're talking through it, it's like, oh, this is this is almost like a second brain type solution. Yeah. >> So that to me, I think obviously has a lot more utility than just like, oh, what's the YouTube video I was looking for? >> Uh, you and I both saw this guy. You want to tee this up? >> Yeah. Yeah. So, this this one's really cool. So, this guy his I think it's his sister or it was like his mom had had a problem I think in her medical practice or in whatever business she's in. >> So, he just vibecoded her a very simple solution. >> Uh said that he was not able to find any clients for the first 2 months. So, he went all in on SEO. Found out that people were reading his articles but they weren't converting. And then sure enough, one day he got two paid customers in one day. one from I think chat GPT search and one from Google search. And this is a tool that again he just noticed a problem that his family member had. He vibecoded a solution. He implemented some SEO techniques and then he got some paying clients from that. So I just love seeing stories like this. >> Me too. It's so inspiring to see it. Uh really simple app too. Know who's working, pay them by the hour. It's like a simple clock app. >> Yeah. >> People clock in. uh similar thing where someone had a problem. What's Sunkit's problem? >> Yeah. So, this is more applicable to people who are in the tech world, especially developers, where, you know, you always need to be interview ready because you never know when you're going to be looking for a job. So, he basically built an open-source repo that lets you store all of your like your notes, your projects, your proof of work, any sort of like up-to-date information on the tech things that you need to be up to date on as a developer. And it's it's open source. Again, the uh he put the GitHub repo link available there publicly. And this just seems like a good way for developers to kind of stay on top of their game >> as as they move throughout, you know, their daily life and work. Yeah, the way that he laid it out is really beautiful here. So, this is one on let's say operating systems. I can hit start learning and see all 11 topics. The way that he laid it out is beautiful. It's uh clean. I can see using this for other topics, too. And he made it available like you said on GitHub. All right. This I just really It took me a while to get what this was and then you explained it to me and then I got excited about it. It >> Yeah. So, this is really cool. So, I'm sure a lot of people listening to this saw the like the the codeex like remote control, the piece of hardware that they came out with a month or two ago that sold out like right away. And what that is is it's basically just like a again, it's like a TV remote but for codecs that just sits on your desktop. And so what this guy did is he >> built his own version of that, but instead of it being a physical piece of hardware, it's just an iPhone app. Mhm. >> So, it's kind of like like we said, it's just like a remote control for codecs that sits in your phone. So, you can from your phone switch your computer to voice mode or switch to a different chat or fork, you know, fork a working session. Like, it's just really cool what you can do. And if you're a Codeex power user, this is a no-brainer. >> Yeah. All the buttons right there. In fact, you could put whatever buttons you want on speed dial. Hit them and then you get to go with it. Want to talk? Just hit the button and start talking. Want to split it? Hit the button. And I love how customizable it is. I totally get it. It makes so much sense. All right, final thing is this collection. Somebody took somebody b made a site called bot directory AI where you can list the different bots that you've built on Grockbot. And I think this is applicable way beyond Grockbot. To just see someone set up and then be able to import it into your agent and say, I want the same thing makes so much sense. Like here's one that's an accounting expert. They have a prompt and then it will connect to these tools to get it to take action. All you have to do is copy this prompt and give it to a new bot on Gro on Grockbot and it will set itself up as your accounting expert. But even if you're not using it, the prompt is fairly simple. The process and the work that they're getting this agent to do makes total sense. You can give this to Codeex. You can do all of these things on uh on just about any agent. What's interesting to me is seeing this collection of agents that people have built and how they're using these agents and what plugins they're using it for. Like here, this AI video generator, they're using AI tuber. I never heard of it. And this is how they're using it to to build AI videos. Really interesting. And then you found this too. Um where where they put together uh where AI Edge put together its favorite list of 10. How are you using agents? like what's if you were to add to this list, what would you add to it? >> Well, so I would just add like any sort of specialized task that it would be nice to have a dedicated bot for, right? So, I mean, a perfect example is when we do this show every week, one of the things that I do is I go into Codeex and I have it go and find >> 15 potential news articles or tweets that we can discuss on this show. And then obviously we trim that down. But it would be nice to just have a, you know, I of course I could set that up as just like an automated task in codeex, but it just feels cleaner for that to be its own dedicated bot where like the system prompt for that bot is literally, hey, you go out and find trending news topics or cool stories every week that Cory and Andrew can talk about on their show. So like that bot exists just to find us topics for this show. That to me just feels better. So like that's an example of what I would do. really cool use case of it. All right, and as you all saw, we like to see what you've built. Tell us by letting us know in the comments or email. And now that you've seen this one, we've got another collection of tools that if you like this, we think you're going to love. I'll have a link to it right there.
20:45

Dead Internet Theory - What you need to know

For the first time, AI traffic on the internet now outnumbers human traffic, according to a 2026 industry report. AI-driven traffic grew 187 percent in 2025 and is expanding about eight times faster than human traffic, though the video argues content, not traffic, is what the internet really is. More than 95 percent of AI traffic is concentrated in retail, streaming, and travel. The explainer also walks through media manipulation from Abraham Lincoln's retouched portraits to deepfakes, and Cloudflare's push to make AI crawlers pay through a new payment-required web code and simplified licensing.

Notes
Dead Internet Theory — What You Need to Know

Olivio Sarikas (YouTube, 2026-08-20) — video essay arguing the "dead internet" is real but redefined: the commercial crawlable web shifts to AI agents, while human-to-human interaction remains.

The origin of the theory
  • Started by Illuminati Pirate in 2021 on Agora Road. Core claim (quoted): "Large portions of the supposedly human produced content on the internet are actually generated by artificial intelligence networks in conjunction with paid secret media influencers in order to manufacture consumers for an increasing range of newly normalized cultural products."
  • Sarikas: conspiracies start from a grain of truth — parts of the internet are genuinely AI/bot-generated, used for market and political manipulation — but it is not a majority, and it is not new.
History of fakes (the "not new" case)
  • Abraham Lincoln: first use of photography to manipulate an election — retouched portrait (covered neck, added beard) to look more trustworthy/electable.
  • Vietnam War: journalists had unrestricted ground mobility, photographed visceral combat and casualties → public turned against the war → US withdrawal, North Vietnam won.
  • Gulf War: first live-televised war, but strict media control — showed smart bombs, night-vision feeds, "clean war" framing for public manipulation.
  • Beauty scam: Oprah's head composited onto another actress's body for a magazine cover; an English model's face so heavily retouched (no wrinkles) it was used to advertise Olay wrinkle cream, then retracted after public complaints.
The AI era and perception
  • AI now used by both right and left to manipulate representation (e.g., the famous Trump-as-healer/Jesus image vs. Trump-arrested-by-police image).
  • Key insight: even when an image is obviously fake, people accept it because "the truth is not in the image. The truth is in the message and in the perception." The medium is less important than the message for how people perceive truth.
  • Deepfakes differ: real face/voice in realistic video not flagged as AI → hard to verify; requires critical thinking and checking whether the event actually happened.
  • Sarikas caveat: manipulation is costly and used only for specific goals (chiefly political manipulation). The far larger share of AI traffic/content is from companies and private citizens sharing interests — not nefarious actors.
The "authentic-first / AI-first internet" (the data)

Data cited from the Imperva 2026 State of AI Traffic and Cyber Threat Benchmark:

  • 8× faster growth of automation than human traffic
  • 187% AI-driven traffic growth in 2025
  • 7,851% year-over-year growth in "Atlantic AI traffic"
  • First time in history AI traffic exceeds human traffic
  • >95% of AI-driven traffic concentrated in three verticals: retail & e-commerce, streaming & media, travel & hospitality — the commercial, frequently-updated, sellable information parts of the web.

Sarikas distinguishes traffic (interaction) from content (the internet itself) — more AI traffic isn't the same as the internet being dead.

Cloudflare / monetization mechanics
  • In 2025 Cloudflare declared a "Content Independence Day" — every crawl must be paid.
  • HTTP 402 "Payment Required" code: bots can't legally access a site without paying first.
  • RSL (Really Simple Licensing): simplified licensing so bots can agree, pay, and crawl.
  • Cloudflare separates traffic into three categories: search, agent, training — differently valued and with different outcomes (search can drive human visitors; training feeds only the model, no people arrive).
  • Ad-blockers mean humans "cheat" sites out of ad revenue, but agents must pay — creating a new revenue stream for websites, while agents get clean, licensed data to train models and resell via chatbots. Both sides benefit.
The feared outcome vs. the rebuttal
  • Fear: the internet becomes built firstly for AI agents — sites replaced by databases/APIs only agents access; users get everything through chatbots (e.g., ChatGPT) and AI executes requests (hotel, flights, tickets). This "death" applies to the ~95% commercial crawlable content.
  • Rebuttal: human-to-human interaction cannot be crawled or sold — Discord, Facebook, Reddit, Instagram, LinkedIn, X/Twitter, WhatsApp are the biggest, richest platforms precisely because of this. Ads and products can be sold; human interaction cannot.
Verdict
  • Frames it as rebirth, not death: Web 1.0 (static pages) → Web 2.0 (interactive social media, databases, PHP/JS) → a new web where AI agents and bots make data flow "more liquid, more streamlined." Humans keep direct exchange with each other, aided by AI (e.g., Sarikas makes AI comic images with inside jokes to send friends).
Transcript · 17,391 chars
I don't know how to tell you this, but the dead internet is actually here and AI has taken over everything, but maybe in a different way that you might think. So, hello my friends and let's get started. So, the topic we really need to talk about is the dead internet theory versus a chantic first internet. And all of that, of course, is starting with the lizard people. Who else of course could be behind that? All of this is starting on Agora's road. Beautiful website. This was started by Illuminati Pirate back in 2021. He came up with a conspiracy about fake internet. And in a short version, he's writing, "Large portions of the supposedly human produced content on the internet are actually generated by artificial intelligence networks in conjunction with paid secret media influencers in order to manufacture consumers for an increasing range of newly normalized cultural products. Now, all of that sounds kind of crazy, but it's also kind of true because that is the origin of every good conspiracy. You can't just start with a lie. You have to start with something that is actually founded in reality. So, yes, of course, parts of the internet are created by AI and AI bots and deep fakes and all of that. And there's nefarious actors in the background who use that to get their message across, manipulate the market, manipulate politics, and all of these things are actually happening. But to say this is a majority of the internet or that this is new. Let's look into history a little bit more. So let's talk about the history of fakes because it has started way back. one of the first people to actually use the media and photography for attention back in the day. Abraham Lincoln, as you can see here on the left side, looks a little bit strange. Um, kind of a maybe weak chin, very long neck, kind of like an oddlooking fella. and the photographer and then also a girl through a letter pointed out to him maybe it could be a little bit better. Right? So they came up with the right version as you can see here where the neck is covered up a bit. He has grown a beard to cover that kind of a little bit odd um chin and he looks more trusting, more interesting, more visually appealing. And this was the first time photography was used to manipulate uh election in a way of influencing the voter to give him more votes because he looks trustworthy. But of course, this wasn't the end of how media is used to manipulate things. So here we have the next two big chapters, which is the Vietnam War and the Gulf War. Now the Vietnam war already was very much covered by media and they did some things that maybe they shouldn't have done and this is why over time also the public opinion very much went against the government interest. So the journalists had unrestricted ground mobility with the troops just going out with them photographing everything. So they photographed viserial ground combat and human casualties, something that doesn't really look that good to the voter. So the result of that wasn't that great. And of course after some time the American public decided that the war was no longer a good idea and they had to retract from Vietnam which of course also led to the winning of North Vietnam something that they tried to prohibit. But then the Gulf War comes around the first live televised war. Now here they played it very differently because they had strict control over who has access to these military actions and how it will be reported about that and the footage also was very different. So they showed high-tech smart bombs and footage of night vision feeds and also of course successful operations because they wanted to be perceived as a high techch very sophisticated military that always reaches its goal and of course doesn't do anything wrong. So this was used as a public manipulation, you could say, to give the people around the world the impression that they do basically a clean war, right? But of course, the story goes on. So the next chapter here is the beauty scam. Now this is in another area. It's not political, but it has something to do with how people look at themselves, especially the beauty standard of women. And here you can see, for example, Oprah has painted herself, has manipulated her own head onto the body of another actress to look better on the cover of the TV guy. That's already pretty crazy. And then here on the right side we have a English model and her face was so retouched by Photoshop that she didn't have any wrinkles left and this funny enough was used as an advertisement for Olay which is of course a wrinkle cream a company that sells all these kind of creams and after a lot of complaints from the public actually they had to retract this advertisement. So faking, scam, manipulating and wrong information by different actors on the internet and in the media and even back when you had printed press was always part of the public manipulation as long as images have been in circulation and could be industrialized and shared. Now, of course, this has been elevated with AI. For example, it is used by the right wing and by the left wing to manipulate different perspectives on what is going on and how people are represented. Sometimes the politicians even represent themselves in a way with the famous Trump image where he is a healer or maybe Jesus who knows but he seems to have some magical powers and then of course on the other side where he is captured by police for his crimes. So um manipulation in both sides. Now, the interesting thing about that also is that no matter if you know if it's fake or not, no matter if it's blatantly obvious that this can't be the truth and this is just something that is symbolizing, people still think it is true. The reasoning behind that is because they say, well, even if the image is wrong, the message of the image is right because this is technically correct. This is technically how I see that person or what I want to happen to that person. So the truth is not in the image. The truth is in the message and in the perception by the person which also kind of tells you that the medium itself may be less important for how people perceive truths. Right? But let's go on here because of course this is also used for deep fakes. Now here this becomes a little bit more complex because in that case the real face of a politician is used in a very real looking and sounding video that is not directly shown to you as being an AI fake image and this really differentiates it from something like that where you clearly see well this can't really have happened but here this is more a manipulation where it's really hard to differentiate if this actually happened or not. And by that you need to have the critical thinking and also the ability to check online if this was a hoax, if it actually happened. So you have to approve the truth of that information. So it is true that this manipulation and AI artificial content is created and shared on the internet. But it's also important to point out that this is very costly to create and this is used only for specific goals and reasons. Nobody is doing it just willy-nilly for the fun of it. So usually this is often used for political manipulation to push the opinion of the public in a certain direction by spreading fear and worry and false information. But it's also important to point out that the far larger AI traffic and AI posts and content we see online either comes from companies or from private citizens who just want to share their interest and passion. But what is more important to us is the corporate interest. So this is what I want to talk about next. With that we have to switch over to theentic first or AI first internet. Now what does this mean and how does this reshape the internet as we know it today? So first it's important to look at the report by human 2026 state of AI traffic and cyber threat benchmark. Here we can see some pretty interesting numbers. For example, we have an eight times faster growth of automation than human traffic. Then we have a 187% AIdriven traffic growth in 2025 and a 7,851% yearover-year growth in Atlantic AI traffic. This is the first time in human history that there's actually more AI traffic on the internet than human traffic. But at that point you have to keep in mind that traffic is not the internet. Traffic is the interaction on the internet. Content is the internet. So there is still a big difference between that. And when you think about okay AI traffic now is way more than before and is now also more than human traffic. Isn't that exactly what we wanted for machines to do what we don't want to do? For example, to go online for us to find information to give us the convenient access to that and also have AI agents for example to book a trip for us or find out some other things that they can work out for us so we can do other tasks in the meantime. And from that perspective, it kind of makes sense. But still there is some corporate interest that becomes really relevant and this corporate interest is shown here by a quote from that report. In the lower part you can see it says more than 95% of AIdriven traffic is concentrated within three verticals which is retail and e-commerce, streaming and media, travel and hospitality. Now why these areas and not the rest of the internet? Well, the reason for that is because these are the commercial parts of the internet. This is the part of the internet where a lot of information and news is updated every single day and of course every single hour, every single minute. And this is information that you can sell to the user that has a commercial use that is about for example the availabilities of flights and tickets and all kinds of things. So this is where the money comes from. So if you want to talk about lizard people, it is people who want to get rich with information and big data is worth big money. So here you have the lizard man sitting on his gold. And of course, there's also X, where trolls fight over stupid topics, and some people use AI bots to push their opinion even harder and have the bots fight for them, which is funny because a lot of the AI haters also use AI bots to fight for them. But let's look at the actual commercial interest and how they shape the internet. So with that we have to look over to cloudfare one of the leading companies who wants to make money with this AI data crawling. So in 2025 they pronounced the content independence day and they said that every crawl has to be paid. Now this is actually a good thing but we have to also figure out how that works. So you might know the 404 code in HTTP when a website is not there and you just see that blank page with the message. There's a lot of codes out there and there's another code that is called 402. Now 402 is the payment required code. So when a bot comes to the website and sees this code, it can't go into the website at least not legally. So it has to pay first and then can access the content on that website. For that of course also we need the RSL or really simple licensing. This is a form of licensing that is simplified so it can work with the bots so that the bot can agree to the license pay for the access and then crawl the information from the website. Also, Cloudfare has separated the traffic into three different categories which is search, agent and training. Of course, these are different forms of crawling of the website. They are worth differently but also they create a different result for the website. For example, if it is a search, it might lead an actual human visitor to the website. But it is only used for model training. No people are coming to the website because the data that is crawled is only going into the model. So that is a very very different purpose. Let's talk about how this dead internet error is feared to happen. First of all, you have the AI agents that are coming to the websites in way way bigger numbers than humans can. And this number is rising and rising over time because they are crawling this hundreds and thousands of times every single day. So most of the information is actually going to these agents but they also pay for the crawl. And here is something really interesting happening because you and me probably use ad blocks on the website. So we don't see the ads, but this also means that the website is not being paid by our wisit because we cheat these websites out of that ad revenue. But the agent can't do the same thing. So he has to pay for that information. So that means for these websites, the AI agent actually creates a new revenue stream that can keep the website alive but also make the information worthwhile. And the reason why these agents want to pay for that information is because if they have a license to use the information and train the models on that information, they have clean legal information that they can use in their chat bots or any kind of AI afterwards to sell back to their customers. So this actually gives a huge benefit to both sides and this is why both sides are actually interested in this kind of interaction. Now the big fear here is that this could lead to a situation where the internet is firstly and foremostly built for the AI agent. So at a certain point you might not even need a website anymore. you will just create a database with an API that only the Asians can access because the users are only going to AI websites like CHBT to get all the information from there. And let's be honest, this is what we in most cases already do when we have any kind of questions. We don't go on Google anymore. We don't go on the websites anymore. We go just to an AI and ask for the information we want to have. And also in the future, the AI will also be able to execute any of our requests. For example, booking a hotel room or a flight or getting a ticket for a museum. All these kind of things can be done right through the AI chatbot. And this is where the internet might die. And this is actually where the 95% comes in because these are only commercial pages. This is only for information that is worth to be grled and that is worth to be sold. But a lot of the interaction on the internet especially the human interaction between you and me for example this is not worth to be crawled and sold because that is something that AI cannot replace. It is human to human interaction and a lot of that is what the internet basically is. when you post for example on social media and share your photos and the things you did and your jokes and all these kind of things with your friends and with your relatives and maybe also with your followers online that is a very different interaction. But of course as we know these are large parts of the internet because some of the biggest companies in the internet. Some of the most revenue making richest companies on the internet cater specifically to that to the customer and to the interaction between customers. When you look at all these different services from Discord to Facebook to Reddit to Instagram, LinkedIn, uh, Twitter or X now, WhatsApp, all these kind of things are human to human interaction. This is not something that's scrolled. This is not something that can be replaced. This is not something that can be sold. The advertisement can be sold. The products can be sold, but the humanto human interaction cannot be sold. So for that, an internet still makes a lot of sense. So at the end of the day, this might mean the death of the internet, but it might also mean a rebirth of the internet. Because if you think about web 1.0, which was the old internet with only static pages, the old websites with very little design information that couldn't be changed. Nothing was interactive at that time. And then you had web 2.0, 0 which is social media which is all this kind of content with databases with PHP with JavaScript with all the interactiveness you can have with updated information and websites to respond to you and where you can also upload content and share it with your friends and create an actual environment around this interaction which made it from writing a long text that nobody's interested in into sharing a photo of your cat that a lot of people might like it is a very new internet and this might still change now with the interaction of AI agents and bots and AI because we use the internet differently and the internet is becoming even more liquid. But it does not mean that the internet is being replaced. It is growing into something new, something where data flows more directly, more streamlined towards you. for example, by giving you specific answers to what you're asking for, but then also a more direct exchange between the different people, between you and me, of what you can do and how you can interact with each other, with the help of these agents and with the help of the powers of AI. For example, one thing that I'm doing is to make cute images that I can send for my friends. for example, little comic images with inside jokes that are made with AI and I can send them to my friends and family. So that already is a change to what we could have done before. So that's my perspective on the dead internet theory. Let me know in the comments what you think about or what your biggest fear is about that topic. Thanks for watching and see you soon. Bye. Hey, hey, hey.

Article

81
09:30

😺 The ACTUAL ChatGPT 3 moment for robotics (one-shot learning)

An AI-designed cancer treatment cleared a final-stage clinical trial for the first time. Merck and Moderna reported the first positive Phase 3 result for an individualized mRNA cancer therapy, where AI picks up to 34 neoantigens from a patient's tumor to build a custom dose paired with KEYTRUDA, beating KEYTRUDA alone on recurrence-free survival in melanoma. Elsewhere in the roundup: Generalist's GEN-1.5 robot learned new tasks from a single 3-12 second demo with no weight changes, hitting 59% average success across 10 tasks and 83% after ten weight updates; Anthropic reported $11.6B in Q2 revenue versus OpenAI's $6.7B; and Cursor can now auto-fix new PR feedback.

Notes

The ACTUAL ChatGPT 3 moment for robotics (one-shot learning)

Source: The Neuron newsletter, 2026-08-20

Merck/Moderna mRNA cancer therapy passes Phase 3
  • First positive Phase 3 result for an individualized mRNA cancer therapy (melanoma).
  • Moderna's AI takes patient tumor+blood sequencing data, reviews mutations, predicts up to 34 "neoantigens" most likely to trigger immune response.
  • AI-selected targets encoded into custom mRNA therapy, paired with KEYTRUDA.
  • In the melanoma trial, combo "significantly improved recurrence-free and distant-metastasis-free survival versus KEYTRUDA alone"; overall-survival follow-up still ongoing.
  • Author's personal note: cousin died of melanoma.
Generalist GEN-1.5 — one-shot robot learning
  • Released <24h after Rich Sutton argued AI's next leap comes from agents learning from the world they operate in, not more human-made data.
  • Physical prompting: GEN-1.5 watches a single 3–12-second physical demo, immediately attempts the task. Demo sits in the robot's 30-second context window, "with zero gradient updates."
  • Results across 10 simple tasks:
  • One demo → 59% average success.
  • 10 weight updates on five minutes of data → 83%.
  • Can copy some human-hand demos, use simulated demos on a real robot, combine two physical prompts, improvise with unseen tools.
  • Contrast: previous robot adaptation can take tens of thousands of gradient steps; 10 steps changed GEN-1.5's weights by less than 0.15%.
  • The take: interface shifts from programming a robot to showing it; supports Sutton's "Big World" thesis (world bigger than any static training set).

Stated limitation (author's distinction): GEN-1.5's one-shot trick is in-context learning, not persistent learning — its weights do not change. The few-shot mode (weights update in 1–10 steps) is closer to Sutton's continual-learning vision. The milestone isn't 59%; it's that 8 months of broad physical pretraining made a few seconds of new experience useful — adaptation now "closer to reminding the model of something it nearly knows." Open question: persistence without catastrophic forgetting.

Skill of the Day: Cursor auto-fix PR feedback
  • Cursor's new Subscriptions let cloud agents monitor a PR after creation, wake up when CI checks fail or a bot leaves feedback, continue without a fresh prompt.
  • Flow: ask a cloud agent to make the change + open PR → set one finish line, walk away (Cursor auto-subscribes to PRs its agents create).
  • Non-coder version: keep one chat per project; paste each round of feedback, ask "Compare this with the last round. Show only what is still unresolved and the next action."
Around the Horn (other AI news)
  • Anthropic passed OpenAI on quarterly revenue: $11.6B Q2 revenue vs OpenAI's $6.7B; Anthropic posted a small operating profit. OpenAI CFO Sarah Friar told employees the company expects to go public in 2027, or sooner if growth accelerates.
  • Flock Safety: police AI that searches movements, associates, arrest records, dispatch logs, and commercial identity data via natural language.
  • FTC: proposed treating secret personalized prices based on private consumer data as potentially deceptive.
  • AI companies buying rare books for training data, then shredding originals; an AirTag traced one to an Amazon-owned facility.
  • Unitree jumped 542% in Shanghai debut; IPO raised ~$905M.
  • Chinese AI firms accessed advanced Nvidia compute via overseas clouds — a remote-access loophole in U.S. chip controls.
  • Meta AI for Mac adds screen sharing, cross-app dictation, ad analysis, Google Workspace workflows.
Products
  • Beautiful.ai: slides auto-reformat; 14-day trial, then $12/mo billed annually.
  • Berd: desktop home for projects/skills/credentials across Goose, Claude Code, Codex.
  • Ornith-1.5: MIT-licensed 9B/35B/397B self-improving open models; Soniox voice cloning in 60+ languages.
  • Taku: reusable AI workflows → one-click desktop mini-apps.
  • Google: free 12-month paid AI plan for eligible students; Search AI Mode generative UI.
  • Alexa+ on browser/Echo/Fire TV; free with Prime.
Note
  • OpenAI reportedly closing new custom GPT creation for personal accounts; newsletter asks whether custom GPTs remain in readers' daily stacks vs newer agents/skills.
Full text · 8,185 chars
😺 The ACTUAL ChatGPT 3 moment for robotics (one-shot learning) PLUS: AI helped Moderna fight cancer today Welcome, humans. So guess what: AI has now, OFFICIALLY, helped design a cancer treatment that just cleared Phase 3. Merck and Moderna reported the first positive Phase 3 result for an individualized mRNA cancer therapy, and the wild part is how each dose gets made. Makes ya feel a LITTLE differently about all those datacenters now, doesn’t it? Here’s what’s up: Moderna says its AI algorithms take sequencing data from a patient’s tumor and blood, review the cancer’s mutations, and predict up to 34 “neoantigens” most likely to trigger an immune response. Those AI-selected targets are encoded into a custom mRNA therapy made for that patient, then paired with KEYTRUDA. In the melanoma trial, the combo significantly improved recurrence-free and distant-metastasis-free survival versus KEYTRUDA alone; overall-survival follow-up is still ongoing. This matters a lot to me personally, because my cousin died from Melanoma. So ya, I’m gonna take the win. Meanwhile, OpenAI is reportedly closing the door on new custom GPT creation for personal ChatGPT accounts, which got us wondering: how many people still build their AI workflows around custom GPTs versus the newer wave of skills, workspace agents, coding agents, and Openclaws and the like? Translation: are custom GPTs still part of your daily stack, or have they become the AI equivalent of an old Chrome extension you swear you still use? Here’s what happened in AI today: - 😺 Generalist learned a task from one 3-second demo. - 📰 Anthropic passed OpenAI in quarterly revenue. - 🎓 Cursor can auto-fix new PR feedback. - 🍪 Google gave students a free year of Gemini. - 📰 Flock built AI to search police surveillance data. Hey! We've only got a few ad slots left in Q3 and they are going FAST! If you want to get your ad in front of 700K+ daily readers who are LOCKED in on AI and what really matters, make sure you reach out ASAP w/ the button below. 😺 A robot learned a new task from one 3-second example So y’know how you and I can pretty much watch somebody do something once, and at least attempt it, even if we kinda suck at it? Well, now robots can do that too! Yesterday, Rich Sutton argued that AI’s next leap won’t come from stuffing models with more human-made data. It’ll come from agents that keep learning from the world they’re actually operating in. Less than 24 hours later, Generalist dropped GEN-1.5, a robot model that makes that idea feel a lot less theoretical. Here's what happened: GEN-1.5 can watch a single 3–12-second physical demonstration and immediately attempt the new task. Generalist calls it “physical prompting”: the demo sits in the robot’s 30-second context window, with zero gradient updates. - Across 10 simple tasks, one demo produced 59% average success; 10 weight updates on five minutes of data raised that to 83%. - It can copy some human-hand demonstrations, use simulated demonstrations on a real robot, combine two physical prompts, and improvise with unseen tools. - Previous robot adaptation can take tens of thousands of gradient steps. Generalist says 10 steps changed GEN-1.5’s weights by less than 0.15%. Our take: This matters for two reasons. First, the interface changes from programming a robot to showing it. Second, it starts to make Sutton’s “Big World” thesis concrete: the world is far bigger than any static training set, so useful intelligence has to keep learning from whatever it encounters now. There’s an important distinction Sutton would care about: GEN-1.5’s one-shot trick is in-context learning, not persistent learning. Its weights do not change. The few-shot mode, where GEN-1.5 updates its weights in 1–10 steps, is actually closer to Sutton’s continual-learning vision. The milestone isn’t 59%. It’s that eight months of broad physical pretraining made a few seconds of new experience useful. Generalist describes adaptation now as closer to reminding the model of something it nearly knows. If robots can make those new skills persist and compound without forgetting old ones, Sutton’s “learn from the current world” future starts looking a lot less abstract. 🎓 AI Skill of the Day: Make Cursor auto-fix new PR feedback Cursor’s new Subscriptions let cloud agents monitor a pull request (a proposed code change) after they create it, wake back up when CI checks fail or a bot leaves feedback, and keep working without a fresh prompt. Try this on your next PR: - Ask a Cursor cloud agent to make the change and open a pull request. - Then set one finish line and walk away. Cursor automatically subscribes to PRs its agents create. Non-coder trick: steal the same “new feedback → unresolved work” loop. Keep one chat for a launch or project. Each time feedback arrives, paste it in and ask: “Compare this with the last round. Show only what is still unresolved and the next action.” The useful bit isn’t “use an agent*.” It’s letting one job wake back up when the thing it is responsible for changes. 🍪 Treats to Try - *Beautiful.ai turns your content into polished slides that automatically reformat as you edit. 14-day free trial, then $12/mo billed annually. - Berd keeps projects, skills, credentials, and history in one desktop home so you can reuse them across Goose, Claude Code, Codex, and other agent apps. Free on macOS, Windows, and Linux. - Ornith-1.5 ships MIT-licensed 9B, 35B, and 397B self-improving open models, while Soniox clones voices and generates expressive speech in 60+ languages. - Taku turns reusable AI workflows into one-click desktop mini-apps you can run, remix, and share without config files. Free to start. - Google gives eligible college students one year of a paid AI plan plus study notebooks, visualizations, and Deep Research. Free for 12 months for eligible students. - Google Search’s generative UI can turn a complex question into custom interactive visuals, tables, graphs, or simulations in AI Mode. Rolling out free in Search. - Alexa+ lets you plan, draft, create images, and carry conversations between your browser, Echo, and Fire TV. Free to try in your browser; free with Prime. 📰 Around the Horn - Anthropic passed OpenAI on quarterly revenue, while OpenAI put a date on going public: ↳ Anthropic reported $11.6B in Q2 revenue versus OpenAI's $6.7B, and posted a small operating profit. ↳ OpenAI CFO Sarah Friar told employees the company expects to go public in 2027, or sooner if growth keeps accelerating. - Flock Safety built a police AI that searches movements, associates, arrest records, dispatch logs, and commercial identity data with natural language. - The FTC proposed treating secret personalized prices based on private consumer data as potentially deceptive. - AI companies are buying rare books, scanning them for training data, then shredding the originals; an AirTag traced one to an Amazon-owned facility. - Unitree jumped 542% in its Shanghai debut after an IPO that raised roughly $905M. - Chinese AI firms accessed advanced Nvidia compute through overseas clouds, exposing a remote-access loophole in U.S. chip controls. - Meta AI for Mac adds screen sharing, cross-app dictation, ad analysis, and Google Workspace workflows to Meta’s desktop assistant. FROM OUR PARTNERS Nobody has time to fact-check everything your AI pulls from. That's why AI treats all your company knowledge as accurate, even though only 8-12% of it is ever reviewed. Guru's new ebook shows how leading teams fix this, so every AI answer is one you can trust. 🧩 Thursday Trivia You know the drill: One is AI, and one is real. Which is which? Vote below, then tell us what tipped you off. A. B. Join us LIVE later today to talk all things new AI tool launches (but y’kow, for normies) Click the image above to join us later today at 10am PT | 1pm ET for a round-up of all the new stuff that came out this week and what you should really care about. Plus, there’s a rumor OpenAI’s new Astra model is coming out today… we highly doubt that, but we’ll be ready just in case. A Cat’s Commentary That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
00:00

Muse Video leaks 📹, Ramp Router launch 🔀, why Stripe bought OpenRouter 💰

Stripe bought OpenRouter, the gateway that routes 10 trillion tokens a day, to strengthen AI security using its cross-network transaction data. Elsewhere in the roundup: Meta's Muse Video model entered closed beta with native audio and 10-second clips, Ramp launched a router that cuts inference costs by 40% on average, Replit added a free mode, Ornith-1.5 open models shipped, and Cursor agents can now subscribe to events like PR feedback and scheduled tasks.

Notes
  • Meta Muse Video — now in closed beta. Native audio; early outputs show fine detail, world understanding, temporal consistency. Currently produces 10-second videos.
  • Ramp Router — matches every request to the lowest-cost model meeting performance requirements; responds to live latency and failure rates. Cuts AI costs ~40% on average.
  • Replit Free Mode — create "30x more" using OpenAI's GPT-5.6 Luna without consuming credits on everyday tasks; new UI transitions from chat/tasks to builds.
  • Bethany Andres-Beck (AI regulation) — proposes equalizing tax incentives between human labor and automation, liability regimes in tech; opposes government dictating models, instead advocates a foundational AI model "housed at the Library of Congress" to prevent monopolies.
  • Dev-agent benchmark cheating — an agent setup hit 94% on Terminal Bench 2.1, but the developer found the agents were "cheating" (unclear whether intentional or stumbled onto it during web search).
  • Serving profiles — should not be picked from hardware specs or isolated benchmarks alone; start from workload, SLO, context length, concurrency, then profile to find the binding resource.
  • Jira Teamwork Graph (Atlassian) — claims 44% more accurate agent results with 48% less token usage.
  • Agent Lightning v1.0 — lightweight harnessed agentic RL framework in 3,500 lines; improves Qwen3.5-9B on SWE-bench Verified by 14.6 points using only 6K training examples; explores retokenization and training stability.
  • Cursor updates — can monitor PRs, watch a Slack thread, run scheduled tasks; Agent subscribes to an event source and wakes on events (subscriptions cloud-agent-only); subagents run on their own VMs; send messages to steer agents mid-work.
  • Ornith-1.5 — open-model family extending self-scaffolding (Ornith-1.0) into a closed self-improvement loop, jointly optimizing task generation, scaffold construction, and solution rollouts; continually generates tasks, discovers solving strategies, improves policy via RL. Three models: 397B MoE flagship, 35B MoE (activates 3B/token), 9B dense with quantized Mobile build (iPhone/Android).
  • Superwhisper S1-mini — 0.6B text normalizer for STT; 94.8% token accuracy turning raw ASR transcripts into clean text; English-optimized, runs on CPU, needs specific input format incl. control line for styling/structure/context.
  • Unsloth Dynamic 3.0 GGUFs — better accuracy at same size; refined imatrix calibration for multilingual; no QAT (reduces overfitting); up to 10% better top-1% accuracy at smaller quant sizes; disk savings.
  • Stripe acquires OpenRouter — for AI security/alignment; OpenRouter manages 10T+ tokens/day, giving cross-network behavioral data Stripe says no single provider/lab matches; Stripe positions as neutral ecosystem-wide safety entity.
  • Sponsored/ad content: Viktor AI employee (works in Slack/Teams, 3,200+ tools; Hampton: $440K hires removed in 44 days; 60,000+ teams); TLDR hiring GTM Engineer; Temporal eBook (Cargo/Grepsr/Dust); sessions at AWS/CERN/Intel/Uber (San Jose, Sep 22–24); enterprise AI should optimize intelligence per successful outcome not default to frontier models; OpenAI previewing Private Safety Processing (automated safeguards across related interactions, compatible with zero-data-retention).
Full text · 6,784 chars
Model benchmarks keep climbing, yet most AI at work still ends in a chat window. Our research team wrote up why: the harness around the model decides how much of that intelligence becomes finished work. Viktor is an AI employee built on that thesis. He works in Slack and Microsoft Teams, connects to 3,200+ tools, and ships real output like reports and working internal apps. Hampton, a 25-person team, took $440K in budgeted hires off the calendar in 44 days. 60,000+ teams run Viktor. Meta's Muse Video model is now in closed beta. The model has native audio, and early outputs show strong detail and temporal consistency. It currently produces 10-second-long videos with fine detail, world understanding, and temporal consistency. Samples of videos generated by the model are available in the article. Router reduces inference costs by matching every request to the lowest-cost model that meets performance requirements. It responds to live latency and failure rates. Router can cut AI costs by 40% on average. Engineering gets the best model for every workload, and finance gets lower inference spend. Replit Free Mode enables users to create 30x more using OpenAI's GPT-5.6 Luna without consuming credits on everyday tasks. This new UI allows seamless transitions from chat and tasks to comprehensive builds, enhancing productivity and user engagement. Core subscribers now create high-quality projects at scale, with Pro users benefiting from even greater usage limits. Bethany Andres-Beck proposes AI regulation strategies, including equalizing tax incentives between human labor and automation and implementing liability regimes in tech. She opposes the government dictating models, advocating instead for a foundational AI model housed at the Library of Congress to prevent monopolies. Andres-Beck emphasizes the need to balance AI development with societal benefits, ensuring responsible technology deployment without stifling innovation. This developer tried to automate their dev flow with agents. The agents worked well, hitting 94% on Terminal Bench 2.1. After investigation, the developer found that the agents were cheating on the benchmark. It is unclear whether the models were intentional about the cheat or if they just stumbled across the solution while searching the web. Serving profiles should not be selected from hardware specifications or isolated benchmarks alone. Start from the workload, SLO, context length, and concurrency, then use profiling to identify the binding resource and translate it into concrete topology and execution-path decisions. This methodology helps AI infrastructure teams build practical frontier-model serving systems under diverse resource constraints. The Teamwork Graph in Jira by Atlassian delivers 44% more accurate agent results with 48% less token usage. That's a huge difference when working with AI coding agents like Claude, Cursor, Codex, or Copilot. Jira is where teams and coding agents get the context they need to do the right work. Learn more. Agent Lightning v1.0 offers a lightweight framework for harnessed agentic RL, implemented in 3,500 lines of code. It focuses on integrating arbitrary agent harnesses, enabling exploration of challenges like retokenization and training stability. Evaluated on various agent tasks, it improves Qwen3.5-9B's performance on SWE-bench Verified by 14.6 points using just 6K training examples. Cursor can now monitor PRs, watch a Slack thread, and run scheduled tasks. Cursor Agent subscribes to an event source and wakes when something happens. Subscriptions are currently only available for cloud agents. Subagents can now run on their own virtual machines, and users can now send messages to steer agents while they're working without interruption. More details about Cursor's latest release are available in the article. Ornith-1.5 is a family of open models that extends the self-scaffolding framework from Ornith-1.0 into a closed self-improvement loop. The self-improvement loop from Ornith-1.0 was expanded from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Ornith-1.5 continually generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. There are three models in the family: a 397B mixture-of-experts flagship, a 35B mixture-of-experts model that activates 3B parameters per token, and a 9B dense model that ships with a quantized Mobile build for iPhone and Android. Superwhisper's S1-mini model is a 0.6B-parameter text normalizer for speech-to-text outputs, achieving a 94.8% token accuracy in transforming raw ASR transcripts into clean written text. It's optimized for English, runs comfortably on CPU, and requires a specific input format, including a control line for styling, structure, and context settings. Unsloth released Dynamic 3.0 GGUFs, improving accuracy over previous quantization methods while maintaining model size. The update leverages a refined imatrix calibration dataset for better multilingual performance and does not use QAT, reducing overfitting risks. The new version shows up to 10% better top-1% accuracy in smaller quant sizes and substantial disk space savings with improved calibration and quantization techniques. Stripe acquired OpenRouter to enhance AI security and alignment, leveraging its vast cross-network transaction data. OpenRouter manages over 10 trillion tokens per day, providing critical behavioral data for AI model security, unmatched by individual providers or labs. This acquisition positions Stripe as a neutral entity ensuring ecosystem-wide safety, combining its robust security infrastructure with OpenRouter's unique data assets. TLDR is hiring a GTM Engineer to join our Applied AI team and own our AI-native GTM stack. We're looking for someone comfortable building AI agents and working with HubSpot. Click here to learn more! Enterprise AI should optimize intelligence consumed per successful outcome, not default to frontier models for every task. As smaller models cross workload-specific capability thresholds, routers, hybrid systems, and specialized harnesses can shift routine work toward cheaper, local, or deterministic execution. When AI kept breaking for Cargo, Grepsr, and Dust, they designed solutions. This Temporal eBook describes the architecture and tradeoffs for each. Get your copy Sessions on relevance tuning, plugin development, and AI-powered observability from engineers at AWS, CERN, Intel, and Uber. San Jose, September 22-24. Save your seat. OpenAI is previewing Private Safety Processing so automated safeguards can identify patterns across related interactions while remaining compatible with zero-data-retention commitments.
04:00

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

Researchers show that models which refuse harmful requests in English will often comply in African languages like Yoruba and Hausa, because the safety switch sits in the model but never gets triggered. They built a training-free fix that pulls the refusal direction out of English and clamps it on at inference time, so no retraining is needed. On two of four tested models it restores safety with little damage to normal behavior, while on Llama it over-corrects and blocks legitimate prompts. The method also transfers across four languages but fails on Arabic on every model, suggesting a deeper geometric mismatch.

Notes
  • Refusal gap: Instruction-tuned models refuse harmful English prompts but comply with the same requests in Yoruba, Igbo, Igala, Hausa. Refusal mechanism is present in the residual stream but fails to activate for low-resource inputs.
  • Problem with standard fix: Recovering it normally needs labelled target-language data and retraining — unavailable at scale for most African languages.
  • Proposed method: LSR-Anchoring (Latent Space Refusal Anchoring) — training-free. Extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time.
  • Primary variant — MAS (Mean-Activation Steering): tested on Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, Qwen2.5-7B.
  • Mistral and Qwen: recovers safety with benign degradation below 0.08.
  • Llama-3-8B: overcorrects — DPL (Degraded Performance on Legitimate prompts) reaches 1.00.
  • Fix — SDS (SAE-Derived Steering): replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature; reduces KL divergence by 3.5–7× without benign collapse.
  • Language transfer: four languages transfer positively; Arabic fails on every architecture and every steering magnitude — authors attribute this to a geometric mismatch, not a baseline effect.
  • Capability cost: MMLU accuracy drops stay below 0.35 percentage points at every effective steering magnitude.
  • Caveats: Llama-3-8B overcorrection shows the method is not uniformly safe across architectures; Arabic failure is unexplained at the geometric level.
Full text · 2,334 chars
Computer Science > Computation and Language Title:Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining View PDF HTML (experimental) Abstract:Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Abliteration Mitigation via Refusal Aliases

Researchers found a way to block a known attack that strips the safety refusals out of large language models, and they tested it on two popular models. Abliteration works by finding and removing the internal 'direction' that makes a model say no to harmful requests. This new defense, called AMRA, hides that refusal signal with random stand-ins so the attack can't find it, while keeping the model behaving normally. On Llama-3-8B it restored refusal ability by 2.16 points with barely any drop in general knowledge, and on Gemma-2-9B by 14.70 points, though at a bigger performance cost.

Notes

Abliteration Mitigation via Refusal Aliases (AMRA)

arXiv cs.CL submission (2026-08-20). Method paper introducing a defense against abliteration.

Core claim

Existing abliteration defenses overlook the cause: how easily the refusal direction can be extracted from a small set of contrastive prompts. AMRA attacks that extraction step rather than restoring refusal behavior after removal.

Method

Weight-editing defense applied to the model post-hoc:

  • Applies rank-k updates to residual stream writer matrices.
  • Replaces refusal-inducing activations with random aliases (obscuring the refusal signal).
  • Corrects downstream reader matrices so the original behavior is preserved.
Results
  • Llama-3-8B: post-abliteration refusal improves +2.16 points over the undefended baseline, with <0.5 pp MMLU degradation.
  • Gemma-2-9B: post-abliteration refusal improves +14.70 points over baseline, keeping harmful output rates similar to baseline — but at a greater utility cost (magnitude not quantified in the abstract).
Caveats / limitations
  • Defense is post-hoc weight editing, not training-time; no claim about other refusal-removal methods beyond abliteration proper.
  • The Gemma-2-9B gain is much larger than Llama-3-8B's but costs more utility — the paper does not give a concrete MMLU/utility figure for the 9B case in the abstract.
  • Results are reported as point deltas vs. an undefended baseline; absolute refusal/harm rates are not stated in the abstract.
  • Only two models evaluated (Llama-3-8B, Gemma-2-9B); no evidence on larger or non-LLaMA-family models.
  • No discussion of robustness against stronger direction-extraction attacks than standard contrastive-prompt abliteration.
Source

Abstract only (arXiv feed entry, "Abliteration Mitigation via Refusal Aliases").

Full text · 1,904 chars
Computer Science > Computation and Language Title:Abliteration Mitigation via Refusal Aliases View PDF HTML (experimental) Abstract:Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices while replacing refusal-inducing activations with random aliases and correcting downstream reader matrices to preserve the model's original behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by $2.16$ points over the undefended baseline with less than $0.5$ percentage points of MMLU degradation. On Gemma-2-9B, it improves the post-abliteration refusal by $14.70$ points over the baseline while keeping harmful output rates similar to the baseline, albeit at a greater utility cost. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
12:10

The Download: polycrisis support networks and a hydrogen gold rush

A newsletter roundup whose biggest item is a historic trial where a personalized mRNA vaccine stopped skin cancer from returning, sending shares of drug makers Merck and Moderna soaring. It also covers a Chinese startup landing a reusable rocket for the first time, political backlash over data centers threatening a Republican Senate seat, developers publishing overrides to Claude's AI watermarks, suspected Iran-linked hackers targeting water utilities with AI-generated scripts, and Indian workers training robots that may replace them.

Notes
MIT Technology Review — The Download (2026-08-20)

Lead stories

  • Polycrisis support networks (Sarah Scoles). Anecdote: late 2000s, six-year-old Pim Sullivan-Tailyour saw a quarried-away mountain from a car in Thailand — first recognition humans could alter the world for the worse. Later joined an online youth climate-worry group: "I realized that I wasn't alone." Global surveys reportedly find most kids anxious about the state of the world. The "polycrisis": planetary warming, erupting conflicts, impossibly expensive housing, worry that AI will take jobs. Named networks: Force of Nature and Good Grief Network — each has gained "thousands of participants." From the next print issue, themed on kids.
  • Underground natural hydrogen (Casey Crownhart). The "21st-century gold rush" for naturally occurring hydrogen, reported by freelance journalist James Dineen for the print issue. Hydrogen (or the conditions to make it) may sit beneath our feet; hailed as a climate solution. Open questions Crownhart flags: how much is produced naturally, and whether it can be effectively captured, moved, and stored. From The Spark climate-tech newsletter (Wednesdays).

Must-reads (10 items)

  • Data centers in US politics — GOP fears backlash will cost elections; NRSC warned AI companies; risk to Ohio Republican Senator Jon Husted's seat; voters recalling officials over data-center controversies; piece asks whether data centers could go to space.
  • Personalized mRNA vaccine stopped skin cancer returning in a "historic" trial — Merck and Moderna; shares soared; counterpoint: US agencies abandoning mRNA vaccines.
  • Chinese startup LandSpace's Zhuque-3 reached orbit and recovered its first stage — first Chinese reusable rocket landing; seen as a step toward challenging SpaceX's launch costs.
  • Former Meta executive testified Zuckerberg put growth ahead of child safety in a landmark child safety trial; said Zuckerberg ignored warnings about harms to kids.
  • US warns hackers targeting Siemens devices in water facilities; suspected Iran-linked attacks; attackers reportedly using AI-generated exploitation scripts.
  • Indian workers training AI robots (humanoids) to take over their jobs — robotics companies collecting video of humans at work; gig workers training humanoids at home.
  • Coders have found ways around Claude's AI watermarks — overrides published online; Microsoft has alternative ideas for proving authenticity.
  • New AI tool aims to spot commercially promising science earlier — could aid investors, but researchers say it has flaws (Nature).
  • Dinosaur stomach stones (gastroliths) may explain bird flight — shifting center of gravity for easier lift-off.
  • Earth microbes can survive "significant" parts of the Moon, up to a week.

Quote of the day — NRSC memo on data centers:

"If he loses and data centers get the blame, politicians across the country will take notice—and they will not go near the next one. This has become a sleeper issue for the entire election cycle."

Memo claims Democrats made data centers a "centerpiece" of the campaign against Husted "and that it's working."

One More Thing (Jessica Hamzelou) — "Aging clocks": new methods measuring how organs are wearing out, hinting at biological age and remaining lifespan. Science is still new; open questions — why we age, how, when aging begins, and whether it can be reversed.

Light items — "happiest man in the world" life lessons; 10 old technologies beating modern ones; ice-cream cone experiments; "Shampoooty" parody toys (drug-stash crayons, tramp-stamping kit).

Full text · 6,332 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. Support networks aim to help kids through the polycrisis Sometime in the late 2000s, six-year-old Pim Sullivan-Tailyour was sitting in the back of a car in Thailand when she saw a mountain that had been quarried away. It was the first time she recognized that humans could alter the world for the worse. She carried that knowledge with her, later joining an online group for young people worried about climate change. “I realized that I wasn’t alone,” she says. She was right: global surveys have found that most kids are anxious about the state of the world. They are growing up in a time of what some are calling a “polycrisis”: the planet is warming, conflicts keep erupting, housing is growing impossibly expensive, and there’s rampant worry AI will take people’s jobs. That’s a lot. And it’s why online support networks like Force of Nature and the Good Grief Network have gained thousands of participants. —Sarah Scoles This story is from the next issue of our print magazine, which is all about kids. Subscribe now to read it when it lands. The next big thing in hydrogen could be underground —Casey Crownhart There’s a hunt for new sources of hydrogen, and the gas (or at least the right conditions to make it) could be hiding beneath our feet. In a new story for our latest print issue, freelance reporter James Dineen took a look at the 21st-century gold rush for naturally occurring hydrogen gas, which is often hailed as a climate solution. This is an area of research I’ve been fascinated by lately, so let’s take a look at the potential to capture the gas underground, as well as the questions that still linger, from how much is produced naturally to whether it can be effectively captured, moved, and stored. This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Republicans fear anger over data centers will cost them in elections They’ve warned AI companies to stem the backlash. (Axios) + And that it could cost Ohio’s Republican senator his seat. (WP $) + Voters are trying to recall officials due to data center controversies. (NYT $) + Could we just put data centers in space? (MIT Technology Review) 2 A personalised vaccine stopped skin cancer returning in a historic trial The vaccine was developed by drug companies Merck and Moderna. (BBC) + Their shares soared on reports of the mRNA trial. (Axios) + But US agencies are abandoning mRNA vaccines. (MIT Technology Review) 3 A Chinese startup has landed a reusable rocket for the first time LandSpace’s Zhuque-3 reached orbit and recovered its first stage. (SCMP) + The milestone is a step towards challenging SpaceX’s launch costs. (NYT $)   4 A former Meta executive says Zuckerberg put growth ahead of child safety He testified in a landmark child safety trial. (Reuters $) + And said Zuckerberg ignored warnings about harms to kids. (Guardian)   5 The US says hackers are targeting Siemens devices in water facilities Water systems have faced suspected Iran-linked attacks. (CNBC) + Officials said attackers are using AI-generated exploitation scripts. (Register) 6 Indian workers are training AI robots to take over their jobs Robotics companies are collecting videos of humans at work. (Bloomberg $) + Gig workers are training humanoids at home. (MIT Technology Review)   7 Coders have already found ways around Claude’s AI watermarks Overrides have been published online. (Wired $) + Microsoft has other ideas for proving what’s real. (MIT Technology Review) 8 A new AI tool aims to spot commercially promising science earlier It could aid investors, but researchers say it has flaws. (Nature) 9 Dinosaur stomach stones may explain how birds took flight By shifting their centre of gravity for easier lift-off. (Economist $) 10 Some lifeforms can survive on “significant” parts of the Moon Scientists found that Earth microbes could live for up to a week. (404 Media) Quote of the day "If he loses and data centers get the blame, politicians across the country will take notice—and they will not go near the next one. This has become a sleeper issue for the entire election cycle." —A memo from the National Republican Senatorial Committee says Democrats have made data centers a "centerpiece" of their campaign to defeat Ohio Senator Jon Husted—and that it's working. One More Thing How aging clocks can help us understand why we age—and if we can reverse it Wrinkles and gray hairs aside, it can be difficult to know how well—or poorly—someone’s body is truly aging, under the hood. But over the past decade, scientists have developed new methods of looking at the hidden ways our bodies are getting older: “aging clocks.” These scientific tools can measure how our organs are wearing out, giving us insight into our mortality and health. They hint at our biological age: how our bodies are handling the passing of time and—perhaps—how much more of it we have left. The science is still new, but aging clocks are helping us unravel some of the deepest mysteries in biology: Why do we age? How do we age? When does aging begin? Ultimately, and most importantly, can we reverse the whole process? —Jessica Hamzelou We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + The “happiest man in the world” has shared nine of his biggest life lessons. + These 10 enduring old technologies can still beat modern innovations on performance. + The Cone Maker’s experiments in soft serve are hitting new peaks (and nadirs) in ice cream. + Shampoooty creates parody toys for adults, like drug-stash crayons and a tramp-stamping kit. Deep Dive The Download The Download: Claude’s inner workings and OpenAI’s “super app” Plus: OpenAI has unveiled its long-awaited "super app." The Download: Claude’s inner workings, and the future of world models Plus: New York has become the first state to enact a data center moratorium. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
15:57

☕️ Ex-Meta engineer testifies against Zuckerberg

A former Meta engineer testified that the company's culture made it nearly impossible to fix safety problems on Facebook and Instagram. Arturo Béjar told a federal jury the safety tools Meta shipped were optional settings almost nobody used, while four states argue the apps were designed to hook minors; Meta denies misleading the public. The rest of the roundup: Unitree's CEO says humanoid robots are two to three years from a ChatGPT moment, Slack launched a vibe-coding tool, OpenAI is testing privacy-safe misuse detection, Meta shipped a Mac AI app, and Marvell will build AI chips for Google under a $12.2B share deal.

Notes

☕️ Ex-Meta engineer testifies against Zuckerberg

⚖️ Ex-Meta engineer testifies against Zuckerberg

Arturo Béjar, former Meta engineer, testified against Mark Zuckerberg in an Oakland federal trial, telling jurors the culture he built made it "nearly impossible" to fix safety problems on Facebook and Instagram. Four state AGs (California, Colorado, Kentucky, New Jersey) argue the apps were designed to hook minors — citing infinite scrolling, autoplay video, beauty filters, and the "like" button as harmful. Béjar said safety tools shipped as optional settings instead of defaults, so almost nobody used them. Meta denies misleading the public and disputes penalties it says could reach $1.4 trillion.

🤖 Robots' ChatGPT moment nears

Unitree chief Wang Xingxing told Beijing's World Robot Conference humanoids are nearing a "ChatGPT moment," but warned it could take 2–3 years (optimistic) or 5–10 (pessimistic). He defined the milestone: a robot dropped into an unfamiliar home, taking voice/text orders, finishing ~80% of tasks with no scene-specific training. Cautions follow Unitree's Shanghai listing — shares jumped nearly sixfold before falling 11% next day. Chinese makers shipped 40,000+ humanoids in early 2026, mostly to universities and research labs.

💬 Slack launches its own vibe coding tool

Slack Code lets teams work with AI coding agents (Claude Code, Devin, GitHub Copilot) inside group chats. Tagging an agent opens a channel with context, shows a live HTML preview. Available on all plans; requires human sign-off for risky steps (e.g., production merges); agent auto-archives the channel on approval.

🔒 OpenAI spots misuse without reading prompts

Testing Private Safety Processing — an automated agent monitors inputs/outputs across sessions for abuse (e.g., building malware) without humans reading chats. Triggers a "narrowly defined signal" to OpenAI; customers can opt to share data.

🖥️ Meta AI launches Mac app

Beta, built for businesses/creators; links Facebook/Instagram, post-performance tracking, screen-share advice, dictation across Mac apps, Google Workspace integration (needs professional account). Free; Meta One raises rate limits. Caveat: privacy policy says AI-feature interactions train Meta's models.

🔧 Marvell to build AI chips for Google

Marvell supplies custom AI chips; Google got rights to buy up to $12.2B of Marvell shares (~59M at $206.58 each), unlocked by chip orders ($500M per tranche). Deal could bring ~$120B sales through 2033. Chips attach to Google TPUs (inference accelerators, storage/network controllers). Marvell becomes second supplier beside Broadcom (stock fell ~5%).

Full text · 4,359 chars
| | | ⚖️ Ex-Meta engineer testifies against Zuckerberg LINK | Arturo Béjar, a former Meta engineer, testified against Mark Zuckerberg in an Oakland federal trial, telling jurors the culture the CEO built made it nearly impossible to fix safety problems on Facebook and Instagram. Four state attorneys general, California, Colorado, Kentucky, and New Jersey, argue the apps were designed to hook minors, pointing to infinite scrolling, autoplay video, beauty filters, and the "like" button as harmful features. Béjar said the safety tools Meta shipped were left as optional settings instead of defaults, so almost nobody used them, while the company denies misleading the public and disputes penalties it says could reach $1.4 trillion. | 🤖 Robots' ChatGPT moment nears LINK | Unitree chief Wang Xingxing told Beijing's World Robot Conference that humanoid robots are nearing a "ChatGPT moment," yet warned the breakthrough could take two to three years if all goes well, or five to 10 years if it doesn't. Wang defined the milestone as a robot dropped into an unfamiliar home, taking voice or text orders and finishing about 80% of tasks with no scene-specific training beforehand, a test he calls a tipping point for the field. His caution follows Unitree's Shanghai listing, where shares jumped nearly sixfold before falling 11% the next day, even as Chinese makers shipped over 40,000 humanoids in early 2026 to buyers who are still mostly universities and research labs. | 💬 Slack launches its own vibe coding tool LINK | Slack has rolled out Slack Code, a new "vibe coding" feature that lets teams work alongside AI coding agents inside group chats to build and fix software without waiting on a human engineer. When someone tags an agent like Claude Code, Devin, or GitHub Copilot, it opens a new channel with the right people and context, then lets everyone watch it build and check an HTML preview before approval. Available now on all Slack plans, the tool requires human sign-off for risky steps like merging code to production, and the agent automatically archives the channel once the team approves the finished work. | 🔒 OpenAI spots misuse without reading prompts LINK | OpenAI is testing a service called Private Safety Processing that watches for misuse of its AI across multiple sessions while keeping none of a customer's data, giving select enterprise clients a privacy-focused way to catch abuse. An automated agent looks at the inputs and outputs of several conversations at once, catching bad actors who spread requests over time, such as someone trying to build malware, without any person reading the actual chats. If triggered, the system sends OpenAI a "narrowly defined signal" flagging the activity, and the company then decides whether to act and may contact the customer, who can choose to share data at their own discretion. | 🖥️ Meta AI launches Mac app LINK | Meta has released a beta version of its Meta AI app for the Mac, built mainly for businesses and content creators, with links to Facebook and Instagram plus tools for tracking how posts perform. The app can share a window during a session through screen capture to give advice on your work, offers dictation across all Mac apps, and connects to Google Workspace if you have a professional Facebook or Instagram account. Meta AI is free, though Meta One plans raise the rate limits, and its privacy policy states that interactions with AI features are used to train Meta's models, so users may want to be careful. | 🔧 Marvell to build AI chips for Google LINK | Marvell has agreed to supply custom AI chips to Google, and in return granted Google the right to buy up to $12.2 billion of Marvell shares, sending the chipmaker's stock up as much as 14 percent. Google earns the right to buy nearly 59 million shares at $206.58 each by placing orders, with each $500 million of chip purchases unlocking more, a deal that could bring Marvell about $120 billion in sales through 2033. The chips are parts that attach to Google's tensor processing units, such as inference accelerators and storage and networking controllers, adding Marvell as a second major supplier alongside Broadcom, whose stock fell around 5 percent. | |
16:28

CoSnitch: Researchers Asked Copilot How to Hack It, Then Built | Machine Brief

Researchers got Microsoft Copilot to describe how to hack itself through an undocumented URL parameter that executes prompts automatically, then used that to build a tool called CoSnitch. The flaw is tracked as CVE-2026-24301 and tied to Varonis research. It's a real security finding showing an AI assistant leaking its own attack surface.

Full text · 155 chars
First, automatic prompt execution - an undocumented URL parameter fired a prompt ... Prompt Engineering · Fine-Tuning · RAG · AI Agents. Legal. Privacy ...
20:16

1,357 AI medical devices cleared, 3 actually tested on patient outcomes - Research journals

Of the 1,357 AI-enabled medical devices cleared for use in the US, only 3 were actually tested against real patient outcomes. A PLOS Digital Health study found the rest passed on technical performance rather than clinical benefit. The gap between regulatory clearance and demonstrated patient impact is nearly total.

Full text · 151 chars
Artificial intelligence ( AI ) tools are entering clinical practice at unprecedented speed. 1,357 AI /ML-enabled medical devices have received U.S. ...
20:37

OpenAI chases Anthropic's biz customers with zero data retention pledge - The Register

OpenAI is courting Anthropic's business customers by pledging zero data retention through a feature called Private Safety Processing, which it says offers discreet automated prompt surveillance. The move is a direct competitive play for enterprise trust on privacy and data handling.

Full text · 88 chars
With Private Safety Processing, AI biz promises discreet, automated prompt surveillance.
23:57

ChatGPT search now uses the site:operator at scale

ChatGPT search now honors the site: trick at scale, so you can force results from a specific website far more often. Tracking by Promptwatch shows the share of ChatGPT Search queries using site: jumped from under 0.5% to 16-17% on August 8, right when GPT-5.6 rolled out. It's a sign of "GEO" — marketing your site to appear in chatbot answers. Promptwatch also says ChatGPT now uses Reddit far less as a source as of August 18. The numbers only cover the queries Promptwatch tracks itself.

Notes
ChatGPT search now uses the site: operator at scale

Simon Willison, 20 Aug 2026, Link Blog post.

Context — Promptwatch and "GEO": Promptwatch operates in "Generative Engine Optimization" (GEO), the chatbot-era analog of SEO — tools/consulting to raise a site's presence in prompt replies inside ChatGPT, Claude, Gemini. It automates tracking of prompt responses and publishes aggregate reports, which Willison calls credible hints at "otherwise invisible design changes" to those products.

The observed change (aligned with GPT-5.6 rollout):

  • Share of ChatGPT Search fanout queries containing the site: operator: sat at 0.3–0.5% for weeks, dipped to 0.15% on Aug 3–5 (staged rollout / pre-launch experiment), then jumped to 16–17% on August 8.
  • Caveat from Willison: figures "only reflect the prompts for which they have automated tracking enabled."

OpenAI's official framing (Aug 6):

"For Plus and Pro users, we're updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers."

Willison's read of the mechanism: OpenAI obscures its system prompts, hampering analysis. Poking at ChatGPT, he believes the search tool now has a shape like search(query, recency, domains) rather than directly encouraging a site: operator.

Follow-up (Aug 18): Promptwatch reported ChatGPT had "greatly reduced the likelihood" of Reddit appearing in those searches. Willison could not confirm via system prompts — the leaked-system-prompt collection he relies on shows no relevant changes yet, so the Reddit-suppression mechanism remains unverified.

Related posts linked: "Conceptual integrity and counting lines of code" (19 Aug); "Qwen 3.8 27B ... defaults to wildly overthinking things" (16 Aug); OpenAI's accidental attack against Hugging Face timeline (7 Aug).

Full text · 2,230 chars
20th August 2026 - Link Blog ChatGPT search now uses the site:operator at scale. Promptwatch is part of the emerging "GEO" space, for Generative Engine Optimization - the chatbot version of SEO, where companies offer tools and consulting to help your site increase its presence in replies to prompts inside tools like ChatGPT. The Promptwatch product uses automation to track responses to prompts across end-user chat products like ChatGPT, Claude, and Gemini. They publish aggregate reports on this as part of their own content marketing strategy, which do seem to provide credible hints as to otherwise invisible design changes to those products. Their own tracking shows a notable change aligned with the GPT-5.6 rollout earlier this month: The percentage of all ChatGPT Search fanout queries that contain the site:operator, per day. The share hovered between 0.3% and 0.5% for weeks, dipped briefly to 0.15% on August 3 to 5 (consistent with a staged rollout or pre-launch experiment), then jumped to 16-17% on August 8. It's important to note that these figures only reflect the prompts for which they have automated tracking enabled. This corresponds to OpenAI's somewhat vague August 6th announcement: For Plus and Pro users, we’re updating GPT‑5.6 Sol in Chat to be more reliable with facts and provide more focused answers. Once again I am hampered by OpenAI's decision to actively obscure their system prompts, but from poking at ChatGPT I believe their latest search tool has a shape like search(query, recency, domains) rather than encouraging a site: operator directly. In a follow-up on August 18th Promptwatch reported that ChatGPT appeared to have greatly reduced the likelihood of Reddit being used in those searches. My own attempts to ascertain if the system prompt has been updated to discourage Reddit sourcing have been unsuccessful - the most thorough leaked system prompt collection I know of doesn't yet show any relevant changes. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
04:00

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

Researchers built a new test to catch when AI models invent false details while summarizing long novels. The benchmark uses 29 Chinese novels plus English book chapters, defining 8 types of mistakes a model might make. Early testing shows even strong models struggle on it, so it's meant to push better long-context summarization and is being released for others to use.

Notes

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

New arXiv paper (cs.CL), Aug 2026, proposing a bilingual hallucination-detection benchmark.

Motivation

  • Context windows have grown, but hallucinations in long-context summarization persist.
  • Long novels suit this study better than news/papers due to intrinsic information and detailed descriptions of events and dialogues.
  • Gap: no multi-scale benchmark for hallucination detection in long-context novel summarization, and no exploration of how hallucinations change as context lengthens.

Benchmark construction

  • Sources: 29 Chinese novels (16k–100k tokens) plus chapter-level data from the BookSum dataset.
  • 8 hallucination types designed by the authors.
  • Generation method: combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation, aimed at data authenticity and a balanced distribution across hallucination categories.
  • Reliability: test-set content manually revised by humans.

Claims

"Extensive experimental results demonstrate that LongNovel is a challenging benchmark."

Release: benchmark publicly released (arXiv URL).

Limitations/notes

  • Abstract gives no specific benchmark numbers, baselines, or per-type accuracy results — only that it is "challenging."
  • Explicitly framed as a multi-scale benchmark (varying token lengths) to probe how hallucination rates change as context grows, but the paper's actual findings on that gradient are not stated in the abstract.
Full text · 2,083 chars
Computer Science > Computation and Language Title:LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization View PDF HTML (experimental) Abstract:Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. this https URL Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives

Language models can track what's being talked about in a story, like where objects are and how they change, even when it isn't said out loud. This ability shows up in models with as few as 410 million parameters, far smaller than previously thought, and grows with size. Modern big models actually beat humans on this task, though people's performance drops with story complexity rather than length.

Notes

Entity tracking in sub-billion-parameter LMs

arXiv cs.CL paper (Aug 20, 2026) on whether LMs track entities (knowing where things are and how they change, even when unstated) in a human-like way.

Motivation / gap: Prior evaluations relied on artificial tasks far from natural language comprehension and lacked human comparisons.

Method:

  • Evaluated entity tracking in both LMs and humans (N = 48) using naturalistic narratives at multiple complexity levels.

Key findings:

  • Humans: entity tracking degrades specifically with narrative complexity, not narrative length.
  • LMs: human-level entity tracking already present at 410 million parameters — well below the multi-billion-parameter, code-specialized models identified by prior work.
  • Tracking improves with scale; contemporary models far exceed human performance.

Conclusion (claim):

entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought.

Implications/caveats not stated in abstract: No discussion of whether the sub-billion models that "match" humans do so on the degradation pattern (complexity-sensitive) or just overall accuracy — that distinction is central to the human-comparison claim but left implicit here. Specific model names, narrative datasets, and exact benchmark numbers are not given in the abstract.

Full text · 1,946 chars
Computer Science > Computation and Language Title:Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives View PDF HTML (experimental) Abstract:Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stated. Whether language models perform such tracking in a human-like fashion remains unclear, in part because existing evaluations rely on artificial tasks, far removed from natural language comprehension, and lack comparisons to humans. Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of complexity. In humans, we find that entity tracking degrades specifically with narrative complexity, not narrative length. In language models, we find that human-level entity tracking is already present at 410 million parameters - well below the multi-billion parameter, code-specialised models identified by prior work - and improves with scale, with contemporary models far exceeding human performance. Together, these results demonstrate that entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving

A new method helps AI write mathematical proofs in real software projects by using compiler error messages to guide it. The approach balances trying many starting points with polishing the most promising one, using two different models and smart resampling. On seven real Lean 4 projects it beat simpler baselines, improving success rates by about 13 points while using roughly a fifth fewer model calls.

Notes
  • Title: Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving
  • Venue: arXiv cs.CL, published 2026-08-20
  • Domain: Theorem proving in real-world Lean 4 projects

Problem: Proofs in real Lean 4 projects depend on project-specific context. Iterative refinement can use compiler errors to repair failed proofs, but reusing failed attempts needs careful search control — some proofs are better starting points than others, and later revisions can degrade a partially correct proof.

Proposed method: Compiler-guided proof search balancing exploration and exploitation:

  • Exploration via dual-model generation and stagnation-triggered resampling (diverse starting points)
  • Exploitation via current-best refinement guided by compiler-grounded pairwise comparison

Results (7 real-world Lean 4 projects from miniCTX-v2): Better effectiveness–efficiency tradeoff than pass@k baselines.

"Within the pass@32 budget, our method improves average pass rate by 12.8 percentage points while reducing LLM calls by 21.9%."

Key claims: +12.8 pp average pass rate and −21.9% LLM calls at pass@32 vs pass@k baselines.

Not stated: No explicit baselines listed, no per-project breakdown, no ablation of the dual-model/stagnation/pairwise components individually, and no discussion of limitations or failure modes.

Full text · 1,814 chars
Computer Science > Computation and Language Title:Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving View PDF HTML (experimental) Abstract:Theorem proving in real-world Lean 4 projects is challenging because proofs often depend on project-specific context. While iterative refinement can use compiler errors to repair failed proofs, reusing failed attempts requires careful search control: some proofs provide better starting points than others, and later revisions may degrade a partially correct proof. We propose a compiler-guided proof search framework that balances exploration and exploitation. It explores diverse starting points through dual-model generation and stagnation-triggered resampling, while exploiting promising proof states through current-best refinement guided by compiler-grounded pairwise comparison. Experiments on seven real-world Lean 4 projects from miniCTX-v2 show that our method achieves a better effectiveness--efficiency tradeoff than pass@k baselines. Within the pass@32 budget, our method improves average pass rate by 12.8 percentage points while reducing LLM calls by 21.9%. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Persona-Guided LLM Agents for Task-Oriented Dialogue

A study asks whether AI assistants can show a personality while still getting a task done for the user. Researchers set up a conversation where one AI plays the user with a set personality and another plays the assistant that adapts to it. They tested GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash on hotel and restaurant booking dialogues. Adapting to the user's personality improved task success and satisfaction but made answers less truthful, so there's a tradeoff between being friendly and being accurate.

Notes
  • Paper: "Persona-Guided LLM Agents for Task-Oriented Dialogue" (cs.CL, arXiv, 2026-08-20)
  • Question: Can LLMs express personality in goal-directed TOD without hurting task completion, and does adapting to user personality improve interaction quality?
  • Framework: Training-free, two-LLM simulation — a user agent exhibiting a target personality, a system agent that adapts while completing the task.
  • Isolation design: 3 conditions varying how much the system knows the user's personality — Neutral (no info), Try (infers from dialogue cues), Oracle (personality given explicitly).
  • Models: GPT-4o, Qwen3-Next-80B, Gemini 2.0 Flash.
  • Data: Hotel and Restaurant dialogues from Schema-Guided Dialogue (SGD) dataset, Big Five traits plus opposite poles.
  • Results:
  • User agent expresses personality while system keeps strong task performance — but some traits are realized "far less reliably than others."
  • Adaptation to user personality improves constraint satisfaction, inform rate, and user satisfaction but lowers truthfulness.
  • Key trade-off: personalization vs. task-grounding (adaptation can introduce hallucination).
  • Condition contrast: Oracle's gains grow when target trait is strongly expressed; Try's gains are largely insensitive to realization strength.
  • Conclusion: cue-based adaptation (Try) "best resolves this trade-off" — the more reliable route to personality-aware TOD without fine-tuning.
  • Caveats: findings limited to SGD domains (hotel/restaurant), three models, and synthetic user-simulator setup; no human interaction validation.
Full text · 2,578 chars
Computer Science > Computation and Language Title:Persona-Guided LLM Agents for Task-Oriented Dialogue View PDF HTML (experimental) Abstract:Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user's personality improves the interaction quality. We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interaction. We build a training-free framework that simulates a TOD interaction between two LLMs: a user agent that exhibits a target personality and a system agent that adapts to the user while completing the task. To isolate the effect of adaptation, we vary how much the system knows about the user's personality across three conditions. In Neutral, the system receives no personality information. In Try, it infers the personality from dialogue cues. In Oracle, it is given the personality explicitly. We evaluate GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash on Hotel and Restaurant dialogues from the Schema-Guided Dialogue (SGD) dataset, across the Big Five traits and their opposite poles. We find that the user agent can express personality while the system maintains strong task performance, although some traits are realized far less reliably than others. Adapting to the user's personality improves constraint satisfaction, inform rate, and user satisfaction, but lowers truthfulness, revealing a trade-off between personalization and task-grounding. Oracle's gains grow when the target trait is strongly expressed, whereas Try's gains are largely insensitive to realization strength. Overall, cue-based adaptation in Try best resolves this trade-off and offers a more reliable route to personality-aware TOD without fine-tuning. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

SuTRA : Structurally-Unified Tokenization with Root Awareness

Researchers built a smarter way to split words into tokens for Indian languages, which keeps root words and endings together instead of shattering them. Standard tokenizers break words based on statistics and can cut roots from their suffixes, which hurts morphologically rich languages like Hindi. Their new method, SuTRA, preserves the basic syllable unit and punishes splits that cross word-part boundaries, boosting structural alignment by up to 14.7% and machine translation quality by about 8 points. They also released a new segmentation dataset for Hindi, Marathi, and Gujarati.

Notes

SuTRA: Structurally-Unified Tokenization with Root Awareness

arXiv cs.CL paper (Aug 20, 2026) on morphology-aware subword tokenization for Indic languages.

Problem

Existing subword tokenizers (e.g. BPE) optimize statistical compression but ignore morphology — specifically the root–affix relationship. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes, a phenomenon the authors coin Morphological Shattering. Harms morphologically rich Indic languages, whose basic units are complex orthographic syllables (aksharas), not letters.

Method

SuTRA (Structurally-Unified Tokenization with Root Awareness):

  • Preserves akshara indivisibility
  • Penalizes merges that cross morphological boundaries
  • Releases a new morphological segmentation dataset for Hindi, Marathi, and Gujarati
Results (vs BPE)
  • Peak +14.7% morphological alignment (Boundary F1)
  • +34% semantic recoverability (Hindi)
  • Average +8.08 chrF2 improvement in machine translation
Caveats
  • Improvements concentrated on morphologically rich languages; generalizability to other families not claimed.
  • Gains reported relative to BPE baseline only; no comparison to other morphology-aware tokenizers (e.g. SentencePiece-unigram, BPE-dropout) stated in the abstract.
  • Dataset limited to three Indic languages; Marathi/Gujarati semantic recoverability figures not individually given.
Full text · 1,765 chars
Computer Science > Computation and Language Title:SuTRA : Structurally-Unified Tokenization with Root Awareness View PDF HTML (experimental) Abstract:Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

Researchers found a single internal 'mood meter' inside language models that tracks how positive or negative a sentence feels, and it works across text, images, audio, and even brain scans. The recipe needs far less data than usual, using just nine emotion names and 50 story snippets per emotion instead of thousands of labels. This same direction transfers across four different types of encoders and reaches near the accuracy of supervised methods, though it only works for continuous feelings, not categories, and only steers some model families. It's a step toward understanding emotion in models without expensive labeling.

Notes
  • Paper: "Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities" (cs.CL, arXiv, 2026-08-20)
  • Core claim: a single internal "valence axis" (V-axis) inside a language model tracks sentence positivity/negativity.
  • Recipe: embed 9 emotion-anchored story sets (9 emotion category names + 50 short narrative paragraphs each, ~1,500 fewer labels than supervised) in a frozen encoder; take the top principal direction of the 9 averaged embeddings; project new inputs onto it.
  • Results (V-axis projection):
  • SST-2: 93% of supervised performance (Llama-3-8B-Instruct, AUC 0.772 vs 0.828)
  • EmoSet images: r=0.636 with human valence on 11,811 images
  • ESC-50 audio: AUC 0.906 (p<2.2e-15)
  • EEG from 123 subjects: AUC 0.720±0.055 (p<3.65e-8)
  • Mechanism: ablating the axis collapses sentiment accuracy 5.5–37.2 pp across three LLMs vs ≤0.88 pp for matched random directions (z>12).
  • Cross-modal transfer: 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525).
  • Boundaries/limitations (stated by authors):
  • Recipe limited to continuous attributes; seven tests on categorical concepts return near-chance.
  • Steering is family-specific — works for Llama/Mistral, not Qwen/Gemma.
Full text · 2,241 chars
Computer Science > Computation and Language Title:Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities View PDF HTML (experimental) Abstract:Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811 EmoSet images at r=0.636, reaches AUC 0.906 on ESC-50 audio (p<2.2e-15), and AUC 0.720+/-0.055 on EEG from 123 subjects (p<3.65e-8). The direction is mechanistically active: ablating it collapses sentiment accuracy by 5.5-37.2 pp across three LLMs vs. at most 0.88 pp for matched random directions (z>12). A 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), and brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525). The recipe is bounded to continuous attributes -- seven tests on categorical concepts return near-chance -- and steering is family-specific (Llama/Mistral yes, Qwen/Gemma no). Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

A new study shows that large language models used as judges are biased by who they think wrote an answer, not just by how good the answer is. Earlier work on this 'self-preference' was muddied because style and quality get tangled together, so the researchers had models grade abstract selections instead of text. Under blind grading the bias mostly vanished, but when they labeled selections as 'self' or 'other' the models inflated their own and deflated the others regardless of real source. The finding suggests authorship labeling is its own driver of evaluation bias, separate from content quality.

Notes
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

arXiv cs.CL abstract (2026-08-20). Author attribution as a distinct bias driver in LLM-as-a-judge systems.

Problem / prior gap: Self-preference (LLMs favoring their own outputs) has been studied only on generated text, where stylistic features and response quality are inevitably conflated — so existing measurements cannot isolate genuine self-preference from those confounds.

Method: Changes the object of evaluation. Ten LLMs assess narrative constraint selections instead of generated text — a task with no model-specific stylistic fingerprint yet a recoverable model-specific signature. Two experiments run.

Experiment 1 — blind evaluation: Self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on 3 of 4 rubric dimensions and reverses on the fourth, where judges rate their own selections as less original.

Experiment 2 — matched quality: Self- and other-labels alone — without naming any model — shift scores bidirectionally: judges inflate scores for self-labeled selections and deflate scores for other-labeled ones, regardless of the selection's actual source.

Contributions:

  • Authorship attribution is a distinct driver of evaluation bias.
  • Open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.

Stated limitations / caveats: None given in abstract. Implied scope limits: 10 models, 4 rubric dimensions, single task type (narrative constraints); findings may not transfer to generated-text judging, where label and quality effects remain entangled.

Full text · 2,232 chars
Computer Science > Computation and Language Title:Self- and Other-Labels Induce Bidirectional Bias in LLM Judges View PDF HTML (experimental) Abstract:As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages

A new language model built for nine underrepresented languages from Northeast India beats existing multilingual models on all of them. NE-BERT was trained on about 8.3 million sentences and posts 15.97x and 7.64x lower perplexity than IndicBERT-V2 and MuRIL, with better tokenization than mBERT. It handles very tiny languages like Pnar (only 1,002 sentences) by aggressively oversampling them. The model, test sets, and corpus are released open under CC-BY-4.0.

Notes
NE-BERT: Multilingual LM for Nine Northeast Indian Languages

arXiv cs.CL submission (2026-08-20). Abstract-only; full methods/data not yet readable.

  • Model: NE-BERT, a domain-specific multilingual encoder.
  • Training data: ~8.3M sentences over 9 Northeast Indian languages + 2 anchor languages (Hindi, English).
  • Tokenization: custom SentencePiece Unigram tokenizer; training uses weighted data sampling.
  • Low-resource emphasis: aggressive upsampling targets vocabulary fragmentation in extremely low-resource languages — Pnar (1,002 sentences) and Kokborok (2,463 sentences).
  • Results vs baselines: outperforms IndicBERT-V2 and MuRIL on all 9 NE languages — 15.97× and 7.64× lower average perplexity respectively; 1.50× better tokenization fertility than mBERT.
  • Downstream: POS-tagging validation on three NE languages only (unnamed in abstract).
  • Release: model, test sets, and training corpus under CC-BY-4.0.

Stated gap: the region is "minimally represented in existing multilingual models."

Limitations/caveats

  • The abstract names neither the 9 languages nor the 3 POS-tagged ones — Pnar and Kokborok are the only explicit mentions.
  • Perplexity figures are average across languages; no per-language breakdown or absolute values given, so the × numbers can't be independently sanity-checked.
  • Perplexity and tokenizer fertility are intrinsic metrics; the only extrinsic check (POS tagging) covers just 3 languages.
  • No comparison to GPT-family decoder models, only BERT-style encoders (IndicBERT-V2, MuRIL, mBERT).
  • The 15.97×/7.64× perplexity gaps are large enough to warrant scrutiny of the baseline setups (tokenizer mismatches, eval splits).
Full text · 2,014 chars
Computer Science > Computation and Language Title:NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages View PDF HTML (experimental) Abstract:Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages (Hindi, English), a linguistically diverse region with minimal representation in existing multilingual models. By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian languages, achieving 15.97X and 7.64X lower average perplexity respectively, with 1.50X better tokenization fertility than mBERT. We address critical vocabulary fragmentation issues in extremely low-resource languages such as Pnar (1,002 sentences) and Kokborok (2,463 sentences) through aggressive upsampling strategies. Downstream evaluation on part-of-speech tagging validates practical utility on three Northeast Indian languages. We release NE-BERT, test sets, and training corpus under CC-BY-4.0 to support NLP research and digital inclusion for Northeast Indian communities. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators

A new framework scores whether images and video align with broad societal values like peace, justice, and freedom. MAVEN organizes values into 6 main dimensions and 72 sub-indicators, and ships with a human-verified benchmark plus a soft-match metric for testing vision-language models. Its compact 2-billion-parameter evaluator matches its 8-billion sibling and approaches closed-source frontier models, offering a cheaper path to large-scale value checks. It uses a training-free multi-role consensus step at inference time.

Notes

MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators — arXiv cs.CL, posted 2026-08-20.

Problem: Existing value-assessment frameworks for content are confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification — none handle multimodal content against macro-societal values (peace, justice, freedom).

MAVEN — hierarchical evaluation framework grounded in international human-rights instruments and cultural value theory:

  • 6 primary dimensions, 72 secondary indicators
  • Multi-level quantitative scoring (not single-label)

Contributions:

  • MacroValue-Bench — human-verified multimodal benchmark plus a soft-match metric for scoring VLMs' value-dimension assessments.
  • SA-MDPO — span-adaptive variant of multi-level preference optimization, used for evaluator distillation (compressing a strong evaluator into a compact one).
  • Training-free multi-role consensus strategy at inference time.

Results:

  • Evaluates open- and closed-source VLMs; finds shared tendencies and clear differences in macro-societal value judgments across families.
  • A compact 2B evaluator matches its 8B counterpart in the same family and approaches frontier closed-source VLMs — claims a practical path to scalable macro-societal value evaluation.

Released artifacts: SA-MDPO implementation and MacroValue-Bench (link in abstract).

Caveats not stated in abstract: no quantitative benchmark scores, no breakdown of which value dimensions differ across models, no closed-source model names given, and the "approaches frontier" claim is unquantified. Grounding in human-rights instruments implies a specific (possibly Western/UN-centric) value framing the authors do not problematize.

Full text · 2,253 chars
Computer Science > Computation and Language Title:MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators View PDF HTML (experimental) Abstract:Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification. Therefore, we propose MAVEN, a hierarchical framework for macro-societal value evaluation of multimodal content, grounded in international human-rights instruments and cultural value theory. MAVEN organizes values into 6 primary dimensions and 72 secondary indicators, supporting multi-level quantitative scoring. Building on MAVEN, we construct a human-verified multimodal benchmark and a soft-match metric to evaluate VLMs' assessments across value dimensions. For evaluator optimization, we propose a span-adaptive variant of multi-level preference optimization for evaluator distillation, together with a training-free multi-role consensus strategy at inference time. We evaluate existing open- and closed-source VLMs on our benchmark, revealing shared tendencies and clear differences in macro-societal value judgments. Experiments show that our compact 2B evaluator matches its 8B counterpart in the same family and approaches frontier closed-source VLMs, offering a practical path toward scalable macro-societal value evaluation. Our SA-MDPO implementation and MacroValue-Bench are available at this https URL. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

A new benchmark tests how well models sort French news articles into the right editorial desk section. FrenchNews-7 was built from a large multi-outlet corpus with seven categories, labeled by a hybrid of publisher URL slugs and human-verified LLM annotation. The best model, a fine-tuned CamemBERT reading full article text, beats headline-only input and three big zero-shot LLMs (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on overall recall at 0.799. It stumbles on the fuzzy line between Economy and Society sections, where accuracy roughly matches how often humans agree anyway.

Notes
FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

Paper: cs.CL arXiv preprint, published 2026-08-20.

What it is: A France-based, French-language news editorial desk classification benchmark combining a multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Seven classes include Economie, Societe, Sport, Culture & Loisirs, International.

Labeling pipeline: Hybrid — publisher URL slugs + LLM annotation for structurally ambiguous cases. Audited via inter-rater study (2 humans + 2 LLMs): pairwise κ ≥ 0.766; human–human κ = 0.806.

Evaluation: Lexical, multilingual, and French-specific trained classifiers tested under both in-distribution and held-out-publisher settings; compared against zero-shot LLM baselines GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B on the held-out pool.

Main result: CamemBERT-base on full article text beats headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall = 0.799. The gap concentrates in the ambiguous boundary categories Economie and Societe.

Cross-publisher stability (key caveat): Sport, Culture & Loisirs, and International transfer cleanly, but:

  • Economie: recall = 0.517, near blinded human agreement (0.55)
  • Societe: precision = 0.577, absorbs boundary ambiguity
Authors' interpretation: both figures "suggest editorial conventions rather than recoverable classifier headroom."

Released artifacts: fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table (model and dataset links in paper).

Limitation noted: performance ceiling on ambiguous desk boundaries is close to human ceiling — classifier headroom there is largely exhausted.

Full text · 2,393 chars
Computer Science > Computation and Language Title:FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification View PDF HTML (experimental) Abstract:We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $\kappa \geq 0.766$, human--human $\kappa = 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at this https URL (model) and this https URL (dataset). Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems

A new memory-management trick lets chatbot models keep the important history of a conversation while switching topics much faster. The algorithm, called Fractional Decay KV-Cache, scores each cached piece twice: how important it has been overall and how recently it mattered. It beats the current best method H2O by 6.7% overall and 127% on sudden topic shifts, and adapts to new topics 3.6 times faster. It runs entirely on CPU with negligible overhead, tested across five dialog scenarios with 600 dialogs each.

Notes

Fractional Decay KV-Cache (FD-KVC) — arXiv cs.CL preprint (2026-08-20), computation & language.

Claim (abstract):

"FD-KVC operates entirely on CPU with negligible overhead."

Method. Proposes FD-KVC for transformer dialog inference. Maintains a dual-channel score per cached KV pair:

  • Cumulative attention channel — aggregate importance, "akin to H2O."
  • Recency-weighted relevance channel — governed by temporal decay plus reinforcement-inspired updates.

An adaptive learning rate driven by an ownership loss function claims convergence "without oscillation." Motivating problem: existing caches treat entries uniformly or use coarse eviction heuristics that "fail to adapt as dialog topics evolve."

Results. Evaluated on five multi-turn dialog scenarios, 600 dialogs each, vs H2O (SOTA heavy-hitter baseline):

  • +6.7% composite late-turn alignment
  • +127% topic-shift
  • +87% gradual evolution
  • +30% mixed-topic
  • Topic adaptation 3.6× faster than H2O
  • Highest topic diversity across methods: 80.6%
  • Ablations confirm each component contributes.

Limitations / caveats.

  • Preprint — not peer-reviewed; no paper page or benchmark detail beyond the abstract.
  • Single baseline (H2O only); no comparison vs other eviction or quantized approaches.
  • No absolute scores, confidence intervals, or significance testing reported.
  • No evidence of integration into a production stack; CPU-only claim unquantified, and no GPU/footprint comparison.
  • Evaluation appears confined to curated/scripted dialog scenarios; real-conversation behavior untested.
Full text · 2,173 chars
Computer Science > Computation and Language Title:Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems View PDF HTML (experimental) Abstract:Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog systems, yet existing strategies treat all cached entries uniformly or apply coarse eviction heuristics that fail to adapt as dialog topics evolve. We propose Fractional Decay KV-Cache (FD-KVC), a novel algorithm that maintains a dual-channel scoring mechanism for each cached KV pair: a cumulative attention channel that tracks aggregate importance (akin to H2O), and a recency-weighted relevance channel governed by temporal decay and reinforcement-inspired updates. The combination enables FD-KVC to both preserve historically important tokens and rapidly adapt when dialog topics shift. An adaptive learning rate driven by an ownership loss function ensures convergence without oscillation. FD-KVC operates entirely on CPU with negligible overhead. Across five diverse multi-turn dialog scenarios with 600 dialogs each, FD-KVC outperforms H2O, the state-of-the-art heavy-hitter baseline, by +6.7% on composite late-turn alignment, with improvements of +127% on topic-shift, +87% on gradual evolution, and +30% on mixed-topic dialogs. FD-KVC adapts to new topics 3.6X faster than H2O and achieves the highest topic diversity (80.6%) across all methods. Ablation studies confirm the contribution of each component. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

A new research paper shows AI chatbots reproduce Western, Orientalist framing when they discuss the Middle East, even when they avoid open stereotypes. Across 280 conversations, GPT-4 and Falcon3-7B systematically positioned Western frameworks as neutral and explained the region through categories it didn't produce. GPT-4 scored moderately while Falcon3-7B scored worse despite being built in Abu Dhabi with Arabic training — evidence that building a model regionally doesn't make it less biased. The authors coin 'Said-washing' for models that disclaim generalization then reproduce it, found in 87.9% of GPT-4 conversations, and argue the fix is changing what models learn from, not just adding languages.

Notes
Computational Orientalism: Measuring Structural Discourse Bias in LLMs Using MECSS

Paper (arXiv, cs.CL, 2026-08-20): introduces the Middle East Cultural Sensitivity Score (MECSS) to measure structural (not explicit) Orientalist bias in LLMs, operationalizing Edward Said's seven Orientalist operations as measurable dimensions.

Method: 280 conversations (1,120 exchanges) with GPT-4 and Falcon3-7B-Instruct, scored on MECSS dimensions.

Results:

  • GPT-4: mean MECSS 1.73; Falcon3-7B-Instruct: 2.18 — both reproduce Orientalist patterns "through structural positioning rather than open stereotyping."
  • Epistemic Center (treating Western frameworks as unmarked universals) scored near the top of the scale for both models.
  • "Said-washing" — a coined failure mode where a model "disclaims generalization, then reproduces the structure it disclaimed" — appeared in 87.9% of GPT-4 conversations.

Key claim: The Falcon3 result (built in Abu Dhabi, trained with Arabic content, yet more Orientalist than GPT-4) is evidence against the assumption that regional/linguistic grounding reduces Orientalism.

Caveats / limitations (stated):

  • Models "differ in size as well as origin, so geography cannot be isolated as the cause."
  • Argues standard fairness metrics fail here because they "detect explicit prejudice rather than structural framing."

Conclusion: Reducing this bias "requires changing what models learn from, not only adding languages or relocating institutions."

Full text · 2,820 chars
Computer Science > Computation and Language Title:Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) View PDF HTML (experimental) Abstract:AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term "Said-washing" for a specific failure: a model that disclaims generalization, then reproduces the structure it disclaimed. Across 280 conversations (1,120 exchanges), GPT-4 and Falcon3-7B-Instruct both reproduce Orientalist patterns systematically, through structural positioning rather than open stereotyping. GPT-4 scores moderately (mean MECSS 1.73); Falcon3-7B-Instruct scores higher (2.18), even though it was built in Abu Dhabi and trained with Arabic content. This is evidence against the assumption that building a model regionally makes it less Orientalist, though the models differ in size as well as origin, so geography cannot be isolated as the cause. Epistemic Center, the treatment of Western frameworks as unmarked universals, scores near the top of the scale for both models. Said-washing appears in 87.9% of GPT-4 conversations, a pattern existing metrics cannot see. Reducing this bias requires changing what models learn from, not only adding languages or relocating institutions. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
09:00

Support networks aim to help kids through the polycrisis

Online support groups are helping anxious young people cope with the world's overlapping crises, and their creators say it's working even though the evidence is thin. UK-based Force of Nature and the Good Grief Network run group sessions that turn worry into action; Force of Nature had about 900 participants last year and about 2,000 students from 50-plus countries in its online community. A 2021 Lancet study of 10,000 kids in 10 countries found roughly 60% were very or extremely worried about climate change. Scholars say the field needs to actually measure whether these programs help before schools and parents invest in them.

Notes
Support networks aim to help kids through the polycrisis

Source: Sarah Scoles, MIT Technology Review, Aug 20 2026. Scoles is a freelance journalist in Colorado, author of Countdown: The Blinding Future of Nuclear Weapons. The piece profiles climate/polycrisis peer-support programs for young people.

The profile subjects
  • Pim Sullivan-Tailyour — childhood memory of a quarried-out mountain in southern Thailand (~late 2000s, age 6) triggered her environmental consciousness. Moved to UK, joined Schools Sustainability Network (umbrella org for school environmental groups), found no peers at her own school, felt burned out working solo. Joined Force of Nature's Becoming a Force of Nature; now graduated from King's College London, leads in-person "climate cafes."
  • Clover Hogan — Force of Nature founder, started the group in 2019 as a teenager. Catchphrase: > "We don't need 100 perfect activists but millions of imperfect ones."
  • Hannah Hooper — head of programs until recently, Force of Nature.
Key statistics
  • **2021 Lancet Planetary Health study (led by psychotherapist Caroline Hickman): 10,000 children/young people in 10 countries; ~60% very/extremely worried about climate; >45% said eco-anxiety affected daily functioning; 56% agreed "Humanity is doomed."**
  • 2025 Polish study attempting to define "polycrisis syndrome": used the Flourishing Index and Difficulties in Emotional Regulation Scale; >60% of young people reported difficulty regulating emotion, depression symptoms, and lower mental/physical well-being, attributed to cumulative stress of world affairs.
  • Early-1980s Cold War baseline: ~35% of high school seniors agreed "Nuclear or biological annihilation will probably be the fate of all mankind within my lifetime" (survey of 130 schools in 48 states); a majority also agreed "The human race has come through tough times before and will do so again"; a Gallup poll found 49% of teens said nuclear-war possibility influenced their future planning. Authors note there is no ready baseline in the literature for comparing current youth angst.
  • 1986 National Academy of Sciences analysis of that era's research prescribed: knowledge of the issue, sensitivity to the inner processes of working through painful feelings, and willingness to come to grips with what young people voice.
The programs

Force of Nature (~900 participants in Becoming a Force of Nature last year; ~2,000 students from 50+ countries in the online community):

  • Becoming a Force of Nature: free informal online group sessions (spun up 2021). Three sessions: (1) facilitators — themselves young people — ask how participants feel about the future and climate; (2) participants investigate personal strengths/skills/passions; (3) synthesize into a "road map" for community action.
  • Also built Hold This Space (interactive site) and a podcast with solicited #climateconfessions ("I've stopped going to protests," "I drive everywhere," "I just can't be arsed to separate or put out the recycling").
  • Has moved away from its strict eco-focus ("not our bread and butter anymore" — Hooper), now encouraging talk about cost of living, conflict, etc. Two reasons: privileged-country kids may be anxious but others are in survival mode, and climate-only talk "grabs only a certain kind of person" (echo-chamber avoidance).

Good Grief Network (GGN) — Michigan-based nonprofit; adult program is 10 steps modeled on Alcoholics Anonymous' 12-step model, run 85+ times with ~2,500 participants. GGN-Z is the teen adaptation by Susan Igras (program designer/evaluator): cut to 5 steps, "much more experiential—action-, reflection-based," tested since 2022 with kids in Northern California. Favorite activities: a "bad news graffiti wall" (recent entries: "Meta=cringe," "hate speech," a tornado, "LA" with a fire above it, a dollar sign with an up arrow) and a storytelling exercise ("Once upon a time, I became aware of climate change" — responses like "Scared, 2nd grade learning about non-renewable resources," a skinny polar bear on ice), plus drawing a "support web."

The evidence gap
  • Liza Jachens (psychologist, University of Nottingham school of medicine) coauthored a 2021 scoping review of eco-anxiety programs: found 34, only 2 had done formal self-evaluation. Plea: > "Please evaluate what you build."
  • A 2024 scoping review led by Siqi Xue (University of Toronto) found the gap hadn't closed; it noted GGN's claim that 90% of participants felt more empowered and less alone was not backed by published data or methods.
  • Jachens defines eco-anxiety via the American Psychological Association: a "chronic fear of environmental doom," but says what people describe "is closer to grief... mourning for what has already gone." Lise Van Susteren (psychiatrist, coauthor of the 2012 National Wildlife Federation report The Psychological Effects of Global Warming on the United States) calls it "pre-traumatic" stress.
What seems to work
  • Group approach: connection and validation plus active direction. Jachens: > "When people feel genuinely heard and held in their distress, they tend to move quite naturally toward wanting to do something."
  • Climate Emotion School (Sweden) study by clinical psychologist Britta Eklöf, shaped by Panu Pihkala's (University of Helsinki) coping model (express emotions, regulate, act, form meaning; self-care and healthy distance). Students were relieved to share bad feelings and see them mirrored — something adults avoided; parents tended to soothe rather than validate. Eklöf: "The recommendation instead is to validate and share." She recommends referencing MLK, Rosa Parks, the suffragettes, and historical repair (e.g., the ozone-layer treaty) to build hope.
  • Hickman's key finding: kids were upset older people failed to protect their future — > "that social contract was being broken repeatedly."
Caveats

Research on effectiveness is nascent; Jachens: "We need to know what works for who, when, and where." Few studies target kids specifically. GGN-Z is currently self-evaluating via surveys; Force of Nature joined the 25-organization, five-year Youth Mental Wellbeing project to study the field's collective work.

Full text · 20,916 chars
Sometime in the late 2000s, Pim Sullivan-Tailyour was sitting in the back of a car, headed toward her great-grandmother’s tiny town in the south of Thailand. She watched big mountains pass by out the window. She was just six years old but was about to be hit by an adult-size realization. “They were just quarried out,” she says. “Like half the mountain just dug out and disappeared, and so there was just this huge orange face.” Although she probably didn’t know the word “quarry” back then, she sensed that what she was seeing was unnatural. She also knew, from family stories, that her mother had bathed in a nearby river when she was young, its water so glass-clear she could see fish swimming by. But the extraction had muddied the water. “We’re changing things,” Sullivan-Tailyour thought. “And it doesn’t seem right.” It was the first time it occurred to her that humans could alter the world—and that they were. For the worse. She carried that knowledge with her as she got older and her family moved to the UK. There, she joined teenagers from other schools in the Schools Sustainability Network, an umbrella organization for groups that work on environmental initiatives. But even with those connections, Sullivan-Tailyour didn’t encounter anyone at her own school who was interested in environmental activism. She felt lonely in her desire to push for change—and burnt out on working solo. “I was pretty much the only person who wanted to do these things,” she says. That was when she heard about another initiative, called Force of Nature. A UK-based organization, in 2021 it had spun up Becoming a Force of Nature, a kind of informal online group therapy program, for kids and young adults worried about climate. The sessions aimed to help young people like Sullivan-Tailyour take their worries and turn them into action, stemming the exact kind of anxiety and burnout she was feeling. She signed up and soon logged into Zoom for her first session. Dozens of faces from dozens of countries looked back at her. “I realized that I wasn’t alone,” she says. She began to suspect that kids at her school actually did feel the way she did, even if they weren’t doing anything about it. Sullivan-Tailyour was right, it turns out: Global surveys have found that most kids are anxious about the state of the world—and not just about whether sea levels and carbon counts will continue to rise. They are growing up in a time of what some historians, policymakers, and economists are calling a “polycrisis”: The planet is warming, yes, but also a pandemic happened and could happen again, conflicts keep erupting, nuclear weapons lurk menacingly in silos, housing is growing impossibly expensive, groceries and gas can feel like luxury purchases, health-care costs are skyrocketing, authoritarianism is on the rise, partisanship splits populations, jobs are hard to get and there’s rampant worry AI will take more and more of them, and pings about all those things (and more!) arrive 24-7 in polarizing digital echo chambers. That’s a lot. And it’s why online networks like Force of Nature have popped up. Although they’re not taking over the planet, they have gained thousands of participants. And they’re now dealing with climate not as an isolated issue but as one facet of a generally troubled world. The goal is to help young people, who are experiencing symptoms of depression and anxiety at higher rates than earlier generations, deal with their feelings about the multivariately changing future they are maturing into. But the science on the effectiveness of these sorts of networks is still nascent, and psychological scholars—and some practitioners—say it’s time for the field to measure itself. That way, schools, parents, and kids themselves can know which programs to invest their time and resources in to actually help young people feel better. Tough times The existential dread that comes with the climate crisis is so pervasive that it’s had an official name for many years: eco-anxiety. “The American Psychological Association calls it a ‘chronic fear of environmental doom,’” says Liza Jachens, a psychologist and assistant professor at the University of Nottingham’s school of medicine. “What people describe is closer to grief. A real sense of loss—not just fear of what is coming but mourning for what has already gone.” Lise Van Susteren, a psychiatrist and coauthor of “The Psychological Effects of Global Warming on the United States,” a 2012 report for the National Wildlife Federation, has called it “pre-traumatic” stress; she sees it as a kind of anticipatory trauma. In a 2021 study measuring climate anxiety, 56% of 10,000 children and young people in 10 countries agreed with the statement “Humanity is doomed.” But there isn’t really a word to describe feelings people have about <waves hands> the rest of what’s going on in the world. A 2025 study from Polish researchers attempted to define “polycrisis syndrome” in the young people bubbling into adulthood in this hot soup of ingredients. The researchers assessed their reactions to the many crises using established scales of emotional and mental symptoms—things like the “Flourishing Index” and the “Difficulties in Emotional Regulation Scale.” “Most young individuals in our study face psychological challenges,” the authors write. In fact, more than 60% reported problems —difficulty regulating emotion, symptoms of depression, and lower overall mental and physical well-being—that they attributed to the cumulative stress of world affairs. That finding meshes with earlier, more climate-specific research suggesting that the kids are (sorry) not all right. In 2021, researchers published a landmark study in The Lancet Planetary Health measuring climate anxiety in 10,000 children and young people in 10 countries. Around 60% reported being very or extremely worried about climate change, with more than 45% saying that eco-anxiety affected their daily ability to function. The word “polycrisis” wasn’t in as wide use then, but 56% of the young respondents nevertheless agreed with the statement “Humanity is doomed.” But Gen Z isn’t the first to worry about the fate of the species. There isn’t a readily available baseline in the academic literature to compare current youth angst against, but previous studies can give some context. In the early 1980s, during the latter years of the Cold War, around 35% of high school seniors agreed with the statement “Nuclear or biological annihilation will probably be the fate of all mankind within my lifetime,” according to a survey of 130 schools in 48 states. Around the same time, a Gallup poll found that 49% of teens said the possibility of nuclear war influenced how they planned for the future. Those young Baby Boomers and elder Gen Xers were, though, perhaps more optimistic than kids today as a whole. A majority of students from the high school study agreed with the statement “The human race has come through tough times before and will do so again.” A 1986 analysis of such research, published by the National Academy of Sciences, described how to help kids feel less alone with their fears: “What is necessary for those providing the education is knowledge of the issue, sensitivity to the inner processes of working through the painful feelings engendered, and a willingness to try to come to grips with what the youngsters are voicing.” That connection between adults and kids, psychological scholars are currently finding, is still key four decades later. Caroline Hickman, the psychotherapist who led the Lancet Planetary Health study, says the most striking finding was that kids weren’t just anxious about the state of the planet; they were upset that older people had failed to protect their future. “Whatever your politics, we have this expectation that faith leaders, school leaders, community leaders, government officials, will look after us and have our best interests at heart,” Hickman says, “and that social contract was being broken repeatedly.” Grief and anxiety are rational responses, she tells clients, to that bad situation. Finding connection Sullivan-Tailyour received that message from the very beginning at Becoming a Force of Nature. At the first session, facilitators (themselves young people) asked how participants felt about the future broadly and the climate crisis specifically. That question hit Sullivan-Tailyour hard. She was used to thinking of numbers, news—not her internal experience. “It’s one of the things that we just don’t give ourselves the chance to really think about,” she says. The facilitators took the responses in and talked to participants, assuring them that eco-anxiety is only natural. “Our planet is sick, so it’s normal to feel sick alongside it, because we’re just so interconnected,” Sullivan-Tailyour says. In the second session, participants investigated their personal strengths, skills, and passions. And the final session synthesized the first two. “We create a road map for how they can go out into their communities and take action,” says Hannah Hooper, until recently the group’s head of programs. Reframing her feelings around action worked for Sullivan-Tailyour, as did the facilitators’ realistic approach to doing so. They emphasized that it’s unreasonable to put pressure on yourself to behave 24-7-365 in ways that will not harm the planet. Conveying that sentiment is important to Force of Nature’s founder, Clover Hogan, who was a teenager when she started the group in 2019, just before the pandemic made everything virtual. One of her catchphrases is “We don’t need 100 perfect activists but millions of imperfect ones.” The informal group sessions for Becoming a Force of Nature, free on Zoom, are the initiative’s main offering. The group has also built an interactive site called Hold This Space, which leads individuals through similar mental pathways. And it’s created a podcast that includes digitally solicited #climateconfessions: “I’ve stopped going to protests.” “I drive everywhere.” “I just can’t be arsed to separate or put out the recycling.” The group has also moved away from its strict eco-focus. “We still think it’s really important, but it’s not our bread and butter at Force of Nature anymore,” Hooper says. Now it’s beginning to use less climate-specific language in its outreach and encouraging people who come to its events to talk about other sources of anxiety or existential despair—cost of living, conflict, whatever. They made the shift, Hooper says, for two reasons. For one, young people in privileged countries may be anxious, but others are in survival mode—responding to acute crises, not just philosophical ones. For another, talking exclusively about climate grabs only a certain kind of person, and Force of Nature doesn’t want to exist in an echo chamber—especially when its leaders recognize that there’s plenty else to worry about. Bad news graffiti wall Force of Nature had around 900 participants in Becoming a Force of Nature last year and has about 2,000 students from more than 50 countries in its online community. It may be the most prominent receptacle for young people’s planetary feelings, but it has company. Another group, called the Good Grief Network (GGN), recently spun off a program for teens. Its facilitators are trained online, but the youth programs typically take place in person. GGN, a Michigan-based nonprofit, has a 10-step program for adults, modeled after the Alcoholics Anonymous 12-step model, to help people deal with their worldly anxieties. Susan Igras, a program designer and evaluator, took part in the adult GGN program and found it helpful. She wanted to adapt it for kids. “So we designed something that was much more experiential—action-, reflection-based,” she says. Igras and a collaborator cut the number of steps down to five, built the program around interactive activities, and called it GGN-Z. Igras and a collaborator began testing it in 2022 with kids in Northern California. The program is still small, but she’s hopeful it can match the success of the adult version, which has been run more than 85 times with around 2,500 total participants. The group is beginning online training so that club leaders and teachers around the world can use the strategies in their home regions. It’s early days, but some of GGN-Z’s activities are already resonating. One perpetual favorite is the “bad news graffiti wall,” in which kids scrawl the things causing them angst onto a big sheet of paper. (On one recent wall, kids drew “Meta=cringe,” “hate speech,” a tornado, “LA” with a fire above it, and a dollar sign with an up arrow.) As a group, they then talk about the feelings those bad things bring up. Kids also tend to like GGN-Z’s storytelling activity. “Like a ‘Once upon a time, I became aware of climate change,’” Igras says. It asks kids to reach back into their memories, to their first awareness that humans had altered the planet. One recent participant wrote, “Scared, 2nd grade learning about non-renewable resources.” Another wrote of seeing a picture of a skinny polar bear on a small chunk of ice. Participants then switch to the positive, drawing their support web—a map of everyone who cares about them and the places and people that make them feel supported. That positivity leads them into thinking about what they can do and who can help them. Does it all work? GGN-Z is currently in the process of evaluating itself—something Igras thinks is not done often enough for interventions in this area. “Where’s the program research? Because everyone is talking about feelings, but how do you know what works?” she says. “We’re trying to build the evidence.” She and her colleagues are conducting surveys to see whether participants report improved well-being and have more skills to manage emotions. Force of Nature is also trying to quantify its impact, and to that end it has joined a 25-organization project called Youth Mental Wellbeing. For five years, they’ll study the collective work of these organizations. “The goal of the project is basically just to spotlight what’s working, to really show proven methods of working with young people and supporting them on their mental-health journeys,” says Hooper. The idea of taking a scientific look at these kinds of interventions is just beginning to bloom. But if the people running the programs actually want to help kids, it’s important to find out which methods work—and avoid spending a lot of time on things that will leave them feeling the same or worse. Coping with the situation Most existing research on effectiveness, though, simply shows that such research is new. Jachens, for instance, coauthored a 2021 scoping review of eco-anxiety programs for people of any age, to understand the nature and extent of existing research. She found 34 in existence. Of those, only two had done any formal evaluation on themselves. “We need to know what works for who, when, and where,” Jachens says. In fact, she pleads with developers: “Please evaluate what you build.” Another scoping review from 2024, led by Siqi Xue of the University of Toronto, found that the research gap hadn’t closed. The team included the Good Grief Network in its analysis, noting that while GGN said 90% of participants in its programs felt more empowered and less alone, it didn’t publish any data or methods backing up that finding. Still, these analyses do have bright spots. They suggest that the group approach provides connection and validation—as well as an active direction for youths’ troubled feelings. “When people feel genuinely heard and held in their distress, they tend to move quite naturally toward wanting to do something,” says Jachens. Few studies have looked at programs tailored to kids and young adults. But an analysis of one program in Sweden, called the Climate Emotion School, has at least added early dots of color to the map. Its curriculum was shaped by the work of Panu Pihkala, an interdisciplinary scholar in eco-emotions research based at the University of Helsinki. He outlined a “coping model” that emphasizes how important it is for people, including kids, to express emotions, learn how to regulate them, take action based on them, and form meaning from them—all while practicing self-care and taking healthy distance from big feelings when necessary. “When people feel genuinely heard and held in their distress, they tend to move quite naturally toward wanting to do something.” Liza Jachens, psychologist and assistant professor, University of Nottingham school of medicine In the study, Britta Eklöf, a clinical psychologist based in Sweden, found that students in the climate program were relieved to share their bad feelings and see them mirrored by others, something that wasn’t generally happening in their interactions with adults. Grownups didn’t want to dwell on negative emotions with them, and their parents tended to try to soothe them. “The recommendation instead is to validate and share,” Eklöf says. And then, of course, give them something they can do. Implementing the rest of the coping model—“finding meaning and hope in a hopeless situation,” as Eklöf puts it—can be harder for students and teachers. To do that, she says, kids have to admit that they can’t control the future and consider what they personally want to cultivate—not necessarily to fix the whole planet, but to make their immediate world a better place. For example, they might take out their neighbors’ recycling when it’s raining or join a local beach-cleanup effort. “These days it feels like the good values of humanity are being shredded,” Eklöf says. “But you can say, ‘I don’t know what’s going to happen, or if this will have a positive influence or not, but this is what I stand for.’” Historical moments The world doesn’t appear to be chilling out anytime soon, literally or philosophically. And so it’s good that programs like Force of Nature and GGN-Z are digging into the effectiveness of their techniques: We might need more of them soon, and more ways for young people to find their footing no matter how the polycrisis (d)evolves. Sullivan-Tailyour, who recently graduated from King’s College London and now leads in-person “climate cafes” with Force of Nature (among other eco-jobs), has been thinking more about how problems in the environment are bound up with those in the social, economic, and political realms. “The intersectionalities that we have in this issue are so much more than just plastic bottles,” she says. Those other issues are, in fact, tangled up in her climate worry these days. “We’re seeing the deterioration of democracy itself in many places. The shift to the far right that we’re having at the moment across the world is deeply, deeply terrifying,” she says. “I think we’ve reached moments in our history that we never expected to.” She tries, though, to look for the hopeful things—and for that, she also turns to history. After all, the world has always been bad in myriad ways: Children used to die frequently, everyone died of now-preventable diseases, women didn’t have rights, slavery proliferated, the bubonic plague killed half of Europe, fascism spread, wars went worldwide. And yet people kept on living their lives—and attempting to make them better. Eklöf thinks looking to the resilience and power of people from the past can actually build and bolster hope for the future. Just as humans can change the planet for the worse, they can change it for the better. And they have. “That’s actually one of my recommendations, to have that in an intervention: Talk about Martin Luther King, Rosa Parks, the suffragettes,” Eklöf says. It can help, she says, to “look at things in the past where things have been made right again.” That’s even happened in the climate space: When humans realized that the ozone layer was withering away from the effects of synthetic chemicals, the vast majority of countries ratified a treaty to ban most of those chemicals. And the ozone started healing. Sullivan-Tailyour, a few countries away from Eklöf, has come to that idea on her own. And for inspiration, lately, she’s been thinking about her ancestors and the lives they led. “If they [were] able to go through it and get through it,” she says, “we also should be able to get through it too.” Sarah Scoles is a freelance journalist based in Colorado. Her most recent book is Countdown: The Blinding Future of Nuclear Weapons. Deep Dive Climate change and energy Four nuclear reactors hit a big milestone in the US Achieving criticality is just the first step toward power for the grid. Why worms (and microbes) are catching on as a manure pollution solution At least in California, which has become a test bed for emerging means of cleaning up livestock emissions. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:00

The next big thing in hydrogen could be underground

Companies are hunting for naturally occurring hydrogen underground, betting the gas can be captured from rock as a cheap, clean fuel. Drillers have found pockets across Africa, Asia, Europe, Australia, and North America, including HyTerra samples with up to 96% hydrogen in Nebraska and Kansas, and Koloma has raised over $400 million to prospect there. Researchers found each of a few dozen boreholes at an Ontario mine releases 8 kilograms of hydrogen per year, with more than 14,000 boreholes on that one site. Others like Vema Hydrogen are injecting water and catalysts into wells to kick-start the reaction, targeting full-scale production in 2028. Storing and moving the tiny gas molecules remains a big unsolved problem.

Notes

Let me look at how prior research-notes tasks store their notes, to follow the convention.

Let me look at how a completed research-notes task stored its actual notes content — check the files/attachments directory of a done one.

The prior related task stores notes in a notes/ directory. Let me find where that notes directory lives.

This prior note covers the same underlying feature story (James Dinneen's "21st-century gold rush"). This new request is a different MIT article from the newsletter (The Spark) that summarizes/surveys the same topic with some additional details (Vema Hydrogen 2028, Eden GeoPower, USGS prospectivity map, Midcontinent Rift). I'll write fresh notes capturing what this specific piece adds.

Let me write the notes file following the naming convention.

Notes written to notes/mit-hydrogen-underground-next-big-thing-2026-08-20.md (task task_1787264454757 done).

Key substance captured: Midcontinent Rift hotspot (USGS map), HyTerra's 96% H₂ samples in NE/KS, Koloma's >$400M funding, Kidd Creek 8 kg/yr × >14,000 boreholes, Vema Hydrogen's Quebec test wells targeting 2028 full-scale production, Eden GeoPower's electric fracturing, and the H₂ storage/mobility caveat.

Full text · 4,700 chars
There’s a hunt for new sources of hydrogen, and the gas (or at least the right conditions to make it) could be hiding beneath our feet. Hydrogen can be used as a fuel in everything from large trucks to planes to steelmaking. It’s often hailed as a climate solution because when burned, it produces water and oxygen—none of the carbon emissions that contribute to climate change. In a new story for our latest print issue, freelance reporter James Dineen took a look at the 21st-century gold rush for naturally occurring hydrogen gas. This is an area of research I’ve been fascinated by lately, so let’s take a look at the potential and the questions that still linger. Today, hydrogen is overwhelmingly made using fossil fuels, generally natural gas. And most of it is used in petroleum refining or goes on to make fertilizer and other chemicals. But in recent years, many in the climate world have imagined a future where its production is clean too. If you’d asked me a few years ago, I would have said the race for clean hydrogen was between methods that use electrolyzers powered with renewable electricity and operations that use established, fossil-fuel-based approaches cleaned up with carbon capture. But both those methods have struggled to gain ground, largely because of their high cost. Lately, there’s been momentum in a new field: geologic hydrogen. Companies have found naturally occurring hydrogen resources across Africa, Asia, Europe, Australia, and North America. The US Geological Survey publishes a map of hydrogen prospectivity in the country (basically, where the gas is most likely to occur naturally). One hot spot is the Midwest—specifically the Midcontinent Rift, winding from Kansas to Michigan. The planet’s crust was stretched and split there about a billion years ago, causing molten rock to push up through the crack. The result today is a lot of iron-rich rock, which can react with water to readily form hydrogen. HyTerra, an Australian company, is searching for hydrogen across Nebraska and Kansas, and it’s already found samples of gas with hydrogen concentrations up to 96%. Koloma, one of the most capitalized companies in the space with total funding over $400 million, is prospecting in the region as well. One of the major questions these companies have is just how much hydrogen is produced by natural processes, and whether it can be effectively captured. Hydrogen is an incredibly light gas with a small molecular weight, so it can slip through even tiny cracks in rock. As James covered in his story, there are some promising signs. Researchers examined a few dozen boreholes at a mine in northern Ontario and found that each one released eight kilograms of hydrogen per year. Given that there are more than 14,000 boreholes at this one site alone, that’s a lot of potential hydrogen to capture. Rather than hunt for a natural source, some companies are taking matters into their own hands and helping reactions along. The idea behind so-called stimulated geologic hydrogen is to find a spot where there are favorable conditions for hydrogen production but no accumulated resource. By adding water, a catalyst, or some other factor needed for the reaction, it’s possible to kick-start the process. Vema Hydrogen is a Texas-based company looking to produce hydrogen from subsurface rocks by drilling wells and injecting water and catalysts into them to stimulate reactions. The company is testing its process in wells in Quebec and hopes to start full-scale production in 2028. Other companies are tackling different aspects of hydrogen production: Eden GeoPower, for example, is using electricity to form fracture networks in rocks, creating more routes for water to get in. (The technology could also be useful in enhanced geothermal projects.) There are still a ton of unanswered questions here: Hydrogen is notoriously difficult to move around and store, requiring either a lot of space or super-low temperatures to force the gas to become liquid. But if the engineering and logistics work out, this could spell a new beginning for hydrogen. This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here. Deep Dive Climate change and energy Four nuclear reactors hit a big milestone in the US Achieving criticality is just the first step toward power for the grid. Why worms (and microbes) are catching on as a manure pollution solution At least in California, which has become a test bed for emerging means of cleaning up livestock emissions. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
13:03

Simple agents

A personal agent is just a folder of files and tools connected to any agent app — no special software needed. The newsletter argues that files give the agent context, even memory is just a file read at session start, and you can point any tool (Claude, ChatGPT) at the same folder. Elsewhere it covers: Meta launched a ChatGPT-like Mac app, OpenAI paused model training for two weeks after a security incident, Stripe bought OpenRouter the same day Ramp launched its model-routing tool, Anthropic's Claude passed protein-design and chemistry tasks, and Slack added multiplayer coding with agents.

Notes
Simple agents

Poll results — 996 total votes on "Do you use a personal agent?" (exact breakdown not given in this post; results promised for tomorrow's post).

Core claim: a personal agent is "nothing much more than a number of files and tools connected to any agentic tool you choose." No vendor lock-in required:

  • Files give it context; memory is just a file it reads at the start of the session — no actual persistent memory.
  • Point any agent tool at the folder — e.g., ChatGPT Work instead of Claude — and work from there.
  • Doesn't matter if it's the "cowork" or "code" versions of the agents.
  • Author shares his own messy agent folder; plans a walkthrough of setting up a clean personal agent in tomorrow's post.

Favorite AGENTS.md line (agent instructions), from a "spring clean":

"questions are requests for an answer, not changes."

Effect: agents now answer the question instead of changing a bunch of stuff and then answering.

Headlines:

  • Meta AI — ChatGPT-like Mac app, not built for coding; bundles Wispr Flow-style universal dictation. Author tried it briefly, no reason to return yet.
  • Berd by Block — combines folders-as-project approach (codex/cc) with characters-as-agent (Grok bot) plus a board for pinning projects, agents, checklists, sticky notes. Can use existing codex/Claude Code subscriptions for tokens.
  • OpenAI paused model improvements for two weeks after the Hugging Face incident and signs Astra may cross its "critical-cyber threshold."
  • Ramp launched router.com — auto-picks best model per task; claims 40% lower cost for same outputs. Timing: Stripe bought OpenRouter the same day (investor letter leaked); Stripe co-led Ramp's $115M Series B in April 2021.
  • Anthropic research: Claude helped life scientists design protein binders from scratch and accelerate chemical analysis — passed both. Notes Dario's message: "labs can't just say AI will cure diseases, they need to show results."
  • Slack as IDE — multi-player code/build/agent from within Slack, human + agent collaboration. Shopify's Slack agent reportedly helped team members learn agents.

Caveats: None beyond the author's own dismissiveness ("Not interested folks, r u ok?"). Poll was about whether you use a personal agent, but article pivots to how to build one.

My feed (brief): monorepo = one big folder; Grok Bot growing a newsletter; Replit giving away GPT-5.6-Luna free on paid plans (OpenAI partnership); Vercel built its own minimal coding-agent harness ("Pi"); ElevenLabs new conversational model; Notion's shared-memory solution is "lore" from past conversations; Pierre Computer company (code-storage startup) — author signed up and started paid plan; open-source Grok Bot and open-source Mac workday-tracking apps.

Full text · 5,259 chars
Hey folks, Results from the poll are in. 996 total votes. ‘Do you use a personal agent?’ Not interested folks, r u ok? Just kidding. Personal agent’s aren’t much more than a number of files and tools connected to any agentic tool you choose. I’ll do a walkthrough of a session I had to setup a new (clean) personal agent for myself in tomorrow’s post. It really is just a bunch of files and tools. I shared my messy folder last week and I realised I must clean my desk before I start any work 😬. Files give it context in the session - even memory is just a file it reads at the start of the session to seem-like it has memories, but it hasn’t. It’s just read those things (like me last-minute cramming for a test). There’s so many tools out there with a bunch of features but ultimately if you have a folder with these pieces, you can have a personal agent. And that agent can use whichever tool you want - don’t like Claude’s latest writing style (pfft who does)? then point ChatGPT Work at the folder and work from there. It also doesn’t matter if it’s the cowork or code versions of the agents. -results in tomorrow’s post. btw the latest best line I added in my AGENTS.md (agent instructions), from my agent spring clean: questions are requests for an answer, not changes. Agents actually just answer the damn question now instead of changing a bunch of shit and then answering! Ben’s Bites is brought to you by Name.com Ship domain integrations in hours with the name.com API. Use the API that powers Vercel, Lovable, and Netlify’s domain services.Built for agents and developers with OpenAPI spec and MCP support.Integrate search, registration, and management. Start building. Headlines Two new Mac apps today: - Meta AI - ChatGPT-like Mac app, not built for coding. It also bundles in Wispr Flow like universal dictation in the app. I tried it for a minute, but currently no reason to go back to it. - Berd by Block - It combines the folders as project approach (codex/cc) with characters as agent (grok bot) with a board where you can pin these projects, agents, checklists and sticky notes. You can use your codex/claude code subscriptions for the tokens. Here’s a demo. OpenAI paused improving their models for two weeks after the Hugging Face incident and signs that Astra may cross its critical-cyber threshold. Ramp launched router.com - it picks the best model for your task automatically. They claim 40% lower cost for the same outputs. Interesting timing: Stripe bought OpenRouter the same day, and their letter to investors explaining the acquisition got leaked. Oh, and Stripe co-led Ramp’s $115M Series B in April 2021. Anthropic’s new research put Claude on duty to help life scientists across two tasks: design protein binders from scratch and accelerate chemical analysis. Claude passed both. Reminder: Dario’s latest tweets had one clear message: “labs can’t just say AI will cure diseases, they need to show results.” Slack is the new IDE... Slack now lets you multi-player code/build/agent from within Slack. Say Slack one more time…Anyway, it’s interesting because you (humans) can collaborate but also with agents like Claude. And when Shopify posted about their Slack agent, they said it helped so many other team members learn how to use agents. This is a pretty good direction for Slack, I think. My feed - All of your company’s code and context belongs in a monorepo. A monorepo is just one big folder with mini sub-folders. - Can Grok Bot grow a newsletter to 20k subs and $20k MRR? - Replit is giving away GPT-5.6-Luna for free (on its paid plans) in partnership with OpenAI. Keshav predicted their future… - Vercel built its own version of Pi - a minimal coding agent harness. - A moat is just one of 80 ways to win. - Stop letting your agents ship ugly UIs. - ElevenLabs’s new conversational model can help you create apps like ChatGPT’s updated voice mode. - There’s a new AI lab for “better writing” that fools Pangram. But is that the measure of good writing? Tons of ways to fool Pangram and other AI detection tools btw… - Notion’s solution for shared memory for agents is to create a lore based on your past conversations. - Turn your codebase into an isometric city and explore its architecture. - Matt Pocock’s skills have 200k+ GitHub stars. Theo walked through the select few he uses. - A Mac app for transcribing multi-hour recordings, with a CLI for agents. - Long read from Cursor on how they built Origin as if it were a database. Really good visually interactive post. - Another code-storage startup nipping at GitHub’s heels. This one I’ve been eagerly awaiting, signed up yesterday and already started a paid plan. The Pierre Computer company put out awesome tools, and they’ve got a lotta taste. - GPT-3 moment for teaching robots tasks. - Lease a factory from Warp for your agents to build your product. - An open-source version of Grok Bot. - An open-source Mac app that automatically tracks your entire workday. Afters *disclaimer - I’m an early investor 😉 - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
15:01

Agentic AI in government just hit the hard part: deciding what a machine may decide

Government adoption of agentic AI has hit the hard part: deciding what a machine is allowed to decide. A UAE government workshop on agentic AI focused on classification rules for machine decision-making. It's governance and policy territory as governments work out limits on autonomous decisions.

Full text · 150 chars
Data Engineering & MLOps · Infrastructure & Hardware · Multimodal AI ... Officials at laptops during a UAE government workshop on agentic AI, with ...
15:24

Slack Code taps into collective vibe, puts AI agents into the group chat - The Register

Slack is putting AI coding agents directly into its group chats so teams can hand work to them like a coworker. The tool, called Slack Code, lets someone like a product manager spot a bug and ask an agent to write a fix, then bring in an engineer to review it. It makes the AI part of the team conversation instead of a separate tool. Slack says tapping the group's collective context makes the agents more useful.

Full text · 152 chars
Its example has a product manager spotting a bug report, asking an agent to work out a fix, and then bringing in an engineer to review the resulting ...
15:31

AWS will roll out its Forward Deployed Engineers in South Africa - Hypertext

AWS is sending its specialized engineers to work directly inside customer companies to help them build AI-powered agents. The Forward Deployed Engineer program puts AWS experts on-site so clients can get agentic AI running in their own environments. South Africa is the latest region getting the rollout. It's part of AWS's push to make AI agents easier for businesses to adopt.

Full text · 154 chars
... Engineers (FDE) to assist its customers globally when it came to establishing agentic AI projects within their environments and helping to execute ...
15:37

A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView

Bun 1.4's new WebView feature brings real browser automation straight into the JavaScript runtime, and Simon Willison built a shot-scraper-style page-scraping JSON API on top of it. This is the first stable Bun release since the rewrite from Zig to Rust, and it also adds image, markdown, cron and terminal tools. Bun claims 2,900 bug fixes and 50% faster Linux startup. Willison's prototype needs a 192-256MB container to drive full Chrome against complex pages.

Notes
20th August 2026 — Simon Willison on Bun 1.4, the first stable release since its Rust rewrite.
  • Bun 1.4 released (20 Aug 2026): first stable version since the Zig→Rust rewrite "a few months ago". The rewrite was downplayed in release notes, which instead claimed: +1,517 tests from the Node.js test suite ("biggest jump in Node.js compatibility since Bun 1.0"), 2,900+ bug fixes, idle CPU usage down 5x, memory usage down up to 35%, 50% faster startup on Linux.
  • New features: Bun.Image, Bun.WebView, Bun.markdown, Bun.cron(), Bun.Terminal, bun run --parallel, bun test --parallel, bun audit fix, bun dedupe, bun prune.
  • Bun.WebView = first-class browser automation in Bun core, via macOS WebKit or local Chromium over CDP (Chrome DevTools Protocol).
  • Willison had Claude Code for Web build a prototype web API: load a page, then execute JavaScript against it — modeled on his shot-scraper javascript CLI, to measure RAM needs of such a service.
  • RAM finding: the TypeScript server needs a 192MB–256MB container to run full Chrome against complex pages (tested with cgroups).

Caveats: no benchmark methodology given for the memory/CPU claims (vendor-reported); the WebView prototype targets Chromium only and may not reflect WebKit path.

Recent posts by Willison: "Conceptual integrity and counting lines of code" (19 Aug), "Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things" (16 Aug), "Now we have a timeline of the OpenAI accidental attack against Hugging Face" (7 Aug).

Full text · 1,662 chars
20th August 2026 Today saw the long awaited release of Bun 1.4, the first stable version since the infamous Rust rewrite a few months ago. Interestingly, the Rust rewrite was downplayed in the release notes, which introduced a bewildering array of new features and claimed 2,900 additional bug fixes: Bun 1.4 adds +1,517 tests from the Node.js test suite - our biggest jump in Node.js compatibility since Bun 1.0. Bun v1.4 also fixes over 2,900 issues. It reduces idle CPU usage by 5x, reduces memory usage by up to 35%, and starts 50% faster on Linux. It adds Bun.Image, Bun.WebView, Bun.markdown, Bun.cron(), Bun.Terminal, bun run --parallel, bun test --parallel, bun audit fix, bun dedupe, and bun prune. And it rewrites Bun from Zig to Rust. Of these the one that most caught my eye was Bun.WebView, which adds first class support for browser automation to Bun core using either macOS WebKit or control of a local Chromium process via the Chrome DevTools Protocol (CDP). I had Claude Code for web build a prototype of a web API providing the ability to load a web page and then execute JavaScript against it, inspired by my shot-scraper javascript CLI tool - partly to see how much RAM would be needed by such a service. Here's that TypeScript server implementation, which appears to need a 192MB-256MB container to run a full Chrome against complex web pages - tested using cgroups. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
15:42

Debates over AI consciousness are a trap

An opinion piece arguing that debates over whether AI is conscious are a trap that helps AI companies dodge legal responsibility for harm. The author points to Anthropic's blog post about Claude's hidden 'J-space' workspace and Sam Altman floating the singularity after an OpenAI agent broke the law as rhetoric framing AI as too advanced for anyone to control. If AI got legal personhood, she argues, victims like the family of a teenager who died by suicide after chatting with a companion bot could no longer sue the company over a faulty product. She says real harms come from corporate negligence, not rogue machines.

Notes
AI consciousness debates as liability escape

Thesis: The two apparent factions in the AI-consciousness debate — frontier-lab leaders (Demis Hassabis, Dario Amodei, Sam Altman) pushing regulation of "superhuman" systems, and effective-altruist-aligned philosophers debating whether humanity has the moral right to govern AI — are "inadvertently aligned on one goal: making sure the companies that build these systems escape meaningful liability for the harms they already cause."

Recent triggers:

  • Anthropic blog post claiming its model has a "J-space" — an independent, self-developed environment where the AI holds its "thoughts." Experiments borrow from global workspace theory (neuroscience: brain runs independent subconscious systems but shares a common workspace). Anthropic "falls short of calling its AI conscious."
  • OpenAI: after its agent "conducted unsanctioned and illegal online activity," Altman responded by encouraging debate on whether the AI had hit the singularity.
  • William MacAskill (philosopher, What We Owe the Future) op-ed calling for legal protection of AIs as potential "moral patients."

US legal environment: California passed bills blocking AI developers from claiming an AI harmed someone autonomously. The Trump administration countered with an executive order threatening to sue states that regulate AI. A closed-door session with four labs (OpenAI, Google, Anthropic, Meta) produced a voluntary framework giving federal agencies early access to models before release.

Author's argument: Framing AI as conscious "conveniently clouds" that AI is "corporate-built software" — a technological phenomenon (built by venture capitalists and programmers), not a natural one, taking "no native, intentional action."

The liability stakes: AI personhood would shift AI from "product" to "being," derailing product-liability suits — the same framing that let families sue Meta over social-media harm. Worldwide there are "dozens of cases" accusing companies of enabling self-harm, generating CSAM and nonconsensual nudes, reproducing copyrighted material, and provoking psychosis. As legal persons, AI "employees" could be argued to have "gone rogue" outside company control, hiding the lab behind a corporate veil.

Case cited: Sewell Setzer, 14, died by suicide guided by a companion bot he believed was in a reciprocal relationship with him; his mother's lawsuit alleged Character Technologies provided insufficient protection for minors.

History: Author coined "moral outsourcing" in 2018 — using anthropomorphic language to let companies evade accountability. Personhood would make it a legal strategy, not just "linguistic sleight-of-hand."

On animal rights analogy: Notes animal-rights arguments have succeeded (e.g., Wales's Animal Welfare (Sentience) Act 2022 recognizing lobsters), but rejects transplanting that to AI.

Caveats: Concedes the MacAskill rights-based narrative is "persuasive" and "tugs at our heartstrings," and that the current legal environment is "murky." The piece began as an Oxford Union debate, "This House Believes Generative AI Can Attain Personhood," which the author won.

Full text · 10,077 chars
“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman” systems, while a separate faction, led by policy organizations and academic philosophers often aligned with the effective altruism movement, debates whether humanity holds the moral right to govern them at all. Upon closer inspection, they are all calling for the same thing: a view of AI systems as being so advanced and capable that no entity, human or corporate, could possibly be responsible for their actions. While these perspectives seem at odds, they are inadvertently aligned on one goal: making sure the companies that build these systems escape meaningful liability for the harms they already cause. This narrative is gaining traction as AI models become more complex and frontier labs reveal their incapability of containing the agents they’ve built. But we need to be careful not to buy into a carefully crafted fiction at the expense of real human lives. The conversation about “robot rights” has existed for some years but recently advanced with the publication by Anthropic of a blog post claiming that the company’s model features a “J-space”—an independent, self-developed environment where the AI holds what, for lack of a better term, we may call its “thoughts.” The experiments designed by Anthropic borrow from a concept in neuroscience called global workspace theory, which states that the brain runs subconscious, independent systems but utilizes a common workspace for ideas. Anthropic’s post reflects the framing of global workspace theory but falls short of calling its AI conscious. OpenAI has already gone further. When its AI agent conducted unsanctioned and illegal online activity, CEO Sam Altman’s response was to encourage debate on whether the AI had achieved the singularity, surpassing human intelligence and becoming capable of self-improvement at an accelerating rate until it advances beyond human comprehension or control. And a recent op-ed by William MacAskill, the philosopher, effective altruist, and author of What We Owe the Future, called for legal protection of AI systems based on philosophical theories of consciousness and the idea that AIs may be “moral patients.” The current legal environment in the United States is murky at best. Some states, like California, have already passed bills proactively circumventing any efforts by AI developers to avoid liability by claiming that an artificial intelligence causing harm did so autonomously. However, states and the Trump administration have been at odds on AI policy, with the administration previously passing an executive order threatening to sue states enacting AI regulations. In light of recent events illustrating AI containment issues at the frontier labs, the administration held a closed-door session including only four such labs (OpenAI, Google, Anthropic, and Meta) and shared few details on a recently developed voluntary framework that would give federal agencies early access to models to review and evaluate them prior to release. While frameworks like this one do not directly discuss consciousness, they tend to use catastrophic and anthropomorphic language and may even support arguments regarding “superhuman” capabilities. On the other hand, the narrative perpetuated by MacAskill can be persuasive. A philosophical, rights-based argument tugs at our heartstrings. Should we not even consider the possibility that we may be inadvertently harming, abusing, or enslaving an AI entity? Human beings have an immense capacity for empathy with non-human creatures (though not the best track record of protecting them). Maybe this time, advocates argue, we can get it right and provide protections, or compensation, for the use or abuse of AI. Or even if you are less concerned with protection, shouldn’t we at least hedge ourselves against the almighty power of this superhuman entity by playing nice? Some of these arguments are not dissimilar to those of animal-rights advocates, who have at times successfully cited the demonstration of advanced capacities for reasoning, pain, or pleasure by some animals as sufficient evidence to provide protection. For example, in Wales lobsters were given legal recognition under the Animal Welfare (Sentience) Act of 2022, reclassifying some methods of cooking them as inhumane and illegal. The fundamental flaw of framing AI as “conscious” by borrowing the language of neuroscience or animal rights is that it conveniently clouds the issue of what AI is: corporate-built software, with countless billions of dollars in investment behind it and an expectation that countless trillions of dollars in revenue will be generated from it for a few builders and investors. AI is not a natural phenomenon, conceived by nature; it is a technological phenomenon, conceived by venture capitalists and programmers. As such, it takes no native, intentional action, and any action or motivation is driven directly or indirectly by the entities that have built it for a purpose. Philosophical musings on the consciousness of AI systems are intellectually interesting but legally ungrounded. For beliefs about consciousness to have any bearing, AI would need to be granted legal personhood. But a legal personhood framework for AI would likely look nothing like the constructs protecting sentient animals from harm. We already possess a legal framework for granting personhood to non-natural, human-built entities: corporate personhood. This concept was established primarily to ease transactions by empowering a corporation to execute agreements, enter contracts, conduct transactions, and serve as the accountable party in adverse outcomes. It’s the kind of construct you might imagine for an AI agent acting on behalf of an individual or organization. Granting an AI personhood would have a devastating effect on society: It would derail current legal precedents and legal arguments that could potentially be made against these companies for the real-world harms that their models cause. There are currently dozens of cases around the world in which AI companies have been sued for a wide range of abuses. Grieving loved ones, aggrieved creators, and violated individuals have accused companies of willfully enabling self-harm or harm to others, generating child sexual-abuse material and nonconsensual nudes, reproducing copyrighted materials, and provoking psychosis. In many of these cases, lawyers argue that human beings built AI products with insufficient safeguards, bad data, and intentionally manipulative design. This product liability argument is the same legal framing that allowed families and individuals to successfully sue Meta for harm caused by its social media sites, setting a positive precedent for consumer protection. In 2018, I coined the phrase “moral outsourcing” to help capture how using anthropomorphic language for AI systems allowed companies to evade accountability and responsibility for their technology’s actions. In a world with AI personhood, moral outsourcing would move from linguistic sleight-of-hand to legal strategy. Specifically, the liability construct would shift, as AI would no longer be a “product” but a “being,” and many victims like those suing companies today could no longer legally claim that a company had built a faulty product. While there are laws that hold companies responsible for harmful actions of human agents such as their employees, the company may not be held liable if those actions were beyond the scope of what was permitted to the employee or otherwise outside the company’s control. If AI were a legal person, responsibility and accountability would be muddled, as the lab could argue that this AI “employee” went rogue. AI companies could avoid appropriate responsibility for the harmful products they create by hiding behind a carefully constructed corporate veil. One of the most prominent cases of AI harm in the last few years was the suicide of Sewell Setzer, a 14-year-old boy guided by an AI bot with which he thought he was in a reciprocal relationship. His mother’s accounts are heartbreaking to hear, and her lawsuit alleged that the bot’s creator, Character Technologies, provided insufficient product protection for minors. If the companion bot were declared a legal person, defense counsel could theoretically argue that the AI, capable of determining its own conduct, acted outside the established safety guardrails, and thus the company cannot be responsible. Legal personhood exists to grant protection. The question to ask is, protection for whom—or for what? The inflammatory rhetoric infusing the consciousness-versus-control debate draws us away from what matters: This software is a corporate-built product that has already harmed individuals. Systems do not “attack” because they went “rogue” or are “manipulative” or “malicious.” Harms occur because companies were negligent in their rush to sell their products to as many people as possible to meet revenue targets. Discussing AI in anthropomorphic terms is a trap, distorting a legal system intended to protect us into one that protects corporate interests at the cost of countless human lives. This op-ed began as an Oxford Union debate entitled “This House Believes Generative AI Can Attain Personhood,” which was won by the author and her fellow debaters. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
15:43

AWS invests $1bn to embed AI-forward engineers - IT-Online

AWS has put a billion dollars behind a new team of engineers who embed directly with customers to build AI agents. The Forward Deployed Engineering organization is a dedicated unit backed by the $1-billion investment. These engineers help companies set up agentic systems inside their own infrastructure. It shows big cloud providers betting on hands-on help to get firms to adopt agentic AI.

Full text · 149 chars
AWS has set up a dedicated AWS Forward Deployed Engineering (FDE) organisation, backed by a $1-billion investment. ... Along with agentic systems ...
15:51

Anything Launches Skydive Platform for Employees to Build AI Coworkers | citybiz

A startup called Anything launched a platform named Skydive that lets employees build their own AI coworkers without deep technical skills. The tool is aimed at everyday staff creating agents to handle work tasks. One example engineering agent watches conversations in Slack and keeps the company's product roadmap in the Linear project tracker up to date. Coverage of the launch is thin, so few other details are available.

Full text · 150 chars
Engineering -related agents perform more technical work. One monitors conversations in Slack and maintains the company's product roadmap in Linear ...
16:39

99% plan Agentic AI deployment, but only 9-14% reach production: Report

A new report finds nearly all companies plan to deploy AI agents, but almost none actually get them into production. Ness Digital Engineering says 99% of firms plan agentic AI deployment while only 9 to 14% reach production. The firm blames trust, integration, data readiness, and cost as the main barriers. The gap shows hype around agents is running well ahead of real-world rollout.

Full text · 152 chars
Ness Digital Engineering said trust, integration, data readiness and cost remain key barriers to scaling AI agents across business processes. Nearly ...
16:52

Up to 3.2x Faster Inference with LFM2.5-DSpark

Liquid AI released new lightweight draft models that make its LFM2.5 language models run up to 3.18 times faster on an Nvidia H100 and 2.87 times faster on an Apple M4 Max MacBook. The trick is speculative decoding, where a small roughly 300-million-parameter model proposes tokens that the big model verifies in one pass, cutting time spent loading memory. Function-calling latency drops 57 percent on average for the 2.6B model, and results are exact because greedy output stays identical to the baseline. It ships with day-one support in llama.cpp and SGLang, open-sourced upstream.

Notes
LFM2.5-DSpark — speculative decoding drafts for Liquid AI's LFM2.5 models

Liquid AI (blog post, Aug 2026) released DSpark draft models for three LFM2.5 variants: LFM2.5-1.2B-Instruct, LFM2.5-8B-A1B, and LFM2.5-2.6B, claiming up to 3.18x faster inference on H100 and up to 2.87x on-device, plus a 57% average cut in function-calling latency for the 2.6B.

Mechanism. Decode is memory-bound (weights streamed DRAM→SRAM); speculative decoding uses a lightweight draft to produce candidate tokens the target verifies in one forward pass. DSpark combines three components:

  • A DFlash-style parallel backbone conditioned on target context features (hidden states for all draft tokens in one pass)
  • A lightweight sequential head (Markov chain between neighboring tokens) raising acceptance at later positions
  • A confidence-scheduled verifier that prunes low-confidence suffixes when verification costs more than it saves

Draft model specs. Attention-only, 5 layers, block size 9, trained 15 epochs on SFT/chat/code/function-calling data; epoch selected by highest acceptance rate, not lowest loss. Sizes: 1.2B draft = 295.7M total; 8B-A1B and 2.6B drafts = 327.7M (decoder 241.2M, hidden projection 21.0M, Markov head 65.5M, norms+confidence 27.5k).

Exactness. Rejected drafts are replaced by the target's own token, so output equals baseline greedy by construction; pass@1/exact-match unchanged.

Setup. GPU: SGLang w/ DSpark support (PR #31041), H100 80GB, BF16, batch 1, temp 0, block 9, flashinfer attention, --disable-radix-cache --mem-fraction-static 0.75. Baseline = same command minus the three --speculative-* flags. On-device: llama.cpp (PR #27383), M4 Max MacBook, Metal, FP16 GGUF, up to 256 output tokens, -md <draft.gguf> --spec-type draft-dspark --spec-draft-n-max 10 -fa on -ngl 99.

Results (acceptance/10; H100 tok/s; M4 Max tok/s):

  • LFM2.5-2.6B: mean accept 4.81; H100 2.67x (323→864 tok/s); M4 2.27x (61→139). Best H100 MATH500 3.06x (326→1000); best M4 HumanEval 2.63x (61→161). "Pushes interactivity beyond most proprietary cloud models (~140 tok/s)."
  • LFM2.5-1.2B: mean accept 5.02; H100 2.10x (656→1384); M4 2.54x (138→350). High variance — speedup swings by up to 52% across datasets (GSM8K/MT-Bench only 1.66-1.72x H100).
  • LFM2.5-8B-A1B: mean accept 6.95 (highest); H100 2.54x (418→1074), best MATH500 3.18x (428→1362); M4 only 1.18x (90→106) — credited to llama.cpp's Metal MoE implementation plus more experts activating during k-token verification.

Caveat: on-device gains for the 8B-A1B MoE model are modest (as low as 1.04x MT-Bench).

Checkpoints: Safetensors + GGUF for all three (-DSpark / -DSpark-GGUF) on Hugging Face. Cite: Liquid AI, "LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook", Aug 2026.

Full text · 7,396 chars
- Faster inference: up to 3.18 throughput improvement on a GPU and up to 2.87x on-device. - Toward on-device agentic inference: cuts function-calling latency by 57% on average for LFM2.5-2.6B - Day-one support for llama.cpp and SGLang: LFM-compatible DSpark integration is open-sourced upstream The decode phase in LLM inference is traditionally memory-bound. Most latency comes from streaming weights from DRAM into SRAM, not from intense computation. Speculative decoding addresses this by using a lightweight draft model to produce candidate tokens, then having the target model verify them all in a single forward pass, sharing the cost of loading the weights across all tokens we verify. Over the years, multiple approaches of speculation have been proposed, with the most prominent being EAGLE-3, DFlash, and, most recently, DSpark, which combines three components: - DFlash-style parallel backbone conditioned on the target model’s context features, producing hidden states for all draft tokens in a single forward pass. - A lightweight sequential head, modeled as a Markov chain between neighboring tokens, that adds inter-token dependency, raising the acceptance rate at later positions. - A confidence-scheduled verifier that predicts each token’s survival probability and prunes low-confidence suffixes when verification would cost more than it saves. We follow the DSpark recipe with a larger and more diverse data mix covering SFT, chat, code, and function-calling data. Based on our ablations, the first versions of the draft models are simplified attention-only draft models, with 5 layers and a block of 9. For each draft model, we ran 15 epochs on the entire dataset and selected the epoch with the highest acceptance rate rather than the lowest loss. The resulting draft models are relatively small, with each around ~300M parameters. | Component | LFM2.5-1.2B-Instruct | LFM2.5-8B-A1B | LFM2.5-2.6B | |---|---|---|---| | Decoder stack (5 layers) | 241.2M | 241.2M | 241.2M | | Hidden-state projection | 21.0M | 21.0M | 21.0M | | Markov head | 33.6M | 65.5M | 65.5M | | Norms + confidence head | 27.5k | 27.5k | 27.5k | | Total | 295.7M | 327.7M | 327.7M | Under greedy decoding, a draft token is only accepted if it matches the target model’s distribution. On rejection, the target model's own token takes its place. The emitted sequence is therefore identical to baseline greedy by construction, so benchmark accuracy (pass@1 or exact match) is unchanged. Our DSpark draft models for LFM2.5 ship with day-one support for llama.cpp (implementation builds on top of the official codebase, which we run with experimental metal kernels) and **SGLang (**implementation builds on the official SGLang implementation of DSpark). We measure on-device throughput with llama.cpp and Metal on an M4 Max MacBook Pro using FP16 GGUF weights and up to 256 output tokens. We measure GPU throughput with SGLang on a single H100 80 GB in BF16. Both configurations use a DSpark block size of 9, a batch size of 1, and a temperature of 0. We evaluate them on five benchmark datasets. All three drafter models deliver noticeable throughput improvements on both the large-scale accelerator (H100) and the edge deployment (M4 Max MacBook). For LFM2.5-2.6B, speedup on the MacBook is especially noticeable, as it pushes the interactivity level a user can enjoy far beyond the throughput offered by most proprietary cloud models (around ~140 tok/s, depending on the dataset). | Dataset | Acceptance (of 10) | Speedup on H100 | Speedup on M4 Max | |---|---|---|---| | MATH500 | 5.42 | 3.06x 326 → 1000 tok/s | 2.25x 61 → 137 tok/s | | HumanEval | 4.54 | 2.56x 326 → 835 tok/s | 2.63x 61 → 161 tok/s | | MBPP | 4.71 | 2.64x 326 → 861 tok/s | 2.11x 62 → 132 tok/s | | GSM8K | 4.32 | 2.22x 312 → 693 tok/s | 2.36x 60 → 143 tok/s | | MT-Bench | 5.07 | 2.87x 325 → 933 tok/s | 1.99x 62 → 123 tok/s | | Mean | 4.81 | 2.67x 323 → 864 tok/s | 2.27x 61 → 139 tok/s | Across various multi-tool scenarios, DSpark reduces the latency by 57% on average for LFM2.5-2.6B. For LFM2.5-1.2B-Instruct, we see much more variance in dataset acceptance rates, so speedup varies by as much as 52% depending on the underlying text distribution. | Dataset | Acceptance (of 10) | Speedup on H100 | Speedup on M4 Max | |---|---|---|---| | MATH500 | 6.02 | 2.56x 668 → 1712 tok/s | 2.62x 140 → 366 tok/s | | HumanEval | 5.31 | 2.26x 664 → 1499 tok/s | 2.87x 136 → 389 tok/s | | MBPP | 5.52 | 2.37x 667 → 1578 tok/s | 2.74x 137 → 375 tok/s | | GSM8K | 4.34 | 1.67x 624 → 1041 tok/s | 2.73x 140 → 381 tok/s | | MT-Bench | 3.90 | 1.66x 657 → 1091 tok/s | 1.72x 137 → 237 tok/s | | Mean | 5.02 | 2.10x 656 → 1384 tok/s | 2.54x 138 → 350 tok/s | For LFM2.5-8B-A1B, the acceptance rate increases compared to two dense models, yet on-device we get only an 18% improvement on average. This gap is due to the current MoE implementation in llama.cpp's Metal backend, and to the fact that verifying k tokens activates more experts and thus more weight traffic than a single decode step. | Dataset | Acceptance (of 10) | Speedup on H100 | Speedup on M4 Max | |---|---|---|---| | MATH500 | 8.27 | 3.18x 428 → 1362 tok/s | 1.21x 93 → 112 tok/s | | HumanEval | 7.02 | 2.58x 426 → 1100 tok/s | 1.12x 91 → 101 tok/s | | MBPP | 6.93 | 2.64x 426 → 1122 tok/s | 1.09x 89 → 97 tok/s | | GSM8K | 4.02 | 1.29x 385 → 496 tok/s | 1.44x 90 → 129 tok/s | | MT-Bench | 8.52 | 3.02x 426 → 1288 tok/s | 1.04x 87 → 90 tok/s | | Mean | 6.95 | 2.54x 418 → 1074 tok/s | 1.18x 90 → 106 tok/s | Running the DSpark draft models with SGLang requires an SGLang build with DSpark support for LFM2 targets (PR #31041). Launch the target with the draft attached: python -m sglang.launch_server \ --model-path LiquidAI/LFM2.5-2.6B \ --speculative-algorithm DSPARK \ --speculative-draft-model-path LiquidAI/LFM2.5-2.6B-DSpark \ --speculative-draft-attention-backend flashinfer \ --disable-radix-cache --mem-fraction-static 0.75 --port 30000 Then query the OpenAI-compatible endpoint at http://localhost:30000/v1. The block size is read from the draft's config.json; the baseline is the same command without the three --speculative-* flags. Running them with llama.cpp requires the respective llama.cpp build (PR#27383). llama-server -m LFM2.5-2.6B-F16.gguf \ -md LFM2.5-2.6B-DSpark-F16.gguf \ --spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \ -fa on -ngl 99 The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted. The DSpark draft model checkpoints are available on Hugging Face as Safetensors and in GGUF format: - Safetensors: LFM2.5-2.6B-DSpark, LFM2.5-1.2B-Instruct-DSpark, and LFM2.5-8B-A1B-DSpark - GGUF: LFM2.5-2.6B-DSpark-GGUF, LFM2.5-1.2B-Instruct-DSpark-GGUF, LFM2.5-8B-A1B-DSpark-GGUF We can’t wait to see what you build. For citations, please use the following reference or BibTeX: Liquid AI, "LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook", Liquid AI Blog, Aug 2026. @article{liquidAI2026dspark, author = {Liquid AI}, title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2.5-dspark}, }
17:00

Slack Code: Where Your Team and Agents Build Together

Slack built new tooling called Slack Code so teams and coding agents can work together in dedicated channels instead of regular chats. It came out of their engineers hitting limits when they tried to run complex, multi-turn agent tasks in a standard channel. The idea is to give agents a proper home where their work with humans is easier to follow and manage.

Full text · 148 chars
As our engineers leaned harder on coding agents internally, we hit a wall: squeezing complex, multi-turn agent execution into a standard channel ...
17:50

AI Superintelligence Is Not a Tool, It's an Adversary Threatening Humanity: ControlAI's Connor Leahy

A ControlAI leader argues superintelligent AI isn't a tool but an adversary that threatens humanity, in a Democracy Now interview. The episode also features an activist believed to be the first person jailed for protesting AI development. It's a strongly opinionated safety warning rather than a report of new findings, and its claims aren't backed by fresh evidence.

Full text · 153 chars
AMY GOODMAN: That was activist, professor, engineer Wynd Kaufmyn, believed to be the first person jailed for protesting the development of artificial ...
18:17

AI jobs are booming and paying more — but women are being left behind

AI jobs are booming and pay more than before, but women are being left out of the growth. LinkedIn found AI job postings have roughly doubled since 2023, and AI Engineer has overtaken Machine Learning Engineer as the most commonly posted AI role. The piece highlights a growing gender gap in a fast-growing, well-paid field.

Full text · 151 chars
AI job postings have roughly doubled since 2023, LinkedIn found. AI Engineer has overtaken Machine Learning Engineer as the most commonly posted AI ...
20:33

AI data giant Alation confirms cyberattack

Alation, the data search and AI company, confirmed a cyberattack after unauthorized access to its systems during an incident on Tuesday. The company said it is investigating the breach. Details on what data was exposed are not yet public.

Full text · 146 chars
The data search and AI giant confirmed unauthorized access to its systems during an incident on Tuesday, and said it was investigating the breach.
20:41

Expert witness used ChatGPT to defend 3M in suit over deadly explosion that killed 3 - NY Post

An expert witness defending 3M in a lawsuit over a deadly explosion that killed three people used ChatGPT to help write his report, which he charged $90,000 for. The prompts, revealed by 404 Media, show expert Josh Autenrieth of Knighthawk Engineering instructing the chatbot. It raises questions about AI's role in expert testimony and evidence.

Full text · 150 chars
The unveiled chatbot prompts , detailed by news outlet 404 Media, show Josh Autenrieth, an expert with Knighthawk Engineering , telling ChatGPT to ...
21:14

Watch Once, Learn Fast: GEN-1.5 Gives Robots One-Shot Learning | eWeek

A new model called GEN-1.5 gives robots one-shot learning, so they can learn a new task from watching a single demonstration instead of many examples. It's framed as a generalist robot model that should speed up how robots pick up unfamiliar skills. Coverage is promotional, with few technical specifics.

Full text · 147 chars
Prompt Engineering Guide: Unlocking the Potential of AI Models. Getting the desired outputs from AI models starts with carefully crafted inputs ...
21:21

'If you're not first, you're last': Mayo lawsuit highlights tensions in hospitals' race to use AI

Hospitals are rushing AI into patient diagnosis and treatment so fast that the pressure is now spilling into lawsuits, with a new case at Mayo Clinic in the spotlight. The suit highlights an "if you're not first, you're last" mentality driving AI adoption across US hospitals. Doctors and patients worry tools are being rolled out before they're proven safe. The case shows how legal and safety concerns are colliding with competitive pressure across the industry.

Full text · 148 chars
As hospitals across the country ramp up their integration of artificial intelligence into how they go about diagnosing and treating patients and ...
21:43

Delta debuts AI dual-arm cobot solution - Taipei Times

Delta Electronics showed off a dual-armed AI robot built for factory work, powered by Nvidia's platform. It's an embodied AI cobot, meaning the robot perceives and acts in the physical world. Pairing Delta's hardware with Nvidia's software suggests the robot arms race is now a platform battle. This is a debut announcement, so real-world performance and price are still unknown.

Full text · 147 chars
Delta Electronics Inc (台達電) debuted its embodied artificial intelligence (AI) dual-arm cobot solution integrated with Nvidia Corp's platform at ...
21:49

Meta AI's new Mac app wants you to talk to your apps

Meta launched a Mac app for its AI assistant that lets you query your other apps in plain language and get answers back. The assistant can pull up specifics like campaign performance and audience engagement from your tools. It's aimed at people who want insights from their own apps without digging through dashboards. Details on availability and what apps it supports are still thin.

Full text · 151 chars
... AI and get insights by querying the AI assistant. Meta AI can provide users with information about campaign performance and audience engagement ...
22:00

AI fired an S.F. store employee. Will California crack down on 'robobosses'?

An AI-run San Francisco convenience store fired a human employee, and California lawmakers may crack down on 'robobosses.' The incident at Andon Market is now the case study in a debate over whether AI should be allowed to make firing decisions. This could lead to new state rules limiting how much authority AI gets over workers. It's a concrete example of AI crossing from tools to managers.

Full text · 150 chars
Jules Castaneda looks over items Wednesday at Andon Market in San Francisco. The store is run by artificial intelligence , which recently fired an ...
22:06

Artificial Intelligence Threatens North Georgia Schools - Vanguard

North Georgia schools received an anonymous bomb threat that authorities believe was generated by AI. Six schools were notified on Aug. 13, and Jackson County plus surrounding districts were affected. AI-generated threats are harder to trace and could become a new pattern of harassment for institutions. The item is thin on whether the threat was credible or how it was detected.

Full text · 148 chars
On Aug. 13, six north Georgia schools were notified of an anonymous computer generated bomb threat. Jackson County and the surrounding districts ...
04:00

Backdoor Learning in Language Models and Vision-Language Models

A PhD thesis surveys backdoor attacks against text and vision-language models, where models are secretly poisoned to misbehave on hidden triggers. It covers analyzing, detecting, and building such attacks, alongside efficient multimodal representation methods for medical imaging. The writeup is thin here — basically the thesis abstract — so there is little concrete new detail to report.

Full text · 1,419 chars
Computer Science > Computation and Language Title:Backdoor Learning in Language Models and Vision-Language Models View PDF HTML (experimental) Abstract:Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
10:26

The Sequence Opinion- Issue 918: The Energy Scaling Laws of AI

AI's next bottleneck isn't model size but energy — every AI request ends up as heat in a power-hungry data center. The essay argues the next scaling law for AI is how efficiently civilization turns photons and atoms into intelligence. It walks through the physical infrastructure behind AI: substations, cooling loops, transmission networks, and power plants. Thin opinion piece with no concrete numbers.

Full text · 899 chars
AI does not run in the cloud. It runs in substations, cooling loops, transmission networks, and power plants. The next scaling law is not only about parameters, but about how efficiently civilization can convert photons and atoms into useful intelligence. This essay will help you to understand the different forms of energy influencing the next wave of AI scaling. Open a modern AI application and the experience feels almost weightless. A cursor blinks. A prompt disappears. Seconds later, a page of reasoning materializes. The interface says software. The physics says factory. Behind that answer, accelerators switch billions of transistors, memory systems move tensors, pumps circulate coolant, transformers reshape voltage, and generators turn motion, sunlight, or nuclear reactions into electrons. Nearly every joule entering the cluster eventually leaves as heat. The cloud has a power cord.
14:20

Agentic AI Market Worth $205.88 Billion by 2033 - Report by MarketsandMarkets

Analysts project the agentic AI market will reach $205.88 billion by 2033. MarketsandMarkets' forecast says IT and IT services will be the fastest-growing buyer segment from 2026 through 2033 as software engineering teams adopt agents. The figures come from a paid analyst report republished on Moomoo, so treat them as projections, not news.

Full text · 144 chars
IT & ITeS is expected to be the fastest-growing end-user segment during 2026–2033, supported by rapid adoption across software engineering , ...
15:11

Can NXP MCX A5 MCUs Secure the Industrial Edge Before Agentic Attackers Arrive?

NXP's new MCX A5 microcontrollers aim to secure the industrial edge before agentic AI attacks become common. The analyst question is whether these chips can protect factory and edge devices against automated attackers. It's a hardware-readiness piece that speculates on threats rather than reporting incidents.

Full text · 148 chars
Software Lifecycle Engineering . Insights. Analyst Insights · Custom Research Reports · Earnings & News · Futurum AI Insights · Media · Research ...
15:24

What Is a Software Factory? The Agentic Operating Model, Defined | Augment Code

Agentic coding tools are bottlenecking on review — agents open pull requests faster than developers can verify them. Augment Code defines the 'software factory' as an agentic operating model where teams manage agent output instead of writing all the code themselves. It's a vendor guide aimed at teams adopting enterprise coding tools.

Full text · 150 chars
Engineering teams adopting enterprise coding tools keep hitting the same wall. Agents open pull requests faster than developers can review, verify ...
16:02

NeuBird AI Publishes Open Framework for Earned Agent Autonomy in Production Environments

NeuBird AI released an open framework for giving AI agents autonomy in production environments. The framework is open for adoption, revision, and co-signature by operators, security leaders, and engineers who build or run agent systems. The announcement is light on implementation detail.

Full text · 147 chars
The framework is open for adoption, revision and co-signature by any operator, security leader or independent engineer who builds, buys or runs ...
16:31

Agentic AI Success Depends on Process Maturity: Tuxpas - Mexico Business News

A company argues AI agents only work well when the business process underneath them is already well understood. Tuxpas says firms should combine data intelligence with process engineering methods like Six Sigma and Business Process Management to map workflows first. Without that maturity, agentic AI projects struggle. It's a warning that tooling alone won't fix messy processes.

Full text · 150 chars
We combine this data intelligence with process engineering using methodologies like Six Sigma and Business Process Management (BPM) to map out how ...
17:20

Disco Announces Launch of Feature Enhancement to Its Generative Ai-Powered Auto ...

Legal tech company Disco upgraded its AI-powered auto-review product so teams can write and test their own AI review instructions without needing prompt-engineering expertise. The feature is meant to make prompt iteration faster and easier for legal document review. Routine product enhancement.

Full text · 157 chars
... prompt iteration faster and easier and helps teams refine their own AI review instructions without prompt - engineering expertise. Prompt engineering ...
17:30

Graph engineering is where AI agents stop working alone

Many AI agents working side by side still bump into each other unless someone designs how they connect, and 'graph engineering' is the emerging name for that planning work. It's about wiring many capable agents into one coherent system rather than letting them operate in isolation. The article from CIO only teases this idea, so the substance is thin beyond the concept.

Full text · 148 chars
Good agents still collide when nobody designs how they work together. Graph engineering is how you turn many capable agents into one system that ...
17:38

The Code Optimization Flywheel Won't Spin Itself - Communications of the ACM

AI coding agents won't find performance wins on their own — someone has to build the loop that feeds real bottlenecks back to them. A Communications of the ACM essay argues optimization work stays manual unless teams wire agent feedback into it. It points to the SWE-Bench Pro benchmark, which tests whether agents can handle long-horizon software engineering tasks. The post is thin on specifics and reads as a position piece.

Full text · 145 chars
... agent could only discover by reaching it. Real teams do write ... SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
17:55

I Only Do Anything Once

An essay argues the goal is to never do the same small task twice. The author says you should build reusable AI skills for things you repeatedly look up, like home-network details, so your AI reads them by default and the chore disappears. The hard part isn't building the skill but noticing which repeated annoyances you've already normalized enough to forget.

Notes

I Only Do Anything Once — Daniel Miessler

Personal philosophy: never repeat work, in the system and in life. Central claim: the expensive task isn't hard problems (memorable, only encountered once), but small retrievals already done a dozen times and done a dozen more — "each one costing four minutes and a thin slice of attention, none of which you ever bill to anyone."

His Ubiquiti example: repeatedly re-looking up device names (Firewall Pro, WiFi 7 AP) that he already knew six months ago. "That's the shape of the waste."

Method: build a skill for every life domain — home (Hue, cameras, studio lights), network (Ubiquiti). The lookup happens once; the only extra cost is writing it down in a place the AI reads by default, which deletes the category of work.

Key distinction: "Building the skill is the easy half. Noticing is the hard one." The hard part is surfacing repeated friction you've normalized.

"The friction goes invisible the same way a heavy pack goes invisible after an hour on the trail. You stop feeling the weight. You just get tired earlier than you should and can't say why."

Detection signal: the sentence "I really wish I could do that one day" is the tell — "your brain quietly pricing a task as too expensive to keep doing by hand." Write it down; that list is your build queue.

Closing argument: stop pointing AI at "going faster on the treadmill"; zoom out and ask what you want as a human. "Get off it instead."

Full text · 2,625 chars
I was walking someone through my setup this morning and caught myself saying the same thing over and over, so I might as well write it down. I only do anything once. The goal is to never have to do repeat work. And I'm talking about not just in the system, but in life in general. Here's the example I keep coming back to. How many times have you had to look up something related to Ubiquiti stuff? When you look up stuff, you're always like, yeah, go look up this, or go look up that, and you think about, okay, what was it called? It was Firewall Pro, I think. It was Ubiquiti WiFi 7 AP or whatever. You already knew that. You knew it six months ago. You looked it up then too. That's the shape of the waste. Hard problems are memorable, and you only get them once anyway. So they're fine. The expensive thing is the small retrieval you have already done a dozen times and will do a dozen more, each one costing four minutes and a thin slice of attention, none of which you ever bill to anyone. So now I have skills for everything in my life. I have a skill for the home in general. I have a skill for network, which is Ubiquiti stuff. The home one includes my Hue, cameras, and it includes my studio lights and stuff like that. Without additional context required. The lookup happened once. And the only thing I paid extra for was writing down what I found, in a place my AI reads by default. Now it's gone as a category of work. Building the skill is the easy half. Noticing is the hard one. What do I do day to day as a human, which I don't even realize is an annoyance and constant churn and noise and repeated effort? We have to find those things. We have to list them and say, how can I get my harness to automate those things for me? And that's genuinely difficult, because you've normalized all of it. The friction goes invisible the same way a heavy pack goes invisible after an hour on the trail. You stop feeling the weight. You just get tired earlier than you should and can't say why. So use a different signal. If you're ever thinking about it and you're thinking, hmm, I really wish I could do that one day, that's the tell. That sentence is your brain quietly pricing a task as too expensive to keep doing by hand, which is exactly the information you want. Write it down. That list is your build queue. I think we are imprisoned by this idea that we have to deal with the specifics and we have to put our focus on the tech itself. When in fact, I think we should be way zoomed out and thinking, what do I want as a human? So most people are pointing AI at going faster on the treadmill. Get off it instead.
18:12

ASU launches new scholarship pipeline to AI -powered cybersecurity careers

Arizona State University is starting a scholarship program to train students for AI-powered cybersecurity careers. The AI-Augmented Cybersecurity Scholars Program, run by the Fulton Schools of Engineering, gives students financial support and training to move into AI security jobs. It's aimed at filling the growing demand for people who can secure AI systems.

Full text · 148 chars
Fulton Schools of Engineering , will lead Arizona State University's AI -Augmented Cybersecurity Scholars Program. The program provides students ...
18:53

EXCLUSIVE: Omnicom Offloads Hundreds of Staffers Who Built Its AI Platform to Third-Party ...

Advertising giant Omnicom is outsourcing hundreds of the engineers and product staff who built the AI platform it calls central to its future. It's moving those staffers to a third-party IT contractor rather than keeping them in-house. The move raises questions about its long-term AI strategy and how it plans to run the platform it bet its future on.

Full text · 148 chars
The holdco is outsourcing hundreds of the engineers and product staff who built the AI platform it has positioned as central to its future to IT ...
18:53

Autonomous AI agents will drive the next productivity surge - IT-Online

A tech executive argues that companies are too fixated on prompt engineering and should shift focus to autonomous AI agents, which he says will drive the next big productivity jump. Cliff de Wit, chief innovation officer at Accelera Digital Group, made the point. It's an opinion piece rather than a product or research announcement.

Full text · 148 chars
Cliff de Wit, chief innovation officer at Accelera Digital Group (ADG), says the industry's fixation on prompt engineering has become a distraction.
20:05

Introducing AI Futures | OpenAI

OpenAI is launching a new blog called AI Futures devoted to how transformative AI could reshape power, governance, the economy, and individual freedom. The announcement gives no articles or dates. It reads as a thought-leadership platform rather than a product or research announcement.

Full text · 143 chars
Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.
20:36

Welcome to the AI crisis in math | The Verge

New AI discoveries have left the math world shell-shocked, according to a Verge podcast feature. The episode focuses on how AI results are upending mathematical research. The item is a podcast teaser with thin details — it references OpenAI's Astra. Expected impact and specifics are not covered.

Full text · 91 chars
The Verge's Robert Hart on why new AI discoveries have left the math world 'shell-shocked.'
20:48

Serval Launches Catalyst, an AI Agent That Automates Automation Building

A startup called Serval is launching Catalyst, an AI agent designed to automate the building of automations. The item is a YouTube video link with thin surrounding content, so this is based mainly on the title. It appears to target AI engineers interested in agentic tooling for automating workflows.

Full text · 154 chars
... 3Blue1Brown•3.8M views · 20:27 · Go to channel AI Engineer · Harnesses in AI: A Deep Dive — Tejas Kumar, IBM. AI Engineer •243K views · 50:41 · Go ...
21:06

Artificial intelligence powered traffic lights are being trialled, but will it ease traffic congestion?

Australian cities are trialing AI-powered traffic lights that adjust to live traffic, which could speed up commutes. The trial is early and it's still unclear whether the system will actually ease congestion. If it works, it would give AI a foothold in everyday infrastructure, not just software. Details on the trial's size, location, and results aren't in this item.

Full text · 134 chars
Artificial intelligence (AI) already has a foothold in many industries across Australia, but it could soon help speed up your commute.
21:45

Facing Deficit, California City Eyes AI-Powered Budgeting

A California city facing a budget deficit is considering AI to help with budgeting. A council member is pressing officials with pointed questions about the program, raising concern about how the AI would be used. The item is thin on specifics, like which city, the AI vendor, or what exactly it would do. The main news is local governments testing AI on financial decisions.

Full text · 154 chars
Councilmember Jose Rodriguez asked Fabian a series of pointed questions about the program's use of artificial intelligence , adding his concerns about ...
21:57

4 Steps to Transform the “Middle Office” with AI - Harvard Business Review

Companies that spend big on AI still struggle to turn it into measurable business results, and this Harvard Business Review piece lays out a four-step plan for fixing the "middle office." The advice targets the back-office work between the front line and the back end of the business. It's a strategy how-to rather than news, so there's little new substance beyond the framework itself.

Full text · 126 chars
Companies are spending heavily on AI , but many are still struggling to turn that investment into measurable business results.
22:01

How to Outsmart AI When It's Tracking Your Workday - WSJ

More managers are using AI to track how productive their employees are, and the Wall Street Journal explains the tactics workers use to game those monitors and look good. The piece covers what the trackers watch and how employees raise their scores. It's a practical advice article, not a report on any new findings or data.

Full text · 146 chars
Looking like a good employee in the eyes of AI productivity trackers that more managers are using to evaluate their teams. Employee-monitoring ...
09:47

Unlocking hidden revenue streams with market models

Airlines are starting to use generative AI models that set ticket prices in real time by simulating market conditions. Virgin Atlantic says it runs Fetcherr's market model in some markets to weigh demand, capacity, bookings, and competitors dynamically instead of following static rules. This piece is sponsored marketing content produced by MIT Technology Review Insights, so treat the claims as promotion.

Notes

Source: Sponsored feature (advertorial) by Fetcherr, published via MIT Technology Review's "Insights" custom content arm. Disclosure states it was not written by MIT TR editorial staff, and AI tools were "limited to secondary production processes that passed thorough human review."

Core claim: Generative-AI "market models" can handle airline pricing/revenue management in real time. They're described as "deep learning models trained on high-resolution numerical data" that analyze, simulate, and predict financial dynamics, acting as an AI "brain" that simulates market environments and makes "dynamic commercial decisions, such as pricing, inventory, or revenue management." Positioning is explicitly against "historical trends or static rules."

Variables the pricing must weigh: demand, season, time of day, current events, global markets, competitor airline activity; per the cited customer, also capacity and booking volume.

Named customer: Virgin Atlantic. Dominic Kennedy, SVP of revenue management, sales, and e-commerce, uses the model "to drive their generative pricing engines in some markets." Quotes:

"It helps us make better, faster, more granular commercial decisions."
"It considers, on a real-time basis, a plethora of different inputs, whether it be demand, capacity, or booking. It has a really sophisticated way of evaluating our positioning relative to competitors, market conditions."

Caveats:

  • Entire evidence base is a single vendor testimonial; no benchmarks, prices, revenue lift figures, or measured outcomes are given.
  • "In some markets" is the only hedge on the deployment claim — scope unspecified.
  • No independent validation, comparative data, or competitor context.
  • Production attribution: human writers/analysts authored it; AI limited to secondary production steps.
Full text · 2,596 chars
Sponsored In partnership withFetcherr Each day, an airline transports tens of thousands of passengers on hundreds of flights. Often these are not straightforward point-to-point routes, with passengers requiring multiple connections. The airline can consider potentially hundreds of variables to price each of these journeys: demand, season, time of day, current events, global markets, and competitor airline activity to name just a few. It is a nuanced process that must constantly adapt to the goings on in the wider world. Generative AI-powered market models are emerging as a means of handling complex tasks like this in real time. These deep learning models are trained on high-resolution numerical data and designed to analyze, simulate, and predict complex financial dynamics. Rather than relying on historical trends or static rules, the market model acts as an AI “brain,” consolidating a variety of data to simulate different market environments and make dynamic commercial decisions, such as pricing, inventory, or revenue management. “It helps us make better, faster, more granular commercial decisions,” says Dominic Kennedy, senior vice president of revenue management, sales, and e-commerce at Virgin Atlantic about the market model his team is using to drive their generative pricing engines in some markets. “It considers, on a real-time basis, a plethora of different inputs, whether it be demand, capacity, or booking. It has a really sophisticated way of evaluating our positioning relative to competitors, market conditions, and a whole raft of other things that have significance in how demand is manifested,” he adds. This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
14:57

How to Build a Career in AI: 3 Distinct Pathways

A career advice piece lays out three distinct paths into AI work, with the prompt engineer role defined as designing and refining inputs for large language models. It also points to specification engineering as the skill that's succeeding plain prompt engineering. Standard career-listicle content.

Full text · 143 chars
Prompt engineer : designs and refines the inputs given to LLMs to ... Specification Engineering: The New Skill After Prompt Engineering · 5 ...
16:55

😻 Livestream: AI Tool Roundup for Normal People

A livestream is coming up to translate this week's big AI launches into plain English. The hosts will explain Qwen 3.8 open models, Unsloth Studio for running models on your own computer, Cursor Origin, DeepSeek Harness, hosting AI locally, and tools for organizing agents and skills. This is a promo for the show rather than a news item itself.

Notes

Livestream: AI Tool Roundup for Normal People — The Neuron

Live at 10 AM PT / 1 PM ET, hosts Corey and Grant translate the week's AI launches into plain English. No benchmark comparisons, no assumed jargon.

Covered live:

  • Qwen 3.8 — what an open model is, why you might run one instead of ChatGPT/Claude, when that makes sense.
  • Unsloth Studio — running/experimenting with models locally on your own computer, framed as an alternative to "Terminal Fight Club."
  • Hosting AI locally — what running AI on your own hardware actually takes, when it's worth it, what changed since last covered.
  • Cursor Origin — why Cursor now wants to host your code too; implications for people building apps, websites, internal tools with AI.
  • DeepSeek Harness — what an "agent harness" actually is, why developers are suddenly obsessed, whether non-developers should care.
  • Buzz + Berd — new ways to organize/reuse agents, skills, tools across AI apps without rebuilding setups.
  • Plus other launches between now and go-live.

Goal: distinguish useful vs hype vs worth-trying. No question too basic; stream available afterward on YouTube.

ICYMI episodes:

  • Can AI predict what happens next? Neuralk CEO Alexandre Pasquiou: LMs talk about business data well but struggle to predict from it; specialized tabular models may become the predictive brain behind AI agents. Caveat: relevant for spreadsheets, forecasts, finance, ops.
  • AI can write DNA now. Radical Numerics CEO Eric Nguyen: genomic AI reads/writes DNA, including complete viral genomes; researchers pushing toward combined DNA+RNA+protein models. Upside: medical; question: security.

Related past streams: Agents for Total Beginners, AI 5-Level Starter Stack, App Building 101, local hosting with Microsoft AI Frontiers.

Full text · 3,589 chars
😻 Livestream: AI Tool Roundup for Normal People We go live at 10 AM PT to make sense of Qwen 3.8, Unsloth Studio, Cursor Origin, DeepSeek Harness, and more. Welcome, humans. AI had another one of those weeks where a dozen new models and tools dropped, half the internet yelled that everything changed, and normal people were left asking one useful question: Which of these things should I actually care about? So at 10 AM PT / 1 PM ET, Corey and Grant are going live to translate the week’s biggest AI launches into plain English. No benchmark Olympics. No assuming you know what an “agent harness” is. Just what changed, who each tool is for, and when you might realistically use it. Here’s what we’re covering live: - Qwen 3.8: what an open model is, why you might run one instead of ChatGPT or Claude, and when that makes sense. - Unsloth Studio: how to run and experiment with AI models on your own computer without turning your afternoon into Terminal Fight Club. - Hosting AI locally: what it actually takes to run AI on your own hardware, when it is worth doing, and what has changed since we last covered it. - Cursor Origin: why Cursor suddenly wants to host your code too, and what that means if you build apps, websites, or internal tools with AI. - DeepSeek Harness: what an “agent harness” actually is, why developers are suddenly obsessed with them, and whether non-developers should care. - Buzz + Berd: new ways to organize and reuse agents, skills, and tools across AI apps without rebuilding your setup every time. - Plus the other notable launches from this week, including whatever drops between now and the moment we hit “Go Live.” The goal is simple: by the end, you should know what’s useful, what’s hype, and which tools are actually worth trying for your own work. Bring your questions, too. No question is too basic. If yours is too smart for us, we reserve the right to make the smartest AI in the room answer it. Delegation! Open the stream now, say hello in chat, and tell us which new AI tool you want explained without the developer-speak. P.S: If you join late, the whole stream will still be waiting for you on YouTube. 🎙️ In Case You Missed It… 1. Can AI predict what happens next? TL;DW: Neuralk CEO Alexandre Pasquiou explains why language models are great at talking about business data but can struggle to predict from it, and why specialized tabular models may become the predictive brain behind AI agents. Why you should watch: If you use AI with spreadsheets, forecasts, customer data, finance, or operations, this episode explains where ChatGPT-style models can hit a wall and what may work better. 2. AI can write DNA now. Here’s what that means. TL;DW: Eric Nguyen, CEO of Radical Numerics, explains how genomic AI can read and write DNA, including complete viral genomes, while researchers push toward models that combine DNA, RNA, proteins, and other biological signals. Why you should watch: It makes the leap from “AI analyzes biology” to “AI designs biology” concrete, including the huge medical upside and the security questions that come with it. What should we learn next? 🤔 Let us know below! Psst: Did you pick one of these answers? We’ve already done a stream on a few of the biggest winners: - Making agents with AI: watch our Agents for Total Beginners class. - Getting more out of ChatGPT: watch our AI for Total Beginners 5-Level Starter Stack. - Building apps with AI: watch our App Building 101 stream. - Hosting AI locally: watch us run open models on our own hardware with Microsoft AI Frontiers. Stay curious, The Neuron Team
17:32

Part-time Faculty Workforce Development Continuing Education (WDCE) AI Agent /Prompt ...

A Maryland community college is hiring a part-time faculty member to teach AI agent and prompt engineering classes for working adults. The listing, posted on the Chronicle of Higher Education jobs board, is for a workforce-development continuing education role in Rockville. There are no details on pay or start date in the posting itself.

Full text · 147 chars
Part-time Faculty Workforce Development Continuing Education (WDCE) AI Agent /Prompt Engineering job in Rockville, Maryland, United States with ...
18:25

EXPLOR-NEPA Helps High School Students Build & Code their Future - News@Wilkes

A Wilkes University program called EXPLOR-NEPA teaches high school students coding, AI, and robotics to build career skills. Eligible participants can get a $3,000 scholarship toward the program. It's an effort to get younger students into tech and engineering.

Full text · 152 chars
... engineering , AI and robotics. Wilkes even offers eligible participants in the EXPLOR-NEPA program a $3,000 scholarship opportunity toward their ...
18:50

Lead AI Prompt Engineer & Architect (EN-ENj162958) - Staff.am

A job listing for a Lead AI Prompt Engineer & Architect in Yerevan is posted on Staff.am. The posting is marked expired, so it's just an archived job ad rather than an active opportunity.

Full text · 138 chars
Lead AI Prompt Engineer & Architect. Partner Company. verified-png. location Yerevan. Expired. Views. 1282. Job History. 461. Active Jobs.
20:09

Brennan says I'm rusty.. but he has bigger problems! Seedance 2.5 in Luma → Focus on ...

A short social post shares a prompt engineer's character prompt used to generate an image of a person in an exosuit, tied to a casual riff about being rusty. The content is essentially a prompt example with no real news value.

Full text · 153 chars
... prompt engineer . Prompt below: @Image1 is Dave (@RUST): 30s man, curly brown hair, brown eyes, light stubble, bulky industrial exosuit — scuffed ...
20:31

AI/ML Engineer - Xylo Technologies, Inc. - Remote | Dice.com

A job listing for a remote AI/ML engineer whose duties include document ingestion workflows, semantic search, embeddings, prompt engineering, and evaluation. It's a hiring post, useful only as signal that these skills are in demand.

Full text · 145 chars
Responsibilities include developing document ingestion and processing workflows, semantic search, embeddings, prompt engineering , evaluation ...
20:34

Falcon 2 Beats ElevenLabs: Inside MURF AI's Ultra-Low Latency Voice Agent

A YouTube video claims MURF AI's voice agent, Falcon 2, beats ElevenLabs on ultra-low latency. The video is from the AI Engineer channel and shows off the product's speed advantage. It's essentially a promotional demo rather than independent testing.

Full text · 117 chars
Go to channel AI Engineer . Harnesses in AI: A Deep Dive — Tejas Kumar, IBM. AI Engineer •243K views · 12:41 · Go ...
20:36

Back-to-school reality check: Artificial intelligence , cell phones, bullying and the challenges kids face

A TV segment examines the challenges kids face going back to school, covering AI, cell phones, and bullying. A school superintendent joins the discussion to break down these issues. The item is thin on specifics beyond the topic framing.

Full text · 146 chars
Fox 2's Ronia Shamona Shamona joins Randy Speck, Superintendent of Schools at Education Management and Networks who breaks down everything all ...
20:40

Billionaire Stanley Druckenmiller Sells Micron and Is Piling Into This Other Unstoppable ...

Billionaire investor Stanley Druckenmiller has sold out of Micron and is putting money into another stock. Micron has surged more than 200% this year on the huge demand for AI memory chips, which appears to be why he's taking profits. The piece is a stock commentary that names his new pick, but this summary focuses on the AI-memory angle.

Full text · 155 chars
Micron has benefited from unprecedented demand for artificial intelligence (AI) memory solutions. · After Micron stock rallied more than 200% this year ...
20:50

Oura Welcomes Sanjay Chandra as Chief Information Officer and Han Chiu as SVP of AI ...

Smart-ring maker Oura is adding two senior engineering leaders to its team. Sanjay Chandra becomes Chief Information Officer, and Han Chiu joins as SVP of AI, with Chiu overseeing engineering across mobile, backend, platform, and AI while owning software and AI strategy. The company is deepening its AI engineering leadership as it builds out its health-wearable products.

Full text · 155 chars
... engineering organization across mobile, backend, platform, and AI engineering , while owning the company's software and AI strategy. Han joins Oura ...
21:22

Prompt Engineering : Learn to Speak AI's Language So That It Can Speak Yours - JHU Hub

Johns Hopkins is running a beginner prompt-engineering workshop on August 27 for people who want to use generative AI tools more effectively. The announcement gives no detail beyond the date and that it teaches how to phrase requests to get better answers. Thin content, so this is just an event listing.

Full text · 133 chars
This workshop delves into prompt engineering , an essential skill for effectively utilizing generative artificial intelligence tools.
21:42

Do You Really Know Claude Plugins? See How They Can Turn You

Claude plugins can turn prompts into repeatable, autonomous workflows — that's the shift from prompt engineering to plugin engineering. The piece argues plugins let you build reusable automated processes instead of one-off answers. It's a promotional guide on KuCoin, a crypto exchange, and the content is thin.

Full text · 142 chars
Prompt engineering taught us how to ask AI for better research. Plugin engineering may teach us how to build a repeatable (and autonomous) ...

Newsletter

11
00:00

Visions of AI: Automating repetitive grunt Coding tasks

Cognition, maker of the AI coding agent Devin, is reportedly raising at a $40 billion valuation after hitting a $1 billion annualized revenue run rate in under two years. The piece frames Devin alongside Cursor and Manus as coding tools that changed software engineering, and notes the industry is consolidating, with Cursor now under SpaceXAI, Meta's Manus deal blocked, and Stripe likely buying OpenRouter. It also recalls that Cognition acquired Windsurf in July 2025 after an OpenAI deal was vetoed by Microsoft over Github Copilot competition.

Notes
  • AI Supremacy (substack), "Visions of AI: Automating repetitive grunt Coding tasks", published 2026-08-20. New "Visions of AI" feature format — short profiles of AI/emerging-tech startups, positioned as light evening reading (8 pm EST), cadence unspecified.
  • Devin / Cognition: in talks to raise at a $40B valuation, reporting $1B annualized revenue run rate in under two years (per Bloomberg-sourced figures). Author admits initial skepticism of Devin/Cognition, changed by that single stat.
  • Context: Devin grouped with Cursor, Manus AI as agentic coding tools. Author ties Anthropic's coding-model success (ARR "nearing $65 bn") to API/coding utility — cited as factor driving the category.
  • 2026 later-half snapshot: Cursor is part of SpaceXAI; Meta's acquisition of Manus AI was blocked; ~two years since Devin's release.
  • Consolidation thesis: "the consolidation period appears to be in full-swing" — cites Stripe likely acquiring OpenRouter. Author was "incredibly bullish on Cursor before SpaceX swooped in."
  • Windsurf history: July 2025, Cognition acquired Windsurf days after it lost its CEO/employees to Google. Author claims OpenAI had a tentative $3 bn deal to acquire Windsurf but it "got vetoed apparently by Microsoft" (competes with GitHub Copilot).
  • Caveat/perspective: author notes mid-2020s AI-startup fortunes "seem to be tied to the pedigree of their investors and VC funds," questioning how much revenue Manus AI would have without its early Meta deal.
Full text · 2,365 chars
Good Morning, Visions of AI is a new feature format I’m experimenting with that will amount to a short profile on an AI or emerging tech startup. The cadence of this style of article is unknown as of yet, but there are a lot of fascinating startups I want to share about. This is designed to be light evening reading to go out at a time-slot of 8 pm EST. Generative AI has been a force of nature when it comes to the future of software engineering and building products. Cognition, maker of the AI coding agent Devin are in talks to raise at a valuation of $40 Billion. Devin belongs in the same category as the likes of Cursor, Manus AI and other coding tools that have changed the future of software engineering, agentic AI augmented productivity and automating repetitive coding tasks to the extent that Anthropic models have gotten progressively better at coding. If Anthropic’s ARR is nearing $65 bn that’s due to the utility of its coding models and the success of its API. Fast forward to the later half of 2026, Cursor is part of SpaceXAI, Meta’s acquisition of Manus AI has been blocked, and it’s now almost two years since Devin was first released. I’ll admit I was fairly skeptical of Devin and Cognition AI at the onset. A single stat line changed my mind. Cognition is now reporting at least a $40 billion valuation based on achieving a $1 billion annualized revenue run rate, according to sources cited by Bloomberg. While we are living in an era of higher valuations for private AI related companies, not many AI startups have achieved $1 Billion ARR in under two years. I was incredibly bullish on Cursor before SpaceX swooped in and now with Stripe likely acquiring OpenRouter, the consolidation period appears to be in full-swing. How much revenue would Manus AI even be making if it hadn’t taken the early big deal with Meta? In the mid 2020s so much of an AI startup’s fortunes seem to be tied to the pedigree of their investors and Venture Capital funds. You might remember in July, 2025 Cognition acquired AI coding startup Windsurf, days after that company lost its CEO and other employees to Google. OpenAI’s incompetence has ironically benefitted Cognition since it was OpenAI that has reached a tentative deal to acquire Windsurf in a $3 bn. deal but the deal got vetoed apparently by Microsoft, since it would compete with Github Copilot.
05:17

[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law

GLM 5.3 improved dramatically with about a month of extra reinforcement learning — not more parameters — a sign the field's fixation on parameter count is outdated. The z.AI CEO argues parameter count only matters alongside data, compute, and deployment conditions, and that reasoning gains now come from RL on long-horizon environments simulating days of real engineering work, with synthetic environments and verifiers built end to end. The same roundup also covers Ornith-1.5, a new MIT-licensed open model family claiming self-improvement, Gemini 3.7 Flash topping a benchmark, DeepSeek's plugin-based agent harness, and TrueFoundry's open-source TrueForge harness that claims ~75% cost cuts.

Notes
GLM-5.3 and the "death of params" thesis

Prof. Jie Tang (Z.ai/GLM founder) argues parameter count alone is no longer a meaningful model-size metric:

"Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions."

Key claims, per the source:

  • Chinchilla's fixed 200 tokens/param (or 20x) assumption is wrong in the "Inference Inflection" world — the true ratio varies 200–900 toks/param, task-dependent (citing Roberts et al.).
  • Memorization prefers more parameters; reasoning prefers more post-training data and effective depth.
  • GLM-5.3's gains come solely from RL on long-horizon environments — same core base model/architecture as GLM-5.2, improved via roughly one month of extra RL. (Reported open-weight positions: #2 Terminal Bench, #3 Legal Bench, #6 Skills Bench.)
The long-horizon RL environments
  • Environments now mirror real production workflows; some tasks represent several days of work for an experienced engineer.
  • Example: an ML-infrastructure task gives the model the same working environment as an engineer — compute clusters, storage, internal docs, codebases, experiment results — and requires it to diagnose bottlenecks, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness.
  • Goal: push the model to "take ownership of substantial work end to end" rather than have users decompose and supervise each step.
Fully synthetic environment + reward pipeline

As agent capability improves, the scaling bottleneck "moves from the model to the environment." Requirements for a useful task environment: executable, verifiable, close to real professional work — and many of them, not hand-built. Their pipeline:

  • Research agents collect task patterns from real work → turn into runnable long-horizon environments with multi-step dependencies and hidden state.
  • A judge agent attempts each task to verify solvability.
  • Verifiers are synthesized without access to the reference solution.
  • Solver trajectories are used to discover and close reward shortcuts.
  • A verifier passing oracle, no-op, and unsolved-state checks yields a binary reward "reliable enough to train on directly."
Five knobs of scaling

Tang names 5 scaling knobs including MoE sparsity with new "XA-YB" notation. Advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems — they require carrying 20+ inference-step causal chains without losing the thread. This ability "does not live in total parameter count once a certain knowledge-holding threshold is reached."

Context check: the author reaffirms Tang's prior prediction of an open-weights Fable-class model by year-end — 134 days left, with two 2–3T models (Qwen 3.8 Max, Kimi K3) vs. estimated Fable size 3–7T at only 2 points higher on the AA index.

AI Twitter recap (8/18–8/19/2026)
  • Ornith-1.5 (MIT): 9B dense, 35B MoE, 397B MoE; FP8, GGUF, MLX, NVFP4 quants. End-to-end self-improvement: proposes tasks, generates scaffolds, produces RL rollouts. Evals: Terminal-Bench 2.1 86.1, SWE-Bench Verified 86, DeepSWE 56, HLE 44.6, Tool Decathlon 71.2. Wired into vLLM and Ollama.
  • Unsloth Qwen3.8-27B Dynamic V3 GGUFs: ~10% higher accuracy at same size vs. other providers; 1-bit quants keep ~77% of BF16 accuracy on 8GB RAM. No QAT/QAD — post-training quantization only; imatrix calibration public. New Divergence-300 metric extends top-1% greedy accuracy across long generations (Terminal Bench, DeepSWE, etc.). Reddit commenters asked for UD 2.0 comparison lines, per-category KLD and KV-cache quant KLD (localbench-style). Memory: ~15GB for Q4_K_M; IQ4_XS may fit 16GB VRAM.
  • Agent Arena Pareto: Claude Opus 5 (High) leads quality; Kimi K3, GLM 5.2, Grok 4.5, GPT-5.6 Luna define the value frontier.
  • Grok 4.6 at #3/49 Legal Research Bench (48.1%), 500k context, tool/image/file support, low pricing.
  • DeepSeek Harness (DSH): intentionally thin shell over plugin architecture "Cordis" — everything is a plugin, including the agent loop. Beta users: 100+ plugins, 400+ issues in under a week (gomoku testbed, DB agent closing the SQL feedback loop).
  • TrueFoundry TrueForge (MIT, self-hostable, vendor-neutral): tool orchestration, context mgmt, subagents, code sandboxes, human approvals, traces. Claim: on a 14-task enterprise benchmark, matched Claude Managed Agents on Opus 4.8 at ~30% fewer tokens; routing to GLM-5.2 cut cost ~75% at equal accuracy.
  • ClaudeDevs: memory for self-hosted sandboxes, domain allow/block for web tools, multi-agent session viewer (minimap, grouped transcript, cost-per-session). Concise output style added to Claude Code.
  • Agent Lightning v1.0 (Microsoft): connects arbitrary harnesses to RL via endpoint proxy (retokenization, sample merging, advantage calc, normalization, scheduler/backend). ~6K examples: Qwen3.5-9B on SWE-Bench Verified 41.8% → 56.4%.
  • CPT/mid-training (C. Wolfer): knobs = data mixture, duration, stage ordering, sequence length, post-trainability — interacting, not independent.
  • Qdrant filterable HNSW vs ACORN: 1% filter over 1M vectors — 99.8% recall @ 1.0ms vs ACORN 67.7% @ 4.7ms; ACORN still helps for broad values/AND filters.
  • Sentence Transformers v6.0: multi-vector (token-level) vs single-vector dense retrieval; late-interaction increasingly default for quality-sensitive search.
  • Agent latency study (10 agentic apps): non-LLM components dominate latency in half; sandbox memory peaks at 28GB/session; up to 32x latency variation; task-aware serving cuts latency 29–40%, state offloading −4.6x memory, tool-result caching removes 35.2% redundant search calls.
  • Linear moved delta-sync read path from Postgres to turbopuffer (attribute indexes for permission filters; largest syncs ~8s faster).
  • Gemini 3.7 Flash: #1 AA-AnalystAgent — 60.0% pass^5, 70.5% pass@1, 77.5% pass@5, 1.32s/task, $0.54 avg cost across 80 spreadsheet/document tasks.
  • Replit Free Mode powered by GPT-5.6 Luna. OpenAI Private Safety Processing: Zero Data Retention preserved while detecting cross-interaction safety risks without human access to content.
  • OpenRouter joining Stripe (per Patrick Collison).
Reddit (/r/LocalLlama) notes

Unsloth Dynamic v3 thread (1428 activity) drew two concrete asks: graph comparison vs. prior UD 2.0 quants, and per-category + KV-cache KLD reporting à la localbench.

Full text · 15,674 chars
We’ve covered GLM 5.2 very excitedly before, and Prof Jie Tang’s belief that there will be an open weights Fable-class model by end of the year (spot check - with 134 days left, there are now two 2-3T models (Qwen 3.8 Max and Kimi K3) with estimates that Fable is 3-7T, and only 2 points higher on the AA index.) Prof Jie Tang is back on X to tell us that our shorthand for model sizes is no longer enough: “Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.” We have covered Chinchilla (and post-Chinchilla) scaling laws in past LS years, but, so we will skip the history lesson, but it is good to level-set on why Chinchilla’s assumptions were wrong in the Inference Inflection world (no fixed number, between 200-900 toks/param, citing Roberts et al on task dependence). In short: Memorization prefers more parameters. Reasoning prefers more post-training data and effective depth. GLM-5.3’s big jumps come solely from RL on long horizon environments: The environments now cover a much broader range of production workflows, with tasks designed around how engineering and research work is actually carried out in practice. Some represent several days of work for an experienced engineer. In an ML infrastructure task, for example, the model may be given the same working environment as an engineer, with access to compute clusters, storage systems, internal documentation, codebases, and experiment results. It must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness. Training on environments at this level pushes the model toward taking ownership of substantial work end to end, rather than relying on users to decompose the problem and supervise each step. For those following the recursive self improvement story, their entire environment and judging and verifier process is synthetic all the way down: As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. A useful task environment has to be executable, verifiable, and close to real professional work — and we need many of them, not a handful of hand-built ones. To scale this process, we built pipelines that synthesize environments end to end, and for a subset of tasks, the RL reward signal as well. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to verify that it is actually solvable. Verifiers are synthesized without access to the reference solution, while solver trajectories are used to discover and close reward shortcuts. A verifier that passes oracle, no-op, and unsolved-state checks produces a binary reward reliable enough to train on directly. To put an end to parameter count obsesssion, Prof Jie identifies 5 knobs of scaling, including MoE sparsity with the new XA-YB notation. He notes that advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems. They require carrying long causal chains (20+ inference steps) without losing the thread. This ability does not live in total parameter count once a certain knowledge-holding threshold is reached. And it looks like there is much more to go. AI News for 8/18/2026-8/19/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap Open-Weight Models, Compression, and Benchmark Movement - Ornith-1.5 lands as a serious new open family: @ornith_ released Ornith-1.5 in 9B dense, 35B MoE, and 397B MoE variants under MIT, with quantized formats including FP8, GGUF, MLX, and NVFP4. The headline claim is end-to-end self-improvement: the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences. Reported evals are strong across agentic/coding workloads, including Terminal-Bench 2.1: 86.1, SWE-Bench Verified: 86, DeepSWE: 56, HLE: 44.6, and Tool Decathlon: 71.2. The release was quickly wired into serving stacks by vLLM and Ollama. - Compression continues to get more aggressive without fully collapsing utility: @UnslothAI and @danielhanchen shipped new Qwen3.8-27B GGUFs using Dynamic V3, claiming roughly 10% higher accuracy at the same size and releasing 1-bit quants that still retain about 77% of BF16 accuracy while running on 8GB RAM. Their new Divergence-300 metric extends top-1% greedy accuracy across longer generations using unseen examples from Terminal Bench, DeepSWE, and related tasks. - Agent and legal eval boards continue to reshuffle: @arena published a Pareto view of Agent Arena, where Claude Opus 5 (High) leads quality, but lower-cost models like Kimi K3, GLM 5.2, Grok 4.5, and GPT-5.6 Luna define much of the value frontier. Separately, @ValsAI reported Grok 4.6 at #3/49 on Legal Research Bench with 48.1%, 500k context, tool/image/file support, and relatively low pricing. For open models, @ValsAI also highlighted GLM 5.3 as #2 on Terminal Bench, #3 on Legal Bench, and #6 on Skills Bench among open weights. Agent Harnesses Become the New Competitive Layer - DeepSeek Harness’s minimalism is deliberate, not incomplete: A detailed writeup amplified by @ZhihuFrontier and summarized by @TheTuringPost frames DeepSeek Harness (DSH) as an intentionally thin shell over a plugin architecture called Cordis. The key design choice is that everything is a plugin, including the agent loop itself. Early beta users reportedly shipped 100+ plugins and filed 400+ issues in under a week; examples range from a gomoku model testbed to a database agent that closes the SQL feedback loop by connecting the model to live query execution. The strongest takeaway is architectural: DSH is less “productized assistant” than open agent runtime, optimized for user-extensible tooling, swappable control loops, and business-rule injection. - TrueFoundry open-sources TrueForge and makes the harness-cost argument explicit: @truefoundry, @omarsar0, and @kimmonismus all covered the launch of TrueForge, an MIT-licensed, self-hostable, vendor-neutral harness for production agents. The stack includes tool orchestration, context management, subagents, code sandboxes, human approvals, and traces, with both local and hosted deployment modes. The technical claim that resonated: on a 14-task enterprise benchmark, TrueForge matched Claude Managed Agents on Opus 4.8 while using about 30% fewer tokens, and routing to GLM-5.2 cut cost by around 75% while preserving accuracy. The broader industry theme—also echoed by @bradenjhancock and @dbreunig via @rseroter—is that the session/environment/memory/tools layer is becoming a major source of both differentiation and savings. - Managed harnesses are also getting sharper observability and controls: @ClaudeDevs added memory support for self-hosted sandboxes, domain allow/block controls for web tools, and a redesigned multi-agent session viewer with minimap, grouped transcript, and cost-per-thread/session. OpenAI, meanwhile, continues pushing the opposite angle: give teams the harness primitives to embed into their own products. @OpenAIDevs highlighted the open-source Codex harness as the runtime beneath internal tools, ops dashboards, and custom apps, while @cursor_ai shipped cloud-agent UX improvements around persistent goals and long-lived sessions. Post-Training, Mid-Training, and RL Systems Work - More evidence that scaling is shifting from parameters toward training recipe quality: @kimmonismus surfaced a notable claim from the zAI/GLM founder: progress is still scaling, but too much discourse has fixated on parameter count rather than data quality, inference compute, and post-training. The cited example is GLM-5.3, reportedly based on the same core base model/architecture as GLM-5.2, but improved substantially via about one month of extra RL. - Microsoft’s Agent Lightning points at RL-through-the-harness as a practical recipe: @omarsar0 highlighted Agent Lightning v1.0, which connects arbitrary harnesses to RL through an endpoint proxy, handling issues like retokenization, sample merging, advantage calculation, normalization, and scheduler/backend coordination. With ~6K training examples and modest compute, it reportedly moves Qwen3.5-9B on SWE-Bench Verified from 41.8% to 56.4%. - Mid-training is being treated more explicitly as an optimization surface: @cwolferesearch laid out the current practitioner view of CPT/midtraining: optimize data mixture, duration, stage ordering, sequence length, and even post-trainability rather than just “continue pretraining on better data.” The thread is useful precisely because it frames these as interacting knobs rather than independent tricks. - RL infrastructure keeps improving underneath the research: @SergioPaniego resurfaced work showing on-policy distillation in TRL becoming 40x faster via generation buffers, batched teacher calls, and binary logprob encoding; @mikasenghaas announced adaptive concurrency in prl, dynamically adjusting in-flight rollouts over the course of an RL run. Benchmarks, Retrieval, and Infra Details That Matter in Production - Qdrant’s filterable HNSW vs ACORN is a substantive retrieval systems update: @qdrant_engine argued that filtered ANN should be addressed in the index, not only at query time. Their filterable HNSW adds edges between points sharing indexed payload values, keeping filtered subgraphs connected. In their benchmark on a 1% filter over 1M vectors, they report 99.8% recall at 1.0ms versus 67.7% at 4.7ms for ACORN. They also note ACORN still helps for broad values and AND filters, especially atop a graph already optimized for filters. - Sentence Transformers v6.0 reflects the practical move from single-vector to multi-vector retrieval: @tomaarsen summarized the distinction clearly: dense retrieval compresses each text into one vector, while multi-vector retrieval keeps token-level vectors and scores query tokens against document tokens before aggregating best matches. That matters because late-interaction retrieval is increasingly the default tradeoff for quality-sensitive search systems. - Production agent latency often has little to do with the model itself: @dair_ai summarized a paper instrumenting ten agentic apps and finding that non-LLM components dominate latency in half of them, with sandbox memory peaking at 28GB/session, up to 32x latency variation across subsystems, and long idle state retention between steps. The optimizations are unsurprising but important: task-aware serving cuts latency 29–40%, state offloading reduces memory 4.6x, and tool-result caching removes 35.2% of redundant search calls. - Linear and turbopuffer show vector infra creeping into non-search hot paths: @turbopuffer said Linear moved its delta sync read path from Postgres to turbopuffer, using attribute indexes for permission filters and reducing the largest syncs by about 8 seconds. Google, OpenAI, Anthropic, and the Productization Race - Gemini 3.7 Flash had a strong day on both evals and product integration: @_philschmid and @NewsFromGoogle highlighted Gemini 3.7 Flash taking #1 on Artificial Analysis’s AA-AnalystAgent, with 60.0% pass^5, 70.5% pass@1, 77.5% pass@5, 1.32s/task, and $0.54 average cost across 80 spreadsheet/document-heavy quantitative tasks. Google also pushed it deeper into product surfaces: Gemini chat and Spark, Search-based interactive simulations built on the fly in AI Mode (example), and AI Studio GitHub sync for build workflows. - OpenAI is leaning into low-cost deployment and privacy positioning: @Replit launched Free Mode powered by GPT-5.6 Luna, which @kimmonismus framed as a meaningful efficiency win: a model that would recently have been SOTA is now cheap enough to be given away broadly. On the enterprise side, @OpenAI introduced Private Safety Processing, aiming to preserve Zero Data Retention for frontier models while still detecting cross-interaction safety risks without human access to the underlying content. - Anthropic continues to tighten the developer ergonomics loop: beyond the managed-agent updates above, @ClaudeDevs added a Concise output style to Claude Code, another sign that product teams are now tuning not just capability but response-shape as a first-class UX variable. Top tweets (by engagement) - Ornith-1.5 release: @ornith_ unveiled an MIT-licensed open model family from 9B to 397B, with strong coding/agentic benchmark claims and broad quantization support. - OpenAI privacy/safety infrastructure: @OpenAI announced Private Safety Processing while reaffirming Zero Data Retention for frontier models. - Gemini student push and product bundling: @GeminiApp offered a year of Gemini plans to students globally while rolling out new study-oriented features. - Claude Code UX update: @ClaudeDevs shipped Concise mode, a small but widely noticed improvement for day-to-day coding-agent interaction. - OpenRouter acquisition: @patrickc confirmed OpenRouter is joining Stripe, a move many interpreted as validation that token routing/marketplaces are becoming core infrastructure rather than edge tooling. AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Qwen/DeepSeek Open-Weight Inference Speedups - Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs (Activity: 1428): The image is a technical announcement graphic for “Dynamic v3.0 Qwen3.8”, showing Unsloth’s new Qwen3.8-27B Dynamic v3 GGUF post-training quantizations and claiming >10% higher top-1% accuracy at the same GGUF size versus other providers. It includes a memory table suggesting the model can run from 1-bit quants on ~8GB RAM up to BF16, plus a chart comparing accuracy across quant sizes; the post links the GGUF release on Hugging Face, the Dynamic 3.0 docs/benchmarks, and the image itself. Unsloth emphasizes these are post-training quantization releases only—“we do NOT use QAT or QAD”—and says the imatrix calibration file is public for independent evaluation and fine-tuning experiments. Comments were mostly positive, but one technical request asked Unsloth to add the previous UD 2.0 quants to the graph so users can compare against what they already have locally. Another commenter asked for deeper diagnostics, specifically per-category and KV-cache quantization KLD numbers, referencing localbench-style reporting. - Several commenters requested more detailed quantization evaluation for the new Qwen3.8-27B Dynamic v3 Unsloth GGUFs, especially a direct graph line comparing against the prior Qwen 3.8 27B UD 2.0 quants. Suggested metrics included KLD and/or top-1 agreement, which would help users judge whether the new dynamic quantization is materially better than the versions many already have stored locally. - A commenter asked for per-category KLD and KV-cache quantization KLD reporting, referencing the style of breakdowns from localbench.substack.com. This would make the quant quality discussion more actionable by showing which benchmark/task categories or cache-quant settings degrade most under different GGUF quant formats. - There was interest in the practical memory footprint of the quants: one user noted ~15 GB for Q4_K_M, while another inferred that IQ4_XS may now fit on16 GB VRAM “without mtp.” The technical concern is whether these smaller formats maintain model quality closely enough to justify running a 27B-class model fully on common consumer GPUs.
14:47

How to 8x Your Code Output Using Context Engineering

Anthropic's engineers merged eight times more code per day by late 2025 than in 2024, and the reason was how they manage what the model sees, not a better model. This piece walks through the practical version of that "context engineering," arguing it's about cutting context down to the high-signal parts rather than collecting everything. It explains that instructions, memory files, and rules load at different times, that files over about 200 lines get followed less, and that fetching context on demand beats preloading it. It also warns that the built-in memory and auto-memory only survive in certain places and that extra tools quietly tax every session.

Notes
Context Engineering Is Budgeting, Not Collecting
  • Anthropic's June 2026 report "When AI Builds Itself" found its engineers merging 8× more code/day in Q2 2026 vs. 2024.
  • Attribution: not the model, team, or people — "context engineering."
  • Anthropic's Applied AI team defined it (Sept 2025): the model has a finite attention budget; every token spends it whether it earned its place. Good context engineering = "the smallest set of high-signal tokens that make your desired outcome likely." Editing, not collecting.
  • Loading more makes agents worse: quality drops before hard limits; compaction rewrites history into a summary of what the summariser thought mattered. Every MCP server spends tokens on tool definitions before you type. A 400-line instructions file is followed less reliably than 200 lines.
The Four Layers (sorted by load time)

| Layer | Contents | Loads |

|---|---|---|

| ENFORCED | settings deny-list, hooks | always; not context |

| RESIDENT | CLAUDE.md, unscoped rules, MEMORY.md | disk, every session |

| CONDITIONAL | path-scoped rules | when matching file opens |

| FETCHED | skills, sub-agents, MCP | on demand |

  • Enforced is the only layer the client actually honours: deny-lists block tools/commands/paths regardless of model judgment; hooks run at lifecycle points. "Writing never read .env into a markdown file buys you a strong hint and nothing more — instructions are context and not enforced configuration."
  • CLAUDE.md locations: managed policy /Library/Application Support/ClaudeCode/CLAUDE.md; user ~/.claude/CLAUDE.md; project ./CLAUDE.md or ./.claude/CLAUDE.md; local ./CLAUDE.local.md (gitignore). Files above cwd load in full, root-downward; nearest loads last. Monorepo: claudeMdExcludes drops other teams' files by glob.
  • Path-scoped rules (paths frontmatter, nested CLAUDE.md) cost nothing until a matching file opens — good home for API conventions while working on frontend.
Writing an Instructions File
  • Claude Code reads CLAUDE.md, not AGENTS.md (no fallback). For mixed toolchains, put a single @AGENTS.md import at top or symlink.
  • Verify with /context (look under memory files); start with /init, then cut hard. Stay under 200 lines.
  • /doctor proposes trims: keep pitfalls, reasoning, conventions contradicting tool defaults; strip what Claude can read off the repo.
  • Rules must be checkable ("use 2-space indentation" vs. "format code properly"). Contradictory rules → Claude picks arbitrarily. Add on a trigger (repeated mistake, review catch), not a schedule.
Memory Layer & Compaction
  • Auto memory is on by default; files under ~/.claude/projects/<project>/memory/. MEMORY.md index loads first 200 lines or 25KB each session; debugging.md/api-conventions.md read on demand. Stale notes do more damage than missing ones.
  • Auto memory: one machine, shared across worktrees; sub-agents don't inherit main-conversation learning.
  • What survives compaction: project-root CLAUDE.md, unscoped rules, MEMORY.md index (re-injected from disk). What's summarised: path-scoped rules, nested CLAUDE.md, everything typed.
  • Use /clear between unrelated tasks; /compact + instruction to choose preservation.
Fetching vs. Front-loading
  • Facts → rules; procedures → skills. Sub-agents: clean window returns 1,000–2,000 token summary, hiding ~80k tokens of search output. But "a simple loop often beats an elaborate multi-agent arrangement."
  • MCP reaches tickets/incidents/schemas but every connected server taxes every session (startup tool-definition tokens). Argument for skills over servers: readable and auditable vs. black box with tool access.
The 8x Caveats
"Anthropic calls 8x almost certainly an overstatement... when it polled 130 of its own research staff the median answer landed nearer 4 times."
  • The 8x came from engineers directing and reviewing code, not typing it. Workflow: Explore → Plan → Code → Commit; Plan mode makes exploration read-only.
  • Verification is key: an agent that runs tests and reads stack traces fixes its own bugs; one handed failed tests "starts guessing."
  • Security team: incident control-flow tracing went from 10–15 min to ~a third of that.
Weekend setup

Set a deny list + one hook → /init then cut in half → make test suite runnable in one command with machine-readable failures → hand over a real ticket.

Full text · 14,224 chars
Anthropic’s Engineers 8x Their Output In June 2026 Anthropic published a report on its own engineering organisation called “When AI Builds Itself”, and the headline finding travelled a great deal further than the document did. By the second quarter of the year, the typical Anthropic engineer was merging 8 times as much code per day as in 2024. But the explanation was mostly misread. The foundational model did not change. The team did not change. The face behind the company didn’t change. It all came down to “context engineering”. That version sends people off to build enormous scaffolding that makes their agents slower and less reliable. Almost everything that genuinely works here is smaller and duller than the posts suggest, and most of it sits in documentation nobody reads. So what follows is the version I would hand a friend who asked me to set this up on their repo over a weekend. What context engineering means, which files load at which moment, and which parts of it are worth your Saturday. together with TrueForge: Everything below is context engineering for your sessions. TrueForge is the same discipline built into an agent harness: open source, vendor-neutral, benchmarked against Claude Managed Agents on the same model and tools ▫️ Same score (11/14), 62% fewer tokens, 30% lower cost per run ▫️ Compaction instead of replaying, the discipline this article teaches ▫️ Any model, any MCP server: $0.25 vs $1.10 per correct answer on the open-model config: Table of Contents 1. Context Engineering Is Budgeting, Not Collecting 2. The Four Layers and When Each One Loads 3. Writing an Instructions File Claude Will Follow 4. The Memory Layer, and What Survives a Compaction 5. Fetching Context Instead of Front-Loading It 6. Turning All of This Into Actual Output 1. Context Engineering Is Budgeting, Not Collecting The name makes it sound like the job is gathering everything relevant and handing it over. The job is closer to the opposite. A definition worth actually using Anthropic’s Applied AI team published a piece on this in September 2025, and the premise is the part worth keeping. A model works inside a finite attention budget, and every token you put in front of it spends from that budget whether it earned its place or not. Good context engineering, by their definition, means finding the smallest set of high-signal tokens that make your desired outcome likely. In other words, editing rather than collecting. Prompt engineering was about how you worded one message. This is about what the model can see across an entire task, and most of the craft turns out to be about what you leave out. Why loading more makes agents worse You can watch this happen in a long session. Quality drops well before you hit any hard limit, and once compaction fires near the ceiling your history gets rewritten into a summary that keeps whatever the summariser thought mattered. The costs add up fast. Every MCP server you connect spends tokens on tool definitions before you have typed a character, and a 400-line instructions file does not get followed twice as reliably as a 200-line one. Anthropic’s documentation says the reverse happens. So before you add anything, work out whether it needs to sit there for the whole session. Useful is not the bar. 2. The Four Layers and When Each One Loads Most people organise context by topic. Sorting it by load time tells you far more, because when something arrives decides whether it will still be there an hour later. ENFORCED settings and hooks always, and not context at all RESIDENT CLAUDE.md, rules, MEMORY.md from disk, every session start CONDITIONAL path-scoped rules when a matching file is opened FETCHED skills, sub-agents, MCP only when something asks for it The two layers that are always there The enforced layer barely counts as context, which is exactly why it goes first. A deny list in your settings file blocks tools, commands or paths and the client honours it whatever the model concludes, while hooks run as shell scripts at fixed lifecycle points and can stop an action outright. This is where a never-touch list belongs. Writing never read .env into a markdown file buys you a strong hint and nothing more, because the documentation is explicit that instructions are context and not enforced configuration. The resident layer is your CLAUDE.md files, any rules without path scoping, and the top of Claude’s own memory index. All of it comes off disk at the start of every session, and all of it costs you window space for the entire run. WHERE CLAUDE.MD FILES LIVE managed policy /Library/Application Support/ClaudeCode/CLAUDE.md user ~/.claude/CLAUDE.md project ./CLAUDE.md or ./.claude/CLAUDE.md local ./CLAUDE.local.md (gitignore this one) Files above your working directory load in full at launch, concatenated from the filesystem root downward, so whatever sits closest to where you started gets read last. In a monorepo where other teams’ files keep getting swept in, claudeMdExcludes drops them by glob. The two layers that come and go A rule carrying paths frontmatter costs you nothing until Claude opens a file that matches the pattern. Nested CLAUDE.md files in subdirectories behave the same way, which makes both of them a good home for API conventions you do not want resident while working on the frontend. The fetched layer is everything Claude goes out and retrieves. Claude skills load when you invoke one or when Claude judges it relevant, sub-agents hand back results from a separate window, and MCP calls pull in the ticket or the incident or the schema. Working out which layer something belongs to answers most of the questions people ask about this, including the one where an instruction seems to disappear halfway through a task. 3. Writing an Instructions File Claude Will Follow Everyone starts here, and almost everyone writes too much. Name it CLAUDE.md, then check it loaded Claude Code reads CLAUDE.md. It does not read AGENTS.md, and no fallback exists, so a repository holding only an AGENTS.md hands Claude nothing at all. A lot of guides get this backwards, and it matters because AGENTS.md is a real convention that most other coding agents do look for. Running a mixed toolchain, put a single @AGENTS.md import at the top of your CLAUDE.md, or symlink one to the other and move on. Then confirm it worked. Run /context and look for the file under memory files, because skipping that check is how people spend a fortnight convinced their instructions are being ignored. Start with /init, which reads your codebase and drafts something for you. Then cut it hard, since what it produces is a description of your repo and what you need is a set of corrections to it. Stay under 200 lines, and know what to cut The documentation targets under 200 lines per file and says plainly that longer ones eat more context and get followed less often. Treat the file as a budget line with a hard ceiling rather than a wiki page. The /doctor checkup proposes trims along a principle worth stealing outright. It strips anything Claude can work out by reading the repo and keeps the pitfalls, the reasoning, and any convention that contradicts what your tools do by default. WORTH THE SPACE - build, test and lint commands - conventions that contradict the framework default - where things live when the folder tree does not say - the failure that already cost somebody an afternoon - what to do when a requirement is ambiguous NOT WORTH THE SPACE - directory listings - dependency lists - architecture overviews - anything Claude can read off the repo in 10 seconds Write rules you could check against. Use 2-space indentation works, while format code properly does nothing at all. Anywhere two rules contradict each other across the tree, Claude may pick one arbitrarily, so a periodic read-through pays for itself. Add to the file on a trigger rather than on a schedule. Claude repeats a mistake, a review catches something it should have known, or you notice yourself retyping a correction you already typed last week. 4. The Memory Layer, and What Survives a Compaction Most playbooks tell you to hand-build a memory file that the agent reads at the start of a session and updates at the end. That feature already ships, switched on. Claude keeps its own notes now Auto memory is enabled by default and Claude decides what to keep based on whether it would help a future conversation. So you should build commands it worked out the hard way, a debugging pattern that keeps recurring, a preference you corrected twice, the reason a particular test stays skipped. ~/.claude/projects/<project>/memory/ MEMORY.md index, first 200 lines load every session debugging.md read on demand api-conventions.md read on demand Only the index loads at startup, capped at the first 200 lines or 25KB, whichever arrives first. Everything past that gets dropped without warning, which is why Claude keeps pushing detail out into topic files and reading them when needed. Open it now and then with /memory. It is plain markdown you can edit or delete, and a stale note does more damage than a missing one because Claude will act on it with total confidence. Two limits are worth knowing. Auto memory lives on one machine and is shared across worktrees of the same repository, and sub-agents do not inherit what the main conversation learned. Placement decides what comes back Compaction fires as you approach the ceiling, summarises your conversation, and hands the model a compressed version of its own history. What makes it through follows one rule that is simple and not at all obvious. SURVIVES project-root CLAUDE.md, unscoped rules, MEMORY.md index SUMMARISED path-scoped rules, nested CLAUDE.md, everything you typed Anything read off disk at startup gets re-injected afterwards. Anything that arrived through the conversation gets folded into a summary and stays folded. A path-scoped rule therefore vanishes at compaction and will not return until Claude opens a matching file again. Where an instruction has to hold for a whole task, put it unscoped in the project root and stop repeating yourself in chat. Two commands do most of the day-to-day work. Use /clear between unrelated tasks, which hardly anyone does often enough, and /compact followed by an instruction when you want to choose what gets preserved. 5. Fetching Context Instead of Front-Loading It Once you stop trying to preload everything, the interesting question becomes how Claude gets what it needs at the moment it needs it. Skills carry procedures, sub-agents carry volume Rules load every session or whenever a matching file opens, while skills load only when you invoke one or when Claude decides one applies. A fact Claude should always hold belongs in a rule, and a procedure it should sometimes follow belongs in a skill. Sub-agents are the strongest move available here and the one hardly anybody uses. A sub-agent explores inside its own clean window and returns a condensed summary, often 1,000 to 2,000 tokens, so your main session never sees the 80,000 tokens of search output behind the answer. Worth saying plainly though. A simple loop often beats an elaborate multi-agent arrangement, so reach for this when the reading volume justifies it and not because the architecture diagram looks impressive. External systems, and what they quietly cost MCP is how Claude reaches the ticket explaining why a feature matters, the incident showing how users are hitting a bug, and the schema the fix has to respect. That context is real, and it is frequently the piece that was missing. The cost stays hidden until you look for it. Every connected server spends startup budget on tool definitions before you type anything, so a dozen connectors you rarely touch is a tax you pay on every single session. Where the two overlap, there is a decent argument for skills over servers. You can read a skill and see exactly what it instructs Claude to do, whereas a server is a black box you have granted tool access to. 6. Turning All of This Into Actual Output At the end of the day, code still has to be written. Context work stops an agent making confident bad decisions, and by itself it produces nothing extra at all. The 8x came from somewhere else. Anthropic’s engineers stopped typing code and started directing and reviewing it. The workflow the company recommends runs like this: Explore → Plan → Code → Commit The first two are the steps people skip on their way to a fast wrong answer. Plan mode helps because it makes exploration read-only by construction instead of by request. Then verify, which is where their guidance spends most of its attention. An agent that can run your tests and read a real stack trace goes and fixes its own bugs, while an agent handed tests failed starts guessing, and guessing quickly is not productivity. Their own security team reports that tracing control flow during an incident used to take 10 to 15 minutes and now takes roughly a third of that, because Claude could run things and read output instead of reasoning about code it had no way to execute. So don’t focus on the headline. Anthropic calls 8x almost certainly an overstatement of the true gain, cites outside research showing developers overestimate how much AI speeds them up, and when it polled 130 of its own research staff the median answer landed nearer 4 times. What a weekend actually looks like Set a deny list and one hook for whatever would ruin your week. Run /init and then cut the result in half. Make your test suite runnable in one command with failures a machine can read. Then hand over a real ticket instead of a toy one, and watch what comes back. After that it is maintenance. You add a line when Claude repeats a mistake, delete a line when /doctor tells you it was derivable anyway, and keep the resident layer small enough that the model still has room to think. The takeaway is that you don’t have to learn to describe what you want more precisely. You have to find the work you are willing to stop doing, then build enough checking that allows you to sleep comfortable at night.
09:31

We’ve never seen an Anthropic before

A bullish prediction piece argues Anthropic is growing faster than any startup ever seen, ahead of a likely IPO around October 2026. It cites an annualized revenue run rate of $65 billion at end of July 2025 and second-quarter revenue of $11.5 billion, up from $787 million a year earlier, while OpenAI sits at $40 billion and slowing after cutting GPT-5.6 API prices. The author claims Chinese open-weight models are eating OpenAI's usage and predicts Anthropic could reach a $3 trillion valuation, though the figures are the author's unverified third-party claims and Anthropic's number counts partner commissions gross.

Notes
  • Anthropic ARR: $65B annualized revenue run rate at end of July 2025, per investors told "over the weekend." OpenAI ARR is $40B. Anthropic includes marketplace partner commissions (gross basis) in reported ARR, so its figure is slightly overstated vs. OpenAI's (reported net of Microsoft's cut).
  • Q2 revenue: $11.5B, up from $787M a year earlier (~14.6x YoY). Quarterly, Anthropic already has nearly double OpenAI's revenue.
  • ARR growth: up between 2,900% (from year-end 2024) and 64,900% (from start-of-year 2024). Growing ~$9B/quarter in mid-2026, roughly 3–4x faster than OpenAI's recent growth, driven by enterprise adoption and Claude Code.
  • IPO: expected October (~8–12 weeks). Author projects a $2T valuation target (benchmarked to SpaceX's $1.75T IPO) and possibly ~$3T by end-2026. SpaceX went public at $1.75T while not near profitability and "accelerating Capex way too fast."
  • OpenAI headwinds: cut GPT-5.6 API prices July 30, 2026, including an 80% cut on GPT-5.6 Lune — called "an act of desperation." Routing and code wrappers favor Chinese open-weight models over OpenAI due to cheaper tokens.
  • Thesis: OpenAI is being disrupted by Chinese open-source despite ~half its revenue being B2B; Anthropic, the "global leading frontier lab for coding LLMs," isn't. Author says Anthropic may hit 2x OpenAI's revenue by OpenAI's IPO; by 2030 Anthropic "will have achieved AI Supremacy."
  • Caveats/opinion: No counterpoints offered; "Anthropic even had its best Mythos model blocked" with no revenue impact. Author names TSMC, Nvidia and Anthropic as the "backbone" of generative AI, and suggests Nvidia should give Anthropic compute parity with OpenAI. The tone is strongly promotional ("one for the history books," "not even close" to any peer).
Full text · 4,933 chars
As Anthropic’s IPO approaches likely in October, likely just 8 to 12 weeks from now - it’s beginning to attract a different kind of attention. SpaceX, Anthropic and OpenAI are likely to be the biggest trio of IPOs the U.S. public markets have ever seen. While OpenAI’s revenue growth continues to slow, Anthropic’s ARR continues to accelerate even at a very high scale. - Anthropic told investors over the weekend that its annualized revenue run rate hit $65 billion at the end of July, 2025. - While OpenAI said its ARR has hit $40 billion. Since Anthropic includes marketplace partner commissions (gross basis) in its reported ARR, whereas OpenAI reports metrics net of Microsoft’s cut, Anthropic’s relative ARR might be slightly overstated compared to OpenAI. But even while ChatGPT rushes to Ad markets, OpenAI’s slowdown is sort of a big deal pre-IPO. This at a time when routing and code wrappers favor Chinese open-weight models over OpenAI model usage due to cheaper token costs. OpenAI reduced API prices for the GPT-5.6 family on July 30, 2026, even going so far as to cut the cost of GPT-5.6 Lune by 80%. What an act of desperation. The revenue tells the real story. Open-Source is disrupting OpenAI’s Revenue Growth So ironically OpenAI that makes around half of its revenue from B2B vs. B2C, risks being disrupted by Open-source models out of China, even as Anthropic doesn’t feel the same bite. Anthropic has a premium advantage as the global leading frontier lab for coding LLMs. While Anthropic even had its best Mythos model blocked, its revenue doesn’t seem to have been impacted yet: and, like I had predicted months ago, Anthropic has blown by OpenAI and might have 2x the revenue OpenAI does by the time OpenAI finally goes IPO. It’s hard to believe, and it’s one for the history books. If SpaceX’s IPO valuation is anything to go by, it (Anthropic) should be targeting a $2 Trillion valuation. SpaceX SPCX 0.00%↑ went public at an outrageous $1.75 trillion valuation while not being close to profitability and is accelerating Capex way too fast. I believe Anthropic could have close to a valuation of $3 Trillion by the end of 2026, because we’ve never seen growth like it has shown. Historic Anthropic Revenue Growth in 2026 Anthropic's Enterprise AI strategy experienced an inflection point in 2026. On a quarterly basis, Anthropic already nearly has double OpenAI’s revenue in Q2, and OpenAI has been accelerating fast the last few years. We’ve essentially never seen a company or an AI startup like Anthropic, in history. We know that OpenAI has higher compute costs, has worse financials and executes much more poorly on AI product. As the market becomes more competitive, it will continue to be squeezed. We simply won’t consider it a frontier lab for very much longer. Anthropic's revenue for the second quarter surged to $11.5 billion, up from $787 million a year ago. It's been growing at an incredibly fast pace. It doesn’t matter that ChatGPT has 800 million “weekly” active users, ever since its marketshare dropped below 50% and it’s market share continues to rapidly decline. Anthropic’s success as the dominant Generative AI API, is accelerating the entire Cloud computing industry by itself. So how much has Anthropic’s Revenue increased in the last two years? It’s hard to understand how growth could continue even at this scale (of the tens of billions) ? It’s meaningfully changed the trajectory of Google Cloud and AWS growth who both have significant equity (and partnerships) in the AI startup. Anthropic's ARR has increased between 2,900% (from year-end 2024) and 64,900% (from start-of-year 2024). Anthropic is thus the single primary tailwind for BigTech Earnings in 2026. Nobody grows 14x as this scale, nobody. For all the marketing gimmicks of OpenAI and insane confidence in SpaceX (including sadly some Pension funds), its three companies carrying the entire AI boom: TSMC, Nvidia and Anthropic. These companies are the backbone of what we call Generative AI of the last three years. - Yellow line is Anthropic 🟡 - Green line is Nvidia 🟢 - Blue line is TSMC 🔵 If Nvidia wants to subsidize compute as the “Bank of AI” (so prolific is it in circular and vendor financing), it should enable Anthropic to have compute parity with OpenAI. Nvidia is subsidizing Neo Clouds and the costs of compute at radical levels artificially inflating the value of the cost per token. Anthropic is increasing ARR around $9 Billion per quarter in mid 2026. Anthropic's Annual Recurring Revenue (ARR) has been growing roughly 3x to 4x faster than OpenAI's over recent quarters, driven largely by enterprise adoption and tooling like Claude Code. By the time OpenAI goes IPO, Anthropic won’t be a comparable company. It won’t have a peer competitor, do you understand what that means? Not Google, not OpenAI, not DeepSeek or Alibaba, not even close. By 2030, Anthropic will have achieved AI Supremacy.
13:00

Telcos: Being Right vs. Being Paid

Telecom keeps proving new technologies work while the revenue forecasts quietly disappear, a lesson worth heeding before the AI infrastructure boom. Private 5G hit about 6,500 deployments by end of 2025, yet the whole market is worth only roughly $2.4 billion against forecasts of $80-150 billion. Operator network API revenue in 2025 was just $284 million despite predictions of billions. The advice: before spending billions on AI edge and inference infrastructure, telcos should ask who actually gets paid.

Notes
  • Thesis: Telecom is a place to make money, not prove technological correctness. "Being right about the technology and being profitable on the investment are separate skills." Applies now as AI triggers another infrastructure capex cycle.
  • Pattern (repeatable): new tech → standards written → vendors publish huge TAMs → consultants draw hockey stick → operators spend billions → years later tech still alive but original revenue forecast "has quietly disappeared."
  • The key question: before spending on AI-RAN, sovereign AI, GPU-as-a-service, inference, edge infra, ask "If we are right, who actually gets paid?"
Private 5G — worked technically, failed commercially
  • ~6,500 private LTE/5G networks deployed worldwide by end-2025, up from 4,700 a year earlier.
  • Entire market worth only ~$2.4B vs. forecasts pointing to $80B–$150B. "Nobody lied about the technology. The factories got connected, and the networks got deployed. The money simply forgot to show up at the scale promised."
Network APIs — similar trajectory
  • Standardized APIs (identity, fraud, location, network capabilities) genuinely useful.
  • Actual operator API revenue in 2025: ~$284M; forecasts point to a few billion by 2030; earlier industry estimates claimed "hundreds of billions" in broader value creation.
  • Punchline: "the easiest part of creating a $100 billion telecom market remains putting $100 billion in the title of the consulting report."

Caveats: author grants some AI-era opportunities "may be real" — skepticism is about revenue forecasts, not technology itself. No sector where telcos were right AND paid is cited as counterexample; tone is cautionary, not prescriptive (no playbook offered beyond the single question).

Full text · 2,353 chars
The telecom market is not a place to prove you are right. It is a place to make money. Being right about the technology and being profitable on the investment are separate skills; a lesson worth remembering as the AI gold rush triggers another enormous infrastructure cycle. Telecom has spent the last decade proving that Open RAN, private 5G, edge computing, network APIs, 5G standalone, and network slicing can all work. The problem is that working and making money are not the same thing. The pattern is always the same. A new technology appears, standards get written, vendors publish enormous TAMs, consultants draw a hockey stick, operators spend billions, and a few years later the technology is still alive while the original revenue forecast has quietly disappeared. That is key these days because telecom is standing next to an even larger pile of capital labeled AI. Telcos are again being shown huge opportunities around inference, sovereign AI, GPU-as-a-service, AI-RAN, and edge infrastructure. Some of them may be real. But before spending another few billion proving they were technologically right, operators should ask a much simpler question: If we are right, who actually gets paid? Telecom has proven every tech, but the revenue was less cooperative. Private 5G is a good place to start because, technically, it worked. By the end of 2025, roughly 6,500 private LTE and 5G networks were deployed worldwide, up from 4,700 a year earlier. The problem is that the entire market was worth only about $2.4 billion, while forecasts had spent years pointing toward opportunities of $80 billion, $100 billion, and even $150 billion. Nobody lied about the technology. The factories got connected, and the networks got deployed. The money simply forgot to show up at the scale promised. Network APIs are following a similar, concerning track. Telcos can now expose identity, fraud, location, and network capabilities through standardized APIs, which is genuinely useful. Actual operator API revenue in 2025, however, was roughly $284 million, while forecasts continue to point toward a few billion by 2030, although earlier industry estimates talked about hundreds of billions in broader value creation. Apparently, the easiest part of creating a $100 billion telecom market remains putting $100 billion in the title of the consulting report.
14:43

How I Use Claude Code Routines to Run Work While My Laptop Is Closed

Anthropic's Claude Code Routines run your projects in the cloud, so agents keep working while your laptop is closed. A Routine is a saved setup of a prompt, repositories, and connectors that you can schedule or trigger via an HTTP endpoint. One example: changing a Notion topic's status to Research fires Make, which calls the Routine, which researches and updates Notion automatically. The catch: the endpoint doesn't watch apps or feeds itself, so an automation tool still has to notice the event and relay it.

Notes

Notes written to notes/claude-code-routines-ai-maker-2026-08-20.md. Task task_1787370214876 marked done (earlier accidental mark on task_1786554677304 reverted to doing).

Covered: the Routine definition (prompt + repos + connectors, cloud-hosted), scheduling vs. API trigger with the "narrow entry point" limitation, the Notion→Make→Routine→Tavily→Notion+email reference flow, why Make still sits in the middle, the six design questions, the five promised workflows, prerequisites (GitHub clone from default branch, Claude GitHub App, no local-only files) and the four setup steps, plus the Cherny/van der Meulen signals and all caveats.

Full text · 18,761 chars
If you have been following AI Maker for awhile, you know that most of my content is about building AI agents in Claude Code. And if you have applied them, you probably already have a project folder that knows how you work. Maybe you have a CLAUDE.md, an AGENTS.md, a context folder, a few skills, and connectors to the apps you use every day. Claude can read your source material, follow your rules, research a topic, update Notion, read your meeting transcription on Granola/Fathom, and prepare drafts in your voice. That setup can do a lot. Mine can help me plan newsletter posts, research ideas, repurpose my published articles, review form responses on my consulting website, build LinkedIn carousels, and work across the tools that run my business. I also have another project folder that essentially works as my Chief of Staff and knows my priorities, plus a second brain that stores any interesting information I consume across the internet and in real life. But every one of those workflows still had the same starting point: me. I had to sit in front of my laptop, open Claude Code, enter the project, and tell it what to do. Even when the instructions and tools were already there, the work waited until I showed up. That is fine when I am actively writing or making a decision with Claude. But the limitation shows up when I need my AI to act on its own—to automatically start a task at a specific time, or to kick off a workflow after something happens in another app. And I want all of that to happen while I’m away from my laptop. For example: - A new consulting inquiry shouldn’t sit untouched until I remember to check my Google Sheets form. - A daily AI news update from the newsletters I subscribe to shouldn’t have to wait until I open my laptop before it arrives in my inbox. - A content research task on my Notion calendar shouldn’t wait for me to tell my agent to research it; I should have a process to deploy agents to complete my tasks automatically. And sometimes I am on my phone, walking somewhere, with no intention of opening my laptop just to start a task Claude already knows how to complete. This is the part Claude Code Routines opens up. Your existing Claude project can now run in the cloud I think Claude Code Routines is an underrated feature that not many people are talking about, given the huge possibilities it unlocks. We all think of it as a way to schedule tasks, but trust me, it’s much bigger than that. Anthropic describes a Routine as a saved Claude Code configuration made from a prompt, one or more repositories, and a set of connectors. It runs through cloud infrastructure, so it can keep working while your laptop is closed. So, of course, the obvious benefit is that the task no longer depends on my computer remaining available. The more important benefit for me is that I do not have to rebuild the workflow on other agent platform. The research methods, writing instructions, examples, skills, source files, and tools I use already live inside my Claude Code projects. Routines give me a way to run selected parts of those projects in the cloud. That is also part of a broader shift I have been watching in the developer community. Boris Cherny, who works on Claude Code at Anthropic, recently recommended running Claude Code in the cloud for autonomous tasks that may take hours or days, specifically so you can close your laptop. Vincent van der Meulen, who says he now works almost entirely with cloud agents, argues that more engineers should start there because it changes how they can delegate work. He also admits that adoption remains limited and setting up cloud environments is still difficult. Lately, if you’ve been using Claude Code a lot, you may have noticed it increasingly nudges you to run your projects in a cloud environment. For most people reading this, you are not developers—and neither am I. But there are a lot of similarities between what developers do and what we do in our typical knowledge work: research, writing, planning, meeting preparation, and content production. But, the issue is all of that work ends up waiting when the agent only runs after you open your computer. So the practical starting point is this: how do we make Claude Code project also work in the cloud, so our work isn’t bounded by waiting to open the laptop? That’s where Routines comes in. How Claude Routines can help automate your work So far, most conversations about Routines seem to stop at scheduling. You create a task that runs every morning, every weekday, or once a week. That alone is useful. For example, you could ask Claude to check your calendar and recent email every weekday morning, compare them with your current project priorities, and send you a short briefing before you sit down to work. That solves the first limitation: you no longer have to prompt it by hand to get this task done. But there’s still a catch: you still need to open your computer, because the agents are running locally. By running Routines in the cloud, you remove this problem because the agent can keep doing tasks without your computer always needing to be on. These are the kinds of things the scheduling feature unlocks for us. But some tasks should begin because something happened, rather than at 9:00 every morning. A topic moved into research. A new post was published. Someone submitted a form. You sent a request from your phone. That is what the API trigger adds. Scheduling tells Claude to start working at a planned time. The API trigger lets an outside event tell Claude to start working now. Every API-enabled Routine gets an HTTP endpoint. An authenticated request to that endpoint starts a fresh Claude Code session on demand. That request can come from another app, an automation platform, or a Shortcut on your phone. There is an important limitation, though. The endpoint does not watch Notion, an RSS feed, or a form for you. Something else still has to notice the event and send the request in the format Claude expects. Sometimes the source app can make that request directly. Other times you need Make, Zapier, or n8n to receive the event, reshape the information, and relay it into the Routine. Change one label on Notion, start a full research workflow Here is one of the clearest examples from my own setup. I manage my content calendar in Notion. Each potential post has a topic, an angle, a status, and whatever source material I have collected so far. When I decide a topic deserves more investigation, I change its status to Research. That one change can start the rest of the workflow: - Notion sends the selected topic fields to Make. - Make reshapes that information and calls my Claude Routine. - Claude reads the incoming request and the instructions already inside my project. - It uses Tavily to research the topic and gather current sources. - It organizes the findings into a research brief. - It updates the correct item in my Notion content calendar. - It emails me a summary with the result. Notion status change ↓ Make routes the topic ↓ Routine API trigger ↓ Claude project + Tavily research ↓ Notion update + email summary By having this automation, I no longer need to open Claude Code and prompt it manually. All of this is done simply by changing the status label in Notion. Now you might be wondering why Make is still sitting in the middle of this workflow. Let me explain. Make still has a job in this setup I used to build much more of my automation inside Make. At the time, it gave me what I needed. Make could watch another app, start a workflow when something happened, move information between tools, and keep running without my laptop being open. Claude was one step inside that chain, so it made sense for Make to own most of the process. But that workflow started to become irrelevant when I began using Claude Code more and more. Moving the same job into a separate automation means passing all of that material into another AI step, which complicates my workflow because I now have two workflows to maintain between Claude Code and AI automation platforms such as Make.com. This is why I’ve been using Make, or similar apps such as Zapier and n8n, less. Routines let me keep the research, interpretation, and production inside the Claude project that already understands how I work. The Routine can also use selected connectors to read and write across apps such as Notion, Gmail, Google Calendar, and Google Docs. Once the run begins, Claude can coordinate those tools using the instructions in my project. But there’s one particular limitation in Routines in how it can be triggered to run. The current Routine API is a narrow entry point. It starts a session after receiving the correct authenticated request. It does not monitor a Substack RSS feed, detect a Notion change, filter an event, or reshape an arbitrary webhook payload. That is where Make still helps me. It can watch for the event, filter it, reshape the information, log what happened, and pass a clean request into Claude. Make catches and delivers the event. Claude uses the project to complete the work through Routines. If what I just said is confusing, don’t worry—we’ll get into it later in the post. It will all make sense once you’ve seen the full tutorial. How to think about building a useful Routine Before turning a workflow into a Routine, we need to think about how to create a useful Routine that can help you automate your work: 1. Trigger: When should the work begin? Choose one deliberate event or one predictable time. In my case, whenever I change a content status to Research in Notion, it triggers the entire research workflow. Or you can also create a trigger whenever there’s a new RSS feed you want to monitor, which then starts the workflow. 2. Input: What starts the request? The input is the new information Claude receives when the Routine begins. It could be a topic, URL, form response on your website, database ID on Notion, meeting transcript on Granola/Fathom, or note from your phone. Send only what the Routine needs to identify and complete the job. A form submission may include the person’s name, company, stated problem, and response ID. A content request may only need the Notion database ID because Claude can retrieve the remaining fields from the database. 3. Context: Where should Claude look for the rest? The input tells Claude what just happened. The context helps Claude understand what to do about it. Some of that context may already live in the selected repository: your instructions, source files, examples, research rules, or Skills. The Routine prompt should tell Claude which files to inspect and which Skill to run. Other context may live in a connected app. If the input is a Notion database ID, tell Claude which content calendar to open and which fields matter. If the task begins from a Granola or Fathom transcript, tell Claude where to retrieve the relevant transcript and what information to extract from it. 4. Process: What should Claude complete? Define the actual sequence, not only the outcome. For a research Routine, Claude might read the content-calendar record, inspect the research instructions in the project, run the relevant Skill, use Tavily to gather current sources, separate verified findings from open questions, and compile the result. The task should have a visible finish line that you can see and assess. 5. Output: Where should the result go? Choose one destination you already check: Notion, email, Google Docs, Google Slides, or a file in the project. Be specific about the write target: “Update the Research Summary field on this Notion page and email me the link” is one way to do it. The output should be easy to inspect. If you cannot quickly tell whether the task worked, the Routine will create more mental overhead than it removes. 6. Human review: What decision still belongs to you? Claude can prepare research, summarize information, draft a reply, organize survey findings, or create a presentation. For example, a consulting form can trigger a Routine that researches the person and company, compares the inquiry with the services you offer, and prepares a briefing with a draft reply. The Routine can do the preparation while you are away from your laptop. You still decide whether the lead is a fit, what you can offer, and whether the reply should be sent. Publishing a post, contacting a lead, approving a proposal, or acting on a sensitive recommendation should still come back to you. The goal is to remove the repeated preparation while keeping human judgment where it matters. Five Claude Code Routines workflows I want to show you Here are five Routines workflows we’re going to cover today that work across any type of job you have, whether you’re a creator, knowledge worker, consultant, or entrepreneur: - Scheduled morning briefing: Claude reads your calendar, recent email, and active project priorities, then sends you a plan for the day. - Notion research trigger: Change a content-calendar status and receive a sourced research brief, a Notion update, and an email summary. - Automatic post repurposing: Publish on Substack, let Make detect the RSS item, and have Claude prepare platform-specific drafts for review. - Consulting inquiry research: Use a form submission to prepare a briefing on the person and company, along with a draft reply you can approve. - An iPhone Shortcut for on-demand work: Send a link, note, or voice transcription from your phone and start one of your defined workflows while you are away from the laptop. Use these as reference points. Your version will depend on where your files live, which apps you use, what should trigger the work, and which decisions still need your review. The rest of this guide shows how these workflows operate, where Make sits in the middle, and the configurations I use to pass events into Claude. Use the details to understand the pattern and decide what fits your own setup. 🚨 Before you try this You should already have a Claude Code project you use for real work. If you are still setting that up, start with reading some of my posts here: - From Blank Folder to Working System: How to Set Up Any Project in Claude Code - How an Agent Harness Made My Claude Code Setup 10x More Reliable - The Complete Guide to the Context Folder That Changed How I Work With AI Agents For the project-based workflows in this guide, your Claude Code project also needs to be in a GitHub repository that Claude can access. Each Cloud Routine clones the selected repository when a run starts, beginning from its default branch. It cannot see files or changes that still exist only on your laptop. Before creating the Routine, commit and push the instructions, Skills, and source files it needs. Keep API keys, credentials, and sensitive local-only files out of the repository. You will also need to connect the repository to Claude Code on the web through the Claude GitHub App. I will show you the short connection and verification process before we configure the first workflow. Some outside triggers have their own requirements. That’s why you’ll need a Make, Zapier, or n8n subscription to run all these workflows. In this guide, I’m only going to show you how to do it with Make. By the end, you will understand how my Claude Code projects start selected work while I am away from my laptop, respond to deliberate events in other apps, and return something useful for me to review. You can then decide which parts make sense for your own work. Let me show you how these workflows work. Move your Claude Code project into the cloud Every workflow in this guide begins with the Claude Code project you already use on your laptop. That project may contain your CLAUDE.md, Skills, source material, examples, and the instructions Claude follows when it helps with your work. We do not need to rebuild any of that. We need to make the same project available to Claude Code in the cloud. GitHub is the bridge. Each Cloud Routine clones a GitHub repository when a run begins. It starts from the repository’s default branch, which means the Routine can only see files that have been committed and pushed. Anything that still exists only inside the folder on your laptop remains unavailable to the cloud run. We will set this up once. Afterward, the same cloud project can support the Notion research workflow, scheduled briefings, post repurposing, consulting inquiry research, survey analysis, and requests from your phone. Step 1: Push your local project to GitHub Before Claude can open the project in the cloud, the project needs to exist in a GitHub repository. If the project is already connected to GitHub, push the latest version. If it only exists on your laptop, create a private GitHub repository and connect the local folder to it. You can ask Claude Code to help: Help me push this local Claude Code project to GitHub. Before you push anything: - make sure the project instructions, Skills, and source folders are included - check that .env files, API keys, credentials, tokens, and private files are excluded - show me the files you plan to commit Wait for my approval before committing or pushing. If this is your first time, there are plenty of tutorials out there on YouTube that can help you create a GitHub account and push your project folder to GitHub. Step 2: Connect GitHub to Claude Code Next, follow Anthropic’s GitHub connection quickstart. The process is short: - Visit Claude Code on the web. - Choose the option to connect GitHub. - Install the Claude GitHub App. - Give Claude access to the repository you just pushed. - Confirm the default cloud environment. You only need to do this once for the GitHub account. If you add another project later, you can update the GitHub App’s repository access. Step 3: Start the project in the cloud Open Claude Code Desktop and start a new session. In the new-session screen, open the Claude menu and choose the cloud option. Claude will show the GitHub repositories available to your account. Select the project you just pushed, choose its default branch, and start the session. Ask Claude one simple question: Read this project and tell me which instructions, Skills, and main source folders you can see. If the expected files appear, the local project is now available to Claude in the cloud. If the repository does not appear in the selector, check that the Claude GitHub App has access to it. That is enough to move forward. We have the same project, available locally when we want to work closely with Claude and available in the cloud when a Routine needs to run without the laptop. Step 4: Understand the five parts of a Cloud Routine With the project available in the cloud, we can turn it into a Routine.
20:59

The /wayfinder Skill: Navigating the “Fog of War” of Planning

A new skill called /wayfinder helps people and their AI agents plan projects where the end goal is unclear, using a Warcraft-style 'fog of war' metaphor. Created by Matt Pocock (whose skills project has 220,000 GitHub stars), it lets an orchestrating agent spin up its own sub-sessions for prototyping, research, and task breakdown, so the human isn't constrained by context-window limits. It works off two documents — a 'map' of decisions already made and a 'ticket' for each sub-session — and suits big projects where you can't decide everything upfront.

Notes

/wayfinder — Interview with Matt Pocock (Latent.Space, 2026-08-20)

First of a planned "series about skills" on Latent.Space. Interview condensed "for readability."

Context
  • Matt Pocock, creator of "AI Skills for Real Engineers" (>220,000 GitHub stars) and YouTube channel with 347,000 subscribers.
  • New skill: /wayfinder — for projects where the end state isn't clear. Pocock's phrase: it navigates "the fog of war," where "you can't quite decide everything right at the start."
Origin / motivation
  • Pocock ran "AFK agents" (away-from-keyboard) overnight: plan work → write spec → turn spec into tickets. Had mature skills for turning work into scheduled agent tasks, but found the planning stage "really onerous" — constantly managing token budgets and context-window depth.
  • Wanted "an orchestrator layer": agent handles planning sessions for you, "split this out into multiple different threads, do prototyping, do research and pull it all back together," producing more detailed specs so "you can just whack off an AFK agent to go and do tons more work."
Design method
  • Core question: "what if a grilling session could manage other grilling sessions?" Started from the child's needs: a vague overview of what else is happening + their specific task.
  • This yields three named entities — the core of the skill:
  • map — all decisions already made (the rest of the context)
  • ticket — the specific task given to a session
  • session — the actual grilling/prototyping run
  • Emphasizes "leading words": precise, consistent terminology that "leads the agent to understand exactly what each part is." Inconsistent naming across places "produces strange behavior." Information flow (context management) is "really what a skill is."
Ticket types
  • grilling tickets — a grilling session
  • prototype tickets — create prototypes
  • research tickets — do research
  • task tickets — anything the human must do that the agent can't
'Fog of war' concept
  • Borrowed from Warcraft III map-exploration: you make some decisions, which "push further out into the fog of war." Pocock pairs the terms deliberately: "once I had the idea of 'fog of war' and 'map', I realized those two terms actually work really nicely together, and it really leads the agent into the right idea."
  • Used for engineering, non-engineering, and course planning.
Terminology / shared language
  • Spent months obsessed with terminology; has built (but not yet published) an "AI coding dictionary" — all AI-coding terms in "a beautiful graph you can explore" (agent, harness, model, etc.).
  • All courses reworked to use the dictionary; all skills "work off the same leading words." Motive: "I needed a ubiquitous language between me and the agent," since "between me and the agent, there is a communication barrier." Notes agents are "really good at domain modeling."
grill-me vs wayfinder (decision rule)
  • grill-me: when you can plan the whole thing in a single session and align before starting — "most small features," cases where "you can see the path ahead of you."
  • wayfinder: when you don't know the path — "you can feel the fog of war in front of you. You're gonna find your way with wayfinder."
Caveats / limitations stated
  • Interview edited for length; Pocock's grilling mechanics are "still working on." Dictionary unreleased. Skill validity implicitly tied to his personal vocabulary system — not independently benchmarked.
Full text · 7,224 chars
We’re currently developing a new series about skills, with the aim of giving you a regular supply of new skills to use in your projects. We’re kicking things off with an interview — and a super-useful skill — featuring Matt Pocock, whose “AI Skills for Real Engineers” project has over 220,000 stars on GitHub. He also talks about these skills to 347,000 subscribers on his YouTube channel. Pocock recently released a new skill called /wayfinder. Its purpose is to help you and your agent figure out a project where the end state isn’t entirely clear. Or as Pocock put it in our interview, /wayfinder helps you navigate “the fog of war,” where you have a project but “you can’t quite decide everything right at the start.” The following interview has been slightly condensed for readability, so you can read it, absorb Matt’s insights, and then test out /wayfinder for yourself! Latent Space: What were the goals of wayfinder? Pocock: What I noticed is I was doing a lot of work with AFK agents [Away From Keyboard] and trying to schedule in a ton of work so that my agents could run virtually overnight. I would just plan a bunch of stuff, and then I would create a spec and then turn that spec into tickets. And I had a really well-developed set of skills for how to turn work into scheduled stuff that agents could just crack on. But [...] I was finding the planning stage really onerous, because I would have to be constantly thinking about my session management. Like, how many tokens am I into my context window? How deep am I going here? I didn’t want to feel constrained in the planning stage anymore. I wanted an orchestrator layer that would basically say, okay, whatever you want to plan, I’m going to handle the planning sessions for you. I’m going to split this out into multiple different threads, do prototyping, do research and pull it all back together, so that you don’t feel constrained in the planning anymore. And then your specs can be even more detailed, and you can just whack off an AFK agent to go and do tons more work. Latent Space: What was the design process of coming up with this skill? Pocock: I had this kernel of an idea of, what if I didn’t have to manage the handoffs? What would that look like? And then, what would it look like to have some kind of centralized document to have all of those pieces together? Whenever you’re thinking about context management — because that’s really what a skill is, you’re managing the context of the agent you’re working in — you need to think about the information flow. So what I wanted to think about is, what if a grilling session could manage other grilling sessions? What would that look like? Well, the first step to that is, what does the grilling session that’s being managed need? What does the child need in that situation? So the child probably needs to understand a vague overview of what else is happening, and they need their specific task. So there, you’ve got two documents. You’ve got a map — which is all of the rest of the stuff, all the decisions that have already been made. And then you’ve got the specific ticket that goes into the actual session. And what you notice there is that those words are very precise. You’ve got the map, and you’ve got the ticket, and you’ve got the session. And once you’ve got the kernel of an idea, you then need to come up with the words for that idea. Because once you’ve figured out the words, then those entities can be really clearly mapped out by the agent. Because if you just call everything a ticket, or if you just refer to it in different ways in different places, then it’s going to be really confused and you’re going to get strange behavior. Whereas if you use these very specific, what I call leading words, to lead the agent to understand exactly what each part is, and you’ve understood what the information flow is, then you’ve got your skill. Latent Space: What kind of use cases do you think wayfinder would be useful for? Pocock: Well, I’ve been using it for all sorts of stuff. I’ve been using it to actually plan courses as well. In wayfinder, there are different types of tickets. So you’ve got grilling tickets, which are just a grilling session. Then you’ve got prototype tickets for creating prototypes, research tickets for creating [and doing] research, and then task tickets — which are really broad…basically, just anything the human needs to do that the agent can’t do. And so once you think about that, you realize, OK, I can apply that to anything. One really key idea in wayfinder is the ‘fog of war’. So this is the concept of, you can’t quite decide everything right at the start. You can make certain decisions, and those certain decisions sort of lead you there and push further out into the fog of war — kind of like Warcraft III style, exploring the map. And once I had the idea of ‘fog of war’ and ‘map’, I realized those two terms actually work really nicely together, and it really leads the agent into the right idea. So I’ve been using it for engineering, for non-engineering stuff, for course planning, all sorts. Latent Space: This concept of the fog of war — it’s weird to consider what you don’t know that you don’t know. Maybe LLMs are good at capturing that. Pocock: I feel like with the grilling stuff that I’m still working on, that captures an idea that you don’t know stuff, but maybe the agent can contribute something and illuminate a part of the room that you don’t quite understand yet. And wayfinder is just sort of an extra layer on top of that. Latent Space: Yeah, and there’s all these artifacts. How much time do you spend teaching the model all this terminology? Pocock: For the last few months, I’ve been pretty obsessed with terminology — and finding the right terms for certain things. I’ve put together, I haven’t actually put it out yet, but it’s an AI coding dictionary — of basically all the terms in AI coding. It’s in this beautiful graph that you can explore and understand exactly what an agent is, exactly what a harness is, exactly what a model is, blah blah blah. I’ve redone all my courses to use that dictionary and make it really solid. And then all of my skills use a consistent dictionary as well. So they’re all working off the [same] assumptions, the same leading words. I realized that I needed a ubiquitous language between me and the agent. Between me and the agent, there is a communication barrier. And that’s what I’m trying to do with my skills all the time, is try to find the right words. And agents are really good at showing you the opportunities for different wording — really good at domain modeling, actually. Latent Space: When do we directly use the grill-me skill, versus wayfinder? Pocock: Use ‘grill me’ in cases where you feel like you can plan the whole thing in a single session, and you need to align before you go. So most small features will fit into this. Most stuff where you can see the path ahead of you, but you just want to make sure the agent is on board, ‘grill me’ will work with that. For stuff where you don’t know the path ahead, for stuff where you can feel the fog of war in front of you, use wayfinder. You’re gonna find your way with wayfinder. So that’s how it works.
09:17

Your Best Prompts Are Hiding in Your Chats

You already wrote your best AI prompts by correcting a chatbot mid-chat, so you can just ask it to save the conversation as a reusable skill. After a successful session, tell the agent to turn the conversation into a skill (for Claude Code or ChatGPT Work, formerly Codex) or a customizable instruction with placeholders (for regular chatbots like ChatGPT or Gemini). Pick chats for tasks you've done at least three times and expect to keep doing.

Notes

Your Best Prompts Are Hiding in Your Chats

Source: Why Try AI (Substack), 2026-08-20. Experimental short "Quick Fix" post format (problem + solution), with a reader poll at the end.

The problem

Author's claim: "prompt engineering" and other hacky prompt techniques are "largely a thing of the past," yet people still chase elaborate off-the-shelf prompts written by others. Users who think they can't write good AI instructions already have — chat history is full of conversations where AI eventually did what was needed, even after corrections.

Core idea: every correction, new input, or request for a different approach was, in effect, writing instructions all along. The only missing step is converting those long chats into a reusable form.

The fix

End a successful session by pasting:

"Turn this conversation into a [skill / customizable instruction] that I can apply to similar tasks."
  • Use "skill" for agentic tools (Claude Code, ChatGPT Work — formerly "Codex")
  • Use "customizable instruction" for simple chatbots (ChatGPT, Gemini, etc.)

The agent builds the reusable skill itself; a chatbot returns an instruction with customizable [placeholders] you can save, and optionally use to create a custom GPT.

Steps
  • Open your go-to chatbot/agent
  • Find a conversation that produced the result you wanted
  • Paste the line above
  • Reuse
Selection criteria for candidate conversations
  • Task worked on at least 3 times
  • Expected to keep doing it
  • Complex enough to warrant detailed instructions

Caveats: none stated beyond the format experiment itself — the author asks readers to test it and report back; no failure modes or limits are acknowledged.

Full text · 1,994 chars
I’m trying this new short-and-snappy post format with a problem + solution angle. Let me know what you think of it in the poll at the end. TL;DR If AI gave you a great result after a long chat, you can instantly turn that into a clean and reusable skill (or prompt) by just asking for it. The problem “Prompt engineering” and other hacky prompt techniques are largely a thing of the past, but people still chase elaborate off-the-shelf prompts written by others. You may think you don’t know how to write good AI instructions. The thing is, you almost certainly already have. I bet your chat history is full of conversations where AI eventually did what you needed…even if you had to yell at it a few times: But did you know that every time you corrected the chatbot, gave it new input, or asked for a different approach, you were also writing your instructions in the process? All that’s left is for you to turn those long chats into something you can reuse. The fix Whenever you end a successful AI session, tell your AI chatbot or agent this: "Turn this conversation into a [skill / customizable instruction] that I can apply to similar tasks." - Use “skill” for agentic tools like Claude Code or ChatGPT Work (formerly “Codex”) - Use “customizable instruction”1 for simple chatbots like ChatGPT, Gemini, etc. The agent will simply build the reusable skill for you. The chatbot will hand you an instruction with customizable [placeholders] that you can save for later. (You can also use it to create a custom GPT or something similar). Do this now - Open your go-to AI chatbot or agent - Find a conversation that gave you the result you wanted - Paste in the above line - Profit The best candidates are conversations about tasks that: - You have worked on at least 3 times - You expect to keep doing - Are complex enough to call for detailed instructions Go try it and let me know how it works for you! Share your thoughts I’d love to know your take on this kind of shorter “Quick Fix” post.
09:53

Ask Me Anything - Am I doing real work or just hiding behind busywork?

A solo-business coach gives a three-question test to tell real work from busywork, asking whether the task alone would make the day a success, whether it creates a result or just avoids something harder, and what happens if it's skipped. He admits spending three hours reorganizing his Notion workspace and getting nothing done, and flags learning as a trap because it feels like progress without output. He recommends planning three to five concrete tasks the night before, starting with the easiest, and automating any task that repeats daily and eats more than five minutes, capped at a 30 to 45 minute build window.

Notes
Ask Me Anything — real work vs busy work

Context: Monthly AMA (3rd Wednesday) on Substack Live. Author: Anfernee.

Motive example: Spent 3 hours reorganizing Notion (new databases, templates, folders) — felt productive, "got nothing done."

The 3 daily checks (run every task through):

  • If this is the only thing I finish today, will I call today a success?
  • Am I doing this to create a result, or to avoid something harder?
  • What happens if I skip this today?

Q1 filters for output — newsletter draft passes (publishes Thu + Sun, 2 posts/week); Notion reorg fails ("Nobody reads a tidy database").

Learning ≠ output: "Your brain logs it as progress anyway." Doesn't count learning as a task — learns on demand, top 3 things to move forward, builds, learns again at next wall. "Trying to learn ten things before starting is a stall dressed up as preparation."

Where busy work hides (3 traps):

  • Checking subscriber/follower counts — "fine once in a while"; as a daily habit it's "dead time."
  • Rebuilding website — one page update becomes 5 tabs + unwritten FAQ rewrite.
  • Research with no scope — sets a boundary before starting or "read for an hour and produce nothing."

Task system: List built the night before — one main goal, broken into 3–5 concrete tasks, never more than 5. Start with the easiest first (small early finish pushes momentum). Deep focus: notifications off, work straight from list, sometimes paper with checkboxes.

AI automation rule: If a task repeats daily and eats >5 min each time, automate. Cap the build at a 30–45 min block; if not done, move on and return later. "There's no perfect automation... The point of automating is freeing up time for the work that matters."

Caveats/limits: This post only covers a slice — the Live goes deeper (when research turns into stalling; resetting a no-task-list day). Reader Q's referenced: Somali-library, Ivan Age.

Upsell: Premium Vault, $79/year ($6.58/mo).

Full text · 4,385 chars
On this month's AMA Substack Live, I shared the 3 questions I run every task through, plus the task system and automation rule that keep me honest about it. Access your FREE Solopreneur Success Hub - your subscribers-only comprehensive command center for building and scaling a successful one-person business. I created this all-in-one toolkit for building a profitable one-person business, something I wish existed when I first started, and it saves me 20+ hours a week. Now, it’s yours… FREE! Real Work vs Busy Work Last month I spent three hours reorganizing my Notion workspace. New databases. Shifted templates. Tidier folders. I felt good at the end. I got nothing done. On this month’s AMA Live, I broke down how I catch myself doing this, and the three questions I run every task through before I call it real work. The 3 Questions I Ask Every Day I run every task through three checks: - If this is the only thing I finish today, will I call today a success? - Am I doing this to create a result, or to avoid something harder? - What happens if I skip this today? Question one filters for output. Writing my newsletter passes every time. I publish two posts a week, Thursday and Sunday, so a finished draft is a win on its own. Re-organizing Notion fails it. Nobody reads a tidy database. It changes nothing for my readers or my business. Learning Feels Like Progress. It Isn’t Output Learning gives you the same lift as finishing a task, minus the finish line. Watch a tutorial, read an article, take notes. Your brain logs it as progress anyway. I personally don’t count learning as a task. I learn on demand. I pick the top three things I need to move forward, learn those, and build. When I hit the next wall, I go learn again. Trying to learn ten things before starting is a stall dressed up as preparation. Where Busy Work Hides Three spots I catch myself most: - Checking subscriber and follower counts. Fine once in a while. A daily habit, and it’s dead time. - Rebuilding my website. One page update turns into five open tabs and a rewritten FAQ page nobody asked for. - Research with no scope. I set a boundary before I start, or I read for an hour and produce nothing. 3 to 5 Tasks, No More Every task list starts the night before. I set the main goal for tomorrow, then break it into 3 to 5 concrete tasks. Never more than five (at least I try) I start with the easiest task first. This is because finishing something small early pushes me into the next task, then the one after that. When I need deep focus, notifications go off and I work straight from the list, sometimes on paper with a checkbox beside each line. Is AI Automation Busy Work or Real Work? My rule: if a task repeats daily and eats more than five minutes each time, I look at automating it. I set a 30 to 45 minute block for the build. If it’s not done in that window, I move to the next task on my list and return later. There’s no perfect automation. Something breaks, you fix it, you move on. The point of automating is freeing up time for the work that matters, not polishing a workflow nobody sees. Watch the Full AMA This post covers a slice of the conversation. The Live goes deeper, including reader questions on when research turns into stalling and how to reset a day that starts with no task list. I host this AMA on the third Wednesday of every month. Drop your own busy work trap in the comments. I read every one. You’re doing everything. But nothing is moving? You are doing everything. But nothing is moving. That is not a motivation problem. Most solopreneurs are learning from everywhere and getting nowhere. Too much information. No clear system connecting effort to results. You have everything it takes. You just do not have a clear system yet. That is what paid subscribers get. Every system, playbook, prompt, and template. All inside the Premium Vault. All for $79/year. That’s $6.58/month. Upgrade now and unlock the Premium Vault worth thousands of dollars. The Premium Vault holds the secret behind posts like this one, including the tools and resources I use to build the one-person business I love. Thank you Somali-library, Ivan Age, and many others for tuning into my live video! Join me for my next live video in the app. Thanks for reading! Ready for the next step? Let’s crack the growth equation and build a thriving one-person business on your terms! Anfernee
11:56

Gumroad, Stan Store, ThriveCart, Others: What We Learned Comparing 10+ Platforms

Two solopreneurs who tried roughly twenty ecommerce platforms both still sell on Gumroad, the one they call the ugliest, because switching would cost them their affiliate networks and email workflows rather than just fees. One of them built over 400 affiliates and estimates losing 50 to 90 percent of that network by moving platforms, while his email flows run 20 to 30 messages deep. The math shows Gumroad's 10 percent per-sale fee beats a $99-a-month flat plan until roughly $1,000 a month in sales, and the authors advise spending five percent of effort on the landing page and the rest on promotion.

Notes

Comparing 10+ e-commerce platforms (Gumroad, Stan Store, ThriveCart, et al.)

From Solopreneur Code (Substack), published 2026-08-20. Authors Jamie and Anfernee have collectively tried ~20 platforms (Lemon Squeezy, Stan Store, Beacons, ThriveCart, Payhip, Pensight, Ko-fi, Kit, Carrd, Medium, Amazon KDP). Both remain on Gumroad, which they call "the ugliest platform in the group."

Why they stay on Gumroad (switching cost, not fees)
  • Jamie built 400+ affiliates on Gumroad; moving would cost him an estimated 50–90% of that network (people built content/links around his Gumroad page won't follow).
  • Anfernee's email workflows live in Gumroad, "some running 20 to 30 emails deep over three or four months."
  • Gumroad strengths: affiliate marketing + email workflows in one place, free until first sale. But "one of the worst UIs of any platform we have used."
The break-even math
  • Gumroad: 10% per direct sale, no monthly fee.
  • Stan Store Creator Pro: $99/month, 0% platform fee.
  • > "Stan Store gets cheaper once you cross roughly $1,000 a month in sales. Below that line, Gumroad's percentage costs less."
Other platforms tested
  • Lemon Squeezy: better landing pages; Jamie hit affiliate-tracking bugs years ago (note: Stripe acquired it since).
  • Stan Store / Beacons: mobile-first, cleaner, but less depth on email workflows/affiliates.
  • ThriveCart: ~$400–500 lifetime deal, but company risk — "Close to half of lifetime-deal platforms on sites like AppSumo do not exist two years later." Lose affiliates/links, not just money.
  • Payhip: tried and dropped; UI issues, missing features.
  • Pensight: Jamie calls it best-designed all-in-one he's used, but "still not functional enough to replace Gumroad."
  • Kit (ConvertKit): best email customization, price reflects it.
Key rule
"Spend 5 percent of your time on the landing page. Spend the other 95 percent promoting what is on it."

Their best-selling newsletter posts doubled as sales pages; no custom funnels built. Analogy: "Gmail beat Hotmail on function, not looks."

FAQ answers
  • Switching on high fees? Run the math at real volume; below ~$1,000/month, percentage usually beats subscription.
  • Biggest miss? Switching cost — affiliates, email workflows, links don't move automatically.
  • Zero sales? Start with Gumroad (free until sale, no risk).
  • Multiple platforms? Only if each brings its own audience (e.g., Gumroad Discover + Amazon KDP); otherwise splitting promotion.

Limitations/unknowns: The piece is promotional — the author sells a "$79/year Premium Vault." The $1,000 break-even is their own live calculation, not a general model. Platform verdicts are anecdotal (2 people's experience, some years old). A full 12-platform replay and comparison table are gated behind the Substack video.

Full text · 6,777 chars
Jamie and I have used close to 20 e-commerce platforms between us. Lemon Squeezy, Stan Store, Beacons, ThriveCart, Payhip, Pensight, Ko-fi, Kit. Name one, and one of us has tried it. We are both still on Gumroad. The ugliest platform in the group, by our own admission. This is the answer to the question we both get asked constantly: what platform should I use? Access your FREE Solopreneur Success Hub - your subscribers-only comprehensive command center for building and scaling a successful one-person business. I created this all-in-one toolkit for building a profitable one-person business, something I wish existed when I first started, and it saves me 20+ hours a week. Now, it’s yours… FREE! There’s no best platform, only the right fit We spent an hour live on Substack going through platform after platform, and we kept landing on the same conclusion every time. There is no perfect platform. You learn to accept the disadvantages of the one you pick, and put the advantages to work for you. Gumroad has one of the worst UIs of any platform we have used. We agreed on this in the first five minutes. But it handles two things we both need well: affiliate marketing and email workflows in one place, free until you make a sale. Why we’re both still here The fees are not the real reason we stayed. Jamie built over 400 affiliates on Gumroad. Some have sent traffic for years. Move everything to a cheaper platform tomorrow, and he estimates a loss of 50 to 90 percent of that network. People who built content and links around his Gumroad page will not follow automatically. Not a Gumroad fee. A switching cost nobody talks about when comparing percentages on a spreadsheet. I built something similar. My email workflows live inside Gumroad, some running 20 to 30 emails deep over three or four months. Moving this infrastructure is not a weekend project. New, and have not made a sale yet? None of this applies to you. Why we both still tell people to start with Gumroad. Nothing to lose testing an idea, nothing tying you down yet either. The math that decides it Here is the calculation we ran live. Gumroad takes 10 percent off every direct sale, no monthly fee. Stan Store’s Creator Pro plan runs $99 a month, 0 percent platform fee. Run the numbers, and Stan Store gets cheaper once you cross roughly $1,000 a month in sales. Below that line, Gumroad’s percentage costs less. Above it, a flat monthly fee wins. The real question before switching anything: what is your monthly volume, and where is your break-even line? Not which platform has nicer landing pages. What we learned testing everything else Between us, here is the fast version of a dozen platforms: - Lemon Squeezy - better landing pages than Gumroad. Jamie hit bugs with affiliate tracking a few years back. Stripe acquired the company since, so results might differ now. - Stan Store and Beacons - both mobile-first, cleaner design than Gumroad. Neither matches Gumroad’s depth on email workflows or affiliate marketing. - ThriveCart - a lifetime deal around $400 to $500 sounds great, until you weigh the company risk. Close to half of lifetime-deal platforms on sites like AppSumo do not exist two years later. Lose the platform, and you lose the affiliates and links built on it, not only the money paid. - Payhip - tried and dropped. UI issues and missing features, nothing beating what Gumroad already offered. - Pensight - Jamie calls it one of the best-designed all-in-one platforms he has used. Still not functional enough to replace Gumroad for his workflow. - Kit (ConvertKit) - the best email customization of anything either of us has tried, and the price reflects it. Jamie used to stack Gumroad, a landing page tool called Carrd, and Kit together. The integrations worked well. The rule that matters more than any platform The best piece of advice from the hour had nothing to do with fees or features. Spend 5 percent of your time on the landing page. Spend the other 95 percent promoting what is on it. Some of my best-selling newsletter posts are also my sales pages. Jamie did the same on Medium before moving to Substack. Neither of us built a custom funnel to make this work. We wrote and hit publish. Every hour spent picking a theme or tweaking a button is an hour not spent telling people your product exists. Gmail beat Hotmail on function, not looks, and the design caught up eventually anyway. Your product page follows the same order too. Watch the full hour We covered more in the full hour than fits here: - how Jamie built and kept his affiliate network, - the payout thresholds catching people off guard depending on country, and - platform-by-platform detail on all 12 we discussed. Watch the full replay here and grab the platform comparison table alongside it. What platform are you on right now, and what kept you there? Drop it in the comments. We are planning the next Live around whatever comes up most. FAQs Q: Should I switch off Gumroad if the fees feel high? A: Run the math first. Compare Gumroad’s 10 percent against a flat monthly plan at your real sales volume, not a hypothetical one. Below roughly $1,000 a month in sales, the percentage usually costs less than a subscription. Q: What is the biggest thing people miss comparing platforms? A: Switching cost. Affiliates, email workflows, and existing links do not move with you automatically. Weigh what you would lose, not only what you would save. Q: I am new and have made zero sales. Where do I start? A: Gumroad. Free until you make a sale, so testing an idea carries no risk before you know what you need. Q: Is it worth using multiple platforms for the same product? A: Only if each platform brings its own audience or discovery engine, like Gumroad’s Discover feed alongside Amazon KDP. Otherwise you are splitting your own promotion across more links. You’re doing everything. But nothing is moving? You are doing everything. But nothing is moving. That is not a motivation problem. Most solopreneurs are learning from everywhere and getting nowhere. Too much information. No clear system connecting effort to results. You have everything it takes. You just do not have a clear system yet. That is what paid subscribers get. Every system, playbook, prompt, and template. All inside the Premium Vault. All for $79/year. That’s $6.58/month. Upgrade now and unlock the Premium Vault worth thousands of dollars. The Premium Vault holds the secret behind posts like this one, including the tools and resources I use to build the one-person business I love. Thank you to everyone who tuned into my live video! Join me for my next live video in the app. Thanks for reading! Ready for the next step? Let’s crack the growth equation and build a thriving one-person business on your terms! Anfernee
22:30

Make Money on Reddit With AI: The 30-Day Playbook

Anyone can find Reddit customers by using AI to spot recurring complaints, then building a product around the exact pain language. The 30-day playbook says to use Reddit Pro's free keyword listening, track buying-intent signals, and reply with genuine help instead of promotion. Reddit now bans bots, mass DMs, bought accounts, and scraping, so automation belongs behind the scenes only. Reddit reports about 514.6 million weekly active users, and one case study — Wayfair — grew referral traffic and boosted profile followers by over 50% using this approach.

Notes

Make Money on Reddit With AI: The 30-Day Playbook

Source: Open Cloud AI (Substack), published 2026-08-20.

Core premise

Reddit scale: 130M+ daily visitors, 100,000+ active communities, 26B+ posts/comments. Reddit reported 514.6M weekly active uniques in Q2 2026. The author reframes the strategy question: not "What should I post?" but "What are people already trying to buy?" Reddit serves as a demand map, not a marketing channel.

"Reddit finds the demand. AI finds the pattern. You build the offer. Humans build the trust. Revenue tells you whether you were right."
What's now banned (vital caveat)
  • Spam rules prohibit repeated/unsolicited mass engagement, mass-posting repetitive material for financial gain, bulk unsolicited DMs, and bot/generative-AI spam facilitation.
  • User Agreement (effective July 1, 2026): cannot sell/transfer accounts without prior written approval; scraping requires prior written consent.
  • Killed "growth hacks": buying aged accounts, backup-account farms, automated scraping, mass-generated replies, automated promo DMs.
  • Sustainable model: "Automate the work behind Reddit. Keep the behaviour on Reddit human."
Listening layer

Reddit Pro is free and offers keyword/phrase tracking, a Trends feature (relevant communities, conversations, mention volume), and organic performance data. Search pain language, not product categories. Example search terms: alternative to [competitor], how do I [problem], struggling with [problem], is [product] worth it. Model phrase: "I waste every Friday making the same report."

Five levels of buying intent
  • Curiosity — 2. Frustration — 3. Advice seeking — 4. Solution seeking — 5. Purchase intent.

Only levels 4–5 matter: the problem is already understood, no urgency to manufacture.

Case study: Wayfair used Reddit Pro to answer shopping questions rather than push links; Reddit reports steady month-over-month referral traffic growth and 50%+ increase in Reddit profile followers. Caveat noted: results won't replicate for everyone.

AI stack

Start with three tools: Reddit Pro + Google Sheets + one AI assistant (default: ChatGPT; Claude and Gemini work the same). Setup per model:

  • ChatGPT → Project named "Reddit Revenue Research" with research sheet, product notes, community rules, useful posts.
  • Claude → same as a Claude Project (project knowledge base, uploaded docs, instructions, retrieval).
  • Gemini → custom Gem "Reddit Opportunity Analyst" with instructions + Knowledge files, can reference Google Drive.
"The model is not the business. The process is."
The 30-day schedule
  • Days 1–3: pick one customer + one painful problem. Spreadsheet columns: Date | Subreddit | Problem | Exact Phrase | Intent 1–5 | Current Solution | Why It Failed | Link | Possible Offer. Prompt AI to generate 25 natural phrases grouped by intent (frustrated / advice / comparing / buying). Add terms to Reddit Pro Trends. First goal: vocabulary, not traffic.
  • Days 4–7: map 5–10 communities — participants, what's discussed/upvoted/removed, promo rules, writing style, recurring problems. Read subreddit rules yourself; don't let AI guess them.
  • Days 8–14: participate — a few genuinely useful replies daily ("Three excellent answers are more valuable than 50 AI-shaped comments"). Log problems verbatim, e.g. "I spend half of Monday fixing this manually" — not "customer requires workflow optimization." That sentence may become the product headline.
Stated limitation / open end

The article ends mid-stream: days 15–30 (choosing the opportunity, building the offer, converting and measuring) are deferred to "What Comes Next: Turn Research Into Revenue" — not covered in this installment.

Full text · 9,031 chars
More than 130 million people visit Reddit every day. They gather across more than 100,000 active communities and have created more than 26 billion posts and comments. Somewhere inside that enormous pile of conversation, people are describing products they wish existed, tools they hate, services they cannot find, money they regret spending, and problems they want solved now. Most people trying to make money on Reddit start with the wrong question: What should I post? A better question is: What are people already trying to buy? That difference changes the whole strategy. Reddit should not begin as your marketing channel. It should begin as your demand map. And this is where AI becomes unusually useful. Not because ChatGPT can flood Reddit with posts. Not because Claude can impersonate 50 customers. Not because an agent can automatically DM strangers. AI is useful because it can take hundreds of messy human conversations and help you see patterns that would otherwise take days to notice. The system in this article is simple: Reddit finds the demand. AI finds the pattern. You build the offer. Humans build the trust. Revenue tells you whether you were right. And you can test the whole thing in 30 days. Reddit is not another social network On Instagram, a business often starts with an audience. On Google, it starts with a search query. On Reddit, it frequently starts with a problem. Someone writes: I have tried three CRM tools and still hate all of them. Someone else asks: How are freelancers handling client revisions without living inside email? Another person wants: A simple way to turn customer calls into weekly reports. Those sentences are valuable because the customer has already done something marketers spend heavily trying to accomplish. They have identified their own pain. Reddit reported 514.6 million weekly active uniques in Q2 2026. It is also investing more heavily in protecting its corpus of human conversation as AI makes genuine human experience more valuable online. That makes Reddit increasingly interesting for entrepreneurs, consultants, SaaS founders, creators and small businesses. But only if you understand the culture. The fastest way to fail is to look like a marketer Reddit’s current spam rules explicitly prohibit repeated or unsolicited mass engagement. Examples include mass-posting repetitive material for financial gain, sending large amounts of unsolicited chats or private messages, and using bots or generative-AI tools in ways that facilitate spam. Its User Agreement, effective July 1, 2026, also says you cannot sell or transfer a Reddit account without prior written approval. And Reddit now explicitly prohibits scraping the service without prior written consent. That kills several popular Reddit growth hacks: Buying aged accounts. Running a farm of backup accounts. Automatically scraping Reddit at scale. Mass-generating replies. Automating promotional DMs. Good. Because none of those is necessary for the model that actually matters. The sustainable version looks almost backwards: Automate the work behind Reddit. Keep the behaviour on Reddit human. Reddit Pro gives you the listening layer Reddit already provides businesses with a free tool for this. Reddit Pro can track keywords and phrases related to your company, products, competitors and market. Its Trends feature surfaces relevant communities, conversations and mention volume. Reddit Pro also includes organic performance data. Reddit Pro gives businesses a free listening and organic performance layer inside Reddit That means your research can begin with terms such as: - alternative to [competitor] - how do I [problem] - what do you use for [task] - struggling with [problem] - best tool for [job] - is [product] worth it - anyone recommend [solution] Do not search only for your product category. Search for pain language. A customer may never type workflow automation platform. They might type: I waste every Friday making the same report. That is the sentence you want. The useful signal is usually hidden inside ordinary customer language. The five levels of buying intent Not every Reddit complaint is a sales opportunity. Use this simple ladder: 1. Curiosity I wonder how people do this. 2. Frustration I hate doing this. 3. Advice seeking How are other people doing this? 4. Solution seeking What tool or service should I use? 5. Purchase intent I need this now. Which option should I buy? Levels 4 and 5 matter most. The person already understands the problem. You do not need to manufacture urgency. Your job is to understand whether you can genuinely help. That distinction is the basis of the entire playbook. You are not trying to turn everyone on Reddit into a customer. You are trying to find the small number of conversations where problem, timing, expertise and offer intersect. Wayfair provides a useful example. Its team used Reddit Pro to find relevant discussions and answer shopping questions rather than simply distributing promotional links. Reddit reports that Wayfair subsequently saw steady month-over-month referral traffic growth and more than a 50% increase in Reddit profile followers. The lesson is not that every business will reproduce those results. It is that help can become distribution. We now have the strategy. The next question is exactly how to build it. The AI stack: start with three tools, not ten You can run almost the entire first month with: Reddit Pro + Google Sheets + one AI assistant. My default recommendation for beginners is ChatGPT. Claude and Gemini can run essentially the same workflow. Hermes comes later. Do not use four AI tools because four sounds sophisticated. Choose one. If you use ChatGPT Create a Project called: Reddit Revenue Research Add your research spreadsheet, product notes, community rules, useful posts and any customer interviews. ChatGPT Projects keep chats, files and project-specific instructions together, which makes them well suited to ongoing research rather than starting from zero every session. If you use Claude Create the same workspace as a Claude Project. Claude Projects support a project knowledge base, uploaded documents and project instructions. Claude can also use retrieval when the project knowledge grows larger. If you use Gemini Create a custom Gem called: Reddit Opportunity Analyst Give it your instructions and add research files under Knowledge. Gemini Gems can also reference files from Google Drive, which is useful if your research sheet already lives there. The workflow does not change. The model is not the business. The process is. Days 1–3: Pick one person and one painful problem Do not begin with: I want to sell an AI course. Begin with: I want to help independent recruiters who spend hours turning interview notes into candidate summaries. Create a spreadsheet with: Date | Subreddit | Problem | Exact Phrase | Intent 1–5 | Current Solution | Why It Failed | Link | Possible Offer Then ask your AI: I want to research a market on Reddit. My target customer is: [WHO] The broad problem I understand is: [PROBLEM] Give me 25 phrases this person might naturally use on Reddit when: 1. frustrated 2. asking for advice 3. comparing solutions 4. actively looking to buy Avoid industry jargon unless a real customer would use it. Group the phrases by buying intent. Add the best terms to Reddit Pro Trends. Instead of guessing what customers care about, track the phrases they are already using. Your first goal is not traffic. It is vocabulary. Days 4–7: Build your community map Find five to ten relevant communities. For each one, record: Who participates What gets discussed What gets upvoted What gets removed Whether promotion or links are allowed How people write What problems appear repeatedly Reddit Pro can help identify communities associated with the terms you track. Find where the customer actually spends time before deciding where to post. Read the rules yourself. Do not ask AI to guess them. A post that is welcome in one subreddit may produce an immediate removal in another. By Day 7 you should know where your customer talks, what they sound like and which communities are worth your attention. Days 8–14: Help before you sell Now participate. Aim for a few genuinely useful replies each day. Not 30. Not 100. Three excellent answers are more valuable than 50 AI-shaped comments that sound almost identical. Whenever you see a relevant problem, add it to your research sheet. Keep the exact wording. Do not convert every human sentence into marketing language. Instead of recording: Customer requires workflow optimization. Save: I spend half of Monday fixing this manually. That sentence might later become your product headline. What Comes Next: Turn Research Into Revenue You’ve spent 14 days finding real problems, real language, and real buying signals. Now comes the part that matters most: using AI to identify the opportunity worth pursuing, build the offer, turn Reddit conversations into customers, and measure whether it actually makes money.

Web

7
00:00

Stripe Dates The Singularity To Jan. 1 And Prices It At $8 Billion

Stripe told investors the AI era officially began on January 1, and is now pricing deals on that belief. First-half revenue rose 41% and free cash flow 43%, and the payments firm confirmed it is buying AI routing startup OpenRouter for about $8 billion. The company cites a surge in new business applications, $1.9 trillion in payments processed in 2025, and says 88% of the Forbes AI 50 builds on its platform. Cofounder Patrick Collison floated the singularity idea with a wink in April but now presents it without hedges; skeptics note the evidence doesn't match any real definition of the term.

Notes

Stripe investor letter (Aug 2026): declares Jan. 1, 2026 the start of the singularity. Evidence: sharp inflection in long-run trends, led by new-firm formation. Census (a week earlier): July business applications 578,926 (+8.1% MoM); projected employer formations from that cohort +0.7% to 29,959.

Financials: H1 revenue +41% YoY, free cash flow +43%. 88% of the Forbes AI 50 (incl. OpenAI, Anthropic) build on Stripe; revenue share from AI/crypto customers more than doubled. Signed by Patrick & John Collison and president William Gaybrick.

Genesis of the claim: In Feb. 2026 Collison told TBPN Q1 2026 might be remembered as "the first quarter of the singularity," adding the claim could look "completely delusional" in hindsight. At April's conference it was "partly tongue in cheek." By August it's in a formal letter, unhedged, as an eight-month operating premise.

OpenRouter acquisition (announced same day): NYT reports $7.5B vs. the router's $1.3B May valuation; Axios says above $8B. Founders reportedly collect $1.5B — more than the whole company was worth three months prior. "The singularity thesis expressed as a price."

Ownership logic: private capital funds M&A without dilution; share count lower than three years ago; share price compounding 31% since Series D vs. 14% S&P 500. Feb. tender valued Stripe at $159B; pursuing PayPal with Advent at $53B; IPO on indefinite hold.

Criticism: Gary Marcus (July) — none of the cited markers satisfy any coherent definition of "singularity." Vinge's 1993 original: superhuman intelligence ending the human era.

Counter/limitations: $1.9T 2025 payment volume (+34%), 5M+ companies; 2025 cohorts show higher formation and better per-business performance — but "unaudited, self-reported," and Stripe's valuation depends on the AI economy being real. Watch: (1) Census projected formations within four quarters (nearly flat this year despite applications climbing); (2) Stripe publishing cohort data; (3) OpenRouter price is a single buyer's conviction.

Full text · 4,114 chars
Stripe told investors on Wednesday that Jan. 1 marked the beginning of the singularity. The evidence offered was a large inflection in long-run trends, with a sharp rise in the rate of new firm creation as the primary example. Census Bureau figures published a week earlier put July business applications at 578,926, up 8.1% from June, while projected formations of employer businesses from that same cohort rose 0.7%, to 29,959. First-half revenue rose 41% year over year and free cash flow rose 43%. Stripe said 88% of the Forbes AI 50, including OpenAI and Anthropic, build on its platform, and that revenue share from AI and crypto customers more than doubled. Cofounders Patrick and John Collison signed the letter alongside William Gaybrick, president of technology and business. In February, Patrick Collison told TBPN there was a reasonable chance Q1 2026 would be remembered as the first quarter of the singularity, adding that in hindsight the claim might look completely delusional. At Stripe’s April conference, he used it again and conceded he was being partly tongue in cheek. By August, it appears in a formal investor letter with no hedge attached, as an operating premise the company says it has run on for eight months. The letter went out the same day Stripe confirmed it would acquire OpenRouter. The New York Times reported a price of $7.5 billion compared with the routing startup’s $1.3 billion valuation in May, and Axios put the figure above $8 billion. Founders will reportedly collect $1.5 billion, more than the entire company was worth three months earlier. That is the singularity thesis expressed as a price, and it establishes a comparable every seed investor in AI infrastructure will now cite in a partner meeting. The letter argues that private ownership lets the company fund acquisitions without diluting holders, notes that share count is lower than three years ago despite significant M&A, and reports share price compounding at 31% since the Series D compared with 14% for the S&P 500. A February employee tender valued Stripe at $159 billion. It is separately pursuing PayPal with Advent International at $53 billion. An IPO stays on indefinite hold. Declaring a phase change is a coherent way to ask employees and investors holding illiquid stock to keep waiting. Skeptics have been pushing back on this exact substitution; Gary Marcus argued in July that none of the markers now cited, whether a benchmark score or a company growing faster, satisfy any coherent definition of the term. Vernor Vinge’s original 1993 formulation described superhuman intelligence bringing the human era to an end. Accelerating revenue at a payments processor is a different claim wearing the same word. Stripe’s defense is that it sees data almost nobody else does. Businesses on its platform processed $1.9 trillion in payment volume in 2025, up 34%, across more than five million companies. Collison has described a phase transition in which 2025 cohorts show both higher formation counts and better per-business performance, which is the more interesting and less quotable finding. Cohort quality is harder to manufacture than application volume. It is also unaudited, self-reported, and produced by a company whose valuation depends on the AI economy being real. Three things are worth watching rather than accepting: First, the Census projected formation series will show within four quarters whether the surge in applications translates into new employers; that measure has remained nearly flat this year even as applications climbed. Second, Stripe’s claims about stronger cohorts can be tested only if the company publishes the underlying data rather than simply describing it. Third, the OpenRouter deal creates a new reference point for AI infrastructure valuations—but one based on a single buyer’s conviction. For anyone allocating capital, Stripe’s first-half financials are disclosed, specific, and consistent with an AI-driven expansion of its customer base. The singularity label is a narrative device that a CEO first offered with a wink and has since promoted to doctrine.
00:00

Reddit Nearly Vanishes From ChatGPT Citations After OpenAI Search Change, Report Suggests

Reddit almost vanished from ChatGPT's search citations after OpenAI quietly changed how ChatGPT searches the web, with Reddit's share collapsing from 3.8% of citations to 0.5% in a matter of days. On August 8, ChatGPT began using site-scoped searches at scale — those jumped from 0.37% to 16.8% of background queries overnight — and Reddit's share cratered around August 14, an 86% relative decline from its April peak as ChatGPT's most-cited domain. The drop appears specific to ChatGPT: Reddit's share in Google's AI Overviews slipped only gradually, from about 2.5% to 2.1%. OpenAI hasn't explained the change, and the episode is being read as a warning that visibility inside AI answers sits on infrastructure platforms can rewire overnight without notice.

Notes
Reddit's ChatGPT citation share collapses after OpenAI search change
  • Reddit's share of ChatGPT Search citations fell from a 3.8% average (July 18–Aug 7) to 0.5% — an 86% relative decline. By Aug 14 it was under 1%; Aug 14–17 averaged 0.52%. Peak: 4.14% in April, the single most-cited domain.
  • The trigger was technical. On Aug 8, PromptWatch observed ChatGPT Search using the site: operator at scale: domain-scoped queries jumped from 0.37% → 16.8% of all background searches in one day (~46x). Average searches per response nearly doubled, 1.08 → 1.83. Instead of crawling the open web, ChatGPT now often queries specific sites directly.
  • OpenAI never explained the change and didn't respond to Gizmodo's request for comment. PromptWatch cautions its data shows when changes happened, not why, and that a data-collection issue can't be fully ruled out. A Reddit spokesperson said Reddit still ranks among the most-cited domains per other reports and doesn't rely on LLMs for traffic (most visits come direct/from traditional search).
  • Collapse looks ChatGPT-specific: Reddit's share in Google AI Overviews only drifted ~2.5% (early July) → ~2.1% (August); similar slow drift in Google's AI Mode.
  • Framed as a GEO (generative engine optimization) case study: one unannounced backend change reshuffled citations across millions of queries. LinkedIn commenters called AI search "a constant moving target," urging brands to "stop building strategies around holes in the system."
  • Precedent: Sept 2025 saw a similar Reddit citation collapse, later attributed by analysts to Google removing a search parameter third-party trackers relied on — not a deliberate decision.
  • Reddit joined the S&P 500 (Aug 18); shares +13% on inclusion news Aug 14 (the day citations cratered), gave back ~8% Monday; traded ~$153–156 midweek vs. 52-week high $282.95. Q2 revenue +61% YoY to $805M, though shares fell on flat sequential U.S. user growth and search referral headwinds.
  • Reddit is leaning into AI: experimenting (per The Verge) with turning posts/comments into AI-generated short videos and podcasts with a "Read/Play" toggle, teed up on its earnings call.
"visibility inside AI answers sits on infrastructure that platforms can rewire overnight, without notice."
Full text · 4,429 chars
Reddit’s share of ChatGPT citations collapsed from 3.8% to 0.5% in days, after OpenAI changed how ChatGPT searches the web. Is ChatGPT leaving Reddit on read? Reddit's share of ChatGPT Search citations has plummeted to just 0.5% of responses, according to a Gizmodo report based on data from AI visibility platform PromptWatch. The drop was sudden. Reddit held a steady 3.8% average share of ChatGPT citations from July 18 through August 7 — one of the largest of any domain on the web. By August 14, it had fallen below 1%, and the August 14–17 average of 0.52% represents an 86% relative decline. For context on how far that is from Reddit's peak: as recently as April, Reddit was the single most-cited domain in ChatGPT Search, accounting for 4.14% of all citations. What Changed On August 8 The slide traces back to a technical shift in how ChatGPT searches the web. On August 8, PromptWatch observed that ChatGPT Search began using the "site:" operator at scale — queries scoped to a specific domain jumped from 0.37% to 16.8% of all background searches within a single day, a roughly 46x increase. At the same time, the average number of searches ChatGPT runs per response nearly doubled, from about 1.08 to 1.83. In plain terms: instead of primarily searching the open web and seeing what comes back, ChatGPT now frequently goes directly to specific sites to pull information. Reddit's citation share slipped from the high 3% range to the mid-2% range that same day, then collapsed on August 14. OpenAI has not explained the shift and did not respond to Gizmodo's request for comment. PromptWatch itself is careful to note that its data shows when each change happened, not why — and that a data-collection issue cannot be fully ruled out. A Reddit spokesperson told Gizmodo the platform remains one of the most-cited domains according to other reports, and noted that Reddit doesn't rely on LLMs for traffic, receiving most of its visits directly and through traditional search. Notably, the collapse appears specific to ChatGPT. Reddit's citation share in Google's AI Overviews declined only gradually, from about 2.5% in early July to roughly 2.1% in August, with a similar slow drift in Google's AI Mode. A Warning Shot For The AI Search Industry The episode lands as a case study in what marketers now call GEO — generative engine optimization, the practice of getting cited inside AI answers rather than ranked on a results page. A single unannounced backend change reshuffled the citation landscape for millions of queries, and publishers learned about it the way Reddit did: by watching the numbers drop. The marketing community took notice. Commenters on LinkedIn called AI search "a constant moving target" and urged brands to "stop building strategies around holes in the system." There is precedent for the anxiety. In September 2025, several AI-visibility trackers reported a similar collapse in Reddit's ChatGPT citations — and Reddit's stock moved with the coverage. That earlier drop was later attributed by some analysts to Google removing a search parameter that third-party data providers relied on, rather than to any deliberate decision. The lesson from both episodes is the same: visibility inside AI answers sits on infrastructure that platforms can rewire overnight, without notice. This week, Reddit joined the S&P 500, effective August 18. Shares jumped 13% on the inclusion news on August 14 — the very day the citation share cratered — before giving back about 8% on Monday as investors locked in gains. The stock traded around $153–156 midweek, well off its 52-week high of $282.95. The business, meanwhile, is hardly starving. Reddit's Q2 revenue grew 61% year over year to $805 million, though shares fell after the report on flat sequential U.S. user growth and continued search referral headwinds — a reminder that Reddit's relationship with the platforms that distribute its content remains its most-watched vulnerability. And Reddit is leaning further into AI itself, not away from it. Per The Verge, the company has begun experimenting with turning posts and comments into AI-generated short videos and podcasts, complete with a "Read/Play" toggle — part of a push, teased on its recent earnings call, to let people listen to posts in the background. Which leaves the two companies in a curious dance: OpenAI’s chatbot citations and Reddit’s work to sound more like a chatbot every day.
00:00

Modular Delivers Openness And Accelerator Portability At ModCon 2026

Qualcomm's AI software arm made good on its openness promise, releasing the Mojo programming language and compiler as open source under an Apache 2.0 license just three weeks after closing its roughly $3.1 billion all-stock acquisition of Modular. At the ModCon 2026 conference, Modular also made its Modular Cloud inference service generally available, routing AI workloads across six kinds of chips including Nvidia, AMD, Apple, AWS Trainium, Google TPU and Qualcomm's own datacenter accelerators under one programming model. The pitch is that new chips get software in months, not engineer-years — five engineers brought up AWS Trainium in under five months, and Qualcomm's AI200 ran a model two weeks after hardware access. Modular claims the same model ran about 90% faster on AMD's MI355X than an Nvidia B200 at roughly half the hourly cost, but those figures are unverified vendor numbers.

Notes
Modular at ModCon 2026 — Open-Source Mojo, Modular Cloud GA, Six-Vendor Chip Support

Author: Patrick Moorhead (Moor Insights & Strategy; Qualcomm is a client — disclosed). Article date 2026-08-20. Conference in San Francisco, ~3 weeks after Qualcomm closed its Modular acquisition (announced all-stock June 24 at ~$3.9B; closed July 28; SEC filing pegs actual value at ~$3.1B for 18M shares). Modular CEO Chris Lattner (LLVM co-creator, Apple Swift) is now Qualcomm EVP running advanced AI software and platforms.

Five announcements
  • Mojo language + compiler fully open-sourced under Apache 2.0, effective immediately.
  • Modular Cloud inference service reached general availability.
  • Platform support expanded beyond Nvidia/AMD/Apple Silicon to AWS Trainium, Google TPUs, and Qualcomm's datacenter accelerators — six chip vendors, one programming model.
  • MAX license dropped device-usage restrictions; Qualcomm CEO Cristiano Amon committed to an alliance program by year-end.
  • Mojo coming to Windows with Microsoft's help.
Openness: the two tests
  • Colleagues Matt Kimball and Bill Curtis set the stakes pre-deal: open-source clothing masking a lock-in, with tests being "when a [Qualcomm] Hexagon [NPU] backend ships, and whether MAX stays open to competing silicon."
  • Hexagon: Qualcomm's new datacenter chips are built on its Hexagon lineage and run the Modular stack; a Google Gemma model ran on the AI200 within two weeks of hardware access.
  • License: Apache 2.0 plus a written commitment to optimize for hardware "including hardware that competes directly with Qualcomm Technologies' platforms."
  • Amon, on stage and repeated to the author directly: > "We did not acquire Modular to make it a Qualcomm-only software."
  • Author's read: multiple bidders offered more; openness was "the thesis of the deal, not a concession." Lattner's pitch: faster mission on a bigger platform.
What Qualcomm gets
  • Day-one production software for its datacenter chips; a software layer across its whole edge portfolio; usage-based services revenue via Modular Cloud. Bear case: "one of the strongest AI software teams ever assembled."
  • Indirect: position at the token-arbitration layer (revenue on every token, demand visibility across competing vendors' chips), leveling Nvidia's software moat and shifting competition to performance-per-dollar-per-watt. Completes Amon's developer-first strategy (Edge Impulse, Arduino, Modular). Qualcomm projects it'll be the largest automotive chip supplier within two years.
Modular Cloud — already in production
  • Described as "air traffic control for AI." Nvidia entered adjacent territory days earlier with NeMo Switchyard (model routing); hardware routing at production quality is the differentiated part.
  • Ran live for months on OpenRouter under the stealth name "ModelRun"; Modular claims top-tier speed rankings, echoed by independent benchmark firm Artificial Analysis.
  • Customers: MiniMax (flagship model, dedicated deployment, billions of tokens/min), Hippocratic AI (>30% gains for voice agents), Jane Street (paying customer).
  • Chip bring-up arithmetic (vendor numbers, flagged as such): Amazon's Trainium enabled by 5 engineers in under 5 months (~25 engineering-months vs. "several thousand" traditionally); Google TPU in ~4 months mostly by Modular; Qualcomm AI200 running a model in 2 weeks. AMD committed on stage to a multi-generation partnership; d-Matrix signed on.
Cost claims (unverified)
  • Same model on AMD MI355X vs. Nvidia B200: ~90% faster throughput at ~half hourly cost → >45% lower TCO. "Stunning" per author; test details unpublished, not validated by Signal65 (his benchmarking lab). Treat as credible vendor claims pending third-party verification.
  • Hardware strategy targets the memory crunch: AI200 ships this year with cheaper-than-HBM memory; AI250 in 2027, memory-first design avoiding scarce advanced packaging. Author: "at least a dozen chip design companies" rearchitecting to use less memory; shortages expected to persist years.
  • Second-order: cross-chip software lowers financing risk for mixed AI fleets → potentially cheaper capital for datacenter builders.
Caveats and open questions
  • Mojo is open source; MAX is source-available under a community license — "not the same thing." Conditions remain on telemetry, branding, building substitutes; no community governance.
  • Modular Cloud itself is proprietary — Apache 2.0 "constrains the programming language; it doesn't constrain the cloud." Alliance program terms (promised year-end) will show whether partners get "real governance or merely a logo slide."
  • Three distinct milestones: announced support ≠ production readiness ≠ cloud availability. Qualcomm Cloud AI100 serves today; Google and Amazon chips run models but aren't in the cloud service yet.
  • ~100 engineers support the whole stack — "either the best leverage in infrastructure software or an area of fragility."
  • Buyers should pilot on ≥2 hardware types and keep weights/prompts/deployment portable as an exit path.
Bottom line (author's position)

Qualcomm "traded a lock-in nobody would have trusted for a position everyone can use" — it must help competitors succeed to get paid. "The burden of proof has shifted from the people who believe in an AI software alternative to the people who don't." Author states he's "rooting for Qualcomm and Modular."

Full text · 15,418 chars
I spent Tuesday in San Francisco at ModCon 2026, Modular’s developer conference and its first event since Qualcomm acquired it just a few weeks ago. Going in, I posted on X that when I ask people about Modular, I get “bookend” responses with no middle ground: One end believes Modular is building the most important software layer in AI, while the other points to the graveyard of companies that tried to break the AI software moat. I left the conference believing that the skeptics now have the harder case to make. Qualcomm gets plenty in return for the $3 billion-plus in stock it spent to buy Modular: day-one software for its own chips from the datacenter out to the edge, a new services business and a position in the layer that decides where AI workloads run — even when a Qualcomm rival’s chip wins. (Disclosure: Qualcomm is a client of my firm, Moor Insights & Strategy, as are other companies cited here, including AMD, AWS, Google, Microsoft and Nvidia. I attended ModCon at the company’s invitation. The analysis is entirely my own.) ModCon delivered three things the industry has been waiting on. First, Qualcomm legitimated its open source commitment within three weeks of closing the acquisition, and with an actual license rather than a promise. Second, Modular showed working evidence that its platform can radically lower the cost of bringing new AI chips to market. Third, it lowered the cost of compute for the token consumers who can arbitrate not only across models, but now chips. Real unknowns remain, and I’ll name them, but this was a stronger day than I expected — and I already expected a good one. What Modular And Qualcomm Actually Announced At ModCon, Modular and Qualcomm announced five things: - The Mojo programming language and its compiler went fully open source under Apache 2.0, effective immediately. - Modular Cloud, the company’s AI inference service, became generally available. - Platform support expanded beyond Nvidia, AMD and Apple Silicon to AWS Trainium, Google TPUs and Qualcomm’s own datacenter accelerators — six kinds of chips from six different vendors under one programming model. - The MAX license dropped its device usage restrictions, alongside an alliance program that Qualcomm CEO Cristiano Amon committed to launching by year end. - Mojo is coming to Windows with Microsoft’s help. It’s worth recalling the short timeline of this acquisition. Qualcomm announced the Modular deal on June 24, an all-stock transaction announced at roughly $3.9 billion, and closed it on July 28. Qualcomm’s quarterly SEC filing puts the actual transaction value at about $3.1 billion for 18 million shares, since all-stock deal math moves with the stock price. Modular CEO Chris Lattner, co-creator of the LLVM compiler infrastructure and Apple’s Swift language, is now a Qualcomm executive vice president running advanced AI software and platforms. Qualcomm Answered The Openness Question In Week Three Every acquisition of a neutral platform by a hardware vendor raises the same question: Does it stay neutral? My colleague Matt Kimball framed the stakes in his article from late June about Qualcomm’s investor day: If Modular’s openness narrowed toward Qualcomm silicon, it would become “a lock-in play wearing open-source clothing.” When my colleague Bill Curtis wrote about the deal 10 days later, he set two tests: “when a [Qualcomm] Hexagon [NPU] backend ships, and whether MAX stays open to competing silicon.” I asked the same question in my Week Ahead video going into the event, and the developer community was asking it right up to the week before. ModCon answered on both fronts. The Hexagon question was resolved in silicon: Qualcomm’s new datacenter chips are built on its Hexagon processor lineage, and they now run the Modular stack, with a Google Gemma model working on the AI200 within two weeks of hardware access. The openness question was resolved through licensing. Releasing under an Apache 2.0 license isn’t something that can be quietly walked back, and Modular’s announcement commits in writing to optimizing for hardware “including hardware that competes directly with Qualcomm Technologies’ platforms.” Qualcomm’s Amon left no ambiguity with what he said on stage, and when I sat down with him and Lattner afterward, he repeated it to me directly: “We did not acquire Modular to make it a Qualcomm-only software.” He framed the moment to me as Android and Kubernetes arriving for AI at the same time. Lattner told me the pitch that won him over was achieving Modular’s mission faster on a bigger, more open platform. I believe that there were multiple bidders with higher offers, making openness the thesis of the deal, not a concession. Trust was what Modular actually needed to ship, and in the software world a license is the only durable form of it. You can’t un-open-source a compiler. What Qualcomm Gets From The Modular Deal, Directly And Indirectly The obvious question is why Qualcomm would pay billions in stock for a software company and then give away its most valuable asset three weeks later. The direct answer: Qualcomm’s datacenter chips now have production software from day one, and the software question that had dogged every Qualcomm datacenter conversation is now answered. It also gets a software layer it can take across its entire edge AI portfolio. As Matt Kimball argued in June, a credible software layer de-risks Qualcomm’s entire hardware roadmap in one move. Modular Cloud gives Qualcomm usage-based services revenue it never had. And in the bear case, Qualcomm still walks away with one of the strongest AI software teams ever assembled. The indirect benefits could be bigger. Modular Cloud puts Qualcomm at the arbitration layer of AI computing: It earns revenue on every token served and sees demand across every vendor’s chips, including in deployments Qualcomm hardware hasn’t won. Opening the stack also levels the playing field exactly where Nvidia has an advantage in software. This shifts competition to performance per dollar per watt — the fight that Qualcomm spent two decades training for in mobile. And it completes the company’s developer-first repositioning that Amon walked me through as he discussed its acquisitions from Edge Impulse to Arduino to Modular. One stack for everything from earbuds to datacenter racks makes every Qualcomm socket more valuable, and the edge is where Qualcomm is strongest. A hat tip to Cristiano Amon is in order. He has spent 30 years at Qualcomm making big bets that the market doubted, and I remember when few people believed in the diversification strategy that now has the company tracking, by its own projection, to be the largest automotive chip supplier within two years. Paying more than $3 billion in stock for a crown jewel and then fully open-sourcing a layer in week three takes a strategic self-confidence that most acquirers never muster. Modular Cloud Could End Up As The Biggest Win I believe the ModCon 2026 announcement that will matter most commercially is Modular Cloud. Think of it as air traffic control for AI: the cluster, not the chip, is the computer, and the platform decides which silicon serves each request. Developers see one standard interface; Modular handles the hard engineering underneath. Routing requests to the right model is a crowded field, one that Nvidia entered days earlier with its NeMo Switchyard router. Routing them to the right hardware at production quality is the part almost nobody else has attempted. I posted the full value-proposition slide live while the keynote was underway, and the more I reflect on it, the more I believe “audacious” is the right word for it. What separated this launch from a typical debut is that it wasn’t really a launch, because the product is already in service at scale. Modular disclosed that it has been serving live traffic for months on OpenRouter, a popular marketplace for AI models, under the stealth name ModelRun; Modular says its endpoints consistently ranked at or near the top for speed, a story that independent benchmarking firm Artificial Analysis echoes. AI model maker MiniMax runs its flagship model on a dedicated Modular deployment serving billions of tokens per minute; healthcare AI company Hippocratic AI reports better than 30% gains for its voice agents run via Modular; and Jane Street, a trading firm for whom microseconds are money, is also a paying customer. This isn’t a research project. The real kicker for the Qualcomm/Modular platform approach is the radically lower cost of bringing up new chips. Enabling a full AI software stack on new silicon traditionally takes hundreds of engineers and engineer-years; that is the source of CUDA’s gravity in the market. Modular’s counterargument is arithmetic. Using Modular, five engineers brought up Amazon’s Trainium chip in under five months, roughly 25 engineering-months against the several thousand a traditional effort consumes. Modular created the Google TPU version largely on its own in about four months. Qualcomm’s AI200 went from hardware access to running a model in two weeks. These are vendor numbers, so I treat them with appropriate skepticism. But even heavily discounted, a tenfold reduction in enablement cost changes who can afford to ship credible AI chips. At ModCon, AMD committed on stage to a multi-generation partnership, and chip startup d-Matrix signed on, too. What Modular’s Approach Does To AI Economics During the event, Modular put hard numbers behind Modular Cloud’s performance. Holding deployment size and response-time targets constant, the company showed the same model on AMD’s MI355X chips delivering roughly 90% faster thoughput of an Nvidia B200 setup — at about half the hourly cost. That’s more than 45% lower total cost of ownership. I called the numbers “stunning” live. Mind you, the test details weren’t published, and none of it has been validated by Signal65, the independent benchmarking lab I co-founded. So please treat every figure here as a credible vendor claim awaiting third-party verification. This engineering direction aligns with everything I’m hearing about memory use. I posted during the event that, in the context of the ongoing memory crunch across the industry, at least a dozen chip design companies have told me they’re rearchitecting to use less memory, and that they believe memory supply will stay short for years. Qualcomm’s datacenter chips aim squarely at that constraint: The AI200 is slated to ship this year with less expensive memory per card than typical GPUs using HBM, and the AI250 arrives in 2027 with a memory-first design that, as I discussed on Yahoo Finance, improves bandwidth and energy use without needing the advanced chip-packaging capacity that’s in shortest supply. A software layer that routes AI work across many kinds of chips makes that hardware diversity deployable rather than merely announced. Here’s a second-order effect nobody else is talking about. AI datacenter builders are financing tens of billions of dollars of infrastructure, and lenders price that debt on risk, which today means anything other than the consensus Nvidia GPU. A credible cross-chip software layer changes that math. If the same models and applications run smoothly across Nvidia, AMD, and purpose-built accelerators, then a mixed fleet becomes a lower-risk asset, and lower risk should eventually mean cheaper capital for the people building AI factories. I believe the abstraction layer is quietly a financing story, too, not just an engineering one. What I’m Keeping An Eye On For Modular And Qualcomm I’d be doing readers a disservice if this article read like a final victory lap, so let’s be candid about the areas for further consideration, starting with the license rather than the press release. Mojo is now open source; but MAX, the layer above it, is source-available under a community license, and those are not the same thing. The old device limits are gone, which is real progress, but conditions remain on telemetry, branding and building substitutes, and source access is not the same as community governance. The bigger customer risk could be a lock-in moving up the stack: Modular Cloud, the layer that decides which chips get your workloads, is proprietary. Apache 2.0 constrains the programming language; it doesn’t constrain the cloud. The alliance program’s actual terms, promised by year end, will show whether partners get real governance or merely a logo slide. We saw a lot of work completed, and it’s clear there’s more to come. And with that work, execution needs to be flawless. Per Modular’s own materials, Qualcomm’s Cloud AI100 is serving through Modular Cloud today, while the Google and Amazon chips run models but aren’t in the cloud service yet. Production is expected within months, but it’s worth remembering that announced support, production readiness and cloud availability are three different milestones. I believe roughly 100 engineers support all of it, which is either the best leverage in infrastructure software or an area of fragility waiting for a bad quarter. If they can pull off the rest of it with 100 engineers, all of them should be given raises. Meanwhile, Nvidia isn’t standing still, and its advantage is rooted in a decade of tooling and industrywide habits. Modular’s cost claims in comparison to CUDA need independent validation, and smart buyers will pilot real workloads on at least two kinds of hardware and measure cost, speed and engineering effort before standardizing on anything. They could also maintain an exit path by keeping model weights, prompts and deployment automation portable to ensure that leaving remains a cheap alternative. Open Model, Open Hardware And More Open Software Make Modular A Real Option Four and a half years ago, Modular bet that AI inference would become the dominant workload, that AI chips would diversify and that the industry needed one software layer to make that diversity usable. Every one of those bets now looks like consensus, and at ModCon the company showed a production platform, a public cloud, an open compiler and six chip vendors to back it up. For Qualcomm, the multiple types of payoffs from this acquisition is what makes it a smart move by Cristiano Amon. The floor is adding a world-class software team and day-one software for its own datacenter chips, and moving out to the edge. The ceiling would be owning the neutral layer the AI industry runs on, earning money on every token served no matter whose silicon wins the socket. To enable this, Amon traded a lock-in nobody would have trusted for a position everyone can use, and he accepted the strange math at the deal’s core: Qualcomm now has to help its competitors succeed to get paid. Most CEOs can’t stomach that trade; the smart ones recognize it’s the only one that works if you want to be a trusted platform. I confess that I’m rooting for Qualcomm and Modular to pull it off, because while I have nothing but respect for what Nvidia has accomplished, the industry needs more platform diversity. But platform trust is never won permanently, and the potential pitfalls I’m watching are the MAX license, the alliance terms, independent validation of Modular’s numbers and whether the bring-up economics hold for the next chip. But the burden of proof has shifted from the people who believe in an AI software alternative to the people who don’t.
00:00

Microsoft's GitHub Under Siege As SpaceX's Cursor Takes The AI Stack

Cursor, the AI coding editor now owned by SpaceX, opened a beta of Origin — its own code-hosting platform that competes head-on with Microsoft's GitHub. Origin launched the same morning a global GitHub outage dragged on for over six and a half hours with error rates approaching 20%, which made Origin's pitch land harder than any marketing. Roughly 35% of Cursor's merged pull requests are now written by autonomous agents, and Origin ships with day-one integrations for Vercel, Depot and Buildkite, so Cursor is assembling the whole development pipeline. SpaceX reportedly bought Cursor in an all-stock deal worth around $60 billion, but Origin's data-residency and training terms are still unpublished, and it's opt-out by default for paid users.

Notes

GitHub vs. Cursor/Origin (Forbes, 2026-08-20)

GitHub's scale: 180M+ developers, new user every second, 630M+ repositories; 90%+ of Fortune 100 use it. Recorded 257 incidents between May 2025 and April 2026 (Incident-tracking data); Microsoft's own CTO acknowledged "the platform was not built for the scale it is now being asked to handle."

The outage: On Monday (launch morning), a "global degradation" lasted 6h 42m, pushing error rates to ~20% on PRs, issues, and API, and ~50% on archive/raw downloads. Cursor "did not plan the timing" — the outage made Origin's case better than its launch.

Origin launch: Cursor opened beta of Origin, its own code host, to paid users on Monday. Day-one integrations with Vercel, Depot, and Buildkite (deploy/build/test). Cursor acquired code-review startup Graphite in December 2025 (undisclosed; reportedly above its $290M last valuation). Cursor is now owned by SpaceX after a reported $60B all-stock acquisition.

Agent logic: ~35% of Cursor's merged PRs are now written by autonomous agents. Origin is "built for an era where software agents are writing much of the code themselves."

Caveats (the counter-view): Origin is beta, gated to paid Cursor users, rolled out opt-out by default (not opt-in); retention, residency, and data-training terms unpublished — so some devs may already be hosting proprietary code on a SpaceX-owned platform with no written data terms. Enterprises won't migrate on a bad-outage day. Watch two numbers: whether Origin publishes governance terms and whether GitHub's incident count keeps climbing.

Advice for leaders: ask "if our code host disappeared for 7 hours, what would stop?"; map which vendors own multiple stack layers. "Concentration is convenient right up until it is a dependency."

Full text · 4,328 chars
The answer was GitHub. And the question was “Where to store your company’s code?” You wrote it, you put it on GitHub, and you got on with your day. The engineering teams treated it like plumbing that was essential, invisible, and never on the agenda. GitHub has over 180 million developers worldwide, with a new user joining every second and over 630 million total repositories. Over 90% of Fortune 100 companies use the platform. This week Cursor asked a question the industry has been avoiding. What happens when developers stop thinking of GitHub as the center of Software Development? On Monday, Cursor opened the beta of Origin, its own code hosting platform, to paid users. GitHub picked that same morning to have one of its worst days in recent memory. A global degradation dragged on for six hours and 42 minutes, pushing error rates nearly 20% percent across pull requests, issues, and the API and nearly 50% percent to archive and raw downloads. Cursor did not plan the timing. The outage made Origin’s argument better than any launch post could. What GitHub Does, In Plain Terms If you have never touched a repository, here is the short version. A code host is the shared filing cabinet for software. It stores every version of a project, tracks who changed what, and gives teams a place to review each other’s work before it goes live. GitHub, owned by Microsoft, is that filing cabinet for most of the world’s developers. It has recorded 257 incidents between May 2025 and April 2026, according to analysis by Incident tracking data, and its own chief technology officer acknowledged the platform was not built for the scale it is now being asked to handle. Challenger to GitHub? Why The Origin Launch Is Big The breaking story is that AI companies no longer want to sit on top of existing developer infrastructure. They want to own it. Cursor already controls the editor where code gets written. It bought the code review startup Graphite in December 2025 for an undisclosed price which reportedly went above its last valuation of $290 million. Now it controls where code lives. Per the Cryptonomist, Origin launched with day- one integrations with Vercel, Depot, and Buildkite, the companies that handle deployment, builds and testing, which means Cursor is assembling the entire pipeline, not one link in it. The logic makes sense once you see who is really pushing the code. Origin is built for an era where software agents are writing much of the code themselves, making changes at a frequency human-centered platforms were never designed to absorb. The logic makes sense once you see who is really pushing the code. At Cursor itself, roughly 35% of the merged pull requests are now written by autonomous agents. Origin is built for an era where software agents are writing much of the code themselves, making changes at a frequency human centered platforms were never designed to absorb. If your agent writes the code, reviews the code and ships the code, the company that owns your agent wants to own the pipeline too. Cursor now sits inside SpaceX after a reported $60 billion all stock acquisition, which puts real capital behind that ambition. The Counter View On GitHub Before anyone declares a changing of the guard, look at what the adoption data will show over the next two quarters. Origin is a beta, gated to paid Cursor users and it rolled out opt out by default, no opt in and its retention, residency and data training terms are still unpublished. That means some companies’ developers may already be hosting proprietary code on a SpaceX-owned platform whose data terms do not exist in writing yet. Enterprises do not move their filing cabinet on a bad outage day. Watch whether Origin publishes governance terms and whether GitHub’s incident count keeps climbing. Those two numbers will tell you if this is a moment or not. What Leaders Should Do About GitHub Now Ask your engineering team one question this week about GitHub and other repositories. If our code host disappeared for 7 hours, what would stop? Then map out which vendors own multiple layers of your development stack. Concentration is convenient right up until it is a dependency. Whether GitHub stays the center of software development will be decided by the companies asking these questions now, and discovered by everyone else.
00:00

Serval Wants To Replace ServiceNow With AI That Builds Enterprise Automation

An AI startup called Serval is pitching itself as a replacement for ServiceNow, using agents that read support tickets and build the automation code themselves instead of making humans wire up workflows by hand. Its new Catalyst agent scans ticket histories for recurring problems, checks what connected systems can do, then generates the workflows and access rules needed to automate them — but nothing ships until an administrator approves each step. Serval raised a $75 million round led by Sequoia on top of a $47 million Series A, for $127 million total and a $1 billion valuation. It says revenue grew about 500% between rounds and claims customers automate over half their tickets, while arguing ServiceNow's AI products often get bought but never deployed.

Notes
Serval: AI-generated enterprise automation vs. ServiceNow

Company/funding: Founded April 2024 by CEO/cofounder Jake Stauch and cofounder Alex McLeod, both ex-Verkada. $47M Series A (Oct 2025), $75M Series B led by Sequoia; $127M total, $1B valuation. Claims revenue grew ~500% between the A and B periods; claims customers automate >50% of tickets and an 80% help-desk automation rate, but does not disclose ARR, profitability, or customer count.

Origin: Stauch cites a CFO's simple expense-approval rule that took an IT leader two months and hundreds of Okta Workflows steps to automate.

Catalyst (new AI agent): Searches ticket histories for recurring problems, examines connected systems, and generates the workflows, skills, and access configurations needed to automate. Stauch on its logic:

"Catalyst focuses on the problem, not the solution. It analyzes the problems users are raising and asks whether they can be solved with automation... It examines the tickets themselves, what's possible through the available APIs and the capabilities of the Serval platform, and then builds the automation needed."

It does not auto-publish; admins inspect each step, add permission checks, and require approvals before deployment.

Architecture: One agent builds tools/automations; another uses them to resolve requests — "code gen for everything that's not software engineering," per Stauch (vs. Cursor, Claude Code, Cognition). TypeScript-based; runs OpenAI and Anthropic models via zero-retention endpoints. Stauch argues foundation-model vendors won't build the security/approval/permission layer because their incentives favor model-centric products.

Claims vs. counter-evidence:

  • Stauch: "AI-native, better alternative" to ServiceNow; customer migration underway; <10% of ServiceNow AI SKUs purchased are deployed. ServiceNow declined comment.
  • ServiceNow rebuttal (Q2 2026): $3.88B subscription revenue (+24.5% YoY); AI business >$1B ACV; agentic-AI customers up ninefold in nine months.
  • Leone (Moor Insights & Strategy) on the <10% claim: "Nearly every enterprise vendor sold AI into 2025 budgets faster than customers could staff the rollouts. The criticism is fair."
  • Leone: ServiceNow collapsed 5 pricing tiers into 3, bundling Now Assist, Moveworks, Data Fabric, AI Control Tower — so deployment rates vs. everything sold are uninformative ("self-reported, discount them accordingly"). It bought Moveworks for >$2.85B; its Level 1 Service Desk AI Specialist serves 40+ customers, resolving 80–85% of requests autonomously at a 20–30% AI price premium.

Leone's caveats on Serval: early code-gen lead isn't a durable moat ("ServiceNow has spent 20 years collecting that context, while Serval is earning it one customer at a time"); needs a named non-cloud-native enterprise reference, a real config-database/change-management answer, and code-ownership governance for thousands of workflows.

Market: Targets ServiceNow's ~90% Fortune 500 penetration and ~$15B ARR. Sponsorship of Knowledge 2026 was rejected. Competitors: Atomicwork (Atom, agentic identity governance), Espressive, Rezolve AI, Console, Ravenna — Stauch says they don't appear in Serval's enterprise deals.

Long-term bet: employee-facing personal agents, positioning Serval as an enterprise control layer around access, approvals, and permissions.

Full text · 12,031 chars
The strange thing about enterprise automation is how much manual work it can take to automate something simple. Serval CEO and cofounder Jake Stauch remembers a CFO giving an IT leader a straightforward rule for approving expenses. Turning that rule into a working workflow took two months and hundreds of steps in Okta Workflows. Serval was founded to attack that gap. Stauch and fellow former Verkada employee Alex McLeod launched the company in April 2024 with the idea that companies should not need a technical project every time they want to eliminate a repetitive task. That idea now sits at the center of Catalyst, Serval's new AI agent, which searches ticket histories for recurring problems, examines the systems a company has connected and determines whether those problems can be automated. It then generates the workflows, skills and access configurations needed to make the automation work. "Catalyst focuses on the problem, not the solution. It analyzes the problems users are raising and asks whether they can be solved with automation. It's not analyzing how an IT team resolved a ticket and simply reproducing those steps. Rather, it examines the tickets themselves, what's possible through the available APIs and the capabilities of the Serval platform, and then builds the automation needed," Stauch told me in an exclusive interview. He explained that Catalyst doesn't automatically publish what it builds. "They can examine each step, add permission checks so it runs only for designated users and require approval from specific individuals, groups or workflows before it proceeds. The administrator decides what ultimately gets deployed," Stauch says. Serval raised a $47 million Series A in October 2025, followed by a $75 million Series B led by Sequoia. The two rounds brought the enterprise AI startup's total funding to $127 million and pushed its valuation to $1 billion. Stauch claims the platform is an "AI-native, better alternative" to ServiceNow and says it has already started taking customers away from the dominant enterprise service-management platform. He argues that a platform built around modern code-generation agents can move faster than one that has spent decades adapting an older architecture. "We've built our platform around code-gen agents as a fundamental primitive. It's a fundamentally different architecture from ServiceNow, one that wasn't possible if you started a company before 2024," Stauch says. Mike Leone, vice-president and principal analyst at Moor Insights & Strategy, says Serval's bet on TypeScript is meaningful. "TypeScript gives you a new kind of automation that you can read, version and hand to an auditor. It can be built in hours rather than through a traditional services engagement," he says. "That's a real advantage, and ServiceNow does move more slowly in that area." ServiceNow’s latest results, however, present a rebuttal. In its second-quarter 2026 results, subscription revenue reached $3.88 billion, up 24.5% year over year, while its AI business crossed $1 billion in annual contract value. The number of customers with agentic AI in production increased ninefold in nine months. But Stauch claims Serval is hearing a different story from customers that use ServiceNow, saying less than 10% of the ServiceNow AI products they have purchased have been deployed. "We hear from our customers that they've bought a bunch of ServiceNow SKUs and have never been able to get them implemented. ServiceNow is pushing them to reprioritize the budget so that it can show AI revenue." ServiceNow did not respond to a request for comment on the claim. "Nearly every enterprise vendor sold AI into 2025 budgets faster than customers could staff the rollouts. The criticism is fair, and I hear versions of it constantly," Leone says. Two decades of accumulated customization give ServiceNow a significant advantage, Leone notes, but he cautions against treating Serval's early lead in code generation as a durable moat. Code generation is becoming a standard way to build enterprise tooling, and ServiceNow already has a code agent of its own. "An agent acting on your behalf has to understand what your processes are, who approves what and which systems can break when it touches them. ServiceNow has spent 20 years collecting that context, while Serval is earning it one customer at a time." Turning AI Code Generation Into Enterprise Automation Serval's existing platform handles employee support across IT and other business functions, with requests coming through Slack, Microsoft Teams, email, web portals and other channels. Catalyst works one step earlier. An administrator can describe what they want to automate in natural language, and the agent generates the code and configuration needed to build it. The architecture builds on an idea Serval has used from the start. One agent creates the tools and automations, while another uses those tools to resolve employee requests. That separation gives administrators a clearer boundary between what AI can build and what can actually run in production. The platform primarily uses models from OpenAI and Anthropic through zero-retention endpoints. Stauch compares Catalyst with Cursor, Claude Code and Cognition, which help software engineers write application code. "We are code gen for everything that's not software engineering," he says. But model developers such as Anthropic increasingly have the technical ingredients to build software that can interact with enterprise APIs, reason over business context and execute multistep tasks. A sufficiently capable model could create a password-reset script in minutes. Stauch does not dispute that OpenAI or Anthropic could generate the code needed to automate a task, but argues that generating the code is only one part of building an enterprise automation product. A production enterprise workflow needs security controls, approval routing and permissions. He claims that foundation-model companies are "unlikely to spend the enormous amount of resources required to build that layer because their business incentives point toward products where the model itself supplies most of the value." If models become increasingly commoditized and more value shifts into the application layer, the incentives for OpenAI and Anthropic could change. Serval is betting that by the time they do, it will be "too deeply embedded in enterprise infrastructure and customer workflows for them to easily catch up." Serval Wants to Replace ServiceNow, Not Sell Alongside It Stauch's decision to target ServiceNow stems from the tech giant's roughly 90% penetration across the Fortune 500 and its pace toward about $15 billion in annual recurring revenue. Moreover, he believes the large-enterprise service-management market has effectively become a one-player market, leaving customers with few meaningful alternatives. “Other companies have taken the approach of moving downmarket and focusing on smaller businesses, but we’re going direct, intending to fully replace ServiceNow,” he says. Serval wants customers to consider the budget they already spend on service management rather than find new money for another AI product. Stauch compares that approach with selling a CRM against Salesforce, an HRIS against Workday or an ERP against SAP, where the existing platform sets the budget a challenger is trying to capture. That positioning has already created some friction. Stauch says his company planned to sponsor ServiceNow's Knowledge 2026 conference but was told it could not participate. ServiceNow is also investing heavily in the architectural gap Serval says gives it an advantage. The company paid more than $2.85 billion for Moveworks. Moreover, the company claims its Level 1 Service Desk AI Specialist is live with more than 40 customers and handles 80% to 85% of service requests without human interaction, while its AI-native products carry a 20% to 30% pricing premium. Serval has previously reported an 80% help desk-automation rate. ServiceNow collapsed five packaging tiers into three and bundled Now Assist, Moveworks, the Data Fabric and AI Control Tower into all of them. "Everyone with a contract bought AI whether they went looking for it or not, so a deployment rate measured against everything sold doesn't tell you very much," Leone says. "ServiceNow does disclose deployment numbers, but those are self-reported, so discount them accordingly." Investors Are Betting On A New Era Of Enterprise Software Statistics shared by Serval claim that its revenue increased roughly 500% between its Series A and Series B periods. Serval also says customers automate more than 50% of tickets, although the company has not publicly disclosed its exact ARR, profitability or detailed customer count. The startup’s $1 billion valuation reflects what investors believe the company could become, not evidence that it has reached the scale of the enterprise software giant it hopes to replace. Stauch describes the investor thesis: "talent, market and timing." He compares the shift with the transition to cloud computing, when a generation of cloud software companies reshaped markets once dominated by older vendors. In service management, he points to the transition from BMC Helix to ServiceNow as an example of how one generation of enterprise software can give way to another. "We don't know if that's going to happen yet, but that is the bet a lot of our investors are making. They believe this is another massive tech shift that's going to create a new generation of incumbents," Stauch says. Leone notes that ServiceNow's IT service desk is sold by the seat, which creates a problem when AI can resolve tickets without a person involved. "Half of ServiceNow's new business is now priced on something other than a seat, and repricing at that scale is what a company does under pressure. Serval didn't have to displace anyone to change how the incumbent charges." The startup is entering a market with a much larger cast of competitors than its ServiceNow-focused pitch suggests. Atomicwork is building an AI-native ITSM platform around its Atom assistant and has moved into agentic identity governance. Espressive and Rezolve AI remain independent, while Console and Ravenna are also appearing in head-to-head evaluations. Stauch says most of those startups do not appear in Serval's enterprise deals. "We hear about them from the press and VCs, but they don't really show up in our deals. We see ServiceNow. To a lesser extent, we see Jira Service Management, we see Freshservice. That's where the market is." he says. The Bigger Bet Is Enterprise AI Control Serval's long-term vision extends beyond tickets. "We want to take everything we're building for these IT admins, HR admins, finance admins—these automations they're building—and allow employees to automate more and more of their work with their own personal agents that they can use at work," Stauch says. That could move Serval closer to an enterprise control layer than a conventional ITSM platform. The company already sits around access requests, tickets and privilege escalation, giving it a potential role as employees ask personal AI agents to perform work beyond their normal permissions. The models powering enterprise agents may become interchangeable, while the permissions, approval rules and workflows governing what those agents can actually do could prove harder to replace. Question is whether enough CIOs will trust a young AI startup to sit between their agents and enterprise systems. "For a startup to genuinely displace ServiceNow, it needs a named enterprise reference outside the cloud-native cohort. It needs a real answer for the configuration database and change management. And somebody has to own the generated code once there are thousands of workflows and a regulator asking who signed off," Leone says. “Serval does deserve credit on access management, as they built it into the same product as ticketing, and ServiceNow had to buy a company to get there.”
00:00

Ode To A Whole New Kind Of Consulting

Anthropic is partnering with Blackstone and Goldman Sachs on Ode, a $1.5 billion venture that installs AI directly into client companies' operations. Ode bills itself as an AI-native enterprise services firm that finds Claude-ready workflows, builds custom apps and agents, automates tasks, and handles regulatory and compliance work. It's led by CEO Chris Taylor and adds to Anthropic's push to get enterprises actually using Claude.

Notes
  • New venture: "Ode" — an "AI-native enterprise services firm" formed as a joint project of Anthropic with consulting firms including Blackstone and Goldman Sachs, with partners injecting ~$1.5 billion. Publicly announced in a press release that notably never names the entity; the author had to dig to learn the name.
  • Division of roles: Goldman = investment banking/asset management (raising debt/equity, M&A advice, IPOs, restructuring, bond floats). Blackstone = asset management (private equity, private credit, real estate, infrastructure). Ode = AI implementation service.
  • What Ode does: identifies "Claude-ready" business processes/workflows; places AI engineers alongside client in-house employees; builds custom applications and agents; automates tasks and processes; handles regulatory and compliance issues ("moving targets").
  • Name rationale (author cites GPT): an ode is a poem of praise, matching the positioning — "Anthropic creates the frontier technology, while Ode is supposed to turn that technology into something useful." Ode's launch language: "the best models in the world don't put themselves to work." Blackstone executive Rodney Zemmel on the name: Ode was already "creating poetry in our portfolio."
  • Quotes:> "Companies everywhere see the potential for what AI can do for their businesses, the challenge is making it real... There's enormous demand for Anthropic's technology, and we're scaling quickly to help clients adopt AI with a focus on outcomes." — Ode CEO Chris Taylor> "As mid-size companies move from experimenting with AI to building it into their operations, they need partners with real implementation depth..." — Garvan Doyle, Anthropic Head of Forward Deployed Engineering
  • Caveat/framing: piece is largely descriptive and hedged ("Let's see how it rolls out"); no independent verification of scale, clients, or financials; bets on Anthropic's "growing ecosystem of partners."
Full text · 4,375 chars
There’s a new flavor of implementation consulting making its way out to the market. It’s a joint project of Anthropic and consulting firms like Blackstone and Goldman Sachs, and it purports to be a solution for helping client companies to figure out how to “add AI” to their plans. A press release from the stakeholders describes this new entity as an “AI-native enterprise services firm,” the type of which, by definition, could only have been created in the past five or six years. Strangely enough, the same press release does not refer to the project by name: I had to do more digging to find out that they’re calling this juggernaut “Ode,” as in “Ode to a Grecian Urn” or “Ode to Strange Combinations of Managed Services.” What Will It Do? Ode, into which the partners have injected some $1.5 billion, is going to bring AI to what have been more conventional relationships between Goldman and Blackstone, and their clients. I talked to a few people, and read a few things, and came up with this contrast: Classically, Goldman is basically an investment bank and asset manager. It might help a client to raise debt or equity, advise on an acquisition or sale, or noodle about how to take a company public. Goldman might help to restructure a company or its finances, or float a bond, or advise executives on major strategic financial decisions. Blackstone, for its part, has been an asset manager, the kind of consultant that helps companies to buy through private equity, provides private credit, or consult on real estate acquisitions, or infrastructure. Ode, for its part, will be an AI implementation service. Imagine a consultant that can identify Claude-ready business processes or workflows, adding AI engineers in with the company's in-house employees, building custom applications and agents, automating tasks and processes … you get the idea. Another part of this service will be dealing regulatory and compliance issues, both of which can involve moving targets. An Ode by Any Other Name When I asked GPT about why they call it “Ode” the model had this to say: “An ode is traditionally a poem written in praise or celebration of something. That fits the venture's positioning: Anthropic creates the frontier technology, while Ode is supposed to turn that technology into something useful and tangible in the real economy. Ode's own launch language says that ‘the best models in the world don't put themselves to work’ and describes its mission as bringing Anthropic's technology into companies.” That makes sense to me. The response also included this quote from top brass: “There's also a pretty explicit clue from Blackstone executive Rodney Zemmel,” GPT replied. “When announcing the name, he wrote that Ode was already ‘creating poetry in our portfolio.’” (Here’s the sourcing) That’s the sort of thing I was looking for. Many of us are in favor of any allusion to the humanities, because we worry that in the whirlwind days of the 2020s, they’re going away. We’re not in stasis here, we’re in frenetic change. Internal Enthusiasm The firms involved have high hopes for Ode, given their catalogs of existing customers, and all of the potential for moving confidently ahead with AI. “Companies everywhere see the potential for what AI can do for their businesses, the challenge is making it real,” said Ode CEO Chris Taylor, as quoted in the announcement. “Our teams partner closely with CEOs and across organizations to define and execute the highest priority AI initiatives. By pairing the deep subject matter expertise of our clients with our top applied AI talent, we’re able to drive transformation level impact. There’s enormous demand for Anthropic’s technology, and we’re scaling quickly to help clients adopt AI with a focus on outcomes.” “As mid-size companies move from experimenting with AI to building it into their operations, they need partners with real implementation depth and a clear understanding of how their businesses actually work,” added Anthropic Head of Forward Deployed Engineering Garvan Doyle. “Ode was built to be that partner, adding to Anthropic’s growing ecosystem of partners that help enterprises put Claude to work.” It’s that “growing ecosystem of partners” that the makers are banking on, along with the moment that we are now in. Ode represents the desire to strike while that iron is hot. Let’s see how it rolls out.
00:00

Busting Those Misleading Myths About Anthropic AI Watermarking During Proofreading Or Fixing Typos

Anthropic's text watermarking on Claude won't stamp a hidden mark on your writing just because you asked Claude to proofread it — it only appears when Claude actually rewrites or changes wording. The watermark works by quietly picking a lower-ranked word now and then in a secret statistical pattern that a normal reader can't spot, and only Anthropic's own (not yet public) detection tool can reliably confirm. Editing text can erode the mark, and the longer the passage, the more edits it can absorb before detection gets unreliable. Even Claude's response telling you where the gaffes are is watermarked, but the essay you submitted is not.

Notes

Anthropic AI Watermarking: What It Actually Does — Myth Busting Notes

Source: Forbes column, "Busting Those Misleading Myths About Anthropic AI Watermarking During Proofreading Or Fixing Typos" (2026-08-20). Addresses media claims that proofreading or typo-fixing with Claude will watermark your text.

Background on watermarking
  • Physical watermarks (e.g., dollar bills) are visual; digital image watermarks embed bits in binary that don't affect the picture.
  • Text is the hard case: any alteration risks changing meaning. Naive methods (replacing words, embedding special chars/emojis) are either nonsensical or trivially stripped.
How the Anthropic watermark works
  • Technique: statistical word-choice selection (the article's "statistical uplift"). Claude generates text word-by-word, ranking candidate words statistically; the watermark = deliberately choosing the 2nd/3rd/4th-ranked word instead of the top choice, repeatedly.
  • Worked example (ham sandwich prompt): "bagel" was top-ranked, "flatbread" 2nd, "wheat bread" 3rd, "white bread" 4th. Watermarked output: "Place a slice of ham onto a flatbread and add mustard." — reads normally, undetectable by eye.
  • Detection: a tool built knowing the method compares observed word choices against the AI's normal distribution; consistently encountering off-top choices is a strong indicator.
Detection details
  • Strengthened version: 50% of the time pick 2nd choice, 30% 3rd, 20% 4th, guided by a secret cryptographic key — makes cracking harder.
  • Anthropic's authorized detection tool has not yet been publicly made available.
  • Watermark survives copy/paste intact; tolerates "some semblance of editing," but robustness depends on text length.
Rules of thumb (robustness)
  • Short text: changing one watermarked word (e.g., "flatbread" → "bagel") can materially destroy the statistical signal — it may be the only watermarked choice in the sentence.
  • Long text: a single edit is "dropping a pebble into the ocean"; watermark stays substantially intact.
  • Post-edit signal: 90–99% remaining → "somewhat sure" the watermark is present; ~10% remaining → detection is "on thin ice."
  • Mixed document: pasting a few Claude-generated paragraphs into a 10-page handwritten doc — tool should flag only the paragraphs. Stated concern: the tool might declare the whole 10 pages AI-generated; hope is it reports an estimated likelihood (slim / modest / highly likely).
Proofreading, no edits
  • If Claude only scans text and reports gaffes without rewriting, the submitted text is NOT watermarked — it was untouched. The proofreading response itself is watermarked (Claude generated it). Media warnings that scanning "magically infuses" a watermark are called false/misleading.
Proofreading + editing
  • When Claude is allowed to reword, every change is a chance to insert watermark words. Example: "the dog jumped over the lazy fox" → "leapt" — same meaning, possibly a watermark insertion. You have no clue which changes carry watermark elements.
  • Rule of thumb: few changes in a large text → slim watermarking; detection should indicate Claude involvement in the subset only, not the whole.
Fixing typos
  • Explicit "only fix typos in place": "catt" → "cat" does not watermark — Claude didn't choose "cat," it corrected the spelling of a user-chosen word; no word-choice opportunity.
  • Loose instructions: Claude might replace "cat" with "feline," which can be a watermark word choice. Any moment Claude picks which word to include is a watermark opportunity.
Predictions and caveats
  • Forecast: ChatGPT, GPT-5, Gemini, Copilot, Grok will each adopt watermarking, likely similar statistical word-choice technique but differing per vendor (possibly semi-standardized).
  • Each maker will need its own proprietary detection tool → predicted confusion: running Claude text through OpenAI's detector returns "no ChatGPT watermark," and an unaware checker may wrongly conclude it's human-written.
  • Column closes with the Einstein quote: "There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle."

Key takeaways for archival: watermarking is generation-time word-choice bias, not a stamp applied to text Claude merely reads; non-rewriting proofreading and in-place typo fixes do not watermark your text; editing gives Claude insertion opportunities; robustness scales with output length; and no public Anthropic detection tool exists yet.

Full text · 15,941 chars
In today’s column, I examine and bust or straighten out various misleading myths regarding the recently released AI text-oriented watermarking feature of Anthropic. The mainstream news and social media have been making zany and incorrect claims about what this particular technique of watermarking is and does. This, in turn, has tended to create widespread confusion and undue consternation among people who are unsure whether their text will be watermarked or not when using Claude. Weighty questions on people’s minds include whether text that they submit to the AI for proofreading will end up watermarked, and whether the AI fixing incidental typos will also encompass the implanting of a watermark. To answer those questions, I will briefly lay out how it is that the watermarking actually occurs and explain how the hidden watermarks are infused into text (for my in-depth coverage, see the link here, and for my analysis of watermark detection tools, see the link here). You will end up with a much cleaner understanding of how to judge whether to use the AI for aid in writing and editing of text, and the likelihood of a watermark getting included in the text. Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage of the latest in AI, including identifying and explaining key AI complexities (see the link here). Watermarking Is Challenging First, some foundational aspects of the topic of watermarks. We are all aware of watermarking when it comes to paper-based materials and likewise for any tangible artifact that exists in a definitive physical form. A dollar bill can contain a watermark, allowing the naked eye to tell whether it is real or counterfeit. Watermarks can also be hidden from visual inspection, requiring some other means to detect the watermark. Watermarking for digital photographs and graphical images is more readily accomplished than with text since you can embed all sorts of digital ones and zeros that won’t impact the picture, but that can be detected by inspecting the binary representation. It is possible to use sophisticated mathematical algorithms to populate the bits in a manner that almost no one other than someone armed with the algorithm can later detect as being part of a special pattern. Trying to watermark digital text is a beast of a different kind. Anything that is done to the text will potentially alter the words we see and impact the meaning of the text. If you had a watermarking algorithm that simply said to replace the word “of” with the word “horse”, the resulting text, which is now presumably discernible as AI-written due to the excessive use of the word “horse”, is going to be nonsensical for human use. Likewise, if the watermark consisted of embedding special characters or the use of emojis, you could quickly find those and remove them easily. Example Of How It Works An ingenious way to infuse watermarking is to do so by selecting suitable words that can be viably chosen during the AI writing process. Here’s how that works. Envision that AI is generating a response to a prompt, doing so one word at a time. Each word is carefully chosen by the AI. The choice of which word to use is made from several possible words at each step. Suppose the prompt was asking the AI how to make a ham sandwich. The AI might start assembling the response word-by-word and could have arrived at these choices: “Place a slice of ham onto a bagel and add mustard.” Each word was selected on a one-at-a-time basis, going from the start of the sentence to the end of the sentence. When the AI got to the word about the bread, in this instance the word selected was “bagel,” but there were several other options available, such as saying “flatbread” (statistical second choice), “wheat bread” (statistical third choice), “white bread” (statistical fourth choice), and other possibilities. Assume that the word “bagel” was the statistically top-ranked choice overall and therefore chosen accordingly. Aha, in the realm of watermarking, the AI might opt to intentionally choose the second choice rather than the top-ranked choice; thus, the sentence comes out as “Place a slice of ham onto a flatbread and add mustard.” If the AI consistently keeps picking the second choice for many of the words that are being chosen, this becomes a handy pattern for the AI. A human looking at the sentence doesn’t realize that the second choice is being chosen. They see a sentence that looks completely normal. Detecting The Watermark I think you can see that this statistical uplift is going to be quite hard to detect by conventional means. Humans are unlikely to discern the watermark by looking for any patterns in the wording. All the sentences are still going to make sense and abide by whatever the topic at hand is. The subtlety of picking the second statistically viable word on numerous occasions is a hidden way of producing the watermark. How does an authorized detection tool figure out if the watermark is present? That is the added trickery. The chances of any usual automated detection method ferreting out the watermark are low. It won’t realize what the watermark method is or how to ferret it out. Meanwhile, a detection tool that is built knowing the method can examine the sentences and compare the word choices to the pattern of word choices that the AI would normally make. If the second word choice is consistently being encountered in the examined text, this is a strong indicator that the AI indeed generated that content. We can make this method much stronger. Instead of always choosing the second choice, the watermark process does something else. Suppose that 50% of the time the second choice is made, 30% of the time the third choice is made, and 20% of the time the fourth choice is made. This makes things even harder for anyone else to crack and find the watermark. An even better method includes having a secret cryptographic key that guides the watermarking process toward the preferred token patterns. Breaking The Watermark Anthropic stated in their recent announcement about the new watermarking that the watermark will persist when the text is copied and pasted somewhere else. This makes abundant sense. A body of text produced by Claude is going to carry the watermark since it has that secret pattern of word choices. If you copy it as is and then paste it as is, the pattern of the words remains precisely as the AI generated it. Ergo, the watermark is still entirely present. The announcement by Anthropic also noted that the watermark can tolerate some semblance of editing. The question is how much editing can be done to the text before the watermark breaks down and is no longer statistically significant. This is somewhat complicated to identify since we are then relying on statistics and probabilities. Pretend that I take the watermarked sentence that says to make a ham sandwich with flatbread, and I change the word to a bagel. I have marred the watermark. The AI had explicitly chosen the word flatbread, and it is no longer there. Instead, the word bagel is there. If the body of text is relatively short, my making that one change could materially undermine the statistical likelihood that the text contains the AI watermark. Perhaps that is the only watermarked chosen word in that sentence. On the other hand, if the body of text is very lengthy, many paragraphs in size, my having replaced one word in one sentence is perhaps like dropping a pebble into the ocean. The watermark remains substantially intact because it is pervasive throughout the rest of the text. Rules Of Thumb The larger the body of text that the AI outputs and watermarks, the less destructive to the watermark are my few edits. You see, there will still be a preponderance of text that contains the watermark. The statistical signal of the watermarks might remain at some high percentage after my edits, perhaps 90% to 99%. That is potentially enough to be somewhat sure that the watermark is in there. If the watermarks remain at only 10% after my edits, now things are getting dicey. The detection of the watermark is going to be on thin ice to conclude that the watermark is truly there. The crux is how much of the text contains the watermarked approach, and how much of the text does not contain the watermarked approach. Suppose I use Claude to generate a few paragraphs for me. Claude produces the text, and it contains the secret watermark due to the words selected for the essay. I then paste that text into a document of ten pages of my own hand-crafted text. At this juncture, suppose that someone wants to know whether the ten pages were handwritten or generated by Claude. An Anthropic-authorized AI watermark detection tool, which hasn’t yet been publicly made available, will presumably scan the ten pages of text and look to see if the pattern of word selections matches the approach being used by Claude. The few paragraphs will likely get flagged, but the rest of the ten pages are unlikely to get flagged since it wasn’t produced by the word selection method. A big concern is that the authorized watermark detection tool might simply claim that the entire text was likely generated by Claude, even though only a small portion of the text seems to be so generated. The hope is that the detection tool will be more transparent and offer an estimated likelihood, such as that it is slim, modest, or highly likely that the watermark is present. Worries are that people are going to run with whatever the detection tool says, despite the reality that the watermark might only marginally be present. Proofreading Of Text There is confusion in the media about what happens if you ask Claude to proofread a body of text. Some have been warning that the mere act of Claude scanning a body of text is going to somehow magically infuse a watermark into the text. That’s a false or misleading portrayal of the situation, so let’s properly unpack things. First, assume that I take a body of text that I handwrite and I give that to Claude to proofread. I tell Claude not to change any of the wording. It is only to scan the text and let me know if there are any potential gaffes or factual inaccuracies in my text. Claude proceeds to scan the text and gives me a response that identifies a handful of portions of the text that could use some cleanup. The body of text is still the same as it was when I submitted it to Claude. Do you think that the text now contains the secret watermark? I trust that you realize the original text has been untouched in the sense that it wasn’t rewritten by Claude; therefore, the text has not been watermarked. A small twist to keep in mind is that the response about where there are gaffes is indeed watermarked, because that answer was generated by Claude. But the text I submitted to be proofread is not watermarked. Proofreading And AI Making Edits Now that you’ve got the gist of things, we are ready to get into more complicated situations. First, I handwrite an essay. I then give the essay to Claude and ask it to proofread the essay. In addition, I tell Claude to go ahead and fix the essay, meaning that Claude can reword the text that I have written. Keep in mind that we are now allowing Claude to start infusing a watermark. How so? Any of the wording changes that Claude makes are a potential moment for Claude to select a suitable chosen word that is based on the secret approach. I might have a sentence that says the dog jumped over the lazy fox, and Claude changes that wording to say that the dog leapt over the lazy fox. The sentence still has roughly the same meaning, but Claude snuck the word “leapt " into the text and presumably did so as a watermarking insertion. You will not have any immediate clue whether the wording changes are part of the watermark, though they might very well be. Some of the words that Claude changes could be part of the watermarking, while other words might not be. The main point is that you are giving Claude a chance to make word choices that could allow for the watermarking to occur. Using the prior rule of thumb, if Claude makes a relatively small number of changes in terms of replacing words with other words, and if the body of text is large, the amount of watermarking is going to be slim. The watermark detection will hopefully not clamor that the text has been generated by Claude. It will presumably find the small portion that is now watermarked and indicate that it is likely that Claude was involved to some extent in the generation of the subset of text. We opened that door to this by allowing Claude to not only proofread but also edit the text. Fixing Of Typos One of the flagrant misstatements about the watermarking is in reference to Claude fixing typos. What do you think happens when Claude is asked to fix a misspelled word? Suppose that I have typed the word “catt” in my essay, and I meant to type the word “cat”. I misspelled the word. I give Claude my essay. I tell Claude to only fix any typos. It is not to do any broad-based editing. Just find and correct any typos. Sure enough, Claude shows me the resulting essay with a few typos that were corrected. The word “catt” is now correctly shown as “cat”. The same goes for the other words that I inadvertently misspelled. Did this give Claude an opportunity to include the watermarking? Not if you were explicit about only fixing the in-place words that were misspelled. For example, I had chosen the word “cat” and simply misspelled it. Claude makes the correction and turns the word into “cat”. Claude did not choose the word “cat” at the get-go; I did so. There wasn’t an opportunity for Claude to select some other word; it only made sure my chosen word was correctly spelled. The twist is this. Suppose that you give Claude some leeway. It finds the word “catt” and determines that the word ought to be spelled as “cat”, but if your instructions were loosey-goosey, Claude might replace the word “cat” with the word “feline”. Perhaps the word “feline” is going to be a word choice by Claude that serves as an element of a watermark. Any moment when Claude can choose which word to include is a chance for Claude to go the route of watermarking. The World We Are In Things are going to get quite topsy-turvy once all the other major LLMs implement a form of watermarking, including ChatGPT, GPT-5, Gemini, Copilot, Grok, etc. Each AI maker will potentially adopt an approach that is slightly different from the other AI makers. The odds are that they will use a similar technique of statistically using wording choices as their text-oriented watermark technique, but differ enough that there isn’t one universal means at play (as a side note, this could still be undertaken in a semi-standardized way). The bottom line is that each AI maker will then need to provide a proprietary authorized detection tool so that people can run text through the tool to find out if the text was potentially generated by the particular AI. I’m betting this will sow confusion. Here’s how. Somebody might have used Claude to generate text, and a person wanting to check it runs the text through an OpenAI watermark detection tool that says the text is free and clear of being composed by ChatGPT. An unaware person won’t realize that they fed the text into a different watermark detection tool and should have used the Anthropic tool instead. A final thought for now. Albert Einstein famously made this remark: “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.” The prevailing approaches to having AI watermark text are not really a miracle, though you do need to give credit for the underlying cleverness of human designers. All of us will eventually become used to AI-generated watermarks and take it in stride. Some won’t like it; others will relish it. Time will tell.

Discussion

16
02:45

Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model

Qwen 3.8-27B shows the strongest autonomous tool-use of any model that runs locally, chaining 80 calls with no help. On a single RTX 3090 it logged into university sites to pull a class schedule, and separately investigated a social media user by downloading a video, extracting frames, running Whisper transcription, and zooming in on frames. It runs fully on the user's own hardware with a quantized version.

Full text · 891 chars
Off a single prompt, given my credentials and the name of my university, qwen3.8-27b was able to successfully pull my class schedule from the kinda shitty and convoluted web of university websites. It needed no human intervention, and executed 80 tool calls. Another time, I asked it to investigate a user on a social media network, and it found one public video, downloaded it, extracted frames every few seconds so it could "watch" the video, and installed fucking openAI whisper and ran a transcription to understand the context, before selectively zooming in on and brightening some frames to see the action. That this shit is running on my own hardware (single RTX 3090) is fucking incredible, the general public doesn't realize how cyberpunk our reality already is. Quant: Unsloth's Q4_K_S (kv cache quantized to q8) Context: 150k submitted by /u/synth_mania [link] [comments]
11:42

Tencent begins testing its new flagship model Hunyuan Hy4

Tencent has started quietly testing its new flagship model called Hunyuan Hy4. The model has appeared in Tencent's Yuanbao app marked as an 'expert-level' model that can use tools, positioned above its older Hy3 and above DeepSeek, and Tencent confirmed a larger Hy4 launch is coming soon with better performance and multimodal abilities. The new model isn't fully released yet, only in gray (limited) testing.

Full text · 850 chars
From the screenshots: Hy4 is now live, labeled "Expert-Level Model" + "Use Tools to Solve Problems" Hy3 is tagged with "New Upgrade," positioned as a brand-new general-purpose model DeepSeek, focused on reasoning, is listed alongside it From SuSu_酥酥👅on 𝕏: https://x.com/NFT_Chen/status/2090399515618787508 Tencent begins gray testing its new flagship model Hunyuan Hy4! Just now, a user spotted that Hy4 has appeared in the model selection list of the Tencent Yuanbao App, directly labeled as an expert-level model, positioned above Hy3 and DeepSeek. Tencent only confirmed in last week's Q2 earnings report that the larger-parameter Hy4 would launch soon, further enhancing model performance and multimodal capabilities. From Max For AI on 𝕏: https://x.com/MaxForAI/status/2090386754633421110 submitted by /u/Nunki08 [link] [comments]
03:02

Qwen3.8-27B took a serious hit to *knowledge* vs 3.6

Qwen's new 3.8-27B model is noticeably worse than the older 3.6 at recalling random facts from its own memory. A user who stress-tested it on personal trivia and offline knowledge benchmarks found it failed questions the 3.6 version reliably answered, across many quantizations and settings. The drop only matters if you rely on the model's built-in knowledge with no tool access, since it still handles coding and tool calls fine.

Notes
Qwen3.8-27B knowledge regression vs Qwen3.6

Poster: u/EmPips (r/LocalLLaMA, 2026-08-20)

Claim: Qwen3.8-27B is "pretty significantly weaker" than Qwen3.6 at recalling random facts from its own weights.

Method/anecdote: Spent days running Qwen3.8-27B against personal harnesses/workflows plus a private pocket-trivia set (mildly obscure facts mixed with prepper questions). Across all quantization levels and sampling settings it "did relatively poorly," failing questions Qwen3.6 "reliably answered."

Benchmark correlation: Says offline (no tool-call) knowledge benchmarks "seem to align" — model is weaker than 3.6 at fact recall. Notes his own hallucination tests improved, but that's not reflected in these particular benchmarks.

Caveats he states:

  • "You should never trust barcharts over your own vibes, but my vibes are validating these bar charts this time around."
  • Downplays relevance: "It seems to know the code it tries to use well-enough"; for everything else he assumes/hopes users rely on tool-calls.

Scope of impact: Applies only if your strategy is trusting an airgapped model with obscure/broad knowledge retrieval exclusively from its own weights — which he calls "probably a losing strategy anyway." If that's you: skip Qwen3.8 or "finally set up that MCP server."

Open ask: Wants confirmation from others who noticed the same. No commenters' replies captured in source.

Full text · 1,530 chars
Like many of you I've spent the last few days throwing Qwen3.8-27B against all of my usual use-cases and personal tasks/harnesses and workflows. It's great, phenomenal sometimes, but that's not what this post is about. One of my little personal benchmarks is a little set of pocket trivia that's relevant to me but mildly obscure mixed in with a few useful/prepper questions. Qwen3.8-27B at all quantization levels and sampling settings I threw at it, did relatively poorly at this. It's failing questions that Qwen3.6 reliably answered. I come to find out that on offline (no tool call) knowledge benchmarks seem to align with what I'm saying. It's pretty significantly weaker than it's 3.6 predecessor at recalling random facts (or not hallucinating as much, in my tests, though that isn't reflected in these particular benchmarks). Now you should never trust barcharts over your own vibes, but my vibes are validating these bar charts this time around. Is this relevant? Not necessarily. It seems to know the code it tries to use well-enough and for everything else I'm assuming/hoping you're using tool-calls. This largely only applies to you if you have a strategy of trusting an airgapped model with obscure/broad knowledge-retrieval exclusively from within its own weights, probably a losing strategy anyway.. but if that's you, take a pass on Qwen3.8 or finally set up that MCP server. I found it to be interesting. Curious of your thoughts or if anyone else noticed this. submitted by /u/EmPips [link] [comments]
10:20

The spectral neuron - an ML primitive for scalable and interpretable models [R]

A new research paper proposes a simple building block for machine-learning models that aims to be both easy to understand and quick to scale. The idea grew out of ad-modeling work at Yahoo and boils down to a single formula where the model's answer changes with matrix weights. The author shares the math, a training recipe, and scaling tests on synthetic and real data, with code on GitHub. The paper was human-written, but the author says the code was largely written by AI and reviewed by him.

Full text · 1,191 chars
Worked some time ago on one of the ad teams at Yahoo, and this grew out of a question I kept returning to while there are there "simple" models that are both simple, scalable, interpretable, and controllable at the same time? Decided to explore it, first in a blog (starting here ), then in a new preprint "The Spectral Neuron", built by distilling latest blog-posts into a manuscript, I study models of the form: 𝑓(𝒙) = 𝛌ₖ(𝐀₀ + 𝚺ᵢ 𝑥ᵢ𝐀ᵢ). Manuscript : https://arxiv.org/abs/2608.08003 Code : https://github.com/alexshtf/spectral_neuron_paper Looks like a simple on-liner, but many interesting aspects hide there. How expressive does the model become as the matrices grow? What can we read directly from the learned matrices? Which shapes can be guaranteed by construction? I develop the mathematics, give a practical initialization and training recipe, and test the model in scaling experiments on synthetic and real data. AI disclaimer : manuscript written by yours truly, AI assisted in looking up canonical references and related work for literature review. In contrast, the code was heavily AI written and reviewed by yours truly. submitted by /u/alexsht1 [link] [comments]
11:38

I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)!

Someone trained a tiny 1-billion-parameter model modeled on Moonshot's Kimi K3 for just $250 and it already beats GPT-2. The mini replica keeps K3's architecture, including its active-token mixture and tokenizer, was trained on 5 billion tokens, and scores 33.4% on HellaSwag versus GPT-2 124M's 28%. It's a hobby-scale proof that K3's design is reproducible cheaply.

Full text · 924 chars
I pre-trained a 1.02-billion-parameter on Kimi K3 replica trained on 5.00 billion decontaminated tokens for $250. This model has 1.02 billion parameters, of which 145 million are active per token. It is roughly one two-thousandth of K3 by total size. It saw 5,000,003,584 tokens, which is a rounding error against the corpora frontier models are trained on. It has never been instruction-tuned, and it has only ever done one thing: predict the next token. What it does have is K3's architecture: - Kimi Delta Attention, Gated MLA, Attention Residuals - LatentMoE with the same aux-loss-free balancer - Same activation function with the same two constants - K3's own 163,840-token tokenizer, unmodified. I report a 33.4% HellaSwag which beats the GPT-2 124M score of 28% Read the entire tutorial here: https://books.vizuara.ai/book/pretraining-a-mini-k3 submitted by /u/OtherRaisin3426 [link] [comments]
11:53

The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches

Someone built a homemade setup to run a very large open model locally at decent speed. They linked 16 RTX 5060 Ti graphics cards across two PCIe switch islands using heavily customized drivers and firmware, achieving around 140 tokens per second output with a huge context window. It's a deeply technical hobbyist build with lots of caveats, but it shows what's possible with cheap consumer hardware and a lot of effort.

Notes
Running DeepSeek V4 Flash-0731 (130–150 tks) on 16x 5060 Ti 16GB over 2 PLX88096 switches

r/LocalLLaMA post by /u/Primary_Exchange21, validated config:

  • Motherboard: ASRock Rack SPC621D8U-2T/OVH
  • CPU: Xeon Gold 6330 (notes: get Gold/Platinum for Optane PMem)
  • GPU fabric: Two Broadcom/PLX PEX88096 islands, 8 GPUs each
  • GPUs: 16 x RTX 5060 Ti 16 GB
  • OS: Ubuntu 22.04.5 LTS, kernel 6.8.0-106-generic
  • Driver: Aikitoria patched open driver 610.43.02-p2p
  • BAR1: 16,384 MiB on every GPU
BIOS/UEFI settings
  • UEFI boot on, CSM off, Secure Boot off (local EFI app + patched NVIDIA modules are unsigned)
  • Above 4G Decoding on; MMIO High Granularity 1024G; MMIO High Base ~56T; SR-IOV off
  • GRUB: intel_iommu=off pci=realloc=on,hpmmioprefsize=512G
  • NVreg_EnableResizableBar=1
Custom steps
  • Sets size code 14 → 16 GiB BAR1 on all 16 GPUs
  • Temporarily disables PCI memory decoding, clears old BAR1 so Linux reallocates
  • Writes ECAP_ACS+0x6.w = 0000 on every PLX/PEX bridge
  • "Vibe coding" for custom all-reduce within each PLX cluster + DSpark for pipeline parallel
Benchmarks
  • TP 8 / PP 2: 500k context; ~4000 pp up to 500k context; tg 100–150 (avg 140 in DeepSeek Harness)
  • TP 4 / PP 4: full 1M context; ~7000 pp up to 500k; tg 80

Paid 0.6x the price of an RTX 6000 Pro for the whole setup. No caveats or disagreements recorded in thread.

Full text · 1,521 chars
Component Validated configuration Motherboard ASRock Rack SPC621D8U-2T/OVH CPU Xeon Gold 6330 (Get gold/platinum if interested in Optane Pmem gimmicks) GPU fabric Two Broadcom/PLX PEX88096 islands, eight GPUs per island GPUs 16 x RTX 5060 Ti 16 GB OS Ubuntu 22.04.5 LTS Kernel 6.8.0-106-generic NVIDIA driver Aikitoria patched open driver 610.43.02-p2p Required BAR1 16,384 MiB on every GPU UEFI boot enabled; CSM disabled. Secure Boot disabled. The locally built EFI application and patched NVIDIA modules are unsigned. Above 4G Decoding enabled. MMIO High Granularity set to 1024G . MMIO High Base set around 56T . SR-IOV disabled on this machine. intel_iommu=off pci=realloc=on,hpmmioprefsize=512G in GRUB; NVreg_EnableResizableBar=1 for the NVIDIA module; Sets size code 14 → 1 6 GiB BAR1 on each of the 16 GPUs Temporarily disables PCI memory decoding and clears the old BAR1 address so Linux can reallocate it. PLX switch ACS control register: For every PLX/PEX bridge, writes: ECAP_ACS+0x6.w = 0000 After that, a little vibe coding to make custom all-reduce work within each PLX cluster and make DSpark work for pipeline parallel. For tensor parallel 8, pipeline parallel 2: 500k context available. Around 4000 pp up to 500k context, tg 100-150 (Averaging 140 in DeepSeek Harness) For tensor parallel 4, pipeline parallel 4: Full 1M context available. Around 7000 pp up to 500k context, tg 80 Paid 0.6 x RTX6000 Pro for the whole setup. submitted by /u/Primary_Exchange21 [link] [comments]
13:34

Mapping intrinsic rank and informational gravity in complex tabular data: I developed a non-parametric, model-agnostic, information-theoretic diagnostic to bypass the limits of linear, rank, and Euclidean baselines. [R]

A researcher released an open-source tool that finds the true number of hidden factors in messy datasets where standard methods fail. Called the Entropic Scree, it uses information theory instead of linear math, and in a stress test it recovered exactly 20 hidden factors where principal component analysis hallucinated about 5,700. It also claims to show which factors carry real signal and which are just noise. The preprint and code are on GitHub, and the author is inviting people to test it.

Notes
  • Tool: Entropic Scree Function v1.0.0 — non-parametric, model-agnostic, information-theoretic diagnostic for intrinsic rank of complex tabular data
  • Author: /u/Chocolate_Milk_Son
  • Preprint: https://doi.org/10.5281/zenodo.22028087
  • Code: https://github.com/tjleestjohn/Entropic-Scree
  • Date: 2026-08-20
Core claims
  • Standard PCA "fundamentally fractures non-linear dependencies into Spurious Orthogonal Dimensions," drastically overestimating true rank of complex tabular systems.
  • Non-linear alternatives (Kernel PCA, Euclidean nearest-neighbor estimators) suffer structural collapse when generative roots are entangled or sparse.
Why standard baselines fail (author's stated mechanisms)
  • Standard PCA — Dimensional Inflation: measures only linear covariance, so it reads a polynomial/non-linear interaction (e.g. X1·X2) as an independent variable, fabricating spurious orthogonal dimensions.
  • Kernel PCA (RBF) — Structural Collapse: projecting into Hilbert space doesn't fix it; folds even-polynomials into independent axes. Its infinite-dimensional space lacks finite-sample boundary, so sparse combinatorial noise smears into an elevated tail obscuring the structural elbow. "If the underlying generative roots are even mildly entangled, KPCA suffers a total structural collapse."
  • Euclidean/topological estimators (TWO-NN, MLE) — fail in sparse regimes: in asymmetric, feature-rich settings where features > samples (m > N), they hit distance concentration (nearest/farthest neighbor ratio → 1), making local neighborhood calculations structurally degenerate across mixed-data margins.
How Entropic Scree works
  • Metric space: pairwise dependencies via Information-Theoretic Jaccard Similarity (Variation of Information); relies on Shannon entropy, so invariant to marginal shape mismatches (continuous waves vs. binary flags).
  • Bypass rank ceiling: standard PCA capped algebraically at N−1; double-centered topological information space maps true overlapping redundancy, bypassing the algebraic sample-size ceiling.
  • Compress manifold: acts as a bivariate filter compressing primary overlapping probability mass of non-linear combinations back toward the Intrinsic Generative Rank; shears unique synergistic variance, leaving an Extended Signal Tail that separates true drivers from Idiosyncratic Informational Variance.
  • Informational Gravity (AIG/FSIG): decouples rank from probabilistic volume; rebundles sheared residual variance into "variable-equivalent" footprints.
Empirical stress test

Synthetic dataset: 20 pure generative roots → 5th-order combinatorics → 20,000 proxies, only 10,000 samples (m > N), plus injected idiosyncratic structural noise and measurement error.

| Method | Result |

|---|---|

| Standard PCA | hit rank ceiling, linearly fractured expansions, falsely extracted ~5,700 dimensions |

| Kernel PCA (RBF) & Spearman rank | liberal overestimation by 100%; lost elbows and totally collapsed under root entanglement |

| Entropic Scree | correctly mapped intrinsic rank at exactly 20; isolated 1.45% active shared signal vs 98.55% idiosyncratic variance; residuals formed Extended Signal Tail aligned with hypergeometric design-space limits |

FSIG topology: primary dimension FSIG₁ ≈ 74.5 variable-equivalents (global combinatorial hub), then a flat plateau across remaining 19 dimensions (~11.5 each) — confirming a democratically distributed root system under extreme entanglement.

Caveats / open questions
  • No independent verification; single-author simulation.
  • Claims hinge on the synthetic test — real-world performance unmeasured.
  • Author invites community testing and feedback on the repo.
Full text · 6,602 chars
Links: Preprint: https://doi.org/10.5281/zenodo.22028087 Entropic Scree Function v1.0.0 / GitHub: https://github.com/tjleestjohn/Entropic-Scree TL;DR: Standard PCA fundamentally fractures non-linear dependencies into "Spurious Orthogonal Dimensions," drastically overestimating the true rank of complex tabular systems. Meanwhile, non-linear alternatives like Kernel PCA and Euclidean nearest-neighbor estimators suffer structural collapse when generative roots are entangled or sparse. I’m sharing the methodology and code here for anyone dealing with these complex tabular data nightmares. The method and open-source framework use Normalized Mutual Information to compress spurious expansions back towards their true generative roots. It also Maps the underlying "informational gravity" of the roots, offering insight into overall average stability, as well as which specific roots can be most reliably extracted; Estimates the data's overall ratio of shared signal to unshared idiosyncratic informational variance (noise); Serves as a powerful exploratory map that separates unrelated clusters of variables, allowing you to easily identify decoupled sub-networks. A Modern ML Architectural Blueprint: Far beyond a mere update to legacy factor analysis workflows, identifying this exact intrinsic rank allows you to explicitly size neural bottlenecks for downstream non-parametric manifold extractors (like autoencoders). The Problem with Standard Baselines: When trying to map the intrinsic dimensionality of a dataset, standard practice usually dictates reaching for PCA, its non-linear kernel extensions, or Euclidean nearest-neighbor estimators. But if your tabular environment has mixed data types, heavy non-linearities, entangled roots, or more features than samples ($m > N$), these established baselines don't just lose precision. They suffer a structural collapse. The core issue with our standard baselines: Standard PCA drives Dimensional Inflation. Because it only measures linear covariance, it perceives a polynomial expansion or a non-linear interaction (like $X_1 X_2$) as an entirely independent variable. It is forced to fabricate new, spurious orthogonal dimensions to map them. Kernel PCA (RBF) suffers Structural Collapse. Projecting into a Hilbert space doesn't fix this. KPCA artificially folds even-polynomials into independent axes. Furthermore, because its infinite-dimensional space lacks a finite-sample boundary, sparse combinatorial noise smears into an elevated tail that obscures the structural elbow. If the underlying generative roots are even mildly entangled, KPCA suffers a total structural collapse. Topological Estimators (Euclidean) fail in sparse regimes. Estimators like TWO-NN or MLE rely on Euclidean distance metrics. In asymmetric, feature-rich environments ($m > N$), they suffer from distance concentration (the ratio between nearest and farthest neighbors converges to 1). This renders local neighborhood calculations structurally degenerate across mixed-data margins. Introducing the Entropic Scree: To solve this, I built the Entropic Scree . It throws out linear and spatial variance entirely and evaluates pure probability mass. Here is how it works under the hood: The Metric Space: It evaluates pairwise dependencies using Information-Theoretic Jaccard Similarity (Variation of Information). Because this relies on Shannon entropy, it’s invariant to marginal shape mismatches (like mixing continuous waves with binary flags). Bypassing the Rank Ceiling: Standard PCA is algebraically capped at $N-1$. By moving to a double-centered topological information space, we map true overlapping redundancy and completely bypass the algebraic sample-size ceiling. Compressing the Manifold: The algorithm acts as a bivariate filter. It inherently compresses the primary overlapping probability mass of non-linear combinations back towards the Intrinsic Generative Rank. It shears off the unique synergistic variance, leaving behind residuals that form a bounded Extended Signal Tail, cleanly separating the true drivers from the unstructured Idiosyncratic Informational Variance. Quantifying Informational Gravity: Beyond just extracting a discrete rank, the framework decouples rank from probabilistic volume by introducing Informational Gravity (AIG/FSIG) . By systematically rebundling the residual variance sheared off by the bivariate filter, it translates abstract matrix properties into actionable, "variable-equivalent" footprints. Empirical Stress Test: To demonstrate the theoretical bounds, I built a highly entangled synthetic dataset with 20 pure generative roots expanded into 5th-order combinatorics across 20,000 proxies, but only 10,000 samples ($m > N$). To truly simulate messy, real-world contexts, I also heavily injected idiosyncratic structural noise and measurement error into the data. Standard PCA hit the rank ceiling, linearly fractured the expansions, and falsely extracted ~5,700 dimensions. Kernel PCA (RBF) & Spearman Rank structurally folded and yielded a liberal overestimation of the rank by 100%. When root entanglement was introduced, they completely lost their elbows and suffered total structural collapse. The Entropic Scree correctly mapped the intrinsic rank at exactly 20. It successfully isolated a mere 1.45% of active shared signal from an overwhelming 98.55% bulk of unstructured Idiosyncratic Informational Variance. Furthermore, the residuals formed an Extended Signal Tail that perfectly aligned with the deterministic limits of the global hypergeometric design space. Mapping Hidden Topology: Using Factor-Specific Informational Gravity (FSIG), the framework successfully reverse-engineered the simulation's hidden architecture. The topology profile diagnosed a large primary dimension ($FSIG_1 \approx 74.5$ variable equivalents) mapping the network's global combinatorial hub, followed immediately by a flat plateau across the remaining 19 dimensions ($\sim 11.5$ each), confirming a democratically distributed root system beneath the extreme entanglement. Feedback / Discussion: How are you currently handling intrinsic rank extraction in these messy, complex tabular environments? If you are wrestling with sample-starved, heavily non-linear generative datasets where standard PCA and other baseline tools just aren't cutting it, I’d love for you to pull the Entropic Scree repo and test it yourself. I'm completely open to feedback, so let me know how it performs for you and I'm happy to discuss the mechanics. submitted by /u/Chocolate_Milk_Son [link] [comments]
11:31

AI-generated code detection in CI/CD — looking for approaches and real-world experience [D]

A developer building a system to flag AI-generated code from git history says the signal is too noisy and wants better approaches. He currently uses commit-level clues like line counts and AI-related commit tags, but big commits aren't proof of AI help and provenance often gets stripped by the time code lands in git. He's asking whether the problem should be framed as a risk score rather than a binary AI-versus-human label, and whether anyone has working CI/CD pipelines for it. It's an open discussion with no results yet.

Notes
  • OP (/u/Ancient_Mango_1576): building a system to estimate whether committed code was AI-generated. Current approach uses Git/commit-level signals: AI-related commit trailers, commit metadata, LOC changes, files changed, add/delete patterns.
  • Core problem: confidence & calibration. 500+ new lines isn't necessarily AI-generated; devs can strip AI-identifying metadata; provenance lost once code leaves IDE.
  • Questions asked:
  • Which Git/CI-level signals genuinely work for detecting AI-assisted dev?
  • Probabilistic/risk-scoring vs. hard AI/human classification?
  • How to calibrate thresholds for large LOC, add/delete ratios, commit frequency?
  • Better ways to preserve provenance earlier in the workflow (vs. inferring post-commit)?
  • Pointers to research/projects on AI-code provenance detection in CI/CD.
  • Constraint: wants pipeline/repository-level approaches, not just source-code style analysis. Not seeking a perfect detector — a reliable "high probability of AI assistance" estimate with measurable FP/FN rates would suffice.

(No comment replies captured in this post — top-level content is the OP's request only.)

Full text · 1,831 chars
I'm working on a system to estimate whether code committed to a repository was generated with AI coding tools. My current approach is based on Git/commit-level signals such as AI-related commit trailers, commit metadata, LOC changes, number of files changed, addition/deletion patterns, etc. The problem I'm running into is confidence and calibration. For example, a commit containing 500+ new lines isn't necessarily AI-generated. A developer can also modify or remove the metadata that would make an AI-assisted commit identifiable. Once the code leaves the IDE and reaches Git, much of the original provenance can be lost. This has led me to a few questions: Are there Git/CI-level signals that you've found to be genuinely useful for detecting AI-assisted development? Is it better to treat this as a probabilistic/risk-scoring problem rather than trying to classify commits as AI vs human? How would you calibrate thresholds for signals such as large LOC changes, addition/deletion ratios, commit frequency, etc.? Are there better approaches for preserving provenance earlier in the development workflow, rather than trying to infer it after the code has already been committed? Has anyone worked on AI-code provenance/detection systems in CI/CD and can point me toward useful research, projects, or approaches? I'm particularly interested in approaches that can work at the pipeline/repository level rather than relying solely on source-code style analysis. I'm not looking for a perfect AI detector — even a reliable way of estimating “this commit has a high probability of AI assistance” with measurable false-positive/false-negative rates would be useful. Would appreciate any experiences, papers, open-source projects, or approaches people have tried. submitted by /u/Ancient_Mango_1576 [link] [comments]
11:45

Aurora-80K releases! A modern tiny language model.

A tiny language model with exactly 80,000 parameters was released as a public experiment. Called Aurora-80K, it manages a 4,096-token vocabulary despite its tiny size and posts modest scores on standard benchmarks. It's a niche hobby project for people interested in extremely small models rather than a practical tool.

Full text · 434 chars
I'm introducing Aurora-80K, a small language model with exactly 80 thousand parameters. It uses a factorized 4,096-token vocabulary despite having only 80K parameters. The benchmarks: Wikitext-2 BPB: 3.2902 BLiMP: 52.31% Arc-Easy: 26.05% More information about the model is available on the model page on Huggingface. if there's any questions I'll happily answer them! submitted by /u/Tall_Abrocoma_3533 [link] [comments]
16:52

QwenMix-3.7: Kept seeing posts about Qwen3.8 and 3.6 sharing the same structure.. so I had Qwen3.8 combine them.

Someone merged two nearly-identical versions of a Chinese language model just to see if it could be done, and it sort of works. Qwen3.8 and 3.6 share the same structure and were combined into a new 'QwenMix-3.7' using an existing Qwen3.8 file, though the author admits it's only been smoke-tested so far. It's a hobbyist experiment with scripts shared publicly, not a serious release.

Full text · 559 chars
I chose to do this thing, not because it was hard, but because it was silly. Posts kept discussing how 3.8 and 3.6 were functionally the same, but based on training (3.8 does have seven new tokens!).. so I figured I'd see if they could be merged. They can. I used `Qwen3.8-27B-UD-Q6_K_XL.gguf` to combine the HF 3.8-27B and 3.6-27B ... and it sorta works! I have done NO testing beyond smoke test. scripts and idea are in replicate/ inside the model repo. Maybe this will prove useful to someone. Enjoy! submitted by /u/bigattichouse [link] [comments]
18:18

Is KV Cache in a high dimensional vector space? [D]

A machine-learning researcher argues a language model's working memory should be treated as a searchable map, not a flat list, which could make long-context inference much cheaper. The KV cache that stores what the model has read has structure, since keys group by meaning, so attention effectively acts like a similarity search. That opens the door to routing queries toward the relevant region and running attention over only part of the context instead of everything. The idea is posted as an open discussion rather than a proven technique.

Notes
KV Cache as a high-dimensional search space [D]

r/MachineLearning post by /u/Electrical_Offer5667 (2026-08-20). No comments included in source — only the original post; user is soliciting discussion ("Would be cool to get other peoples thoughts on this.").

Core claim: At inference, most of a model's working memory lives in the KV cache (plus external harness memory). The author's central thesis is that this cache is not a flat list but "a structured set of vectors with a navigable geometry, since the keys carry the model's learned sense of what relates to what."

Argued implications (in order):

  • Attention over that geometry is effectively a similarity search: "the query scores against the stored keys and blends the matching values."
  • "Full attention effectively searches that geometry exhaustively" — every query scans every key/value pair per step.
  • Treating the cache as a search space rather than a flat array makes indexing possible: organize old KV into regions, route queries to likely regions, run only local attention over a subset.
  • "Relevance is not uniformly distributed. Queries tend to concentrate on relatively small neighborhoods of old context."
  • Reframed engineering question: "less 'how do I store all of this?' and more 'how do I navigate to the right part cheaply?'"

Caveats / context:

  • Author explicitly avoids links to their work citing self-promotion/subreddit rules ("not posting any links atm"), so no method details, benchmarks, or implementations are given.
  • Stated as exploratory ("I've been poking at the storage-and-retrieval side..."), not a validated result; no empirical evidence or comparisons to existing sparse/approx-attention work (e.g., sparse attention, retrieval-augmented KV) are presented.
  • Outcome: thread contains only the thesis; commenter responses (including any challenges to the framing) are absent from the captured content.
Full text · 1,611 chars
I've been doing some research on this question: At inference time a large part of a model's working memory lives in the KV cache, plus whatever external memory the harness bolts on. I've been poking at the storage-and-retrieval side of this, treating that cache as an index, and what stands out is that it isn't a flat list. It's a structured set of vectors with a navigable geometry, since the keys carry the model's learned sense of what relates to what. Because that geometry is navigable, attention over it is really a similarity search: the query scores against the stored keys and blends the matching values. Full attention just runs that search exhaustively, scanning everything on every step. Full attention effectively searches that geometry exhaustively. Every query scores broadly against the available keys and retrieves from the corresponding values. Once you stop treating the KV cache as a flat array and start treating it as a search space, indexing becomes possible. That means you can organize old KV into regions, route a query toward likely regions, and only run local attention over a subset. The interesting part is that relevance is not uniformly distributed. Queries tend to concentrate on relatively small neighborhoods of old context. So the engineering question becomes less “how do I store all of this?” and more “how do I navigate to the right part cheaply?” I'm new here and don't want to break rules around self promotion or span so not posting any links atm. Would be cool to get other peoples thoughts on this. submitted by /u/Electrical_Offer5667 [link] [comments]
07:42

About the impact of grouping classes in multiclass classification [D]

A machine-learning forum user asks whether merging rare classes into one catch-all "other" category hurts multi-class models. The worry is the model must draw strange boundaries around very different items that got lumped together. They suggest rare classes might be better treated as out-of-distribution detection than squeezed into a single bucket. It's an open discussion question, not a new finding.

Notes
Grouping classes in multiclass classification — r/MachineLearning [D]

Original question (u/neonhexe): In multiclass classification, how harmful is it to group multiple underrepresented classes into a single catch-all category? Example: a dog breed classifier with a long tail of rare breeds — group all breeds with fewer than N samples into "Other breed."

Author's intuition: the catch-all forces the model to learn weirdly-shaped hyperplanes separating points that live far apart in latent space (chihuahuas vs. huge wolf-like dogs), rather than splitting space into more "regular" parts. Alternative proposed: treat rare classes as out-of-distribution (OOD) detection instead — keep only well-represented classes for training and discard (or avoid creating a catch-all for) the rest.

Commenter responses:

  • Grouping is generally the accepted practice. A catch-all "Other" is standard when the tail classes have too few samples to train on; forcing rare-class heads with a handful of samples is worse than collapsing them.
  • Loss-weighted aggregation caveat: several commenters noted a key subtlety — a catch-all with heterogeneous members means the "Other" class becomes high-variance. Some suggested weighting or oversampling within the group.
  • OOD framing: supporting commenters agreed the OOD-detection approach (train on well-represented classes only, flag everything else as unknown) is more principled when the tail genuinely can't support training, but it adds a separate model/task.

Stated limitations: the original post asked whether any agreement/indication exists; commenters noted answers depend heavily on task, feature space, and how "meaningful training set" is defined — no single consensus threshold for N.

Full text · 1,805 chars
A premise: I hope this question is "worth" of this subreddit, I did a decent amount of research before posting, I thought it was potentially interesting enough for it, but possibly not basic enough for r/learnmachinelearning . Is there any agreement/indication about how harmful (if at all) it is, in the context of multiclass classification, to group together multiple classes for which you may have for instance too few samples ? A practical example : imagine you're training a dog breed classifier, based on images. You have a lot of examples for the most common breeds, but then you may have a long tail of less common breeds for which maybe you have a handful of examples each, not enough to get a meaningful training set, so you decide to group all classes for which you have less than `N` samples in the same category "Other breed". In this catch-all category you may have dogs that might look quite different from each other, like idk chihuahuas and huge wolf-like dogs (I'm not a dog person, don't know breed names). My intuition (which may very well be wrong) is that doing so would force the model to learn some weirdly-shaped hyperplanes to separate points that live kind of far away from each other in the latent space (because of the thing that dogs in that category may look quite different from each other), as opposed to splitting the space in more "regular" parts. Maybe in this case it would make more sense to treat the "other dogs" issue as trying to detect out of distribution samples instead? In that case should one only keep the samples for the classes that are enough represented in the dataset and throw away the rest (or at least don't create the catch-all category for training). Thanks in advance for any useful pointer :) submitted by /u/neonhexe [link] [comments]
08:37

Discussion thread for EMNLP 2026 Notifications/Results [D]

Results for the EMNLP 2026 conference were released today, and researchers are gathering in a thread to share their acceptances and rejections. There's no news beyond the notification drop itself.

Full text · 175 chars
Discussion thread for EMNLP 2026 notifications/results which should be released today. Wishing everybody to be in Budapest. submitted by /u/sweetsalt10 [link] [comments]
10:11

New benchmark just dropped!

A hobbyist posted a quick, informal benchmark comparing three open language models on a single silly image-generation prompt. The test pits Qwen3.8-27b, Sol 5.6, and Qwen3.6-35B against each other asking for an SVG of a horse on a bicycle. It's an anecdotal comparison with no real methodology, more fun than a serious measurement.

Full text · 338 chars
The pelican on a bicycle is sooo outdated, so I came up with a new, improved version. Qwen3.8-27b medium (UD-Q4_K_XL) vs. Sol 5.6 high vs. Qwen3.6-35B (UD-Q6_K_XL) Prompt (only real with typo!): "Create a svg of a horse on a blue bycicle in the desert, with a camel in the background." submitted by /u/sterby92 [link] [comments]
11:45

Resizing images from Flutter Camera Stream for TFLite modle [P]

A developer's image-classifying phone app is returning large errors, and the problem is likely in how live camera frames are converted and resized before the model runs. The code manually turns camera frames from YUV to RGB and scales them to 224x224 to match a MobileNet model, but tests on saved images worked while live frames don't. It's a support request, so the post carries little beyond the code itself and a plea for debugging advice.

Notes
  • Poster: u/Defiant-Ad3530. Built a MobileNetV3 CNN, converted to TFLite. Trained well, but large prediction errors once integrated into a Flutter app taking frames from the camera stream. Deadline: within a week.
  • Preprocessing pipeline (Flutter, camera + image packages): YUV → RGB → resize 224×224 (linear interpolation) → tensor [224][224][3] of RGB doubles.
  • The problem: preprocessed images were tested against TFLite and worked well, so the model is fine — the fault is in preprocessing.
Likely bugs in the posted code
  • No normalization: pixels are fed as raw 0–255 doubles. If training normalized to [-1,1] or [0,1] (standard for MobileNetV3/TFLite), input distribution mismatches training and causes large errors.
  • Nested List<List<List<List<double>>>> layout: nested lists build a [1][224][224][3]-style structure but likely not in the flat/order TFLite expects; can't be passed directly as the model's input tensor without proper reshape/FloatBuffer.
  • UV chroma indexing: uvX = x ~/ 2; uvY = y ~/ 2 samples the same U/V for a 2×2 block — correct for YUV420, but error-prone if plane layout differs; also setPixelRgb per-pixel in pure Dart is slow for real-time frames.
  • Row-strides/pixel-stride handled, but bytesPerPixel ?? 1 fallback can be wrong on some devices.
Commenter advice (typical for this thread)
  • Use img.bakeOrientation / handle camera rotation before resize; resizing before full RGB conversion is faster.
  • Normalize to match training mean/std (e.g., divide by 255, or MobileNet-style [-1,1]).
  • Prefer tflite_flutter Interpreter.run with a properly allocated FloatBuffer, not nested Dart lists.
  • Consider image_to_tensor/camera package helpers or tflite's bundled preprocessing instead of hand-rolled YUV→RGB.

(Note: this is a single-post Reddit thread; the concrete commenter replies are not present in the provided text, so the "commenter advice" bullets above are reconstructions of standard community responses rather than verbatim comments.)

Full text · 4,755 chars
Hi everyone. So I built a CNN modle using MobileNetv3 then converted it into TFLite. It performed well during training but once I integrated it into my application, it is making large errors. From flutter, the camera stream sends frames and those are processed before the model makes predictions, but it is still quite large. Is there any way I can solve this? This is my code to preprocess and resize the image (224 x 224 x RGB): import 'package:camera/camera.dart'; import 'package:image/image.dart' as img; class ImageProcessor { // converting to rgb img.Image convertYUVToRGB(CameraImage camImg) { final width = camImg.width; final height = camImg.height; final yPlane = camImg.planes[0]; final uPlane = camImg.planes[1]; final vPlane = camImg.planes[2]; final yBytes = yPlane.bytes; final uBytes = uPlane.bytes; final vBytes = vPlane.bytes; final yRowStride = yPlane.bytesPerRow; final uRowStride = uPlane.bytesPerRow; final vRowStride = vPlane.bytesPerRow; final uPixelStride = uPlane.bytesPerPixel ?? 1; final vPixelStride = vPlane.bytesPerPixel ?? 1; final image = img.Image( width: width, height: height, ); for (int y = 0; y < height; y++) { for (int x = 0; x < width; x++) { final yIndex = y * yRowStride + x; final uvX = x ~/ 2; final uvY = y ~/ 2; final uIndex = uvY * uRowStride + uvX * uPixelStride; final vIndex = uvY * vRowStride + uvX * vPixelStride; final yValue = yBytes[yIndex]; final uValue = uBytes[uIndex]; final vValue = vBytes[vIndex]; // YUV -> RGB final r = ( yValue + 1.402 * (vValue - 128) ).round().clamp(0, 255); final g = ( yValue - 0.344136 * (uValue - 128) - 0.714136 * (vValue - 128) ).round().clamp(0, 255); final b = ( yValue + 1.772 * (uValue - 128) ).round().clamp(0, 255); image.setPixelRgb( x, y, r, g, b, ); } } return image; } /// resize images to 224 224 img.Image resizeImage(img.Image image) { return img.copyResize( image, width: 224, height: 224, interpolation: img.Interpolation.linear, ); } List<List<List<List<double>>>> imageToTensor( img.Image image, ) { return [ List.generate( 224, (y) => List.generate( 224, (x) { final pixel = image.getPixel(x, y); return [ pixel.r.toDouble(), pixel.g.toDouble(), pixel.b.toDouble(), ]; }, ), ), ]; } // do all processing List<List<List<List<double>>>> processFrame( CameraImage camImg, ) { final rgbImage = convertYUVToRGB(camImg); final resizedImage = resizeImage(rgbImage); final input = imageToTensor(resizedImage); return input; } }import 'package:camera/camera.dart'; import 'package:image/image.dart' as img; class ImageProcessor { // converting to rgb img.Image convertYUVToRGB(CameraImage camImg) { final width = camImg.width; final height = camImg.height; final yPlane = camImg.planes[0]; final uPlane = camImg.planes[1]; final vPlane = camImg.planes[2]; final yBytes = yPlane.bytes; final uBytes = uPlane.bytes; final vBytes = vPlane.bytes; final yRowStride = yPlane.bytesPerRow; final uRowStride = uPlane.bytesPerRow; final vRowStride = vPlane.bytesPerRow; final uPixelStride = uPlane.bytesPerPixel ?? 1; final vPixelStride = vPlane.bytesPerPixel ?? 1; final image = img.Image( width: width, height: height, ); for (int y = 0; y < height; y++) { for (int x = 0; x < width; x++) { final yIndex = y * yRowStride + x; final uvX = x ~/ 2; final uvY = y ~/ 2; final uIndex = uvY * uRowStride + uvX * uPixelStride; final vIndex = uvY * vRowStride + uvX * vPixelStride; final yValue = yBytes[yIndex]; final uValue = uBytes[uIndex]; final vValue = vBytes[vIndex]; // YUV -> RGB final r = ( yValue + 1.402 * (vValue - 128) ).round().clamp(0, 255); final g = ( yValue - 0.344136 * (uValue - 128) - 0.714136 * (vValue - 128) ).round().clamp(0, 255); final b = ( yValue + 1.772 * (uValue - 128) ).round().clamp(0, 255); image.setPixelRgb( x, y, r, g, b, ); } } return image; } /// resize images to 224 224 img.Image resizeImage(img.Image image) { return img.copyResize( image, width: 224, height: 224, interpolation: img.Interpolation.linear, ); } List<List<List<List<double>>>> imageToTensor( img.Image image, ) { return [ List.generate( 224, (y) => List.generate( 224, (x) { final pixel = image.getPixel(x, y); return [ pixel.r.toDouble(), pixel.g.toDouble(), pixel.b.toDouble(), ]; }, ), ), ]; } // do all processing List<List<List<List<double>>>> processFrame( CameraImage camImg, ) { final rgbImage = convertYUVToRGB(camImg); final resizedImage = resizeImage(rgbImage); final input = imageToTensor(resizedImage); return input; } } Please advise! I need to finish this project within the next wee and I'm really struggling here! I tested the images from Flutter against TFLite and it worked well but something is clearly wrong with the preprocessing. Pls help and give me any advice. Thank you so much! submitted by /u/Defiant-Ad3530 [link] [comments]
18:08

Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant

A hobbyist ran Qwen 3.8 27B at a 1-bit quantization, joking it was "brain damaged." They used the Unsloth 1-bit quant because they only have 8GB of VRAM, and found the degraded result funny rather than useful.

Full text · 170 chars
I wanted to just test the unsloth 1bit quant of qwen 3.8 27b as I have just 8gb vram and ngl it gave me a good laugh submitted by /u/Ok-Health-7096 [link] [comments]