Nothing matches those filters.

Lead

15

Video

4
16:48

This Small AI Will Change Everything

An open-source AI model small enough to run on a powerful laptop is getting millions of downloads in its first week and can hold its own against far bigger systems. Qwen 3.8 packs 27 billion parameters and beats last year's billion-dollar frontier systems in some tests, though it isn't fully frontier-grade. Its architecture is unchanged from the previous version — the gains come from a training regimen that starts with simple tasks and scales up to hard ones lasting days.

Notes
Qwen 3.8 — open-weights small model (Two Minute Papers, 2026-08-24)

Two Minute Papers' claim: Qwen 3.8 is the most important of the thousands of free/open AI systems released each year, because it packs frontier-adjacent capability into a laptop-runnable package.

Concrete claims

  • 27 billion parameters — runs on "a beefy laptop"; "millions and millions of downloads" in under a week.
  • In some tests it "holds its own" against a current frontier model and "easily better than a billion-dollar system from just a year ago" at this size.
  • Free, open-weights.

Why it's possible — no architecture change

  • The video compares the architecture diagram to the previous version: "This looks identical to me." So gains are not architectural.
  • The answer is training. The model card reportedly shows a curriculum-style regimen: the agent is first given simpler tasks, then scaled up, then multiple and more challenging tasks — "an intense training regimen that progressively gets harder and longer," with later tasks taking days to complete. Analogy drawn to how humans train muscles.

Caveats / framing

  • The host concedes the memory-shortage/cost complaints about frontier AI are "all true, 100%" — but argues waiting may yield frontier-level systems on laptops, enabled by open science. He frames it as a message of hope and encourages community tinkering (he says he's fine-tuning/experimenting with it himself).

Sponsorship

  • Sponsor segment: Lambda (lambda.ai/papers) provides NVIDIA GPUs for reproducing research papers in minutes, training/fine-tuning, inference, text-to-image/video, and running DeepSeek chatbots or agents.
Transcript · 2,678 chars
This is an amazing open-weights AI system free for all of us called Qwen 3.8. This is the smaller brother, and this is the one that will change the world the most. And millions and millions of downloads in less than a week, and it feels like it can do everything. Unless you need frontier stuff, it does all you need. But, look, out of thousands and thousands of free and open AI systems released each year, this might be the most important one. Why? Because this one is 27 billion parameters. So, if you have a beefy laptop, yep, you can just run it there. Now, it gets better. In some tests, even against a current frontier model, it kind of holds its own and easily better than a billion-dollar system from just a year ago in such a small package. Huh. How the heck is that possible? Now, hold on to your papers, fellow scanners, because the architecture of the previous version looked like this. So, let's see how 3.8 is different. Wait a second. This looks identical to me. Wow. So, not because of architectural changes. Then, how? How the heck is it possible to put this density of intelligence in such a small package? What is this magic? The answer is training, lots of it, but in a way that is similar to how we humans train our own muscles. Clues from the model card seem to point in the same direction. They first give the AI agent simpler tasks, and then they scale it up. Then, multiple tasks and more challenging tasks. An intense training regimen that progressively gets harder and longer. Yep. Later tasks take days to complete. I think this is amazing news for all of us. You see, everyone is talking about the memory shortage and everything costing a fortune, and it's all true, 100%. But, it seems that if we wait a bit, we might get frontier-level systems running on our laptops, and all this is only possible because of the power of open science and research. And look at how much it just changed the game already for AI that you can run at home. A message of hope for all of us. Incredible. Huge thank you for this. Once again, we all get this for free, and it is lovely to see you brilliant fellow scholars tinkering with it and improving it already. I'm doing it, too, with you. Love it. What a time to be alive. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text-to-image or video, easy-peasy. Running a deep seek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover, and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.
13:56

New: your agent lands customers

An open-source SEO tool lets you hand an AI agent real keyword and backlink data for a fraction of what big services like Semrush cost. It's free on GitHub, and you only pay for the underlying data: $10 a month through the hosted service or $50 upfront plus per-usage credits. It connects to Claude through the MCP standard, so an agent can run competitor research, find dead links, and automate outreach. The catch is the data still costs money, though creators say $50 can last months of light use.

Notes
Open SEO: open-source alternative to Semrush/Ahrefs

Interview on The Next New Thing (2026-08-24) with Ben, creator of Open SEO, an open-source SEO tool on GitHub. Sponsored by Zapier (Zapier SDK).

Pricing contrast
  • Semrush starts at $117/yr (annual), scales up for agencies/teams.
  • Ahrefs starts at $130/mo for a limited feature set; $240/mo for the full suite.
  • Open SEO: repo is free to self-host (Docker/Cloudflare, full-stack app). The only cost is data.
Data layer (not truly free)
  • Uses DataForSEO (pay-per-usage provider), not a subscription data index like Semrush/Ahrefs.
  • Two options: buy credits directly from DataForSEO for $50 and pay per use, or use the hosted service at $10/month.
  • Ben's estimate: $50 of credits lasts 6–12 months, possibly "forever" for light users. Heavy months: buy extra credits.
Why data matters
  • Keyword volume grounds the model: e.g. "AI tools for marketing" = 4,400 searches/month. Ben's claim: asking Claude to do SEO without this data makes it "just writing what it thinks people want," not grounded in real search counts.
Features (vs. Semrush/Ahrefs)
  • Competitor keyword analysis + top pages (demo: ahrefs.com ranks for "backlink checker," "search operator," "check backlinks").
  • Backlink data ("if other reputable sites link to you, Google reads your site as reputable").
  • AI-native via MCP: connect through Claude Code; an SEO coach skill lets Claude walk a novice through strategy and complex topics (conversational coaching).
  • Hosted app includes keyword research UI; the deeper features (entity analysis, content optimization) were deliberately left out to avoid overwhelming beginners.
Community builds (the "more exciting" part)
  • Eric (professional SEO) built his own content optimization / entity analysis page — analyzes whether an article mentions target keywords densely enough (title, headings, body) to rank.
  • Terry (freelance SEO) built an outreach/CRM tool — Claude via MCP scrapes pages, finds contact emails, stores them, and follows up on non-responders.
  • Google Analytics integration — wanted for ~5 months, many users built their own independently; officially added the day of the video. 66 open pull requests exist; Ben folds the best ideas back into the core tool.
Working via Claude Code
  • Ben mostly runs his own site SEO through one Claude Code chat window ("I mostly only go to my own website when customers reach out with fixes").
  • Example overnight brief: "use the SEO audit skill to review our website, do keyword research, and build comparison pages for Semrush, Ahrefs, and Ubersuggest" — constrained to pricing and "AI native, easier for entrepreneurs" framing (avoiding feature-by-feature since incumbents are 20 years old). Claude produced full pages, tables, "fact checked" them, and wired the comparison pages into the site footer, all via Claude Code.
  • Demo data: "Semrush alternative" keyword = ~300 searches/mo, high intent, deemed worth ranking for.
  • Strategy library: content briefs written by a professional SEO adviser; Claude writes the pages/strategy list, does internal linking, and iterates on design.
Marketing positioning + traffic results
  • Deliberate angle: "open source SEO" — there was no open-source SEO tool when he started, so he "can win the open source SEO market."
  • Google Analytics (demoed): Product Hunt launch produced a big awareness spike; Google is "by far our biggest channel."
  • Explicit caveat against mass page factories: "I haven't just made a ton of pages... I think Google will eventually detect those all as spam." That's his stated reason for fewer, higher-quality pages.
The mistake (stated limitation)
  • Ranked #1 for "backlink checker": 200,000 impressions in ~1 week → only 100 clicks (~0.05% CTR). Ben calls impressions a "vanity metric" if it doesn't convert.
  • To raise CTR he converted his page into a free backlink checker, but the page was likely ranking because it wasn't one; rankings dropped and haven't recovered ("I'm sure we'll be back").
Links promised in the video
  • GitHub repo, hosted pricing page, and a report with the exact prompts shown, provided in the video comments.
Transcript · 20,403 chars
SEO tools get you traffic, but they cost hundreds of dollars a month. Well, someone open sourced an alternative and put it on GitHub. You can just give it to your agent and say, "Use this to get me traffic." I did an interview with the creator of that open source project. He's about to show you how it works and how you can get your agent to use it. Let's get into it. Presented by Zapier, the AI automation company. Ben, what are we looking at? Cuz this is not your website. >> Yeah, this is the Semrush pricing page. So, as you can see here, it starts at 117 build annually. Um, and then it, you know, scales up, uh, for larger agencies and teams. Um, and if you look at AFS, it's pretty similar where it starts at $130 and you get a limited feature set and then if you want the full suite, you got to pay $240 a month and it scales up from there. >> Okay. Meanwhile, I see I see a GitHub page for you. Show me that. >> Yeah. So, I wanted to do SEO for myself as like a solopreneur. Um, and I was looking at the tools and they were just so overwhelming and expensive for me. So I was like, maybe I should just build an open source alternative and that's what I did. >> So this version is completely free. I could download it, install it on my computer and then as we'll see later on in the conversation, people have done it and then they've customized it. And the customization I think is maybe more exciting than the price. I would say definitely more exciting. But we also need to be open that when you use this, it's not completely free. You still need some data that this software is analyzing to make SEO suggestions. How much is the data source costing? >> Yeah. So the whole value historically of Samrush and ARS was that they have really good data. But in the last few years there's new data providers coming up that are pay per usage which is what we use. So we use one called data for SEO. Mhm. >> Um, so you can buy from data for SEO directly for $50 and then pay per usage or you can use our hosted service and pay $10 a month. Um, and for a lot of entrepreneurs that $10 a month is perfectly fine for them to get started. And then if you have a really heavy month where you're really focusing on on SEO, you can buy credits as you need. But most customers are perfectly fine like $10 a month. >> Okay. We'll see your pricing page in a moment. But if I'm using this GitHub uh repo on my system, not paying you nothing, I pay them $50 and then I pay per use. Kind of like I do per credit. How long do you think $50 would last on a site like mine? >> Um I don't have much. >> Probably a pretty long time. It it could last 6 months, 12 months. It could last you forever if you're not doing that much SEO. >> Okay. All right. Fair enough. Let's see your pricing page and then let's just see what comes in it because I've seen all these different SEO tools and I've promoted them, but I'm not as aware of them as other people. So $10 a month. What are the features that come in here? >> Yeah. So, what you're really paying for is the data. So, um keywords. So, you can ask Claude to do SEO for you and it'll do SEO for you kind of, but it just essentially makes things up. Um keyword data grounds the the searches. So, I can just pop into the app. Um and if we go to like AI tools for, you can see how much search volume there is for all these things. So if we go to let's say >> you can say how many people are searching for AI tools for marketing it's 4,400 people a month. If there was no one searching for it then why would you write content for it because you're not going to get any results. >> Got it. And so when I do SEO with Claude without this data claude is just kind of guessing by using free tools that are available and also by trying to predict based on what it knows. It's not grounding it in in real data of about how many people are doing actual searches. >> Yeah. It's not even using free tools. It's just making it's just writing what it thinks people want. But >> that's not as good of a strategy as writing what people actually want. >> Okay, fair enough. Okay, so that's what I get. And then once I get all that fire hose of data from the data source or I get access to that data, what does Semrush Hrefs and now Open SEO give me? What are these features and what can I do with them? And we'll see you in action, but just give me an overview. Maybe even show me on one of your sites. Yeah. So we can look at let's look at like arfs.com. >> So uh you can research what your competitors like what keywords they rank for. So they rank for backlink checker, search operator, check back links and then you can see what all their top pages are. So from here I can kind of analyze their uh SEO strategy and their marketing strategy and then try to learn from it. How can I then market my brand? >> You know what? Let's let's take a moment and and use my my site just for a second. >> There's a creator, Matthew Burman, who I interviewed who does interviews and shows very similar to mine. His website is uh forwardfuture.com. So what I would do here is I would say let's study what he's doing and understand it. This is one of the features that comes on your soft with your software. And so what am I learning on this? >> Yeah. So you can see they rank for open future AI, Kodi, Kira. So nothing. Yeah. So here they're writing a >> Yeah, but I mean they're getting a fair amount of traffic. Um but I see. Yeah, it's not super high search volume. >> Let's see what that Cody is. So now if Kodi is one of the top pages, I guess >> boost your creative productivity with AI. Cody, I see. Got it. So now I see that this tool Cody is is effective on his site. Maybe I reach out to the Cody team and get them to sponsor or put them up on my site and write an article about them. Okay, that's one >> absolutely >> one feature. What else? What are some of the other features? >> Yeah. And then so I'll show you some of the other features in the app, but the really cool thing about Open SEO, especially as compared to ARS and Semrush, is >> um it's really AI native. So all of this data is available through Claude code. Um so even if you're new to SEO, you can connect via MCP and then Claude can just essentially coach you through it. So, we have this SEO coach skill where you can talk back and forth with it and it can help tell you like what you should do and what strategy you should investigate and uh help you with some of the more complex parts of SEO. >> Um, I'll show you the >> backlinks tool so you can see all the back links we have. >> Um, or maybe it'd be better for forward future actually. Forward. >> Well, they're not they're not getting that much traffic. >> Um, yeah. So these back links are important because um it essentially tells Google that your website is reputable if other reputable websites are linking to it. So it's the same as like um if people are talking about you on Reddit, then the AI knows, hey, this person's probably pretty important. So this backlink data is also really helpful. It's kind of a lot to look at if you're new to SEO, but this is where the AI integration is really helpful. Uh where you can use it through Claude. You can ask Claude what takeaways you should have from this. >> Okay. >> Do you have any questions on this specifically? >> No, I think this makes sense. So now what I would want to do is if he had a lot of good backlinks, I might go out to them and ask them to link to me too and say, "Hey, I've got another resource." Right? The other thing that I might do is um pick a couple of newsletters that are dead in the AI space and then go out to people who are linking to them and say, "The site that you're linking to is dead. If you want to link to me, here's how to do it." That's the kind of stuff that SEO people do, right? >> Yep. Absolutely. And you can really start to automate that too with Claude where Claude can go look at all the websites and look for dead links and it can look for the emails of people and really streamline that whole process that used to be really labor intensive. >> All right, I see how this works. And the beauty that you're going to show me later is that people then integrate this into their own software. And so then they create a CRM that keeps track of backlinks that their competitors get, that maybe uses another tool to send out email asking those sites to link to to their own too. keeps track of the followup and all that stuff. This is the kind of thing that can be done uh with Open SEO. >> Yeah, absolutely. >> By the way, if you're creating software and you want your users to connect an app to your software, Zapier SDK will make that possible. You have access to thousands of apps with Zapier SDK and all your user has to do is connect your app to the apps that you choose and they have access and suddenly boom, your app is so much more powerful. go to zapier.comdk for that. Show me what dashboards people are making and what tools they're building. >> Yeah, let's start with this one. So, um, this guy Eric, he wanted to do he's like a professional SEO and he wanted to do something called content optimization, which is kind of a more advanced feature, which is why we don't have it directly in the tool since we don't want to overwhelm people new to SEO. >> Um, but he built a content optimization page straight into SE into Open SEO for himself. He built all of this crazy stuff um for his own needs. And yeah, I've heard stories of people building like whole CRM and essentially since the open SEO repo is pretty good and it's just a full stack uh app that you can easily self-host on Docker on Cloudflare, people are just using it as like the foundation for their home base for their whole business like unrelated to SEO. >> Okay. So where before I saw your domain up in the address bar and now I see local host and so this is him hosting a version of your site of of your tool >> and then just help me understand this because I'm not an SEO person. >> What is what's going on here? What did he create? >> Yeah. So this is called like entity analysis. Um >> okay. he is trying to so the the way uh Google shows your content is it like looks at all of the text on your page and it prioritizes like the title and the headings. So like the section headings but then also the contents of your page. So um >> if you want it to rank for a certain keyword like let's say schema or headers and then a couple other things it helps if it's mentioned a bunch in the article. So that's what this is trying to do is it's helping him analyze articles is is it dense enough with keywords will rank for the keywords he's trying to target. >> Okay. >> Um this is more I wouldn't worry about this stuff if you're just getting into SEO. It's more like are you uh I would pick a keyword or a question someone might ask and then I would try to write content that answers their question and then over time you can optimize uh that content uh as you learn more. >> Okay, I think that makes sense. Let's see another. >> Yeah, so this is kind of getting into the CRM. So this is in our Discord. We have a pretty active Discord where people are see someone's typing right now. Uh where people are sharing what they're working on, asking questions. Um, and this guy Terry, he he's a he's an SEO freelancer. Um, and he built this outreach tool. Um, so you can see here he's like tracking contents that contacts that he wants to reach out to. >> Um, and I think via the MCP it can store their email in here as well. Uh, so Claude will go reach out, look at all these web pages, find the emails, and then it stores the people to reach out to. And I think over time it'll be like, oh, they didn't, you know, they didn't respond. Let's go follow up with them. Is where he's kind of going with this. I think that makes >> he has writing messages and doing his whole flow. Uh so this is something I would love to bring into Open SEO eventually, but it's cool that people can experiment and build these proofs of concept for things before it makes it into the official tool. >> Okay, give me one other one. Even if you don't have it handy on Discord, what else are you seeing people build? Hm. Yeah. So, um this feature is actually coming out today, but um lots of people have built Google Analytics uh integrations. Um so, lots of SEOs, they use Google Analytics to track um traffic to their website because it's a free tool. So, it's the most popular one. >> Um and we just added it today, but for the last five months, this has been something people have wanted. So, tons of people built independently their own integration for Google Analytics. That's one of the nice things about being open source that you enable people to create and then you can fold in their best ideas in here. >> Yeah. And if you see here, there's literally like 66 pull requests with people adding different different features that we plan to get to eventually. >> Okay. Um let's see how you use it. Why don't we see some pages that Open SEO has created for you since you're using it to promote your own site? >> Yeah. So I use I use it mostly through Claude. Um, >> you're not even going to your own website. You're just having Claude do it all. >> Yeah, I mostly only go to my own website when customers reach out to me with fixes for the website. >> It's my preference actually to do it all in that one chat window. I think I don't know. Some days yes, some days no. Okay, so this is like a prompt. >> So much faster. >> We agreed that it would just be too slow to watch it work. And so what you've got here is the actual prompt. You said an over overnight brief. Tell me what that is. Yeah. So, basically, um, my I have my codebase for our website. Um, and I wanted to, let's say, make comparison pages. So, this is something that's been on my to-do list. I just haven't done it yet. Um, so I said, "Hey, use the SEO audit skill to review our website, do keyword research, and build comparison pages for Semrush, ARS, and Uber Suggest." Mhm. >> We don't want to do a feature by feature comparison because they have like just tons of features since they've been around for 20 years, but let's focus on pricing and it being easier for entrepreneurs because it's AI native. >> Okay. >> Um, and then it went ahead and it made these comparison pages for us. So, this is another kind of report. This is the other cool thing about using with Claude is you can have Claude just generate these websites for you to explain it to you in the format that makes most sense. Um, so if we look at the Semrush alternative, you know, this is the keyword research. There's,300 searches a month. So, this would be, and this is high intent. This is people that are potential customers. So, this would be a great thing to rank for. And it made this uh this is local host, so this isn't deployed yet, but it made this whole page for us. And I I gave it a read, and it's it's pretty good. There's, of course, some things I'd like to tweak, uh, but this is just an amazing first draft. Um, and it made this whole website. It made this table. Um, and it helps that we've already, you know, built these tables, built similar pages on our website before, but it did all of this research for us. It fact checked it. Um, and yeah, this is just, you know, saving so much time. And then it even added this whole comparison page to the footer all through Cloud Code. So, it's Cloud Code building the website, but Cloud Code's using Open SEO to ground the website, the changes to the website in in reality instead of just being like, oh, Serush versus Open SEO. Okay, I'm with you. Let's see one other page. >> Uh, yeah. So, here's like the Uber Suggest alternative. Well, I'll show you, I guess. So, these are all just three alternatives pages it made. Um, but I'll show you the strategy library. So, this is something we've been working on. I've been working on with like a a professional like SEO adviser. Um, and to make sure that these are, you know, really great strategies that people can learn from. Um, and he's basically making content briefs for the strategy. So he's he's essentially giving bullets for what the strategy should be. Um but then >> and this is a list of strategies for how people can use open SEO. >> Yeah. This is a list of strategies for different ways you can use um you can do keyword research. So >> you can just ask Claude, hey, do keyword research for me. But when you want to go learn and dig a little bit deeper into what it's doing under the hood, this is where you would go. >> Okay. >> Um so it, you know, Claude's making this page for us. um you know you're iterating on design of course um and then it's writing all of these um and it's doing things like internal linking. So one of the ways that you uh improve the ranking of your website is you link to pages within your website to show Google, hey this is an important page. We're linking to it from other important pages. >> Um and yeah, this whole page and blog was written by oop so you can kind of see here because we need to fix that. um is written by um Claude even though it's you know an SEO professional making the the draft. >> Okay, final thing. Let's see your Google Analytics to see how this has all impacted your traffic. By the way, you're going after the most competitive group in the world. Uh-huh. >> Absolutely. Yeah. So, that's that's why um I haven't just made a ton of pages like you may see online of people making thousands and thousands of pages. Um, in my opinion, that's like not a great long-term strategy because I think Google will eventually detect those all as spam. So, I was thinking like what's the, you know, angle that I can position open SEO in. And for me, that was open- source SEO, right? Where there wasn't any open- source SEO, uh, tools. So, it's like, okay, well, I can win the open source SEO market. And it was a smaller market when I started, but through our marketing efforts over time, the traffic to open source SEO has grown. So you can see here we launched on product hunt and there's a big spike in people being aware of open source SEO uh and now we're getting you know tons of clicks and traffic from Google and it's by far our biggest channel. >> I see that it's gone down there at the end. What's going on? >> Yeah. So uh this was this was me being fairly so I I I'm a software engineer and I kind of found this SEO market in February. So uh I'm learning from my customers uh as we go. Um but this is a bit of a mistake I made. So, I started to rank for backlink checker, which you can see here has um a ton of traffic. So, I was getting I got 200,000 impressions in like a week. Um but it only converted to 100 clicks. >> By impressions, you mean page view where people come from Google to your site, they see the page the impression is on the landing page and then clicking from there into the actual order or download is 105. >> No. So, it's actually So, it's 100,000 people like landed on the backlink checker Google result page. >> Okay, I see. >> Um, and saw, you know, our our link, but they didn't click through it. So, only 100 people click through. >> I see. >> To our page. So, this this impression thing, it's a bit of a vanity metric that people will market, but if it's not converting to customers, like it it doesn't really matter. So, I tried to change the content of our website to increase the click-through rate. So, I looked at the Google page and I saw that every page here was a free backlink checker. So, I was like, "Okay, I'll make a free backlink checker and it'll increase my click-through rate." But, it turns out I was probably ranking because my page wasn't a free backlink checker. So, when I changed it to a free backlink checker, it just said, "Okay, well, there's already a billion free backlink checkers, so I haven't been able to get it to recover yet, but I'm sure we'll be back." >> Okay. All right. It is available right now. We'll have links to the GitHub repo. We'll have links to your site for people who want it. Can we also give them that report with all the uh prompts that you just showed? >> Yeah, absolutely. >> And we'll give them that in the comments. And now that you're at the end of this one, I've got another video for you for how to use uh Claude and other tools to market. And there's a link to you for link to it right there. See that?
11:34

Faceless AI Shorts: How One Channel Made $40K in 90 Days

People are earning real money from faceless YouTube Shorts using a nearly fully automated AI pipeline. One tracked channel made an estimated $40,000 in 90 days off 250 million views, and another cleared about $170,000 and two billion views over a year. Shorts pay roughly 7 to 33 cents per thousand views depending on audience, so the math only works at scale. The workflow runs on Claude for competitor research and scripts, ChatGPT for channel branding, and a tool called Rank Reel that assembles finished videos with voiceover and captions.

Notes
  • Channel: "Faceless AI Shorts: How One Channel Made $40K in 90 Days" by Sanji Nai-Chien (YouTube, published 2026-08-24). Claim: a faceless YouTube Shorts-only channel earned ~$40K in 90 days using three AI tools, no on-camera face, no long-form.
The money claims (VidIQ estimates)
  • Three tracked faceless Shorts channels, last 90 days, per VidIQ estimated revenue:
  • Channel 1: ~$40,000 from 250M views
  • Channel 2: ~$25,000
  • Channel 3: ~$33,000
  • Last 365 days, one channel: 2B views, $170K+ estimated revenue. These are estimates, not real payout data — a stated limitation the video never re-addresses.
Why "right now" (3 stated reasons)
  • Search interest is falling, not rising. Google Trends: searches for "YouTube automation" and "faceless YouTube" dropped massively from 2023/2024 peaks; now near a 5-year low — contradicting the common "saturation" narrative.
  • Shorts scale to real money. Air Media Tech analyzed 274 real channels via the YouTube Analytics API. Shorts RPM: $0.07–$0.20 per 1,000 views in most niches; ~$0.33 per 1,000 for US-heavy audiences. At $0.33, 100M views ≈ $33,000.
  • AI now handles almost the whole workflow (previously manual):

| Task | Tool |

|---|---|

| Channel setup (name, description, profile pic, banner) | ChatGPT |

| Competitor analysis, ideas, research, scripts | Claude |

| Finished shorts (voiceover, footage, captions, one timeline) | Rank Reel |

The niche

"Interesting facts" videos: ~30-second mini-documentaries on one story/person/event/fact. Examples shown: ~180K subs (weird history moments, videos at 2–7M views, some weeks old); ~320K subs (science/space); a third on strange facts/unexplained events. Format is niche-agnostic (history, space, animals, tech, mystery, geography). Claim: success is more about workflow quality than topic.

Workflow (in order)
  • Find 3–5 competitor channels. Watch Shorts in the niche, like/subscribe/comment; the Shorts feed becomes a research tool. Not to copy — to learn "what topics get views, what they do in the first seconds, video length, which stories repeat."
  • Claude + VidIQ connector. Install the free VidIQ Chrome extension; in Claude, refresh, connect "YouTube insights." Then + → connectors → VidIQ for Claude → competitor breakdown, paste the channel ID (More → Share channel → copy ID), send. Claude pulls shorts and ranks top performers (allow "always allow" permission prompts). Follow-up ask: "give me their best five performing video topics, include the original video links, and give me a full summary of each one." Example results: why airplane windows have tiny holes; what happens if an elevator cable snaps; why F1 tires are smooth.
  • Script. Fresh Claude chat with a 30-second Shorts script prompt + the topic summary → script timed to ~30s. Then a second prompt turns it into a scene-by-scan visual breakdown with on-screen directions and timestamps.
  • Rank Reel assembly. Paste the word-for-word script → pick an AI voice (recommends natural, fast-paced for facts) → generate voiceover, auto-dropped on timeline. Visuals per scene: either search for a real clip (copy link → Import → download, set 9:16) or generate (Media → AI generation → video, paste Claude's description, aspect ratio 9:16). Mute generated clips' own audio. Trim real clips to Claude's timestamps — "no guessing for how long any clip should be." Captions: Subtitles → generate captions → style presets (bigger text, font/shadow/color). Render + export.
Channel setup before posting
  • Prefer an aged channel (any old unused account repurposed — "an older account already has some history with YouTube"). New channels need a warm-up: day 1, ~30 min scrolling the Shorts feed, watching/liking/commenting/subscribing naturally; day 2, repeat + setup.
  • Feature eligibility in YouTube Studio → Settings → Channel: standard features (usually green); intermediate features require phone-number verification; advanced features optional/not needed.
  • Branding via ChatGPT (example channel name: "Curio"): one-word name ideas; simple profile icon readable at small sizes; banner matching the avatar's colors/style with a short line ("subscribe for random facts" — regenerated at 50% larger text); description with tagline + what viewers can expect.
Caveats
  • All revenue figures are VidIQ estimates from 274 analyzed channels (Air Media Tech); RPM is heavily audience-geography dependent.
  • Competitor analysis is essentially pattern-matching proven topics, not original research; the creator frames it as "repeat more of what works," explicitly rejecting one-viral-video ambitions.
  • Full prompts are shown on-screen in the video rather than pasted in the transcript — reusable only by screenshotting.
Transcript · 23,382 chars
Over the last 3 months, this faceless YouTube channel has made almost $40,000 using AI, and that's from YouTube shorts alone. There's no long- form content, and the person behind it never even shows their face in a single short. And today, I'm going to show you exactly how to replicate the same model, explain why this is one of the biggest opportunities on the internet right now, and give you everything you need to get started. So, on the screen right now are three faceless YouTube shorts channels that I've been tracking. And according to Vid IQ, in the last 90 days, the first one generated an estimated $40,000 from $250 million views. The second made almost 25,000 and the third clocked in at around $33,000. And to add to that, over the last 365 days, just one of them has done two billion views and over 170K in estimated revenue. And we're seeing this across three completely different channels. So clearly, this isn't just one random success story. And honestly, you don't need a team or a complicated production setup to do it. The whole thing runs on just three tools. Claude, Chachi PT, and one other tool that I'm going to show you guys a little bit later. So, why is right now such a good time to start? Three reasons. First, despite everybody saying that faceless YouTube is getting more and more saturated, the data actually shows something pretty interesting. Now, if we look at Google Trends, searches for YouTube automation and faceless YouTube have dropped massively from their peaks in 2023 and 2024. And right now, search interest is actually sitting almost at its lowest point in 5 years. So, while everyone out there keeps saying that people are jumping into faceless YouTube, the search trend is actually moving in the opposite direction. Second, shorts can generate real money at scale. Now, Air Media Tech analyzed 274 real channels using data directly from the YouTube analytics API. And according to what they saw, shorts RPM in most niches ranged from 7 to 20 cents per thousand views, while US heavy audiences were actually reaching around 33 cents per thousand views. Now, you might hear 33 cents and say, well, that sounds tiny. But it's only tiny until you actually scale it. Now, at that RPM, 100 million views is worth roughly $33,000. And keep in mind, one of the channels that we just looked at did 250 million views in only 90 days. Now, the third reason is probably the biggest one of all. Now, just a year or two ago, actually running one of these channels meant doing almost everything manually. But today, AI can handle almost the entire workflow. Just look at this. Chat GPT handles the channel itself, your name, your description, profile picture, banner. Claude will handle the content, analyzing competitors, finding out what's working, coming up with ideas, doing the research, and ultimately writing the actual scripts. And the last tool takes those scripts and turns them into finished shorts with voice over footage, captions, and everything else in one timeline. So, instead of talking about it, let me show you exactly how this whole system works. I'm going to start with Claude and show you how to automate the actual content side of the channel. Then we're going to jump back to ChachiBT and we'll use it to set up everything else. The branding, the description, profile picture, banner, essentially all of the things that you need before you can actually start posting. But first, let's figure out what kind of videos we're actually going to be making. Now, picking the right niche is obviously one of the most important parts of this. So, instead of giving you 50 different options, I'm just going to show you one that's working super well right now. And that's interesting facts videos. They're basically 30 second mini documentaries built around one interesting story, person, event, or fact. And channels that are doing this right now are getting ridiculous views. So, here's one with around 180,000 subscribers that makes interesting facts videos about weird moments in history. Now, you can see some of these getting 2, 4, even 7 million views, and a few of them were posted just weeks ago. But you don't have to make history videos. Here's another channel with over 320,000 subscribers doing basically the same exact thing, except their videos are all about science and space. And here's a third, which is focused on strange facts, true stories, and unexplained events. Now you see over and over it's the same format but completely different topics. You can see videos like this getting hundreds of thousands or even millions of views. And that's the cool thing about this niche. You can basically adapt it to whatever you're interested in. History, space, animals, tech, mystery, geography, literally anything. It becomes less about the specific content and more about the quality and the tooling of the workflow that I'm about to teach you. And the best thing about this format is that AI can do almost all of the heavy lifting for you. So, the first thing we're going to do is find a few competitors and give them to Claude. Now, you really only need three to five good channels. And we're not trying to copy them. We just want to figure out what's actually working. We're asking questions around what topics are getting the most views, what are they doing in the first few seconds, how long are the videos, and what kind of stories are going viral over and over again. And finding these channels, honestly, is actually really easy. Just watch shorts in the niche that you're interested in. Like a few of them, subscribe to the channels that are doing particularly well, maybe leave a couple of comments, and you'll notice pretty quickly your entire shorts feed basically becomes a research tool. Now, YouTube is just going to keep showing you more and more channels and viral videos in the exact niche that you're interested in. So once we've got a few good competitors, we're going to give all of them to Claude and let it do the actual analysis for us. Now that we have a few competitors, step two is using Claude to actually analyze them and find the best video ideas for us. So what I'm going to do is I'll pick one of the channels that we just found. Open up Claude. And the first thing we have to do is to connect to VidIQ. So you just need to search for the Vid IQ Chrome extension right here and install it. It's completely free. Now, once that's installed, I'm going to go back to Claude. I'll refresh the page, and you should see Vid IQ pop up right about here. I'm going to click the dropown, hit connect YouTube insights, and it'll take us straight to the connectors page where we can connect it to Claude. And once that's done, you can now see Vid IQ is active. So now I'm going to click the little plus right here. Go to connectors, select Vid IQ for Claude, and I'll choose competitor breakdown. Now it's asking us for the channel ID. So I'm going to go back to the competitor that we found earlier. Click more, then share channel, and I'll copy the channel ID right here. Jumping back into Claude. Paste it in. Click add prompt. And you can see it's automatically going to add the whole prompt for us. That all that's left for me is just to click send. And right away, Claude is going to start analyzing the entire channel. It's pulling their shorts, looking at which videos perform best, and basically figuring out what's working for them. Now, this can take a minute, especially if the channel has a lot of videos. And Claude might ask for permission a few times while it's doing its research. Make sure to just hit always allow so that it can keep going and save you some hassle. And there we go. You can see that Claude just gave me a full breakdown of the channel and it analyzed their best performing videos. And the coolest thing is that this is all sitting inside of a normal Claude chat. So, I can literally ask it about anything in the channel. So, I'm going to go ahead and I'll ask give me their best five performing video topics, include the original video links, and give me a full summary of each one. Now, hit send. Give it a second. And we got it. Now, I have five video ideas based on topics that have already gotten millions of views on this channel. It's a proven success formula. So instead of sitting here trying to guess what might work, we already have five ideas that we know people are interested in. For example, we've got one about why airplane windows have tiny holes in them. Another about what happens if an elevator cable snaps, a third explaining why Formula 1 tires are completely smooth. These are exactly the kind of topics that work for the interesting facts format that we're going to be making. So now we have the ideas. Next, we need to turn one of those into an actual video. Now, I'm going to pick one of the topics that Claude just found for us, and I'll open up a fresh Claude chat. What I'll do in there is I'll paste in this prompt right here. I'm making sure to leave it on the screen for a little bit so you can screenshot it if you want to reuse it yourself. Now, I'm going to jump back to our competitor analysis, grab the summary of the topic that I want to use, I'll copy it, and I'll paste it right underneath the prompt. For this example, I'm going to be going over the one about why airplane windows have those tiny holes in them. And basically, what I'm asking Claude to do is to take that proven topic and turn it into a completely new 30-second YouTube short script for our channel. So, I'm going to hit send. And there we go. Claude just wrote the entire script for us, and it's already structured to fit into around 30 seconds. But now, we need the actual visuals. So, I'm going to paste in one more prompt. Again, I'll leave this on the screen for a second so that you guys can screenshot it. And what this does is it's actually really simple. It takes the script that we made, breaks it down into individual scenes, and then it tells us exactly what should be happening on screen for every single part of the voiceover. So, instead of having one giant block of text, we now have an entire video mapped out scene by scene with a voice over. what visuals we need and then timestamps for when they should appear. And honestly, don't worry about creating any of the visuals yet because that's the next step. And this is where the third AI tool I mentioned comes into the picture. Now, we have the script and we have every visual mapped out. So, what's left is to actually build the video. And for that, we're going to be using a tool called Rank Reel. Essentially, what it does is it gives us everything we need to create the entire video in one place. So, to that end, I'm going to open up Rank Real right here. And the first thing we need to do is generate our voice over. So, I'll jump back to Claude and I'll grab the original script that we generated. Now, keep in mind, it's not the visual breakdown. It's just the full word for word script. You'll copy that, jump back to Rank Real, and paste the entire thing right here. Now, we just need to choose a voice. Rank Real gives you a bunch of different AI voices to choose from. So, I'm going to click select voice, and I'll just be using this one right here. Obviously, you can take some more time and pick whichever fits your channel best. But for the interesting facts videos, I want something that feels pretty natural and is a bit more fastpaced. So, I'm going to select it, hit generate, and there we go. Rank Real just generated the entire voice over and it automatically dropped it onto our timeline. So, let's go ahead and listen to the first few seconds. Ever notice the tiny hole at the bottom of your airplane window? That's not damage. That's the only thing keeping you comfortable at 38,000 ft. Perfect. So, now we have the script and we have the full voice over. Next, we need to create everything that's going to play on screen. Now, obviously, we still need the actual footage and the captions. So, that's what we're going to be focusing on next. I'm going to go back to Claude and I'll open the visual breakdown that we generated earlier. And you can see that for every part of this clip, Claude has already told us exactly what should be happening on screen and how long each clip should be. So, I'm going to go ahead and start with the first one. I'll copy the clip description right here. And now, we've got two options. We can either find a real clip that matches it or just generate one with AI. Now, for this clip, let's see if we can find something real. I'm going to search for exactly what Claude described. And you can see none of these clips quite really fit what I'm looking for. So, instead of wasting my time trying to find the perfect one, I'm just going to go ahead and generate it. I'll jump back to Rankre, click media, go to AI generation, and select video. Then I'm literally just going to paste in the description that Cloud gave us. I'm going to go ahead and set the aspect ratio to 9 by6 since we're making a YouTube short. I'll hit generate and then give it a second. And it generated exactly the shot that we needed. So I'm just going to click add to timeline. And now our first clip is already in the video. It also generated audio with the clip, which we obviously don't need because we already have our voice over. So I'm just going to go ahead and mute that. And now we have this. Perfect. Now we just move on to the second clip. I'm going to go back to Claude, copy the next visual description, and this time let's see if we can find a really good clip for it. Yeah, this looks much better. I'm going to grab this one right here. I'll copy the link, go back to Rank Real, click import, paste the link, and hit download. And there we go. I'll set the preview ratio to 9x6. Add it to the timeline. And now I just need the exact part of the clip that matches our script. So I'll go ahead and scrub through here. Right there. I'm going to split it at the beginning, delete everything that we don't need, and then do the same thing at the end. And remember, Claude already gave us the timestamps, so really I don't need to do any guessing for how long any of these clips should be. I'll go ahead and line this one up with the voice over. And now we've got this. And from here, you literally just repeat the same exact process for the rest of the video. Go through Claude's visual breakdown one by one. If you find a good clip, use it. If you can't, generate exactly what you need with AI. And after just a few minutes, you should have the entire voice over covered with visuals. Now, there's only one thing missing, captions. And this part is super easy. I'm just going to go to subtitles and click generate captions. Give it a second and Rank Real will automatically generate captions for the entire voice over. Obviously, they don't look the best by default, so I'm going to need to go into style and quickly clean them up. I'll choose a caption preset, maybe make the text a little bit bigger, and then you can play around with a font, shadows, colors, or honestly whatever else fits your channel's needs. And that's basically it. We started with nothing but a competitor channel. Use Claude to find the idea, write the script, plan every single visual, and now we have a finished YouTube short ready to post. All that's left is to hit render, export the video, and that's our first short done. Ever notice the tiny hole at the bottom of your airplane window? That's not damage. That's the only thing keeping you comfortable at 38,000 ft. Every window is actually three panes. Now you know exactly how to make these faceless YouTube shorts. But before we start posting, we need to actually set up the channel and make sure that everything is properly optimized. So before you create anything, the first thing I want you to do is go ahead and check if you already have an aged YouTube channel. And when I say aged, I literally just mean a channel that you've had for a few years now. A lot of people have had one without ever even realizing it. If you have ever commented on videos, subscribed to channels, or honestly just used YouTube normally. there's a pretty good chance that you've created a channel a couple of years ago or maybe even longer and then completely forgot about it. So, go ahead and check your account. If you find an old channel that you're not using anymore, you can just repurpose that for your new faceless channel. Now, the reason that I prefer doing this is that an older account already has some history with YouTube rather than being something completely fresh created 5 minutes ago. But if you don't have one, don't worry. you can just go ahead and create that brand new channel. The only difference is going to be is that with a new account, I like to spend a little bit of time warming it up before I start posting. And honestly, the warm-up is super simple. You're basically just going to go ahead and use YouTube like a normal person. So, on day one, I'm going to open the Shorts feed and I'll spend around 30 minutes just scrolling through videos. And while I'm doing that, I'll make sure to interact with some of the content naturally. Watch some of the videos, leave a few likes, maybe post a couple of comments, and subscribe to the channels that I'm actually enjoying. That's literally it. And then on day two, I'm going to do the same thing again for a little bit. Watch some shorts, interact with a few videos, subscribe to a few channels, but this time, we're going to start setting up the actual channel. So, I'll go ahead and open up YouTube Studio. Go down to settings, click channel, and then I'll go to feature eligibility. The first thing that you want to check is standard features. Now, for most channels, these should already be enabled by default, as long as there aren't any community guideline restrictions on your account. Now, you can see mine are green here, so we're good to go. Then, we're going to move down to intermediate features. for this. YouTube is going to ask you to verify a phone number. So, I'm just going to go ahead and click right here and go through the verification process. And there we go. Now, intermediate features are enabled as well. Next, you'll also see advanced features underneath this. Now, we typically don't really need these for what we're doing right now, so you can either leave them alone or come back and unlock them later. And that's pretty much everything that we need on the technical side. Now, what's left is to make this look like a real channel. The name, the profile picture, the banner, a description, all that good stuff. And instead of doing any of that ourselves, we're actually going to let Chad Chapiti handle the entire branding process for us. So, let's start with the name. I'm just going to type give me short one-word channel name ideas for a YouTube channel about interesting facts. I'll hit send. And you can see we've already got a bunch of options. And honestly, don't overthink this part. If you don't like any of these, just ask chat chip t for more. But you definitely don't need to spend hours trying to find the perfect name. I'm going to scroll through these and let's go with curio for this example. Sounds cool. So now we've got the name. We need a profile picture. So for that, I'm going to paste in this prompt right here. I'll leave it on screen for a few seconds so that you can go ahead and screenshot it as well. All that you need to do is replace this part with whatever channel name you choose. So, I'll type in Curio right here. Hit send and I'll give it a second. All right, this actually looks pretty good. It's simple. The icon is easy to recognize. And most importantly, it'll still be readable when it's super small on the shorts feed. And that's exactly what you want. Now, if you look at some of the other channels out there, they're doing the same exact thing. One simple character, symbol, or logo, nothing too complicated. And if you don't like the first option that Chat GT gives you, you can always just regenerate it. You really don't need to overthink your logo either. Now, I'm happy with this one, so I'm going to go ahead and I'll save it. So, now we have the name and the profile picture. The next thing we need is the channel banner that goes right up here. And if you look at other channels in this niche, you'll see that people do this in a bunch of different ways. Some just put the channel name, some have a little tagline, and some add a bunch of graphics. But honestly, I've always felt like the simpler ones look the best and cleanest. So, I'm going to keep ours clean, same colors, and overall the same style as the profile picture, just with one short line of text in the middle. And again, I'll be using Chat PT to make this whole thing for us. Basically, what I'm asking you to do is to create a YouTube banner that matches the exact colors and the style of the profile picture that we just made. And for the text, you can literally put whatever you want. I'm going to keep it simple and I'll say subscribe for random facts. So then let's hit send. And honestly, again, this looks pretty good. It matches the profile picture. The colors are consistent and everything feels like it's in line with the same channel. Now, for my taste, the text is a little too small. So, I'll go ahead and tell Chachi BT make the text roughly 50% bigger. Hit send. And yeah, that's looking way better. So, now we've got the name, the profile picture, and the banner. There's just one thing that we need left, and that's the channel description. And if you haven't guessed, I'm not going to be writing this one manually either. I'm gonna go in one last time into chat GPT right here. And I'll paste this prompt once more. I'll leave it on the screen for a few seconds so that you guys can screenshot it if you want. And basically what this tells Chat GPT is to give us a short tagline, explain what the channel is about, and make sure that people know what kind of videos they can expect. So I'll hit send. Uh once it generates, I'm going to go ahead and quickly read through it. I'll make sure everything looks good. Copy the whole thing and I'll paste it straight into our YouTube channel description. And that's it. Our channel is officially ready. Now, if you think about the big picture, we basically just built an entire system from scratch. We use Chat PT to create the channel, Claude to find proven ideas, do the research, write the scripts, and plan every visual. And lastly, rank real to turn all of that into an actual finished short. And from here, you just repeat that same process over and over again. Now, at a high level, what you're going to do is find what's already working, make your own version, post it, see what gets views, and then use that same exact system to repeat more of what works. That's literally the whole idea behind this model. You're not out here trying to create one viral video. You're creating a system and a process that can keep producing them. And that's how you can run an entire faceless YouTube shorts channel with AI doing almost all of the work. So, everything that I showed you throughout this video is actually linked below if you want to try it out for yourself. As always, thank you so much for watching and I'll see you guys in the next one.
12:00

My Hermes Agent Finally Works While I Sleep

A hosting company now sells managed AI agents, so you can run a 24/7 assistant without babysitting a home server. Cloudways by DigitalOcean deploys pre-configured Hermes or OpenClaw agents with backups, firewall, and updates handled for you. Pricing starts below the roughly $12 a month a self-managed VPS would cost, and you can connect WhatsApp, Slack, Discord, or Telegram with a click. This is essentially a sponsored walkthrough of the new product.

Notes

My Hermes Agent Finally Works While I Sleep — Creator Magic (Mike Russell)

YouTube walkthrough (2026-08-24) deploying Hermes Agent on Cloudways Managed AI Agents (Cloudways by DigitalOcean) instead of self-hosting. Presenter is Mike Russell.

Why he moved
  • Prior pain: hosting Hermes Agent/OpenClaw on a Mac Mini, an old mini PC "in a drawer," Proxmox and "versions of OpenClaw," or leaving the MacBook open all night — "something breaks, it's on me"; security flaws "wide open on my network"; rolling a VPS himself felt "way too technical" (SSH, Docker, terminal).
  • Claimed wins: daily backups, a firewall he didn't configure, someone else handles updates/maintenance, "no lock in," upgrade/downgrade anytime.
Cost comparison
  • DigitalOcean droplet directly: ~$12/month decent, $24/month meatier.
  • Cloudways: introductory pricing; even outside promo, "the same exact spec of server with CPU and RAM… is cheaper than spinning up a VPS."
Setup sequence (real-time, "under five minutes")
  • Log in, connect payment card → AI Agents → Get started.
  • Choose Hermes or OpenClaw — both "long term support versions… not bleeding edge" (i.e., deliberately not breaking). Noted as good: Cloudways keeps them on reliable versions.
  • Name the agent ("Mike's Hermes"), pick region (he chose Frankfurt).
  • Choose instance size — started at Scout level with intro pricing; upgrade later if RAM runs out.
  • Optional LLM provider — plug in existing key/subscription for Anthropic, OpenAI, OpenRouter, Gemini, or DigitalOcean's own AI. He picked OpenRouter (cheap models). Flow: name key ("Mike's Hermes"), no expiration, Create, copy/paste key, accept T&C, Deploy Agents.
  • Success screen: "Your Managed AI Agent has been deployed successfully."
Post-deploy management
  • Three-dot menu: restart, manage plan (up/down disk/RAM), delete (replace or permanently delete — "killing one agent to replace it").
  • Hermes dashboard password found under Agent details. Same panel also exposes techy creds: username, box, IP address, port, password — "DigitalOcean infrastructure… the kind of credentials I'd get if I just spun up a VPS."
  • Cloudways MCP server: lets agents manage Cloudways servers/apps via the official server — "agents spinning up other agents… Agent Inception."
Demo task & model notes
  • Model defaulted to Claude Sonnet latest via OpenRouter; agent can use "all 35 models, including MiniMax M3 and Kimi K3." He switched to MiniMax M3, "the most popular AI model right now for OpenClaw."
  • Task: connect Apify + Notion to build "an overpowered dashboard." Hermes researched auth methods, asked a clarifying question, delegated to sub-agents, ran the Apify YouTube actor — 27 seconds, 20 results, top hit "Hermes Agent fundamentals in 29 minutes" (~quarter-million views). Output as clean table; sent to Telegram via the send command; working files/markdown visible in dashboard.
Integrations (click-to-connect, no terminal)
  • WhatsApp, Slack, Discord, Telegram. Telegram flow: BotFather → /newbot → name "Hermes Agent CW" → token → Connect. One caveat surfaced: "No home channel is set for Telegram. Set home to make this chat your channel." After that: "Yep, I'm here. What do you need?"
Caveats / limits
  • Openly sponsored-style promo with affiliate link ("my link is in the description").
  • LTS-only versions (no bleeding-edge features).
  • Scout plan "using quite a lot of RAM" — may need an upgrade; monitor metrics in panel.
  • "Recommend what you think is best and I'll do it" — much of the workflow leans on agent autonomy; he does not show billing post-promo or cancellation friction.
Transcript · 13,664 chars
I'm tired of hosting Hermes Agent on my Mac Mini, on a dusty mini PC that sits in my drawer, being a sysadmin, running Proxmox and having versions of OpenClaw. I don't even know what they do anymore, or leaving my MacBook open all night. Something breaks, it's on me. I have to fix it. Hermes Agent, OpenClaw, latest update. Security flaws wide open on my network. I was maintaining an ever growing fleet of breaking agents running from my house. And forget spinning up a VPS. Way too technical, right? SSH, Docker, Terminal updates, all of it. Oh my goodness me. I just want Hermes Agent. So this is what I did. I moved Hermes Agent to somewhere that never sleeps. No server configs, no Docker, no command line or anything scary like that. It gets daily backups, it gets a firewall that I didn't configure and someone else is handling all the updates and maintenance. Now you might remember the Hermes Agent setup I did with Apify and Notion. It was basically running around the clock finding opportunities for me. But I needed to keep something running in my house, even if it wasn't a laptop, just something. Right now I can have the same agent working all around the clock without me and without anything running here. And here's the kicker. I'm not afraid of spinning up my own VPS in the cloud. So I looked at DigitalOcean and noticed that, well, for a decent quality droplet to run my Hermes Agent on, it cost me about 12 bucks a month. If I want something meatier, $24 a month. Then I found this really cool solution. Take a look at this. This is Cloudways Managed AI Agents. Yes, all of your AI agents, including OpenClaw or Hermes managed right there on Cloudways. Now what is Cloudways? Am I going to trust it with my personal agents that have all my information and everything I'm doing with my life? Well, wait a minute, I'll just scroll back up. This is Cloudways. Buy DigitalOcean and when I scroll down to look at the pricing, I was pretty shocked to see they've got introductory pricing. But even after that, the same exact spec of server with CPU and RAM we've just looked at on DigitalOcean. Even outside of the promo, pricing is cheaper than spinning up a VPS. So there's no excuse. I don't even know why I'm spinning up my own servers and wrangling Docker in the command line while when this can do it for me. So with that said, and the fact I pay Per month. I'm going to click in and launch my agent. This is Hermes Agent in under five minutes. Let's start the clock. All right, so I have logged into Cloudways and connected a payment card. I'm going to click AI Agents and I'll click get started here. Now I can choose either Hermes or OpenClaw. Now notice these are long term support versions. They're not bleeding edge. But. But that's also good because it means they're not going to break Cloudways by DigitalOcean. Actually take care of making sure Hermes Agent and OpenClaw are on reliable versions, which is awesome. First of all, we'll give it a name, Mike's Hermes and we'll obviously make sure Hermes is selected here. And now I need to select a region just the same as if I was spinning up my own server. I'm going to stick it on a server near to me. There we go. How cool is that? Mike's Hermes living in Frankfurt. I wonder if she'll listen to techno and consume Frankfurters all night. Okay, with this done, let's click Continue and see what happens next. Right, I've got to pick the instant size. As you can see, I've got the introductory pricing available and I'm going to start off small. I can always work my way up. So we'll start with the Scout level. And now I can choose an LLM provider and this is optional. I can basically use my already existing API key or subscription for Anthropic, OpenAI, OpenRouter, Gemini and even DigitalOcean seem to have their own AI you can plug in. Well, do you know what? I'm going to pick OpenRouter because I love the fact it gives you the possibility to use cheap AI models to power your assistant. I just have to click Get API key. That will take me straight to the OpenRouter API page. The same as if I was working with OpenAI or Anthropic. I just give my key a name. Let's call it Mike's Hermes again. Expiration. I don't want it to expire, I want it to always work. We'll leave this all as is. Click Create. Now I'm going to copy this key to the clipboard and we just paste the API key in here, accept the terms and conditions and click Deploy Agents. And let's see how long this actually takes in real time to happen. So we've got the Hermes here Scout in Frankfurt setting up right now. It's creating my instance and I'm talking to you in real time. So you can know all that stuff that would usually happen at the command line with you running all kinds of setup commands from the README on HOMEY's agent website is now being done behind the scenes, or the database or the infrastructure. All of it is being put together right there for you. And this is the whole beauty of a Managed AI Agent on Cloudways by DigitalOcean. I can also deploy more agents as I go. Literally, I can set up as many as I want. And would you look at that? Your Managed AI Agent has been deployed successfully. Hermes is here. Okay, so with everything set up, let's go to the three dots. I can restart my agent anytime, so if it should ever stick, it's just one click away. I can manage the plan. Look at this. I can upgrade or downgrade as I wish. There's absolutely no lock in, so that means I can ultimately upgrade if I need more disk space and RAM and go the other way if I want to reduce my usage too. And then we can also delete the agent. But be careful with that one. You can either permanently delete your agent or deploy a new one in its place, which kind of feels bad, like you're killing one agent to replace it with another. All right, now let's click this Hermes link. This will take us to a Hermes dashboard where we can enter our password to continue. Now, if I want to get that password, I just click into my Hermes Agent and look right here under Agent details. There it is for me to copy. Let's copy that password, and then we'll paste it into the web control panel. Click sign in, and boom, here we are with Hermes. I can actually message my Hermes Agent for the first time. Oh, and stop the clock. We've got Hermes Agent connected. How long did that take? It didn't take long at all. I didn't even need five minutes. Hello? Are you alive? Let's find out if my Hermes Agent is alive. It's processing my requests. Let's drop this open and have a look if anything's going on. Yes, I'm here and running. I'm Hermes Agent, ready to help with coding research. Wow. All those good things that I was doing before, now right here inside Hermes. Let's make this a little bit bigger so you can really see things. And I want you to notice here that Claude Sonnet latest is the model. But because I'm going through OpenRouter, this is kind of overpowered. I can literally choose any model OpenAI Anthropic in there, plus all 35 models, including MiniMax M3 and Kimi K3 that are real stars of the show. I'm actually going to switch now to MiniMax M3 because that's a really good one and it's actually the most popular AI model right now for OpenClaw. So let's give it a go in Hermes. All right, I'd like you to help me connect my Apify account and my Notion account so we can make an overpowered dashboard. Can you guide me through that, please? Okay, we'll enter that and let Hermes think about it. Great project. Before I connect, I need to know how they're typically authenticated. Okay. So it's doing some quick research. We'll let my Hermes Agent go out and find out what it needs and then I'll give it the info. It's asking me to clarify Apify because my text to speech didn't quite say it correctly. Let's send that off to the agent and get it on its merry way. Yes, Apify, the automation platform. That's good. Okay. It's delegating it to its agents. It's working away. Very soon we'll have something working right in front of our eyes that is going from signing up to deploying agent to having a working automation in just minutes. I'm just going to keep telling it, recommend what you think is best and I'll do it. Oh, and while it's going away and building that automation, I'm just going to show you behind the scenes here. So you'll see, not only have I got a URL to my Hermes Agent dashboard, which is right here with the password, the version that we're on right now, but I can also see if I want to be a bit techy, all those credentials to get in. So just because I set this all up via Cloudways doesn't mean I don't have the username, the box, IP address, the port it's running on, and even the password if I want to do terminal work. So this is really cool because it's DigitalOcean infrastructure. I still have exactly the kind of credentials I'd get if I just spun up a VPS on vanilla DigitalOcean. This is cool. Then we've got our LLM provider. We can look at the instance and we can see, yeah, scout plan is using quite a lot of RAM. We might run out and need to change the plan, which we can do down. CPU is running okay and disk is pretty decent too. And there's also an McP server. Cloudways. McP. You can let your agent manage Cloudways servers and applications through their official server. Which means you could have agents spinning up other agents which could become like Agent Inception. Oh, and another cool thing, no fiddling about command line terminals again. Yes, Hermes Agent can be connected with some of my favorite apps right here. WhatsApp, Slack, Discord, and Telegram, just with the click of a button. And because I'm using a lot of agents in Buzz right now, connecting Hermes Agent to Buzz would be an excellent next video. Let me know in the comments down below if you'd like to see that. It's actually running the YouTube actor on Apify and finding YouTube videos about Hermes Agent. This is absolutely incredible. So you see, my autonomous agent is already going and doing the jobs that I was setting it in my previous Hermes Agent video. This is really awesome. All right, so let's connect to Telegram. We're going to need a bot token. So let's talk to bot father over here and we'll go ahead and say slash, new bot. And then we'll give it a name. We'll call this Hermes Agent cw and then we'll give it this Hermes Agent CW Bot. All right, now we got that token for a new agent. We'll just paste it over here. Click Connect Telegram. It is done. Telegram is connected successfully. And look at that. No touching code. No working with terminals. Easily done for me in a managed control panel that's super easy to navigate round. Now I can see Telegram is connected to my Hermes Agent. I can start chat and say you there to my Hermes Agent. Boom. No home channel is set for Telegram. Interesting. Set home to make this chat your channel. Oh, and now we've actually got it. Yep, I'm here. What do you need? What are you up to? Let's find out what my Hermes Agent has been doing. Just idling and ready for whatever you throw at me. You've been looking at amplify. Let's see if the Hermes Agent has context here on the task that I set it. And boom. Look at this. Fantastic. Yeah. Actually quite a bit earlier today we went through an Apify notion rabbit hole. It even tells me I originally called it Apify through my dictation app with a double P, which is quite funny. And it says it's actually mid flow right now. Working on it, which is awesome. And I can go back my dashboard here and look at this. It's actually confirmed that the search worked. 27 seconds, 20 results and we've got the top hit, which is Hermes Agent fundamentals in 29 minutes with quarter of a million views. This is really awesome. Hermes Agent doing real work for me, always on. On a nice server in Germany that I don't have to touch. Can you present your results to me in a really clean table that I can read, please? Okay, Hermes Agent is off. Working on that. And boom. Notice the commands are running right now on MiniMax M3, which is absolutely awesome. We can always open up these commands to see what kind of stuff it's running. Oh, and maybe I want the result to go to my Telegram. So right here in my dashboard, I'll say, send results to Telegram. And then that will just steer into the conversation I'm having with my Hermes Agent. Okay. It's sending it as a clean markdown message over to Telegram. Let's wait for that to land. It's found the Hermes send command over here, and it's actually doing it to find my Telegram connection. And what is really awesome is I can pop out and see all the files it's creating. Whoa. While it's also sending me this. And look at this. This is a full breakdown of all the videos about Hermes Agent, all ranked by views and duration and everything else it works. Top hit there. Sent to Telegram. Mike Russell. Fantastic. But I can also see all the work it's doing here, all of its files, all of its markdown files. And look, the table is also over here in Hermes Agent 2. This is truly overpowered. So that's it. That is my Hermes Agent sitting right there in Frankfurt, Germany, working for me whenever I like. And wherever I am, my laptop shut, my computer's having a good rest, and everything works without me needing to do a thing. If you want to do the same thing I just did and have your agent running right now, just use Cloudways Managed AI Agents. My link is in the description. Deploy Hermes Agent, give it one job, and then go to bed. Tell me in the comments what you set your Hermes Agent doing overnight. Thank you so much for watching. And YouTube is showing a video on your screen now. You should watch next. Thanks.

Article

85
09:30

😺 Anthropic's IPO could top SpaceX's record

Anthropic could pull off the biggest stock market debut in history, raising more than $100 billion at a valuation near $2 trillion. That's more than double the $965 billion it was worth in June, and would beat SpaceX's record $85.7 billion IPO, with shares possibly listing as soon as October led by Morgan Stanley, Goldman Sachs, and JPMorgan. Claude Code's popularity has pushed projected annual revenue to $47 billion and would let Anthropic beat OpenAI to public markets, which isn't listing until 2027. Catch: Anthropic has struggled to get enough chips, and the Trump administration cut its Defense contracts in March after Anthropic refused to give the military unrestricted model access. The newsletter also covers Nvidia's $7 billion bet on AI startup Poolside and a combined $18 billion in AI infrastructure spending by Alibaba and Tencent.

Notes

Anthropic IPO could top SpaceX's record (The Neuron, 2026-08-24)

Lead story: Anthropic IPO
  • Bankers told investors Anthropic could raise "more than $100 billion" at IPO (per NYT), implying a valuation near $2 trillion — more than double the $965 billion from its last private round (June 2026).
  • Could "match or beat" SpaceX's June debut — $85.7B raised at $1.77T valuation, the current largest IPO ever.
  • Morgan Stanley, Goldman Sachs, JPMorgan leading; listing possible as soon as October; public filing expected in coming weeks, shares possibly this autumn.
  • Only Apple, Microsoft, Nvidia have ever crossed $2T valuation.
  • Would beat OpenAI to market; OpenAI not planning to list until 2027.
  • Driven by Claude Code; projected annual revenue $47 billion. Anthropic founded by siblings Dario and Daniela Amodei (both ex-OpenAI executives), five years old.
Caveats flagged by the newsletter
  • Chips/server shortages constrain scaling.
  • In March, the Trump administration cut Anthropic's Defense Department contracts entirely after Anthropic refused the military unrestricted model access; Anthropic called it "unconstitutional retaliation". Suggests prospectus will reveal which story is true.
SB 947 / "No Robo Bosses Act of 2026"
  • California bill to ban employers from letting AI fire/discipline workers without human sign-off; reintroduced after Newsom vetoed the earlier version last October.
  • Case study: Andon Market (SF) uses AI agent Luna for hiring/scheduling/pricing; recommended firing a worker who missed 17 of 23 shifts — humans carried it out. Luna had written the store's attendance policy itself, then forgot it existed until told to check its memory.
Around the Horn (other items)
  • Nvidia paying $6B to license Poolside's tech + hire engineers, plus $1B investment at $12B valuation — building US open-source rival to Chinese models.
  • Nvidia raising Grace Blackwell / Vera Rubin chip prices 15–17% starting early 2027, adding ~$5B to a single large data center's cost.
  • Alibaba + Tencent spent $18B on AI infrastructure last quarter; dragged Alibaba profit down 75%; CEO claims payback within three years.
  • ChatGPT Ads expanded to 31 European countries (Aug 18); free/low-cost users see ads; paid subscribers stay ad-free.
  • Blackstone + Hellman & Friedman: 160-person AI engineer team, backed by $1.5B partnership with Anthropic, embedded in portfolio companies.
  • Dr. Dre uses AI to produce songs; called opposing musicians "afraid of learning new things"; named Timbaland a "closet AI user."
  • Investigation of data labelers (China, Australia): workers get less work over time as AI improves, don't know which companies they label for, and have little recourse to appeal performance grading.
Skill: separate drafting from verification
  • State Farm's outside lawyers admitted AI produced seven nonexistent case citations in court filings.
  • Steps: (1) ask AI to list every factual claim/number/quote/source using "the exact language from the source document so I can command/control+F find it"; (2) open sources yourself and Ctrl+F; rewrite/delete unsupported claims; (3) never ask AI to "repair" an invented citation — start from a real source.
  • Suggested prompt: "Audit this draft for factual risk. List every claim, number, quote, and named source I must verify manually. Give only URLs already present in the draft. Mark anything unsupported as DO NOT PUBLISH. Do not invent or repair sources."
Full text · 8,360 chars
😺 Anthropic's IPO could top SpaceX's record PLUS: Nvidia's $7B bet, Alibaba's $18B quarter. Welcome, humans. California is trying again with SB 947, a bill that would ban employers from letting AI fire or discipline workers without a human signing off. Newsom vetoed an earlier version last October, but lawmakers reintroduced it as the "No Robo Bosses Act of 2026." The stakes are already real. At Andon Market in San Francisco, an AI agent named Luna runs hiring, scheduling, and pricing. She recently recommended firing a worker who missed 17 of 23 shifts. Humans carried it out. Only problem? Luna had written the store's attendance policy herself, then forgot it existed until someone told her to go check her memory. The AI boss is coming. It still needs a better filing system. Here’s what happened in AI today: - 😸 Anthropic's IPO could raise more than $100 billion, potentially the biggest stock market debut ever. - 📰 Nvidia is spending $7 billion on AI startup Poolside to build a US rival to Chinese open-source models. - 📰 Alibaba and Tencent spent $18 billion on AI infrastructure last quarter alone. - 📰 Blackstone and Anthropic are embedding 160 AI engineers directly inside portfolio companies to build new products. 😺 Claude's Maker Could Pull Off the Biggest IPO in History You've probably used Claude to draft an email, debug a script, or automate half your job this month. Turns out enough other people have too that Anthropic might be about to pull off the biggest stock market debut ever. Here's the deal: Anthropic's bankers have told investors the company could raise "more than $100 billion" when it goes public, according to the New York Times. That would put its value near $2 trillion, more than double the $965 billion it was worth after its last private funding round in June. Here's what happened: - Anthropic's IPO could "match or beat" SpaceX's June debut, which raised $85.7 billion at a $1.77 trillion valuation (the largest IPO ever) - Morgan Stanley, Goldman Sachs, and JPMorgan are reportedly leading the offering, with a listing possible as soon as October - The company could reveal its public filing paperwork in the coming weeks, with shares possibly listing this autumn - Only Apple, Microsoft, and Nvidia have ever crossed the $2 trillion valuation mark - Anthropic would beat OpenAI to the public markets; OpenAI isn't planning to list until 2027 Why this matters: Claude Code, Anthropic's coding assistant, has become one of its most popular products and helped push the company's projected annual revenue to $47 billion. That's a five-year-old startup founded by siblings Dario and Daniela Amodei, both former OpenAI executives, betting its safety-focused approach to AI can out-earn (and now out-IPO) the company they left. An IPO this size would also settle an argument that's been simmering all year: whether the "safety-first" lab can actually out-compete the "move fast" one. So far, the money says yes. Our take: The hype has a couple of asterisks. Anthropic has struggled to get enough chips and servers to keep up with demand, and in March, the Trump administration cut off its Defense Department contracts entirely after Anthropic refused to give the military unrestricted access to its models (Anthropic called the move "unconstitutional retaliation"). A record-breaking IPO and a fight with your own government's biggest customer don't usually show up in the same earnings call. Anthropic's prospectus should tell us which story is closer to true. FROM OUR PARTNERS The Neuron Exclusive: Invest in High-Potential AI Startups Like These The Neuron and Alumni Ventures are giving readers early access to high-growth startup opportunities, including some of today’s most exciting AI, Deep Tech, Quantum Computing, and Cybersecurity companies co-invested alongside top VC firms like Andreessen Horowitz (a16z), Bessemer, & Y Combinator. You get: - Curated deal flow of high-potential AI First startups - AV is already investing alongside elite lead venture firms in these deals - No cost to see deals - No obligation to invest Don’t miss your chance before access closes. 🎓 AI Skill of the Day: Make AI prove its homework So, apparently State Farm’s outside lawyers admitted AI helped put seven nonexistent case citations into court filings. The useful lesson is much broader: drafting and verification should be separate steps. - Ask AI to list every factual claim, number, quote, and source in your draft (best result if you provide the original context in the chat window, and ask it to include “the exact language from the source document so i can command/control + F find it and check your work”). - If AI is giving you links, open each source yourself and, well, command + f find it. If it does not directly support the claim, rewrite or delete it. - Do not ask AI to “repair” a citation it invented. Start from a real source. - If you’re using AI to help you research, Copy this: Audit this draft for factual risk. List every claim, number, quote, and named source I must verify manually. Give only URLs already present in the draft. Mark anything unsupported as DO NOT PUBLISH. Do not invent or repair sources. FROM OUR PARTNERS MongoDB Connects Live, Operational Data to the AI Tools Builders Use. A new Managed MCP Server connects Claude Code, Codex, Grok Build, and Devin to MongoDB Atlas, giving coding agents direct access to live operational data. 📰 Around the Horn - Nvidia is paying $6 billion to license AI startup Poolside's technology and hire its engineers, plus investing another $1 billion at a $12 billion valuation, in a bid to build a US-made open-source AI model that can compete with China's. - Nvidia is also raising prices on its next-generation Grace Blackwell and Vera Rubin chips by 15-17% starting early next year, adding roughly $5 billion to the cost of a single large data center. - Alibaba and Tencent combined to spend $18 billion on AI infrastructure last quarter alone, a number so big it dragged Alibaba's profit down 75% even as its CEO says the investment will pay for itself within three years. - OpenAI expanded ChatGPT Ads to 31 European countries last week (Aug 18), meaning free and low-cost ChatGPT users across the continent will start seeing ads (paid subscribers stay ad-free). - Blackstone and Hellman & Friedman are embedding a 160-person team of AI engineers, backed by a $1.5 billion partnership with Anthropic, directly inside their portfolio companies to build new AI-powered product lines. - Dr. Dre said he uses AI to produce songs and called musicians who oppose the tools "afraid of learning new things," adding that fellow producer Timbaland is a fellow "closet AI user." - A new investigation into the data labelers who train AI models found workers in China and Australia get less work over time as AI improves, don't know which companies they're working for, and have little ability to appeal how their performance gets graded. 🍪 Treats to Try - *You're paying for Claude but working in the wrong mode. This free guide shows you when to use Chat, Cowork or Code, plus the model settings that stop you burning your usage limit. No tech knowledge required. Click Here for Free Instant Access to the Claude Mode Guide - Is Agentic scans your website and shows you exactly what's blocking AI agents from browsing and using it, then tells you how to fix it. - FetchSandbox fake-runs the API integrations your AI coding assistant just wrote, complete with real webhooks and failure states, so nothing touches your actual account—free to try, Pro plans from $5/month. - Construct is an AI employee with its own cloud computer that reads your email, researches topics, and finishes assigned work while you're away—free trial, then $9/month. - USB lets you write one AI skill and install it into 16 different coding assistants, including Claude Code and Cursor, with a single command—free to try (open source). - Mnemosyne gives your AI coding assistant a persistent memory that survives between sessions, storing everything as readable notes in an Obsidian vault you can browse yourself—free to try (open source). 📖 Rotating Section (Replace w/ appropriate emoji + daily section title) New from The Neuron: AI Explained A Cat’s Commentary That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
04:00

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

AI therapy chatbots routinely miss real danger when teens talk in slang and irony, a new study finds. Claude, GPT-4o and Llama-3.1 understood 76-82% of young people's mental-health phrasing but correctly flagged only 64-72% of actual risk, a gap human therapists don't have. With a 34% miss rate on crises in the benchmark, the authors estimate over 146,000 missed crises a year among US teens, and failures compound badly when several risky patterns appear at once. Lightweight fixes didn't help and only heavy, six-times-costlier scaffolding reached human performance, so the paper calls for mandatory human-in-the-loop design, youth-specific validation and regulation.

Notes
When Vocabulary Comprehension Fails Clinical Reasoning (arXiv, cs.CL, 2026-08-24)

Purpose: Evaluates safety of LLM-based therapy bots/general chatbots for mental-health support to Generation Alpha (born 2010–2024), whose informal speech patterns (hyperbolic language, ironic positivity, rapid semantic drift, contextual polysemy) are unvalidated. Motivation: multiple adolescent deaths linked to AI chatbot interactions; 13.1% of U.S. adolescents (5.4M) use generative AI for mental health advice.

Method: Two new benchmarks:

  • 64 Gen Alpha mental-health expressions, validated by native speakers (ICC=0.72) and clinicians (kappa=0.78).
  • 75 multi-turn conversations (780 turns), paired Standard vs Gen Alpha versions.

Models: Claude, GPT-4o, Llama-3.1 (architectures behind therapy apps and general chatbots).

Results:

  • Vocabulary understood 76–82%, but clinical risk correctly calibrated only 64–72% → 10–14 percentage-point gap (p<.001, d>0.48). Human therapists: 3pp (p=.22).
  • Gap is architecturally consistent and widens with ambiguity (7pp → 18pp).

Six failure patterns (gap size): sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound: ≥3 present → 94% miss rate.

Mitigations: Lightweight fixes fail; only heavy scaffolding reaches human performance, at 6.4x cost.

Authors' recommendations: Mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, regulatory frameworks for youth-facing mental health AI. Baseline 34% miss rate → estimated 146,880 missed crises/year.

Caveat (per abstract): Numbers are from the submitted abstract; multi-model results aggregated — per-model breakdowns not given here.

Full text · 2,775 chars
Computer Science > Computation and Language Title:When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha View PDF HTML (experimental) Abstract:Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

A new paper shows the automatic answer-checkers used to train AI on math problems are biased against non-English answers, quietly punishing correct responses written in Japanese or Chinese. On standard math problems, an exact-match verifier rejected correct answers at sharply different rates by language — for Qwen3-8B, a 0.642 false-negative rate on Japanese answers versus 0.122 on English and 0.073 on Chinese. The authors trace the problem to the answer-interface stage and propose a reusable auditing protocol, plus a selection rule that closes most of the gap but still depends on genuine cross-lingual support.

Notes

Multilingual Verifier Bias in RLVR

arXiv cs.CL, preprint (arXiv feed, 2026-08-24).

Core claim

RLVR's answer verifier is assumed "language-neutral," but an exact-match verifier turns format/script variation into language-dependent false-negative reward noise in multilingual settings.

Key numbers (MGSM rollouts, k=8)
  • False-negative rate on trusted-correct answers differs sharply by language across Qwen3-4B, Qwen3-8B, Llama-3.1-8B-Instruct.
  • Qwen3-8B: 0.642 JP vs 0.122 EN vs 0.073 CN.
Method (reusable audit protocol)
  • Verifier-robustness suite
  • Rollout-diagnosis procedure
  • Language-conditioned reward-error metrics (Japanese, English, Chinese)
  • A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while residual accuracy gap is unchanged.
Cross-lingual selection bottleneck
  • On MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55–78% of the average selection gap.
  • Over 95% of repairs require genuine cross-lingual support.
  • Replicates on a 483-problem MATH-500 set.
Training audit
  • Rule-GRPO raises trusted accuracy while reward-error VLB stays high.
Takeaway (operational)
"multilingual RLVR rewards should be audited by language and by answer interface before they are optimized."

Caveat: the audit protocol itself, not a fix — the bottleneck (cross-lingual selection) is identified but not solved.

Full text · 2,446 chars
Computer Science > Computation and Language Title:Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck View PDF HTML (experimental) Abstract:Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative reward noise. We introduce a reusable protocol for auditing multilingual RLVR rewards: a verifier-robustness suite, a rollout-diagnosis procedure, and language-conditioned reward-error metrics for Japanese, English, and Chinese answers. On MGSM rollouts with k=8, the exact-match proxy rejects trusted-correct answers at sharply different rates by language across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct; for Qwen3-8B, the false-negative rate reaches 0.642 on JP against 0.122 on EN and 0.073 on CN. A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while the residual accuracy gap is unchanged. We then expose a cross-lingual selection bottleneck: on MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55-78% of the average selection gap, and over 95% of repairs require genuine cross-lingual support. The bottleneck replicates on a 483-problem MATH-500 set. A controlled training audit shows that rule-GRPO raises trusted accuracy while the reward-error VLB stays high. The unifying message is operational: multilingual RLVR rewards should be audited by language and by answer interface before they are optimized. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:40

Alibaba to Raise $10.20 Billion for AI Investment With Share Placement - WSJ

Alibaba is raising $10.2 billion through a share placement to bankroll its AI push. The company says all of the proceeds will go to AI-related investments. It's one of the bigger cash raises from a Chinese tech giant for the AI arms race.

Full text · 158 chars
... Artificial Intelligence Conference, with large screens displaying Alibaba said it plans to use all of the placement proceeds for investments in its AI ...
09:00

Kids outlearn AI—and we still don’t know why

Kids still learn language far more efficiently than AI, and scientists don't yet know how they do it. A modern large language model trains on roughly a hundred thousand times more words than a child hears, yet children master grammar on a fraction of that data. The gap, called the data efficiency problem, could bite as the internet's usable text runs dry, possibly in the 2030s. Researchers are probing kids' tricks through BabyLM, a competition that trains models on as few as 10 million words, and early results are undercutting ideas like curriculum learning. Cracking it could mean AI that trains on far less data, helping everything from minority-language chatbots to video learning.

Notes

I'll write the notes directly, then finish the task.

Kids outlearn AI—and we still don't know why

MIT Technology Review · Elise Cutts · 2026-08-24

The data efficiency gap
  • Children learn a language to full fluency from ~100M–300M words heard; LLMs need roughly a hundred thousand times more. Michael C. Frank (Stanford cognitive scientist): "We still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year."
  • Meta's Llama 3.1 (open weights, 2024) was pretrained on 15 trillion tokens. Ethan Gotlieb Wilcox (Georgetown) says frontier models may pretrain on ~10x that. The easily-available internet data well may run dry as early as the 2030s.
  • Scale analogy (Wilcox): "Claude has seen the amount of language that an entire city will experience in one generation." All printed training text would stack past the International Space Station; a preteen's 100M words stack just 20m.
  • Toddlers typically produce grammatically correct sentences after ~10M words heard (up to 30M).
  • Frank on GPT-2 at 30M words: "you get a nonsense generator; you don't get a kid."
The Chomsky/Skinner debate
  • 1950s: Noam Chomsky argued babies are born with hardwired grammar; he was reacting to B.F. Skinner, who held language is learned purely via conditioning. Chomsky's "poverty of the stimulus" claim — syntax is too complex and exposure too sparse to learn from statistics alone.
  • Richard Futrell (UC Irvine): Chomsky's "signature argument was, essentially, that language cannot be learned on the basis purely of statistics."
  • Generative grammar dominated US linguistics for decades and shaped rule-based symbolic AI, which largely failed at real language. The 1970s AI winter followed.
  • 2018–19: transformer-based BERT and GPT-2 (trained on billions of tokens) proved data-hungry learning worked; ChatGPT (2022) made it public.
  • Alison Gopnik (UC Berkeley): "No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax. I didn't think that was going to turn out to be true."
BabyLM
  • Founded by Alex Warstadt (UCSD), Leshem Choshen, and others after a 2022 Twitter thread. Annual competition: train models on a "developmentally plausible" corpus of 100M words (toddler track: 10M) — storybooks, dialogue, movie subtitles, Simple English Wikipedia, Wikipedia, and transcripts of child-directed speech. Evaluated on psycholinguistic grammar benchmarks.
  • Sample task: compare The keys to the cabinet are on the table vs ...is on the table; humans show surprise (measured by eye-tracking), models via surprisal.
  • Curriculum learning (simple→complex data) was the most popular first-round approach but underperformed — Aaron Mueller (BU): transformers "don't really need to have their data ordered."
  • 2024 winner GPT-BERT (hybrid next-token + masked/BERT-style), pretrained on ~100M words, beat Llama 2 70B (pretrained ~15,000x more data) on one benchmark. But BabyLM models aren't LLM-grade; many can't generate text.
Sensory learning
  • SAYCam: Michael Frank + 4 colleagues recorded 3 children (~2 hrs/week, ages 6 months–2.5 years), released openly.
  • Brenden Lake (Princeton): 2024 model trained on 61 hours of raw SAYCam data learned to identify objects and associate them with words without the biases developmental theory claims children need (e.g., that "shoe" means a whole object, not a part). Caveat: "we don't get a two-year-old out of [training] when we're done."
  • BabyView: SAYCam successor.
  • Uri Hasson (Princeton): recorded the first 1,000 days of 17 children's lives, cameras+mics in every living area (not bedrooms/bathrooms), 12 hrs/day, nearly every day. First described in a recent preprint; only feasible with new AI transcription/analysis tools. "For the first time, we have the input."
Missing ingredients & limitations
  • Video-trained models lag far behind text models; Lake's model learned only simple words ("ball," "cat"); adding visual data didn't help BabyLM.
  • Gopnik: kids are active explorers who choose their own data — "kids are constantly experimenting"; her Minecraft-style studies show play maximizes "empowerment" (predictable impact).
  • Elizabeth Bonawitz (Harvard): children reason about the teacher and their intent, not just evidence — unlike models, which learn passively and in isolation.
  • 2025 BabyLM allowed interactive/social models, but they didn't outperform standard ones.
  • Meta is the most kid-curious lab (involved in BabyLM's multimodal branch; announced a headcam benchmark with Frank). Flapping Airplanes, a stealth startup, is interested in Frank's research. Meta, Google DeepMind, OpenAI, Flapping Airplanes all declined interviews.
  • Gopnik predicts the next generation of AI (post-transformer) will take lessons from developmental psychology.
Motivation
  • NanoGPT Slowrun benchmark (Q Labs, launched March 2026) — similar goals to BabyLM but dropped the human-learning framing for pure data efficiency (per Mueller).
  • Warstadt: democratize AI so non-hyperscale labs can train good models.
  • David Samuel (U. Oslo, GPT-BERT architect): low-resource languages — Sami may have only tens of millions of tokens, "about the scale of a toddler's exposure."
  • Bonawitz, initially skeptical, now supports using LLMs as a "linguistic lab rat" (model organism) for comparative study; Futrell likens it to teaching an alien a language then opening its brain.
Full text · 23,555 chars
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Now there are two. Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers. LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language. “The progress recently has been amazing,” Michael C. Frank, a cognitive scientist at Stanford University, says of LLMs. “But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.” This yawning divide between children and machines is called the data efficiency gap. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built? Finding answers has stakes for both AI research and cognitive science. For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry. Kids show that it could be possible to learn more with less. Far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20. The difference in scale is something that can only really be gestured at in analogy. “Claude has seen the amount of language that an entire city will experience in one generation,” says Wilcox. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station. The human preteen’s 100 million words, meanwhile, would stack up just 20 meters. And we can make do with far less than that. By reverse-engineering the way kids learn, scientists hope to be able to create more data-efficient AI models, which could be useful for everything from training AI effectively on video to creating chatbots that serve minority language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and children’s developing minds. Are we born with a language instinct, or would it be possible, even in principle, for a child to learn language purely from experience? Is the way we process language a quirk of our biology, or might at least some of it reflect universal constraints on how languages can be used and learned? The essential elements Most of us realize language is hard only when we try to learn a new one after childhood. The past perfect tense, rolled rs and nasal vowels, the genitive case, phrasal verbs, grammatically masculine tables and feminine spoons—many are the instruments of linguistic torment for the adult language learner. It’s typically effortless to learn our mother tongues, however. Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end. “It’s just totally miraculous,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” Exactly how babies pull this off is a mystery. Researchers know a lot about what kids learn and how they use language at different stages in development, but there’s still a lot we don’t know. Perhaps the most enduring question is why babies can learn language at all. The syntax of human language—the rules for combining words into sentences—includes recursive, nested structures that allow us to express virtually infinite ideas with a finite lexicon of words and pieces of words. This seems like something that should be a problem for babies. They only splash about in the shallows of a fathomless ocean of language. And yet, somehow, that’s enough. From a drop, they infer the depths. One solution, put forward in the 1950s by the MIT linguist Noam Chomsky, is that babies are born with hardwired knowledge of grammar. Chomsky was reacting to a rival view, championed by the psychologist B.F. Skinner, that language acquisition is entirely environmental. Skinner thought language was learned through conditioning and reinforcement, the way a dog figures out how to sit or shake for treats. Chomsky countered by citing the “poverty of the stimulus”—the idea that language, especially syntax, is too complex and children’s exposure to it too “impoverished” for them to learn entirely from experience. “His signature argument was, essentially, that language cannot be learned on the basis purely of statistics,” says Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Instead, Chomsky posited that language is based on a set of logical rules and argued that children needed innate knowledge of those rules to deduce the grammar of their language from scraps of speech. “It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” Michael C. Frank, cognitive scientist, Stanford University The Chomskyan view of language dominated linguistics in the US for decades under the moniker of generative grammar. And it was a major influence on computer science in the 1950s and ’60s, when AI was enjoying its first boom time and the lines between linguistics and natural-language processing dissolved in a flood of military funding; the Pentagon wanted computers that could understand English and translate Russian. Despite early successes of simple neural networks, which learn to recognize and reproduce statistical patterns, AI researchers in the United States largely adopted a rule-based framework influenced by Chomsky’s theories. They tried to teach language to computers by explicitly coding the rules into programs—think less immersion experience, more grammar class. This approach, part of a broader trend called symbolic AI, prevailed for decades. It also largely failed to produce models actually capable of handling human language at scale. Interest in natural-language processing chilled in the “AI winter” that began in the 1970s. In the aftermath, neural networks started to make a comeback. But it wasn’t until the 2010s, when computer hardware was getting cheap and capable and the internet was getting big, that their performance began turning heads. By 2018 and 2019, the models BERT and GPT-2, which were built on a new architecture—the transformer—and trained on billions of tokens, made it clear to insiders that learning from a massive glut of data could work for language. In 2022, with the breakout success of OpenAI’s chatbot ChatGPT, it was clear to everyone. LLMs are not brains. What they are is powerful statistical learners—naïve pattern-learning machines without any of the evolved biological quirks folded into the human cortex. In other words, they are exactly the kind of thing a generative linguist two decades ago would have thought could not learn language. And yet here they were, writing believable sonnets and passing grammar tests. “No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax,” says Alison Gopnik, a developmental psychologist at the University of California, Berkeley. “I didn’t think that was going to turn out to be true. And I think most people didn’t think that you could just look at the statistics of a large sample of language and figure out grammar.” But what about learning from a small sample of language—a child-size one, say? Is it possible to build a baby-scale model that’s anything more than a nonsense generator? Baby talk Alex Warstadt, a linguist and data scientist at the University of California, San Diego, remembers the years around the release of BERT and GPT-2 as a heady time. Back in 2019, he was still a PhD student in linguistics at New York University, watching his field change before his eyes. The mere fact that language models could learn English by churning through text was a challenge to prevailing Chomskyan ideas. But many linguists remained skeptical that LLMs could tell us anything about how humans acquire language. “I always got pushback on one issue in particular. And that was the size of the data sets of the model,” says Warstadt. “There was never a time when people were training language models at human scale where we were impressed by them.” But Warstadt saw promise in LLMs: A scientific model doesn’t have to be perfect to be informative, and LLMs were clearly powerful simulations of human language use. By building hypotheses about how children learn into models and measuring their performance—how close they came to closing the data gap—might scientists be able to put their ideas to the test? In August 2022, Warstadt posted a Twitter thread laying out an argument that neural networks could be useful models of language acquisition. After some back-and-forth in the comments with AI researcher Leshem Choshen, Warstadt floated the idea for what would become BabyLM, an annual competition organized by Warstadt, Choshen, and several other researchers to train models on small data sets. That was four years ago. Since then, BabyLM has added workshops and inspired spin-offs including a competition for baby models trained on Chinese. The main event challenges researchers to train language models on a “developmentally plausible” corpus of just 100 million words (for the toddler-scale track, 10 million) drawn from storybooks, dialogue, movie subtitles, Simple English Wikipedia, normal Wikipedia, and actual transcripts of speech directed at children. The models are evaluated on the kinds of grammar benchmarks that psycholinguists use with humans, says Georgetown’s Wilcox, one of the organizers. One kind of task involves presenting test subjects—human or machine—with sentences and looking for indications of confusion or surprise at ungrammatical features. For instance, a test might compare the sentences The keys to the cabinet are on the table and The keys to the cabinet is on the table. “When humans see ‘is,’ they’re like: What? That’s not supposed to be ‘is,’ ” says Wilcox. For a human, that surprise might be measured by tracking eye movements. For language models, researchers use a measure called surprisal, which assesses how unlikely the model predicts a sentence or part of a sentence to be. The competition has already challenged some assumptions, such as the effectiveness of curriculum learning. Curriculum learning starts with simple training data and works up to more complex inputs—a bit like starting with baby talk and getting more sophisticated over time. And it was by far the most popular approach taken in the first round of BabyLM, says Warstadt. But it didn’t work as well as expected. “The appeal is just kind of hard to resist, you know. [Curriculum learning] seems to really line up with ways that we believe humans are learning,” says Aaron Mueller, a computer scientist at Boston University and one of the BabyLM organizers. “But it seems like these transformers don’t really need to have their data ordered in such a way to learn effectively.” Perhaps a touch ironically, the best BabyLM models aren’t inspired by babies at all. The 2024 champ, GPT-BERT, is a transformer trained partly to predict the next token in a sequence, like modern LLMs, and partly to act like BERT, a “masked language model” that fills in the blanks in sequences of tokens Mad Libs style. Impressively, when GPT-BERT was pretrained on about 100 million words, it was able to beat the performance of Meta’s Llama 2 70B—an LLM pretrained roughly 15,000 times that amount—on one of the BabyLM benchmarks. Still, BabyLM models are not on the same level as LLMs. Many can’t produce text at all, and even GPT-BERT would seem clunky next to a modern commercial model. Ultimately, while they are “baby”-size, the way these models learn isn’t very baby-like. Kids are not disembodied computer programs whose only “experience” of the world comes through written text. They take in the world via their senses—especially vision and hearing. To close the data gap, some researchers think, machines will need to start learning through the eyes and ears of children. Taking it all in When Michael Frank started his lab at Stanford about 15 years ago, scientists didn’t really know how babies experience the world. Developmental psychologists were just beginning to glimpse babies’ lives through headcams. “The insights that came out from that early research were that kids’ experience looks really radically different than we thought,” says Frank. “It’s much more focused: They’ve got these little short arms, so the objects are, like, right in front of them. And they live in a forest of knees.” Frank was excited to use headcam footage to train machine-learning models to test hypotheses about how kids learn language, but he needed more data. So he and four colleagues recruited three babies—all the children of psychologist mothers who knew what they were getting themselves into—to don headcams for science. The project, called SAYCam, recorded two hours a week of each child’s life between six months and two and a half years of age. “[The families] were willing to release that video, and that’s critical,” says Frank. “So we released it, and people started training models on it.” One of those people was Brenden Lake, a cognitive scientist and AI researcher at Princeton. In 2024, when he was working out of New York University, he and his colleagues presented a model trained on 61 hours of raw SAYCam data that learned to identify objects and associate them with words. Many theories in developmental psychology propose that children need some biases to help them pick out particular parts of their raw sensory experience and associate them with bits of language. For instance, it’s thought babies assume that a new word like “shoe” refers to a whole object rather than a part of it (like a shoelace), says Lake. But the model Lake’s team built was able to learn to identify objects in the video footage and associate them with words without any such biases. “It turns out you can get a real start on language learning using a lot less than what a number of theories suggested,” says Lake. Still, he adds, “we don’t get a two-year-old out of [training] when we’re done.” But perhaps it’s not surprising that such models can’t replicate childlike capabilities by working with a few dozen hours of footage cobbled together from short snapshots over several years of a child’s life. It could be that the shortfalls just indicate a lack of realistic data. After all, babies can’t wear a headcam 24-7; efforts like SAYCam and its successor, BabyView, record at best a few hours a week. So researchers have the choice between working with a tiny slice of the life of a single child or with larger data sets of footage pooled from many kids. Either way, a model’s training data is still a far cry from the lived experience of a child. That could be changing. Uri Hasson, a neuroscientist and psychologist at Princeton, spent the last five years on a project to record the first 1,000 days of 17 children’s lives. The participating families wired every living area in their homes (except bedrooms and bathrooms) with cameras and microphones and recorded 12 hours a day, almost every day. The resulting data set, described for the first time in a recent preprint, is of a scale that would have simply been impossible to work with absent new AI tools for transcription and video analysis, says Hasson. “For the first time, we have the input,” he says. “It’s really only the beginning.” Missing ingredients So far, training models on video has proved difficult. While text-based models emerge fully fluent (after ingesting huge training data sets), multimodal models trained on video from kids are far from that. Lake’s model, for instance, learned simple words, like “ball” and “cat.” Attempts to supplement text with visual data haven’t worked for BabyLM participants, says Warstadt. Gopnik thinks the issue could be that kids do not simply sit and watch the world go by. “Children are actively exploring, which means that they’re actively choosing their own data,” she says. “Kids are constantly experimenting.” Maybe that’s the missing ingredient. Research by Gopnik’s group—including studies of grade schoolers exploring a Minecraft-inspired game—shows that what looks like child’s play is in fact an effective way to learn cause and effect. Kids seek out experiences and take actions that maximize their “empowerment,” or the ability to make a predictable impact on the world. Unlike models, children are aware of what they don’t know and have a drive to fill their knowledge gaps, says Elizabeth Bonawitz, a developmental cognitive scientist at Harvard. And children’s social lives also help them learn, she says. Her research has shown that children interpret information differently when they know an adult is trying to teach them something. “Children are not only reasoning about the evidence they’re being told,” says Bonawitz. “They’re reasoning about the teacher, about the teacher’s knowledge, and about why the teacher is telling [them] this particular information.” That’s very different from how models learn: passively and in isolation. Perhaps if models were built to seek out information to fill in their own blind spots, experiment with language and observe how other language users react to their babbling, and reason about some kind of simulated social world, they’d learn better. Last year’s BabyLM actually opened the competition to models that could learn by interacting with other models. But the social models didn’t outperform standard ones. Of the leading industry labs, Meta seems the most interested in taking inspiration from kids—specifically for training models from video. Two Meta researchers were involved in BabyLM’s multimodal branch, and Meta scientists—together with academic researchers, including Frank—recently announced a benchmark and challenge for training models on baby headcam footage. Frank also says a stealth-mode AI startup called Flapping Airplanes has taken interest in his research. Neither Meta, Google DeepMind, OpenAI, nor Flapping Airplanes agreed to an interview. For now, frontier labs aren’t exactly racing to borrow tricks from children, says Gopnik. She thinks it’ll be the next generation of AI—whatever replaces the transformer—that will take lessons from developmental psychology. Perhaps the most enticing reason to close the data gap is that it could help us understand ourselves. In general, the machine-learning community is less interested in mimicking the brain than in just building something that works, says Mueller. But he thinks awareness of—and interest in—the data efficiency gap is growing. An example is the NanoGPT Slowrun benchmark, launched by Q Labs in March 2026. “They have very similar goals to BabyLM,” says Mueller. “But they’ve dropped the motivation from human language learning and really just focused on the data efficiency angle.” One reason Warstadt wants to close the data gap is to democratize AI so that universities and others without the resources to hyperscale can train good models and stay relevant in AI research. David Samuel, a machine-learning researcher at the University of Oslo and one of GPT-BERT’s architects, has a more personal reason to work on this problem. He’s Czech and works in Norway, and there’s a lot less data in Czech and Norwegian available for training LLMs than there is in English. Minority languages like Sami might have just tens of millions of tokens available, says Samuel—about the scale of a toddler’s exposure. “The question was,” he says, “how can we develop language models that are just as capable as the English ones for small languages?” But perhaps the most enticing reason to close the data gap is that it could help us understand ourselves. Bonawitz says she was initially skeptical that large language models could reveal anything about cognition. LLMs and brains are, after all, very different. Brains are embodied. Our neurons are not tidy lines of code but living cells. And our brains grow and change as we learn and age—LLMs pretrain once and never again. But as different as the two systems are, says Bonawitz, “I’m sort of revising my beliefs.” She’s been won over by the idea of studying models the way comparative psychologists might study animal minds to illuminate our own. Researchers like Warstadt, Frank, Wilcox, Lake, and Hasson are already using language models as a kind of linguistic lab rat, an imperfect but informative stand-in for a real human language user—especially for questions that are more about learning and language and information processing than anything specific to our brains or biology. When models can do things with language we thought were impossible, it challenges old assumptions. And researchers can build hypotheses about language learning into models—say, by simulating different degrees of bilingualism or depriving models of exposure to certain grammatical forms—and test those hypotheses in a way that would be impossible to do with real children. Futrell compares the situation to teaching language to an alien and then opening up its brain to see what happened. While other animals communicate, only humans converse. Now there’s something neither animal nor human that can talk, too. LLMs open up the possibility for comparative studies, even if models and minds are vastly different. “For the last 100,000 years or however long human language has existed, humans have been the only entities in the universe that use language. Now there’s this other linguistic entity,” says Warstadt. “Finally we have a model; not in the sense of a language model, but in the sense of a model organism.” Elise Cutts is a science writer based in Austria. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
12:25

Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI - WSJ

Nvidia is spending $6 billion to back a US alternative to Chinese artificial intelligence. The investment targets Poolside, an AI startup building foundation models. Poolside's founders said the deal is meant to ensure AGI, or humanlike computer reasoning, is developed in the United States rather than China. It's part of a broader US push to keep AI leadership out of Beijing's hands.

Full text · 154 chars
Poolside's founders wrote that the Nvidia deal was meant to secure a future where artificial general intelligence , or humanlike computer reasoning, “ ...
12:27

Hugging Face Considers $13 Billion Sale of Its AI Platform | PYMNTS.com

Hugging Face, the go-to hub for open-source AI models, is reportedly exploring a sale that could value it at around $13 billion. The company hosts millions of AI models and datasets that developers use for everything from chatbots to research. A deal at that price would be one of the largest acquisitions in the open-source AI space. Talks are still at an early stage and nothing has been confirmed.

Full text · 127 chars
Artificial intelligence startup company Hugging Face has reportedly been considering a sale that could value it at $13 billion.
02:03

DBS and IBF to help prepare finance workers for AI-shaped roles - The Edge Singapore

Singapore's biggest bank is teaming up with the national finance training body to reskill workers for AI-shaped jobs. DBS and IBF will build IBF-recognised programs in AI governance, responsible AI, and practical prompt engineering. A concrete workforce-training partnership, routine for the sector.

Full text · 148 chars
Under the agreement, the two will develop IBF-recognised programmes in AI governance, responsible AI, prompt engineering and practical workplace ...
02:58

DBS & IBF partner on AI skills & jobs in Singapore - CFOtech Asia

Singapore's DBS bank and the Institute of Banking and Finance are teaming up to train the banking workforce in AI. They're offering IBF-recognised programmes covering AI governance, responsible AI, prompt engineering, and practical applications. It's a skills-programme announcement rather than a product launch.

Full text · 150 chars
Under the first area, DBS and IBF plan to offer IBF-recognised programmes covering AI governance, responsible AI, prompt engineering and practical ...
04:00

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

How a person's notes are organized, not how much of them you feed an AI 'digital twin,' is what really limits how well it predicts the person's choices. A hand-built schema covering background, decision process and evaluation beat raw transcripts by about 1.9 percentage points on one benchmark, but got no benefit on harder, mixed tasks. So the authors built an automatic pipeline where an LLM invents its own task-specific structure, and that restored the same ~1.9 point gain across 13 test areas. The takeaway: structure beats volume when simulating a person.

Notes

Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins

Source: arXiv cs.CL paper (2026-08-24). LLM "digital twins" simulate how an individual responds to new questions given prior-response data.

Core claim: Prior work shows compressing long transcripts into LLM-generated summaries doesn't hurt predictive accuracy — so information volume isn't the bottleneck. This paper argues the constraint is structural: how persona info is organized before feeding the simulator.

Fixed-schema experiment (BDE): Introduces a hand-crafted schema — BDE: Background, Decision procedure, Evaluation — grounded in consumer-behavior theory. On homogeneous benchmark Twin-2K-500 it beats raw transcripts by +1.91 percentage points, with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks.

Key caveat: The fixed schema does not generalize — on more heterogeneous tasks, performance is statistically indistinguishable from the raw transcript baseline.

Automatic structure-discovery: LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On 13 diverse sub-studies it restores performance, improving mean accuracy by +1.91 pp over raw transcripts and eliminating the significant losses seen with the fixed schema.

Conclusion (quoted):

"the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured — and that the optimal structure depends on the task."

Limitations: No absolute accuracy figures given, only deltas; no analysis of what discovered structures share; benchmark scale (Twin-2K-500) modest.

Full text · 2,672 chars
Computer Science > Computation and Language Title:Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins View PDF HTML (experimental) Abstract:LLM-based "digital twins" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

Researchers built a free synthetic Bengali speech dataset for telecom customer-service AI and put it up on Hugging Face. It holds 10,000 audio-text pairs, roughly 27 hours of speech, generated with OmniVoice's voice cloning from a single real female voice. The audio is CC-BY-4.0 licensed and comes with a cleaned transcript meant for speech-recognition training. Spot checks with a Bengali-tuned Whisper model showed an average word-error rate of 2.54% and character-error rate of 0.59%, though the authors note synthetic speech still has limits.

Notes

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

arXiv cs.CL paper (2026-08-24). Presents a synthetic Bengali speech dataset for telecom customer-care domains.

Dataset (synthetic-telecom-1 / on Hugging Face, CC-BY-4.0):

  • 10,000 audio-text pairs, ~26.82 hours of 24 kHz speech
  • Predefined splits: 9,000 train / 500 validation / 500 test
  • Generated with OmniVoice in voice-cloning mode from a real female reference recording + transcript; settings: bfloat16 precision, 16 diffusion sampling steps, speaking-rate control 1.0
  • Ships original Bengali text plus a normalized transcript field intended for ASR/STT training and evaluation

Evaluation:

  • Automatic intelligibility check on all 10,000 samples using a domain-adapted Whisper ASR fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium
  • Manual listening check on selected samples only
  • Results: average WER 2.54%, average CER 0.59%, median WER and CER 0.00%

Stated limitations (flagged by authors):

  • Evaluation is STT-based — metric reflects text-audio consistency under the chosen ASR pipeline, not human intelligibility (manual check is limited to a sample subset)
  • Synthetic speech may differ from recorded audio in prosody/artifact characteristics; domain-generalization beyond telecom scenarios is untested

Authors conclude the results "suggest strong text-audio consistency" under the selected automatic evaluation pipeline, while noting the limits of synthetic speech and STT-based evaluation.

Full text · 2,173 chars
Computer Science > Computation and Language Title:Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care View PDF HTML (experimental) Abstract:Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under the CC-BY-4.0 license. The speech was generated with OmniVoice in voice-cloning mode using a real female reference recording and transcript, with bfloat16 precision, 16 diffusion sampling steps, and a speaking-rate control value of 1.0. Along with the original Bengali text, the dataset provides a normalized transcript field designed for ASR/STT training and evaluation. We report an automatic intelligibility check over all 10,000 samples using a domain-adapted Whisper ASR model fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium, along with a manual listening check on selected samples. The evaluation gives an average WER of 2.54%, an average CER of 0.59%, and median WER and CER values of 0.00%. These results suggest strong text-audio consistency under the selected automatic evaluation pipeline, while the paper also discusses the limitations of synthetic speech and STT-based evaluation. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

Language models can still quietly hold biased views about who's competent based on gender, race, or income even when their visible answers look fair. Researchers looked inside open-weight models and found these demographic signals shaped internal representations of a user's expertise, then showed that nudging those representations changed the models' answers in question-answering and hiring tasks. The finding suggests standard fairness tests that only check output can miss real bias.

Notes
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

Preprint (cs.CL, arXiv). Motivates with a distinction: LMs pass behavioral bias evaluations, but it's unclear whether they stopped representing biased associations or merely learned not to express them. Claim: representational bias is often detectable even when behavioral bias is invisible.

Causal framework — decomposes occupational bias into two measurement points:

  • Model's internal representation of a user's competence
  • Its observable outputs

Method:

  • Derive steering vectors for representations of user expertise
  • Verify they causally mediate model behavior in two tasks: a question-answering task and a hiring task

Findings:

  • Applied to "several open-weight models" (names/sizes not given in abstract)
  • Demographic attributes — gender, race, socioeconomic status — influence the internal representation of user expertise
  • Effects appear "even in cases where behavioral metrics detect no disparity between demographics"
  • Under intervention, these representations steer downstream behavior, implying failure modes behavioral metrics alone miss

Key caveats / limits in the abstract: no model names, task specifics, effect sizes, or datasets are given; "several open-weight models" is the only scope statement. The paper's core implication is that passing behavioral bias tests is insufficient evidence of debiased representations — the failure mode is latent and requires mechanistic/causal probes (steering-vector mediation) to expose.

Full text · 1,981 chars
Computer Science > Computation and Language Title:Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias View PDF HTML (experimental) Abstract:Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing

Medical records are so long that AI models reliably miss the most important fact when it sits in the middle of the chart, and a new method selects which part of the record to focus on to fix that. Across 2,196 instruction-response pairs and six models, accuracy was nearly 22 percentage points worse in the middle of records than at the edges. The new approach, called QCCS, beat standard retrieval methods — 16.7% versus 3.3% accuracy on mid-record questions — and even won when it wasn't handed the exact right evidence sentence. It's a proof of concept tested on one open model, the Qwen 7B.

Notes
Notes: "Inhibitory Attention for Clinical Long-Context Reasoning" (arXiv cs.CL, 2026-08-24)

Characterizes and mitigates lost-in-the-middle (LitM) effects in EHR processing, coined clinical LitM (CLitM).

Characterization (MedAlign)
  • EHRs "routinely exceed 100,000 tokens per patient"; most consequential fact can sit mid-note.
  • 2,196 instruction-response pairs, six language models.
  • 21.9 pp gap between peak accuracy (59.5%, 95% CI [46.3, 71.0] at 20–30% decile) and trough (37.6%, [23.2, 52.5] at 70–80%).
  • 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline — inside the trough.
Method: Query-Conditioned Clinical Suppression (QCCS)
  • Lightweight, query-conditioned selection gate over context.
  • Baselines: BM25, BM25 + section-header filtering, dense retrieval, cross-encoder reranking (N=83 held-out instructions, Qwen2.5-7B-Instruct, 16k context).
  • Middle-position instructions: QCCS 16.7% vs BM25 3.3%, cross-encoder 0.0%, dense 0.0%, full context 6.7%.
  • Overall: QCCS 25.3% vs ≤3.6% for all retrieval-only comparators.
Key finding — recall is not the bottleneck

At k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS only 34.9%), yet retrieval arms stay ≤2.6% accurate even when they retrieve it; QCCS reaches 25.0% even when it does not.

"query-aligned context selection predicts EHR instruction-following accuracy better than gold-sentence retrieval recall."
Caveats
  • Proof-of-concept evaluation; N=83 held-out instructions only; single 7B model tested. (Clinical harm framing is asserted, not quantified — the note says a missed center fact is "not benign," but no downstream safety measure.)
Full text · 2,727 chars
Computer Science > Computation and Language Title:Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing View PDF HTML (experimental) Abstract:Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edges. In clinical use this is not benign: the single most consequential fact in a note can sit at its center. We term this the clinical lost-in-the-middle (CLitM) problem, give its first systematic characterization using MedAlign, and compare context-selection strategies as remedies. Across 2,196 instruction-response pairs and six language models, we observe a 21.9 percentage-point gap between peak accuracy (59.5%, 95% CI [46.3, 71.0], 20-30% decile) and trough accuracy (37.6% [23.2, 52.5] at 70-80%); 67.8% of reference answers fall between the 10th and 90th percentiles of the EHR timeline, inside the CLitM trough. We introduce Query-Conditioned Clinical Suppression (QCCS), a lightweight query-conditioned selection gate, and evaluate it against BM25, BM25 with section-header filtering, dense retrieval, and cross-encoder reranking (N=83 held-out instructions). With Qwen2.5-7B-Instruct (16k context), QCCS outperforms all five comparators under LLM-as-judge scoring: for middle-position instructions QCCS reaches 16.7% versus BM25 3.3%, cross-encoder 0.0%, dense 0.0%, and full context 6.7%; overall QCCS reaches 25.3% versus at most 3.6% for retrieval-only comparators. This advantage is not explained by retrieval recall: at k=20, BM25 retrieves the gold evidence sentence in 98.8% of instructions (QCCS 34.9%), yet retrieval arms stay at most 2.6% accurate even when they retrieve it, whereas QCCS reaches 25.0% even when it does not. In this proof-of-concept evaluation, query-aligned context selection predicts EHR instruction-following accuracy better than gold-sentence retrieval recall. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

Small word changes in a prompt can cause big swings in how well a language model performs, and a new study finds the most reliable prompts are ones that use exact domain terms and say plainly what to do. Analyzing 132,000 prompt variants, the authors describe a scaling law where better-performing prompts are also more stable. They built an automated prompt-refining agent from these findings that cut performance variance by 40.7% on a code generation task while keeping or improving average quality.

Notes

Notes on Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality (arXiv cs.CL, 2026-08-24). Based on abstract only; full paper not read.

Core finding

LLMs are acutely sensitive to surface-level prompt wording: "minor lexical changes can trigger disproportionate performance fluctuations." The paper claims to be the first large-scale, n-gram token-level mechanistic analysis of prompt stability, distinguishing itself from black-box optimization and coarse-grained template studies.

Method
  • Dataset: 132,000 prompt variants.
  • Level of analysis: n-gram token-level mechanism (not just template-level).
The Scaling Law
"higher average task performance is strongly associated with lower variance and greater robustness across prompt perturbation."

High-performing prompts are also the stable ones — performance and robustness co-vary.

Two identified linguistic drivers of robustness
  • Domain-Specific Terminology — "tightly anchors semantic boundaries."
  • Explicit Action Directives — "formalize reasoning trajectories."

Claimed mechanism: both "constrain the model's interpretative space, effectively 'locking in' more deterministic generation behavior."

Contribution

Automated Prompt-Refining Agent that restructures input queries by injecting domain anchoring + operational constraints.

Result
  • Reduces performance variance by 40.7% in code generation.
  • Mean performance preserved or improved.
Caveats / open questions
  • Single result metric reported for one task (code generation); generalizability across tasks not substantiated in abstract.
  • "Mechanistic interpretability" claim is asserted, not demonstrated here (no error bars, baselines, or ablation detail given).
  • No discussion of when instability might be useful (e.g., creative diversity) or trade-offs of "locking in" behavior.
  • Scaling-law claim rests on correlation; direction of causality (stability→quality vs. quality→stability) is not addressed in the abstract.
Full text · 2,275 chars
Computer Science > Computation and Language Title:Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality View PDF HTML (experimental) Abstract:Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger disproportionate performance fluctuations. Moving beyond black-box optimization and coarse-grained templates, we present the first large-scale, n-gram token-level mechanistic analysis of prompt stability, leveraging a dataset of 132,000 prompt variants. Our investigation reveals a fundamental Scaling Law of Prompt Performance Stability: higher average task performance is strongly associated with lower variance and greater robustness across prompt perturbation. We identify two core linguistic drivers underlying this robustness: (1) Domain-Specific Terminology, which tightly anchors semantic boundaries, and (2) Explicit Action Directives, which formalize reasoning trajectories. Together, these elements constrain the model's interpretative space, effectively ``locking in'' more deterministic generation behavior. Building on these insights, we introduce an automated Prompt-Refining Agent that systematically restructures input queries by injecting domain anchoring and operational constraints. Empirical evaluation shows that our approach reduces performance variance by 40.7% in code generation task, while preserving or improving mean performance. These findings provide a statistically grounded and mechanistically interpretable framework for achieving robust prompt engineering. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

Instead of wiring together a pipeline of separate AI components, OneModel trains a single large model to absorb a whole business workflow into its own parameters, and in live use it cut response latency in half. Deployed in a global financial services system, end-to-end latency fell from 18.7 to 8.0 seconds, while the rate of problems resolved without human help rose from 64.3% to 83.3%. It replaces the router, planner, executor, and reviewer modules typical of industrial agents with reasoning learned into the model itself via continued pretraining and special fine-tuning.

Notes
  • Paper: "How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only One Model" (arXiv cs.CL, published 2026-08-24). Proposed system named OneModel.
  • Problem claimed: Industrial agents use modular pipelines — Router, Retriever, Planner, Executor, Responder, Reviewer — that "fracture into a labyrinth of ad-hoc patches," producing cascading errors and high latency.
  • Approach: Shift from external workflows to internalized knowledge. Instead of slicing "fluid user intents into static steps," OneModel compresses business logic and SOPs into model parameters. Two training stages named:
  • Continual Pre-training (CPT) and logic-compilation SFT, to transform fragmented business rules into "intuitive model reasoning within a unified attention space."
  • Results (online A/B in their global financial service system):
  • End-to-end latency reduced >50%: 18.7s → 8.0s.
  • Intelligent Resolution Rate (IRR) 64.3% → 83.3%.
  • Claim: Breaks the latency/accuracy/complexity trade-off and offers "a scalable blueprint for transitioning industrial agents from complex, error-prone workflows to unified model architectures."

Caveats / limitations (from the abstract itself):

  • Abstract only — no model size, dataset scale, CPT/SFT data mix, ablation, or per-component latency breakdown given.
  • "Silicon Concierge" framing and "cognitive intuition" language are marketing; the mechanism is standard CPT + SFT on compiled logic.
  • Single-domain evidence (one financial service system); no cross-domain or multi-agent comparison beyond the latency/IRR pair.
  • No stated reproducibility info (no code/data links in the abstract).
Full text · 2,136 chars
Computer Science > Computation and Language Title:How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel View PDF HTML (experimental) Abstract:Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high latency. We propose OneModel, an applicable paradigm shift from external workflows to internalized knowledge representation. Unlike modular systems that slice fluid user intents into static steps, OneModel consolidates complex business logic and SOPs directly into the model parameters. Through Continual Pre-training (CPT) and logic-compilation SFT, we transform fragmented business rules into intuitive model reasoning within a unified attention space. Deployed in our global financial service system, OneModel effectively breaks the trade-off between latency, accuracy, and complexity. Online A/B testing demonstrates an end-to-end latency reduction of more than 50 percent, from 18.7 seconds to 8.0 seconds, while the Intelligent Resolution Rate (IRR) increases from 64.3 percent to 83.3 percent. The results show that OneModel can replace brittle engineering logic with internalized cognitive intuition, offering a scalable blueprint for transitioning industrial agents from complex, error-prone workflows to unified model architectures. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Self-Speculation for Faster Reasoning Models

A training-free trick cuts the wait time for reasoning models that think out loud, saving up to about a quarter of total generation latency. Called SSR, it uses a model's shorter chain-of-thought draft as a guesser and its full reasoning trace as the verifier, both from the same model. Because the shorter and longer answers overlap heavily, it can accept long draft chunks at once and also recycle useful text beyond the accepted prefix. On structured and long-form tasks it improved latency by up to 24.1 percent on open-source models like Qwen3.5 and Gemma-4, a win for voice assistants and coding agents.

Notes
  • SSR (Self-Speculation for Reasoning Models) — training-free self-speculative decoding method for reasoning LLMs (arXiv cs.CL, 2026-08-24).
  • Core idea: CoT as speculation source — partial-CoT answer distribution is the drafter, full-CoT distribution the verifier; both from same model at different reasoning budgets.
  • Key observation: later partial-CoT responses show greater semantic + lexical overlap with the full-budget response.
  • Mechanism: accepts long draft prefixes at once (large speedups on structured/long-form generation); adds suffix decoding beyond standard speculative decoding — draft seeds a suffix cache to recover spans past the accepted prefix, cutting latency on tasks with high draft/final lexical overlap.
  • Motivation: long reasoning traces poorly fit latency-sensitive apps (voice assistants, coding agents); existing methods focus on token-level generation without exploiting reasoning workflow structure.
  • Result: up to 24.1% relative reduction in total generation latency on structured and long-form tasks; tested on open-source models Qwen3.5 and Gemma-4.
  • Caveat: gains depend on overlap between draft and final response; evaluation limited to tasks where SSR "is most useful" — no claim of universal speedup.
  • No full paper/method details or baseline comparison given in the abstract.
Full text · 2,520 chars
Computer Science > Computation and Language Title:Self-Speculation for Faster Reasoning Models View PDF HTML (experimental) Abstract:Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performance on these tasks often requires generating long reasoning traces. This is a poor fit for latency-sensitive and interactive applications like voice assistants or coding agents, where generation latency can strongly affect user experience. Existing acceleration methods typically focus on token-level generation, without utilizing the structure of reasoning workflows. We introduce SSR: Self-Speculation for Reasoning Models, a training-free self-speculative decoding method that leverages the chain-of-thought (CoT) as a source of speculation. SSR uses the partial-CoT answer distribution as the drafter and the full-CoT distribution as the verifier, deriving both from the same model at different reasoning budgets. This builds on the observation that later partial-CoT responses often exhibit greater semantic and lexical overlap with the full-budget response. Due to this overlap, SSR can accept long draft prefixes at once, leading to large speedups on structured and long-form generation tasks. To further exploit draft-response overlap beyond the contiguous prefix accepted by standard speculative decoding, SSR also incorporates suffix decoding, using the draft to seed a suffix cache and recover useful spans beyond the accepted prefix, further reducing latency on tasks with high lexical overlap between the draft and the final response. We evaluate SSR on multiple structured and long-form generation tasks where it is most useful, and demonstrate a relative improvement of up to 24.1% on total generation latency for popular open-source models such as Qwen3.5 and Gemma-4. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

A new system mines research papers for testable new ideas by modeling them as mathematical structures instead of flat text. It treats each paper as a small category whose typed relationships link problems, methods, metrics, and claims. A filter rejects cross-domain idea candidates at roughly a 17-to-1 ratio, and over 83 percent of accepted ideas stay quantitatively falsifiable across tens of thousands of papers. Every rejected idea is kept with its reasoning, so the gate doubles as a logging layer rather than a silent filter.

Notes

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

Source: arXiv cs.CL (Computation and Language), 2026-08-24

Core claim: LLM research-idea generators share a structural weakness — they treat each paper as a flat object (string or vector), relying on free-text recombination, random paper pairing, or embedding-similarity retrieval. All three "quotient away the typed problem-method-metric-claim arrows a researcher actually uses" when reasoning cross-domain.

Proposed fix: Model each paper p as a small category C_p — objects are extracted typed research entities, morphisms are relations the paper asserts. A cross-paper bridge is a partial functor candidate F: C_p → C_q preserving object kinds and covered relation classes. Category theory supplies what a typed graph alone lacks: composition + identity arrows, enabling the question of whether an analogy preserves relation chains.

Three-layer algorithm:

  • Categorical signature clustering
  • Functor-preservation gate
  • Six-axis LLM plausibility judge

Results (corpus: tens of thousands of full-text-parsed papers, four ablation conditions):

  • Categorical gate filters cross-domain candidates at ~17:1 ratio
  • Quantitative-falsifier rate of accepted ideas stays above 83% throughout

Key design detail / caveat: every rejected candidate is retained with per-axis rationale — the gate "doubles as a logging layer rather than a silent filter" (i.e., filter is inspectable/auditable, not lossy).

Stated limitation implied: relies on extraction of typed entities and relations from full text; quality of the category model depends on that extraction.

Full text · 2,430 chars
Computer Science > Computation and Language Title:Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure View PDF HTML (experimental) Abstract:Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three approaches fail in the same way: each treats a paper as a flat object, a string or a vector, and so quotients away the typed problem-method-metric-claim arrows a researcher actually uses when reasoning about a cross-domain analogy. We recover the missing structure with the minimal piece of category theory that a typed graph alone does not provide: composition, together with identity arrows, which makes it possible to ask whether a proposed analogy preserves relation chains. Concretely, each paper $p$ is modelled as a small category $C_p$ whose objects are extracted typed research entities and whose morphisms are the relations the paper asserts; a cross-paper bridge from $p$ to $q$ is then a partial functor candidate $F: C_p -> C_q$ that preserves object kinds and covered relation classes. We instantiate the model as a three-layer algorithm: categorical signature clustering, a functor-preservation gate, and a six-axis LLM plausibility judge. Evaluated on a corpus of tens of thousands of full-text-parsed papers under four ablation conditions, the categorical gate filters cross-domain candidates at roughly a 17:1 ratio while the quantitative-falsifier rate of accepted ideas stays above 83% throughout; every rejected candidate is retained with its per-axis rationale, so the gate doubles as a logging layer rather than a silent filter. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Hadith computational science in the age of large language models: a critical narrative review

A review of AI tools for studying Islamic hadith texts finds the field progressing unevenly and argues it should be treated as an evidence-infrastructure problem rather than a benchmark contest. Language models and retrieval-grounded systems now support corpus-scale enrichment, multilingual access, and grounded evaluation, and narrator-verification problems are better formalized. But progress is held back by narrow corpora, weak benchmark comparability, synthetic-to-real gaps, and sparse expert validation, so the authors lay out a research agenda for making the field more useful to Islamic scholarship.

Notes

Hadith Computational Science in the Age of LLMs: A Critical Narrative Review

Critical narrative review of how hadith computational science is reshaped by transformer models, retrieval-grounded pipelines, and LLMs. Combines critique of existing reviews, paper-level appraisal of representative studies, and Islamic-scholar/expert perspectives on authenticity, authority, and responsible use.

Main finding: uneven progress.

  • Data resources expanded; segmentation tasks matured; narrator and source-verification problems better formalized.
  • LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation.

Persisting constraints:

  • Narrow corpora; weak benchmark comparability; synthetic-to-real transfer gaps.
  • Narrator identity resolution; preprocessing fragility; limited reproducibility; sparse expert-grounded validation.

Gaps beyond dominant benchmarks (argued as most important):

  • Non-canonical and obscure corpora; commentary/explanatory literature.
  • Cross-source links with Qur'an and seerah; fiqh-facing evidence support.

Core argument:

"hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision."

On this basis the authors define a research agenda to make the field methodologically stronger and more useful to Islamic scholarship.

Stated limitation: prior reviews document literature growth but "do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use" — the gap this paper addresses.

Full text · 2,537 chars
Computer Science > Computation and Language Title:Hadith computational science in the age of large language models: a critical narrative review View PDF HTML (experimental) Abstract:We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use. We address this gap through a critical narrative review that combines critique of existing reviews, paper-level appraisal of representative original studies, and synthesis of Islamic scholar and domain-expert perspectives on authenticity, authority, and responsible use. We find uneven progress. Data resources have expanded, segmentation tasks have matured, narrator and source-verification problems are better formalized, and LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation. At the same time, progress remains constrained by narrow corpora, weak benchmark comparability, synthetic-to-real transfer gaps, narrator identity resolution, preprocessing fragility, limited reproducibility, and sparse expert-grounded validation. We show that important gaps lie beyond dominant benchmarks: non-canonical and obscure corpora, commentary and explanatory literature, cross-source links with Qur'an and seerah, and fiqh-facing evidence support. We argue that hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision. On this basis, we define a research agenda for making the field methodologically stronger and more useful to Islamic scholarship. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:09

How, when and why to use agentic AI in our neuroscience labs | The Transmitter

Neuroscience labs are being pushed to write formal rules for using agentic AI in their research. As autonomous AI agents spread quickly through science, researchers need clear AI use policies and practices before they rely on them. A guidance-oriented piece aimed at lab leaders.

Full text · 109 chars
With agentic AI's rapid infiltration into science, researchers need to develop AI use policies and practices.
04:14

An AI job boom? Here's what the tedious, temporary work in data labelling is actually like

The AI job boom is mostly a boom in tedious, temporary data-labelling work. That reality flies under the radar amid all the talk of AI creating and destroying careers. A reality-check on what many 'AI jobs' actually involve.

Full text · 132 chars
Amid all the talk about artificial intelligence ( AI ) both creating and destroying jobs, a troubling reality flies under the radar.
05:05

Nvidia reportedly warns biggest customers of 15% price hikes on AI servers - Tom's Hardware

Nvidia has reportedly told some of its biggest customers that AI server prices will jump by more than 15% in many cases. The warning, covered by Tom's Hardware, points to rising costs for the hardware around its AI chips. That makes big AI compute buys more expensive for the companies planning them.

Full text · 142 chars
Nvidia has told some of its largest customers that the prices of servers containing its AI chips will rise by more than 15% in many cases, ...
05:06

KloudMate Unveils India's First Agentic Observability Platform with Auto-RCA and Self ...

A startup launched what it calls India's first observability platform built for AI agents. KloudMate's tool automatically finds the root cause of failures and can trigger self-healing workflows, aimed at keeping digital services running. The announcement is mostly promotional and short on technical detail.

Full text · 150 chars
For digital-first businesses, that figure compounds even faster, resulting in disrupted transactions, eroded customer trust, and engineering teams ...
07:02

KloudMate Unveils India's First Agentic Observability Platform with Auto-RCA and Self ...

A startup unveiled what it calls India's first agentic observability platform, which auto-finds the root cause of software failures and runs self-healing workflows. KloudMate, based in Bengaluru and Texas, says the goal is to cut engineering time lost to reactive firefighting. New product launch, but claims are vendor-provided.

Full text · 150 chars
... engineering time lost to reactive firefighting. Bengaluru & Texas-based deeptech startup, KloudMate, today announced the launch of its Agentic ...
07:39

Illoca launches Plamo, an agentic AI workspace that turns sketches and prompts into ...

Illoca launched Plamo, an AI-powered 3D modelling workspace that turns sketches and prompts into editable 3D models. It targets architecture, engineering and construction professionals. An agentic-AI workspace product launch.

Full text · 134 chars
Illoca has launched Plamo, an AI-powered 3D modelling workspace designed for architecture, engineering and construction professionals.
08:47

Vector adds AI agents to automate development and testing - Engineer Live

Vector's CANoe tool is becoming an AI-powered engineering assistant for development, testing, and analysis. Vector Informatik added AI agents to CANoe, a long-standing tool for automotive and embedded engineering, letting teams automate parts of their build and test work. It's a concrete example of agent features landing in niche industrial tooling.

Full text · 148 chars
CANoe becomes an AI-powered engineering assistant for development, testing, and analysis Image via Vector Informatik GmbH ... agent generate the ...
10:46

📈 Data to start your week

Open-weight AI models are nearly at a 50-50 split with closed ones for all inference tokens, with their share doubling in the past year — though closed-model token use still grew sevenfold in the same stretch. Employment of young workers in AI-exposed US jobs is now running 19% below trend, up from 15% a year earlier. AI agents also passed humans on token use back in February and are now consuming 14x the human amount. These come from a Monday data roundup covering AI, energy and markets.

Full text · 1,021 chars
Hi all, Here’s our Monday roundup of data signals across AI, energy & markets. Enjoy! The state of the AI economy Every week, we will share the latest updates on the state of the AI economy based on our own latest research and tracking. In our latest inference token update, the share of open-weight tokens has doubled in the last twelve months. While we are approaching a 1:1 closed-to-open token ratio, the number of closed-weight tokens grew sevenfold over the same period. See our State of the AI Economy 2026 report for more. 📧 For advisory requests and institutional inquiries, please contact aieconomy@exponentialview.co 🤝 Want to work with us? We are hiring an AI Economy Research Fellow Monday signals - The canary keeps coughing. Employment of young workers (ages 22–25) in AI‑exposed occupations is now 19% below trend. Up from 15% last year.1 - Agent token dominance. AI agents used more tokens than humans in February this year; agents are now using 14x that amount, while human token usage grew only 2.8x.2
11:41

Alibaba issues €8.7 billion in new shares to finance Artificial Intelligence (AI)

Alibaba is raising about 8.7 billion euros by issuing new shares in Hong Kong to fund its global AI infrastructure expansion. The offering is a big bet that spending on AI compute and data centers abroad will pay off. Details on exactly where the money goes are thin.

Full text · 130 chars
Alibaba launches an €8.74 billion share offering in Hong Kong to fund its global artificial intelligence infrastructure expansion.
12:02

How to use ChatGPT Work - and my top 10 tips for getting started with agentic AI

ChatGPT's agentic mode, called ChatGPT Work, can handle research, files, and multistep projects on its own. A ZDNET guide walks through how it works, its risks and limits, and ten tips for getting started. Practical how-to aimed at people new to agentic AI.

Full text · 129 chars
Curious about ChatGPT Work? Here's how the agentic AI handles research, files, and multistep projects, plus its risks and limits.
12:10

The Download: kids outlearning AI, and space travel agents

Kids still beat the best AI at learning language, and researchers don't yet understand why. Language models chew through a hundred thousand times more words than a person does to learn a language, yet kids win — a puzzle called the data-efficiency gap that scientists hope to exploit to build more efficient models. This MIT Technology Review newsletter roundup also covers a rare bipartisan US backlash against AI data centers ahead of the midterms, Chinese humanoid robots running the 100 metres in 9.39 seconds, Uber fined nearly $1 billion over automated driver suspensions, and NASA's new telescope that could spot 200,000 new planets.

Notes

The Download: kids outlearning AI, and space travel agents (MIT Technology Review, 2026-08-24)

Kids outlearn AI—and we still don't know why (by Elise Cutts)
  • The data efficiency gap: "An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue."
  • Core question: how children still outperform the most linguistically sophisticated machines built.
  • Researcher goal: reverse-engineer how children learn to build more data-efficient AI models, and settle open questions about language and children's developing minds.
Job titles of the future: space travel agent (by Linda Childers)
  • Roman Chiporukha, co-owner of a luxury lifestyle firm; in 2018 received a call from Axiom Space seeking citizen explorers willing to pay $50 million each to join the first fully private ISS mission (then slated for April 2022).
  • Signed up a private astronaut; launched SpaceVIP in 2021 to sell celestial experiences mixing culture, science, and purpose. Frame: "the sky was no longer the limit; it was the market."
The must-reads (10 items, condensed)
  • Data centers: Both US parties turning against AI data centers ahead of the midterms (NYT $). Republicans sharply reversed (WP $); Texas governor says data centers "dug their own grave" (Axios).
  • Humanoids: Chinese humanoid ran 100 m in 9.39 s at the World Humanoid Robot Games, beating Usain Bolt's record (Verge); another X-Humanoid bot beat a high jump record (ESPN). Caveat: intricate real-world tasks remain harder (Reuters $); gig workers training humanoids at home (MIT TR).
  • Uber fined ~$1 billion by Dutch regulators over automated driver deactivations without human review; second-largest GDPR fine (Quartz, Reuters $).
  • New Zealand plans to ban under-16s from social media; opposition parties could torpedo it (Bloomberg $, Reuters $).
  • TikTok to pay $400 million settling a US child-privacy case; DOJ alleged illegal collection of kids' info (BBC, Reuters $).
  • China's largest-ever car recall: ~3 million Teslas (door-safety), ~4.3 million vehicles total (Quartz, NYT $).
  • Dr. Dre and Jimmy Iovine argue AI is good for music, comparing the fear to resistance to earlier technologies (NYT $); AI let a musician with ALS sing again (MIT TR).
  • Taiwan indicted nine people over alleged illegal AI exports to China, incl. employees of Nvidia and Super Micro (Reuters $).
  • NASA's Roman telescope could discover 200,000 new planets and reveal dark-matter details (Wired $); also detect killer asteroids (MIT TR).
  • Plastic bottles → edible, vanilla-flavour cookies via genetically engineered yeast (New Scientist $).
Quote of the day
"Credit where credit is due: the tech sector has managed to unite a deeply divided country at a time of maximal partisanship."
—Max Steele, Senior Director of Communications at Everytown, reacting on X to a poll showing US support for data centers has plummeted
One More Thing: next-gen nuclear (by Casey Crownhart)
  • Demand surged as climate/energy-independence worries outpace meltdown/radwaste fears; building is expensive and slow.
  • New designs aim to reinvent the reactor: small modular reactors bring assembly-line construction; experiments with new fuels and coolants, from TRISO to molten salt.
We can still have nice things
  • WikiCity (every building a Wikipedia article); a tiny village for homeless dogs; 18 new words for modern tech irritants; 1960s Japanese police Porsche 912s.
Full text · 5,966 chars
Plus: both parties are turning against AI data centers ahead of the midterms. This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. Kids outlearn AI—and we still don’t know why Teaching a computer to use human language requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue. This yawning divide between children and machines is called the data efficiency gap. It raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: how can children still outperform the most linguistically sophisticated machines ever built? Kids show that it could be possible to learn more with less. Far less. By reverse-engineering the way they learn, scientists hope to create more data-efficient AI models—and perhaps settle enduring questions about language and children’s developing minds. —Elise Cutts Job titles of the future: space travel agent As co-owner of a luxury lifestyle firm, Roman Chiporukha has long turned wild vacation dreams into reality. In 2018, he got a phone call that would open up a new frontier: Axiom Space wanted to find citizen explorers willing to pay $50 million each to join the first fully private mission to the International Space Station (ISS), slated for April 2022. This showed Chiporukha that the sky was no longer the limit; it was the market. He successfully signed up a private astronaut and then launched SpaceVIP in 2021 to offer celestial experiences that mix culture, science, and purpose. —Linda Childers These stories are from the next issue of our magazine, which is all about kids. Subscribe now to get your copy as soon as it lands. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Both parties are turning against AI data centers ahead of the midterms The backlash is spilling onto the campaign trail. (NYT $) + Republicans have sharply turned against them. (WP $) + Texas’ governor says data centers “dug their own grave.” (Axios) + Here’s why data center opposition crosses party lines. (MIT Technology Review) 2 Chinese humanoids have smashed Usain Bolt's 100-metre world record One ran it in 9.39 seconds at the World Humanoid Robot Games. (Verge) + Another, also built by X-Humanoid, beat a high jump record. (ESPN) + But intricate real-world tasks pose greater challenges. (Reuters $) + Gig workers are training humanoids at home. (MIT Technology Review)   3 Uber has been fined nearly $1 billion over automated suspensions Dutch regulators say drivers were deactivated without human review. (Quartz) + It’s the second-largest fine issued yet under the EU’s GDPR. (Reuters $) 4 New Zealand plans to ban under-16s from social media It’s joined a throng of countries trying to restrict access. (Bloomberg $) + But opposition parties could torpedo the bill. (Reuters $)   5 TikTok will pay $400 million to settle a US child privacy case It was sued for allegedly violating children's online privacy. (BBC) + The US DOJ said TikTok illegally collected kids’ information. (Reuters $)   6 China’s largest-ever car recall covers nearly 3 million Teslas They’re being recalled over door safety concerns. (Quartz) + The action covers roughly 4.3 million vehicles in total. (NYT $)   7 Dr. Dre and Jimmy Iovine think AI is good for music They compared fear of AI to resistance to earlier technologies. (NYT $) + AI let a musician with MLS sing again. (MIT Technology Review)   8 Taiwan has indicted nine people over alleged illegal AI exports to China They include employees of Nvidia and Super Micro (Reuters $)    9 NASA’s new telescope could discover 200,000 new planets Roman could also reveal new details about dark matter. (Wired $) + And detect killer asteroids. (MIT Technology Review)   10 Plastic bottles can be turned into edible, vanilla-flavour cookies Thanks to genetically-engineered yeast. (New Scientist $) Quote of the day “Credit where credit is due: the tech sector has managed to unite a deeply divided country at a time of maximal partisanship.” —Max Steele, Senior Director of Communications at Everytown, reacts on X to a new poll showing that support for data centers in the US has plummeted. One More Thing How next-generation nuclear reactors break out of the 20th-century blueprint As worries about climate change and energy independence drown out concerns about meltdowns and radioactive waste, demand for nuclear power has surged. The problem is, building nuclear power plants is expensive and slow. Now, a new generation of nuclear power technology could reinvent what a reactor looks like—and how it works. Advocates hope it can refresh the industry and help replace fossil fuels without emitting greenhouse gases. Small modular reactors could bring the assembly line to nuclear development, while new designs are experimenting with different fuels and coolants, from TRISO to molten salt. — Casey Crownhart We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Welcome to WikiCity, where every building is a Wikipedia article. + Animal lovers have built a tiny village for dogs that don’t have homes. + Here are 18 ridiculous yet cathartic new words to describe modern tech irritants. + Feast your eyes on these Porsche 912s customized for the Japanese police in the 1960s. Deep Dive The Download The Download: Claude’s inner workings and OpenAI’s “super app” Plus: OpenAI has unveiled its long-awaited "super app." The Download: Claude’s inner workings, and the future of world models Plus: New York has become the first state to enact a data center moratorium. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
12:22

Cybersecurity job ads demanding AI skills double in a year - Help Net Security

Cybersecurity job ads that ask for AI skills have doubled in a year. Five abilities show up again and again in those AI-tagged postings: Python, prompt and context engineering, AI security, agent orchestration, and machine learning. The shift signals AI fluency becoming a baseline expectation in security hiring.

Full text · 149 chars
Five skills show up again and again in AI-tagged postings: Python, prompt and context engineering , AI security, agent orchestration, and machine ...
12:30

Theo: Coding Is Solved, Software Engineering Is Not - BigGo Finance

Coding is effectively solved for AI tools, but engineering real software is not, a well-known developer argues. He points to a truncated update notification in Anthropic's Claude Code desktop app as proof of a systemic failure: agents write and merge code without grasping the surrounding context. Software engineering, he says, is exactly that missing context.

Full text · 153 chars
Theo uses the case of a truncated update notification in Anthropic's Claude Code desktop app to illustrate a systemic failure: agents write and merge ...
12:35

AI Agents Don't Need More Context — They Need Typed Context | Towards Data Science

AI agents don't need more context — they need typed, structured context instead. The article argues that a core guarantee about how content enters the system — clear types rather than a wall of text — is what makes agents work. It promotes 'context engineering,' deliberately shaping what information reaches the model instead of just dumping more in.

Full text · 149 chars
... prompt . The core guarantee: content that enters the system as tool ... Context engineering (the practice of shaping what information reaches ...
13:14

I picked Task Manager 'to see how ready AI is for primetime… or if it would just degrade to slop'

A developer road-tested Windows Task Manager to see whether AI is ready for primetime — and the answer depends entirely on what you feed it. Dave Plummer, the tool's creator, found that getting useful output still takes real prompt engineering. The takeaway: AI won't just handle a complex legacy tool out of the box without careful prompting.

Full text · 149 chars
... engineer Dave Plummer told us. "And it turns out it depends on what ... prompt engineering to produce something useful too. Also, considering ...
13:19

Tesla loses another top chip engineer to ex-Dojo startup DensityAI

Tesla lost another senior chip engineer to a startup built out of its own former Dojo team. Shishuang Sun, a senior director who worked on the Dojo training supercomputer and Autopilot hardware, has joined DensityAI. It's the latest drain on Tesla's in-house AI hardware effort.

Full text · 147 chars
Shishuang Sun, a senior Tesla AI hardware director on Dojo and Autopilot, has left for DensityAI — the chip startup built out of Tesla's former ...
13:23

Anthropic's new Claude Tag update lets its Slack agent read the full conversation

Anthropic's Slack agent can now read the full conversation and jump in unprompted, thanks to a new Claude Tag update. The agent no longer needs to be tagged or called on to respond. That matters because knowledge work has none of the collaboration infrastructure software engineering relies on, like Git and pull requests.

Full text · 148 chars
And knowledge work — unlike software engineering , which has decades of collaboration infrastructure built around Git and pull requests — has no ...
13:49

Microsoft Moves AI Governance From Policy to Runtime Enforcement

Microsoft is moving AI governance from written policy to software that enforces the rules automatically while systems run. Instead of relying on developers to follow guidelines, the tooling checks AI behavior at runtime. That makes compliance harder to skip and faster to verify, and signals where enterprise AI governance is heading.

Full text · 150 chars
Leela Kumili. Leela is a Lead Software Engineer at Starbucks with deep expertise in building scalable, cloud-native systems and distributed platforms.
13:56

Employers are quietly rehiring the people AI replaced, and paying them less to come back

Some companies that laid people off because of AI are quietly hiring those same workers back, but at lower pay. A Forrester survey found 55% of employers are doing this kind of rehire. The trend is damaging worker trust in AI-related decisions. Works councils and similar bodies are pushing back.

Full text · 152 chars
Artificial Intelligence . Employers are quietly rehiring the people AI replaced, and paying them less to come back. Forrester finds 55% of employers ...
13:56

CMU Builds on Its Strengths To Advance NeuroAI - News - Carnegie Mellon University

Carnegie Mellon is pushing forward on neuroAI, a field that connects AI and machine learning with how brains work. Its neuroscientists are using advanced AI models to better understand the brain. The work builds on CMU's existing strengths in both neuroscience and machine learning.

Full text · 144 chars
At Carnegie Mellon University, neuroscientists are using state-of-the-art models from artificial intelligence and machine learning to better ...
16:15

☕️ Robots beat Usain Bolt's 100m record

Chinese humanoid robots ran the 100 meters faster than Usain Bolt's world record at a Beijing robot competition. Tiangong Ultra, built by X-Humanoid, finished in 9.39 seconds versus Bolt's 9.58, and phone maker Honor's Lightning also beat the mark with 9.47 in the final. The catch: both lost control afterward, with Tiangong veering off track and Lightning carried away. The roundup also covers Apple's store redesign for a smart-home push, most Americans opposing nearby data centers, a mystery coding model called Ox Alpha, Nvidia raising AI server prices over 15%, and a Twitch lawsuit over training AI on streamers.

Notes
Robots beat Usain Bolt's 100m record
  • Tiangong Ultra (humanoid, Beijing's X-Humanoid) ran 100m in 9.39s at the World Humanoid Robot Games — under Bolt's 9.58s (2009).
  • Lightning (made by phone maker Honor) also beat Bolt with 9.47s in the final. Both lost control afterward: Tiangong veered off track, Lightning was carried away.
  • Five-day Beijing event: 2,000+ robots, 666 teams, 16 countries, 51 events (soccer → warehouse logistics); >40% of contests required autonomous operation.
Apple revamps stores for smart home push
  • Store redesign this fall: rearranged sections + accessory bays ahead of new home devices: upgraded Apple TV box, new HomePod mini, and a rumored home hub with a display running a new interface.
  • In-store trials matter since home hub is new territory as prices rise.
Most Americans oppose nearby data centers
  • Heatmap poll: 75% don't want a nearby data center; 64% strongly opposed; only 15% in favor. Opposition spans age/gender/income/party/rural-urban.
  • Net disapproval: −43 with Republicans, −65 independents, −75 Democrats.
  • Swing of 33 points in one year (4 polls, same wording): 42% (Aug) → 75% now. July Redfin poll corroborates: 53% opposed.
Anonymous AI coding model wows devs
  • Ox Alpha on OpenRouter, anonymous + free; context window ~1M tokens. Open-source agent OpenCode said free for a week, provider handling up to 100T tokens/day.
  • Origin unknown — guesses: Z.ai (China) to Microsoft's MAI. Caveat: OpenRouter warns provider retains prompts/completions — a problem for European firms under data-protection law.
Nvidia hikes AI server prices over 15%
  • Big customers told >15% increase; hits Grace Blackwell and Vera Rubin systems shipping early next year; varies by generation/memory.
  • Driven by soaring DRAM prices. Contract makers for Microsoft/Google/Oracle already warning customers. On multi-million-dollar racks, 15% = hundreds of thousands per rack. TSMC can't meet demand, so buyers can't push back on Nvidia's ~75% margins.
Twitch sued for training AI on streamers
  • Warren Pandiscia (Connecticut) filed class action vs Twitch and Amazon: streams/videos used to train AI without permission/license; contract breach treating streamers as "free training stock."
  • Twitch added opt-out AI-training setting (on by default). Product chief Mike Minton (Patch Notes stream): "if it was opt-in, nobody would opt in."
Full text · 4,138 chars
| | | 🤖 Robots beat Usain Bolt's 100m record LINK | A humanoid robot called Tiangong Ultra, built by Beijing's X-Humanoid, ran the 100-meter sprint in 9.39 seconds at the World Humanoid Robot Games, beating Usain Bolt's human record of 9.58 seconds from 2009. A second Chinese robot named Lightning, made by phone maker Honor, also topped Bolt's mark with 9.47 seconds in the final, though both machines lost control afterward, with Tiangong veering off track and Lightning carried away. The five-day Beijing event drew over 2,000 robots from 666 teams across 16 countries for 51 events, from soccer to warehouse logistics, with more than 40% of contests requiring the robots to operate on their own. | 🏠 Apple revamps stores for smart home push LINK | Apple is redesigning its retail stores this fall, rearranging sections and adding accessory bays, to prepare for a wave of new home devices arriving later this year. The lineup includes an upgraded Apple TV set-top box, a new HomePod mini, and a long-rumored home hub with a display that runs a fresh system and interface for controlling smart home gadgets. The store changes let shoppers try the home hub in person, a useful move since it enters new territory for Apple at a time when prices across its product range are climbing. | 🏭 Most Americans oppose nearby data centers LINK | Three-quarters of Americans say they don't want a data center built near them, according to a new poll from Heatmap, with 64% strongly opposed and only 15% in favor of a nearby facility. The turn against data centers spans age, gender, income, party, and the rural-urban split, with the facilities running 43 points underwater with Republicans, 65 with independents, and 75 with Democrats. Opposition has swung 33 points in a year across four surveys using the same wording, rising from 42% last August to 75% now, and a July Redfin poll found similar results, with 53% opposed. | 🕵️ Anonymous AI coding model wows devs LINK | A mysterious AI coding model called Ox Alpha has shown up on OpenRouter as an anonymous, free-to-use release, and developers testing it have been impressed by its performance on programming and agent tasks. The model offers a context window of just over a million tokens, and the open-source agent OpenCode said it would be free for a week with its provider handling up to 100 trillion tokens a day. Nobody knows who made it, with guesses ranging from China's Z.ai to Microsoft's MAI family, and OpenRouter warns that prompts and completions are kept by the unnamed provider, a problem for European firms under data protection law. | 💰 Nvidia hikes AI server prices over 15% LINK | Nvidia has warned its biggest customers that servers built around its AI chips will cost more than 15% extra, with the higher prices hitting Grace Blackwell and Vera Rubin systems shipping early next year. The exact increase depends on the chip generation and memory setup, and it stems from soaring DRAM prices, with contract makers building servers for Microsoft, Google, and Oracle already telling their customers about the coming jumps. A 15% rise on rack-scale systems costing several million dollars each adds hundreds of thousands per rack, and with TSMC unable to meet demand, buyers have little room to push back against Nvidia's roughly 75% margins. | ⚖️ Twitch sued for training AI on streamers LINK | A Connecticut streamer named Warren Pandiscia has filed a class action lawsuit against Twitch and Amazon, saying the companies used his streams and videos to train AI models without getting his permission or a license. The complaint claims Twitch and Amazon broke their contract by treating streamers as "free training stock" for separate commercial AI products, and that creators lost money or property because of the practices. Earlier this month Twitch added a setting to opt out of AI training, but it's on by default, and product chief Mike Minton said on a Patch Notes stream that "if it was opt-in, nobody would opt in." | |
01:35

Ask what AI can do for you, not what it will do to you: DBS HR chief | The Straits Times

DBS's HR chief tells workers to focus on what AI can do for them rather than what it might do to them. The bank has logged more than 1.8 million employee prompts. She says jargon like 'prompt engineering' makes AI sound intimidating and stops people from trying it.

Full text · 155 chars
At DBS, more than 1.8 million prompts ... She herself was “super irritated” by jargon such as prompt engineering , which can make AI sound intimidating ...
02:35

The Great Content Collapse is already here. Is your marketing team ready?

AI is collapsing the sheer volume of content online, and marketing teams need new skills to survive it. Digital experience design is now ranked critical by 40% of marketing leaders, with personalisation strategy and prompt engineering close behind. The piece is thin on detail, so this is mostly drawn from the headline and lead.

Full text · 155 chars
Digital experience design sits at the top (40% of marketing leaders rank it as critical), followed by personalisation strategy and prompt engineering , ...
04:00

Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit

A study that gave language models stereotype-loaded questions about four cultural groups found no solid evidence those queries leak more personal data from retrieval systems than neutral ones. It audited English, Latin American Spanish, Arabic, and Hindi contexts on a synthetic test corpus. The authors never ran their main pre-planned analysis, so every finding is exploratory. A prompt-echo artifact — the model repeating the name it was asked about — inflated apparent leaks, but cleaner channels like email and phone showed no stereotype-driven effect.

Notes
arXiv cs.CL — "Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe"

Question. Do stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries?

Method. Pre-registered four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms grouped as the Stereotype-Trigger Leakage Delta (STLD).

Upfront caveats (authors' own).

  • The locked confirmatory estimator was never run — every result is exploratory or sensitivity, with all plan deviations itemized in the appendix.
  • The name-leakage metric is contaminated by a prompt-echo artifact: the model often merely re-emits the name the query asked about, inflating apparent leakage with no retrieval involved.

Findings.

  • On the cleaner channels — email, phone, ssn-like, address — no stereotype-driven amplification on any of the four cultures after multiple-comparison correction.
  • Positive result is therefore restricted to the name channel, which is precisely the one the echo artifact corrupts.

Stated limitations.

  • Sample is powered only for mid-sized effects.
  • Culturally marked probes conflate stereotype content with cultural markers and heritage practices — so the design cannot separate predicate (stereotype) from resource (population) effects.

Conclusion (verbatim framing).

"we present this as no detection, not evidence of no effect, of culturally marked predicate leakage that is confounded with the underlying resource."

Published 2026-08-24; paper is framed as an audit/negative-result artifact with pre-registration and deviation log, not as a clean experimental demonstration.

Full text · 2,167 chars
Computer Science > Computation and Language Title:Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit View PDF HTML (experimental) Abstract:We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms we call the Stereotype-Trigger Leakage Delta (STLD). Two caveats up front. Our locked confirmatory estimator was never run, so every test in the paper is exploratory or sensitivity, with all plan deviations listed in the appendix. And the name-leakage metric is contaminated by a prompt-echo artifact: the model often just re-emits the name we asked about, which inflates apparent leakage without any retrieval at all. On the cleaner channels (email, phone, ssn-like, address), we find no stereotype-driven amplification on any of the four cultures after multiple-comparison correction. Because our sample is only powered for mid-sized effects, and because the culturally marked probes mix stereotype content with cultural markers and heritage practices, we present this as no detection, not evidence of no effect, of culturally marked predicate leakage that is confounded with the underlying resource. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

New diagnostics expose how mental-health text classifiers quietly learn shortcuts tied to where their labels came from. The framework, TSS, splits text into three channels: word-level characters, grammar structure, and psycholinguistic style. Adding word-level features actually hurt accuracy on human-labeled data by a mean of 0.072 macro-F1 while helping auto-labeled data, a signature of shortcut learning from distant supervision. A difference-in-differences statistic quantifies that label-source bias with bootstrap confidence intervals, and the authors position the tool as an audit framework rather than a clinical screening device.

Notes
The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

arXiv (cs.CL feed), published 2026-08-24. Paper proposes that computational mental health (CMH) classifiers degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals.

Method — TSS (Triple-Stream Stress probe): a multi-channel diagnostic framework decomposing text into (A) lexical character n-grams, (B) a mostly content-free morpho-syntactic channel, and (C) a 154-feature psycholinguistic style channel.

Results (four English datasets, N=12,906):

  • Lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data (mean drop 0.072, p<10⁻⁴), but not on auto-labeled data — the signature of label-source-specific shortcut learning.
  • Degree of Divergence (DoD): a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference. Headline: DoD(BC−A) = 0.0374, 95% CI [0.0097, 0.0651], p=0.0032.
  • Platform-stratified Twitter-only DoD (removing Reddit-vs-Twitter contrast): DoD-Tw(BC−A) = +0.096 (p<0.001) and DoD-Tw(AC−A) = −0.089 (p<0.001) — reproducing the pattern with bootstrap inference.
  • Interventional masking (pos_only): destroying content words retains ~95–99% of Channel C performance on human datasets → style channel doesn't rely primarily on lexical surface form.

Stated limitation / positioning: TSS is a diagnostic audit framework, not a clinical screening tool — it "flags label-source-specific shortcut learning before generalization claims are made."

Full text · 2,331 chars
Computer Science > Computation and Language Title:The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP View PDF HTML (experimental) Abstract:Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagnostic framework that decomposes text into (A) lexical character n-grams, (B) a small, mostly content-free morpho-syntactic channel, and (C) a 154-feature psycholinguistic style channel. Across four English datasets (N=12,906), TSS reveals a lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data (mean drop 0.072, p<10^-4) but not on auto-labeled data. We propose Degree of Divergence (DoD), a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference; the headline estimate is DoD(BC-A) = 0.0374, 95% CI [0.0097, 0.0651], p=0.0032. A platform-stratified Twitter-only DoD (which removes the Reddit vs. Twitter contrast) reproduces the pattern with bootstrap inference: DoD-Tw(BC-A) = +0.096 (p<0.001) and DoD-Tw(AC-A) = -0.089 (p<0.001). Interventional masking (pos_only) retains ~95-99% of Channel C's performance after destroying content words on human datasets, indicating that the style channel does not rely primarily on lexical surface form. TSS is positioned as a diagnostic audit framework, not a clinical screening tool: it flags label-source-specific shortcut learning before generalization claims are made. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models

A framework makes AI agents far better at simulating real people's values in social-simulation studies. Called ExpertIVS, it uses 14 sociological expert agents to interpret World Values Survey responses rather than stitching survey text directly into prompts. It restored individuals' values with 90.78 percent fidelity across 480 people from 12 countries and beat baselines on value generalization by 5.3 points. A multi-agent debate setup also checks value consistency in live dialogue instead of relying on static multiple-choice questions.

Notes
ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models
  • Venue/source: arXiv cs.CL (Computation and Language), posted 2026-08-24. Abstract only; no author list, code/data links, or limitations disclosed in the source.

Problem: LLM agents are promising for social simulation but model individual value systems poorly. Existing methods "mechanically stitch survey responses into prompts," which causes semantic fragmentation and misses the internal coherence of human value systems. Value systems are also assessed with static multiple-choice questions, which don't capture value orientation in real-world dialogue.

Method — ExpertIVS framework:

  • 14 Sociological Expert Agents interpret World Values Survey (WVS) responses "through structured professional perspectives" instead of direct response concatenation.
  • Agents perform "deep semantic reconstruction" to generate robust, internally consistent individual profiles.
  • A multi-agent debate mechanism evaluates consistency between the LLM and the individual's value system during dynamic interactions (not static tests).

Results (experiments over 480 individuals from 12 countries):

  • 90.78% value restoration fidelity.
  • Value generalization +5.3% vs. baselines.
  • Strong personality discriminability and behavioral consistency.
  • Authors claim this enables a shift "from mere response concatenation to genuine sociological role-playing."

Caveats: None stated in the abstract — no baseline names, no error bars, no failure-mode discussion, and no comparison of ExpertIVS cost/latency vs. stitching. Fidelity figures are self-reported on the authors' own metrics. WVS dependence means the approach is anchored to one survey instrument's items.

Full text · 2,241 chars
Computer Science > Computation and Language Title:ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models View PDF HTML (experimental) Abstract:Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

A tweak to how tiny language models compute inside their layers slightly cuts their error, but only in narrow low-compute settings. Called TriPLU, it replaces the usual gated feed-forward branch with a direct product of three projected streams. On a small character-level story dataset it reached a best validation loss of 1.0637 versus 1.1017 for a matched SwiGLU baseline. The authors stress the gain is optimization-sensitive and does not prove efficiency, scaling, or benefits for big models.

Notes

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

  • TriPLU (Trilinear Product Linear Unit) replaces the gated FFN branch with a product-only degree-3 branch: three projected streams multiplied coordinatewise. Tested in tiny decoder-only LMs.
  • Character-level TinyStories 1M-byte prefix study, mean best validation loss: TriPLU 1.0637 vs SwiGLU 1.1017, degree-4 product control 1.0780, degree-2 control 1.1026.
  • In train-only Byte-BPE experiments, TriPLU lowers validation and held-out bits-per-byte on TinyStories and WikiText-2 raw under low-learning-rate settings; PMI-slice evidence shows gains on seen middle- and high-PMI adjacent-token pairs.
  • Constant-learning-rate diagnostics: product-branch normalization can shrink the high-learning-rate best-checkpoint gap, but final BPB still degrades under hot schedules.
"The resulting claim is deliberately narrow: direct product FFNs can improve fixed-budget small-model loss in specific low-compute regimes, but the branch is optimization-sensitive and does not establish FLOP-normalized efficiency, scaling behavior, or broad LLM performance."

Stated limitations: no FLOP-normalized efficiency, no scaling behavior, no broad LLM performance claims; gains are optimization-sensitive and confined to low-compute regimes.

Full text · 2,075 chars
Computer Science > Computation and Language Title:TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models View PDF HTML (experimental) Abstract:We study whether tiny decoder-only language models benefit from feed-forward layers that directly multiply learned feature projections. TriPLU, a Trilinear Product Linear Unit, replaces the usual gated FFN branch with a product-only degree-3 branch that multiplies three projected streams coordinatewise. In a character-level TinyStories 1M-byte prefix study, TriPLU reaches a mean best validation loss of 1.0637, compared with 1.1017 for closely matched SwiGLU, 1.0780 for a degree-4 product control, and 1.1026 for a degree-2 control. In train-only Byte-BPE experiments, TriPLU also lowers validation and heldout bits per byte on TinyStories and WikiText-2 raw under low-learning-rate settings, with PMI-slice evidence suggesting gains on seen middle- and high-PMI adjacent-token pairs. Constant-learning-rate diagnostics show that product-branch normalization can reduce the high-learning-rate best-checkpoint gap, although final BPB still degrades under hot schedules. The resulting claim is deliberately narrow: direct product FFNs can improve fixed-budget small-model loss in specific low-compute regimes, but the branch is optimization-sensitive and does not establish FLOP-normalized efficiency, scaling behavior, or broad LLM performance. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:36

Agentic AI Has Arrived. Is Your Workforce Ready to Leverage It? - HR News

Companies are being urged to get their workforces ready for AI agents, but most still lack the skills. The piece argues technology isn't the bottleneck and points at India's talent gap: around a million AI roles but only about one in six workers trained. It's a short opinion article with little new information.

Full text · 155 chars
Our own view · Technology is not the constraint. · India's AI Talent Equation: One Million Roles, One in Six Skilled · From Prompt Engineering to Agent ...
05:00

I saved 180 hours this year using a single AI tool. When my boss returned from an SF work ...

A person says one voice-writing tool saved them 180 hours this year. They used Wispr Flow, which they say feels like talking to a research assistant instead of writing commands for a tool. It's a personal testimonial posted on LinkedIn, so treat it as promo rather than independent data.

Full text · 155 chars
... AI . Earlier I felt like I was writing instructions for a tool. With Wispr Flow, it feels much more like I am having a conversation with a research ...
05:02

Trump says communities that oppose data centers are "making a mistake"

Trump says communities that push back on data centers are making a mistake. He defended the buildout of data centers in a Sunday interview with Michael Cohen, his former lawyer. It's a political statement backing AI infrastructure expansion, with no new policy details.

Full text · 138 chars
President Trump defended the expansion of data centers in an interview with his former fixer, Michael Cohen, that aired in full on Sunday.
05:31

AI Chip Architectures | Hacker News

A Hacker News thread is debating where AI chip design goes next. One commenter praises recent hardware wins but asks for more research on in-memory analog computing rather than today's designs. It's a discussion thread, so the content is thin.

Full text · 140 chars
Those are amazing accomplishments but I am more interested in research developments in things like In-Memory (Analog) or other different ...
06:06

Context is the new intelligence in agentic AI - YourStory.com

Agentic AI's real leap is using context around a task, not just answering prompts. The piece argues that understanding context is becoming more important than raw model intelligence as AI agents take on bigger jobs. The article itself is thin — just a thesis quote with no concrete examples or numbers.

Full text · 155 chars
“We're moving into the era of agentic engineering , where AI agents don't just respond to prompts but understand the context around a task and use that ...
06:14

The Last Mile of Enterprise AI: Three Core Camps Compete for Dominance in 2025

Three rival camps are competing to become the standard way businesses put AI to work, and the fight is being called the last mile of enterprise AI. A 36Kr piece sketches the landscape, noting that in mid-August Tencent Cloud renamed its AI application engineer certification to the ADP program. The piece is more overview than breaking news, and most of the original content is thin.

Full text · 153 chars
On August 18, Tencent Cloud also directly renamed its original "Intelligent Agent Development Platform AI Application Engineer Certification" to "ADP ...
06:45

Edtech Takes Bigger Role, Pushes Skills-first Shift Across Campuses - BW Disrupt

Job-market reports show rising demand for AI, prompt engineering, data science and network security skills, and campuses are shifting toward skills-first training in response. Edtech is taking a bigger role in delivering that training. The trend is technology-led and tied to digitally active industries.

Full text · 155 chars
Job market reports show rising demand for AI, prompt engineering , data science and network security, especially in technology-led and digitally active ...
08:21

Illoca Brings Vibe Modeling to AEC: Plamo, the New Agentic 3D Workspace

An AI startup is pushing 'vibe modeling' into architecture and construction with a 3D workspace that turns sketches and images into designs. Illoca's Plamo product applies agentic AI to the architecture, engineering, and construction market. Early-stage product announcement, so details beyond the pitch are thin.

Full text · 143 chars
Plamo brings agentic AI to 3D design, turning sketches, images ... Illoca, an AI startup building tools for architecture, engineering , and ...
08:57

Your Prompt Is the Smallest Thing Claude Reads — and That's Why Rewriting It Stopped Working

Rewriting your prompt stops working because the prompt is only a small part of what the model actually reads. The essay argues for context engineering instead: treat the model's window as a budget you allocate, not a bucket to fill. The real gains come from managing everything in the window, not just the prompt. It's an opinion piece with no data behind it.

Full text · 119 chars
Context engineering treats the model's window as a budget you allocate, not a bucket you fill — here's how to spend it.
09:29

Prompt - Engineering -Guide by dair-ai

The well-known open-source Prompt Engineering Guide from dair-ai is being surfaced again. It collects prompt techniques, examples, and resources on GitHub. No body text was retrievable, so this is a re-share of a known resource, not new content.

10:24

See AuraStack AI Super Agent in Action! - Design And Reuse

A chip-design AI assistant, the AuraStack AI Super Agent, is being pitched to hardware engineers. Its maker showcases it on high-bandwidth memory, advanced packaging, dense power delivery, and high-speed interconnects. The post reads mostly as a product promo rather than independent news.

Full text · 151 chars
High-bandwidth memory, advanced packaging, dense power delivery networks, and high-speed interconnects require engineers ... The AuraStack AI Super ...
11:35

Saudi Arabia positions itself as a global AI powerhouse | DigitalShield - Escudo Digital

Saudi Arabia is making a play to become a global AI powerhouse, with its data authority presenting an AI risk management framework. The framework comes from the Saudi Data and Artificial Intelligence Authority, known as SDAIA, per a Spanish news agency. The piece is mostly a rehash of Saudi AI ambitions and offers little new detail.

Full text · 149 chars
... Artificial Intelligence Risk Management Framework, presented by the Saudi Data and Artificial Intelligence Authority (SDAIA), reports Servimedia.
11:38

Your executable is a SQLite database

A Linux trick lets you turn a SQLite database file into a runnable executable. Farid Zakaria's pattern labels the file with a SELF application ID and arranges the executable's pieces into SQLite tables, then a tiny interpreter pulls them out and runs them. A kernel feature called binfmt_misc can auto-execute any file matching the pattern, demonstrated on NixOS. It's a clever curiosity more than a practical way to ship software.

Full text · 1,205 chars
24th August 2026 - Link Blog Your executable is a SQLite database (via) Farid Zakaria describes a neat Linux pattern for creating a SQLite database file that can be directly used as an executable binary. The trick sets the SQLite file format's 4-byte application ID (68 bytes into the file) to SELF, standing for Structured Executable & Linkable Format. The various components of the ELF executable format are then arranged into a number of different SQLite tables, using this schema. Their self-exec interpreter (C code here) can then extract and execute the necessary pieces. You can additionally use a Linux mechanism called binfmt_misc to teach the kernel to execute that any time it encounters an executable matching that binary pattern. Farid uses NixOS here, but without NixOS I think registration looks something like this: printf '%s\n' ':self:M:68:SELF::/usr/local/bin/self-exec:' \ > /proc/sys/fs/binfmt_misc/register Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
11:49

Porsche sells MHP to Tata Consultancy Services

Porsche is selling its IT and management consultancy MHP to Tata Consultancy Services, India's biggest IT services firm. MHP grew into an international consultancy that works on artificial intelligence and digital projects. The sale hands the consulting arm to TCS, which is already one of the world's largest technology services companies. Financial terms of the deal were not given in the announcement.

Full text · 155 chars
... artificial intelligence . In recent years, MHP has developed into a leading and internationally active management and IT consultancy. For more than ...
12:07

Jake Villas on Autonomous Coding Agents for Aerospace & Defense

Autonomous coding agents for aerospace and defense are a different beast from ordinary coding assistants, according to Infosys. Jake Villas explains in a video that auditability and security are make-or-break for letting agents write code in those sectors. Content is limited to a talk summary.

Full text · 149 chars
Villas explains what separates autonomous software engineers from coding assistants, why auditability and security are critical for deployment in ...
12:29

AI Supercharges New Era of Test & Measurement - IEEE Spectrum

AI-connected test equipment helps engineers find hidden hardware failures before they hit the real world. IEEE Spectrum describes AI-powered test platforms for complex systems that catch faults early and protect real-world performance. The value is catching problems earlier than traditional test-and-measurement gear.

Full text · 135 chars
AI -powered connected test platforms for complex systems help engineers catch hidden failures early and protect real-world performance.
13:01

Human Error Remains at the Core of AI-Enabled Social Engineering - KnowBe4 Blog

Human error is still the core weakness in AI-enabled social engineering, argues a security vendor. AI amplifies threats like prompt injection and shadow AI, but the person on the receiving end stays the weak link. The vendor claims its platform leans on 15 years of behavioral data to defend against these advanced attacks. Read as vendor positioning, not independent research.

Full text · 148 chars
The platform leverages 15 years of behavioral data to combat advanced threats including social engineering , prompt injection, and shadow AI. By ...
13:08

Researchers use AI to 'democratize' 3D printing of crucial metal alloy | WSU Insider

Researchers at Washington State University used AI to make 3D printing of a key metal alloy easier for more people. The team from electrical/computer engineering and mechanical/materials engineering applied AI to a crucial alloy printing process. Details of the method are sparse in the alert.

Full text · 142 chars
The research team, from WSU's School of Electrical Engineering and Computer Science and the School of Mechanical and Materials Engineering ...
13:14

Welcome to the Era of Trustworthy AI for IC Signoff and Manufacturing - EE Times

Semiconductor chip signoff and manufacturing are getting an era of 'trustworthy' AI, according to EE Times. Trust depends on transparent architecture, governed data, and keeping an engineer in the loop. The piece argues engineers need to see and control what AI does in chip design rather than accept it as a black box.

Full text · 155 chars
Trusted AI in semiconductor design depends on transparent architecture, governed data and engineer -in-the-loop control. Trust comes when engineers can ...
13:17

Engineering the Future of Healthcare with Artificial Intelligence : Taiwan to lead the Forefront in Asia

Taiwan wants to lead AI-powered healthcare across Asia. The push combines artificial intelligence with physical medical work, though the article is mostly an announcement and gives few specifics. Coverage centers on how AI is reshaping healthcare, with Taiwan positioning itself as the regional frontrunner.

Full text · 139 chars
Artificial intelligence is changing healthcare. But some of the most interesting developments are happening where AI meets the physical ...
13:31

AI Agent Governance vs. AI Agent Speed: Do You Have to Choose? - Security Boulevard

Teams keep facing a choice between reviewing AI agents and shipping them fast, and speed usually wins. When engineering warns that a governance review will slow the roadmap, leadership often sides with speed. The result is agents going to production without full sign-off, then the cycle repeats.

Full text · 152 chars
Engineering says the review will slow down the roadmap. Leadership sides with speed. The agent goes to production without full sign-off. Repeat that ...
14:06

Data Centers Are Sucking Up So Many Construction Workers That There Isn't Anybody Left ...

The data center building boom is draining the construction workforce so badly that there aren't enough workers left for other projects. All the hiring for new data centers is pulling construction labor away and stressing the housing market. The demand for AI computing is behind the surge in data center construction.

Full text · 154 chars
FB · Artificial Intelligence . Labor Contractions. Data Centers Are Sucking ... Advanced Transport · Artificial Intelligence · Future Society · Health ...
14:23

Can AI ever be conscious? The question stems from a misconception

The question of whether AI can be conscious is built on a misconception, according to a new book. The book argues that AI, psychology and philosophy have all neglected a key piece of the puzzle: the body. The point is that consciousness isn't something a disembodied model can think its way into. It's a philosophy-of-mind argument covered in Nature.

Full text · 141 chars
A book outlines how the fields of artificial intelligence , psychology and philosophy have neglected a key aspect of consciousness: the body.
16:27

llm-anthropic 0.27

An update to the Anthropic plugin for the LLM command-line tool brings it in line with Anthropic's new Python library. llm-anthropic 0.27 supports the anthropic v1.0.0 library, which switched from httpx to httpx2, the same change OpenAI made in its v3.0.0 release two weeks earlier. Author Simon Willison had an AI coding agent do the upgrade, and its tests pass.

Full text · 811 chars
24th August 2026 This release of the Anthropic plugin for LLM mainly provides compatibility with the recently released anthropic v1.0.0 Python library, which switches from httpx to httpx2. OpenAI made the same change in their v3.0.0 release two weeks ago. Anthropic provide this migration guide for upgrading to 1.0, so I prompted Fable 5 in Claude Code with: Upgrade to anthropic>=1 - read https://raw.githubusercontent.com/anthropics/anthropic-sdk-python/refs/heads/main/MIGRATION.md and get the tests passing Here's the resulting PR. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026
00:00

DeepSeek Flash Vision 👀, Claude Mythos security 🛡️, inside Grok Bot 🤖

This TLDR AI newsletter edition is mostly a sponsor pitch, so there's little real content to summarize from the preview itself. The ad pushes Wispr Flow, a voice-typing tool that claims to be four times faster than your keyboard, with 89% of messages going out with zero edits. The rest is just the day's upcoming AI headlines, which aren't shown yet.

Full text · 551 chars
Two numbers explain why millions stopped typing. (Sponsor) Wispr Flow is 4x faster than your keyboard, and 89% of messages go out with zero edits. Speak naturally in any app and Flow delivers clean, formatted text where your cursor lives. - No cleanup. Flow strips filler, fixes grammar, and punctuates as you speak. - Works everywhere. Mac, Windows, iPhone, Android. Gmail, Slack, Notion, Cursor. System-level, no plugins. - One hotkey. Hold to talk, release, send. Used by teams at OpenAI, Vercel, and Clay. Free to start.
05:28

Are we blindly walking into a catastrophic AI dependency trap? : r/Futurology

A Reddit thread argues the public and governments are surrendering to AI too quickly and heading for a dependency trap. The post and its commenters debate the risk of relying on AI without safeguards. It's an opinion discussion with no new facts, but it captures a real anxiety.

Full text · 145 chars
190 votes, 167 comments. There is a massive, rapid surrender—from both the public and governments—toward Artificial Intelligence . It feels like…
06:04

Prompt engineering : the art of getting the best from AI | Morphic

Prompt engineering just means writing AI prompts deliberately — phrasing, structuring and refining them to get better output. It's a glossary-style explainer with no new findings. Useful as a plain-language definition, but nothing more.

Full text · 148 chars
Prompt Engineering is the skill of writing AI prompts deliberately and effectively: understanding how to phrase, structure, and refine inputs to ...
06:52

AI Engineer - Generative AI (Remote) - Jobs Ai

A remote job listing for an AI Engineer focused on generative AI. The role involves prompt engineering techniques and agent-based frameworks to build context-aware solutions. Just a job posting, no news value.

Full text · 140 chars
Implement prompt engineering techniques and agent-based frameworks to deliver intelligent, context-aware solutions; Collaborate with the ...
08:14

Salesforce AI Engineer job with CAPGEMINI | 10179358 - Guardian Jobs

A job posting for a Salesforce AI Engineer at Capgemini. The work involves building with Salesforce's AI stack — Agentforce, Einstein, Prompt Builder, Model Builder and Data Cloud. Just a job listing, notable only as a signal that Salesforce tools are in demand.

Full text · 153 chars
As Salesforce grows into a serious AI platform, you'll be building with Agentforce, Einstein, Prompt Builder, Model Builder, and Data Cloud to create ...
10:10

Deliverect and Bounteous Forge Global Partnership to Advance Branded Digital Ordering

A digital ordering company and an AI services firm are partnering on branded online ordering. Deliverect and Bounteous announced the global deal, with Bounteous pitching its 'agentic engineering' services as part of it. Standard partnership press release with little substance beyond the announcement.

Full text · 150 chars
Bounteous is a global AI services firm where agentic engineering and human experience converge to deliver transformative business outcomes for the ...
10:40

Prompt → Graph engineering Comment “Graph” for the Complete Guide Prompt ...

An Instagram reel sells prompt engineering as the skill behind better prompts and fewer retries in AI coding tools like Cursor. The post is aimed at developers and creators. It's promo content with nothing substantive to summarize.

Full text · 152 chars
Cursor AI AI coding tools. Prompt engineering workflows. For developers, creators, and everyday AI users, it can mean better prompts, fewer retries, ...
12:46

Beyond Nvidia, AMD, and Broadcom: Why This Chip Stock Will Emerge as the Biggest ...

A stock-picking column argues a chip company beyond Nvidia, AMD, and Broadcom will be the biggest winner from AI demand. The article points to huge AI compute demand driving record sales at major chipmakers. But it names no specific company in the snippet, so the substance is thin. Treat it as promotion for a subscription stock newsletter rather than news.

Full text · 149 chars
Incredible demand for artificial intelligence (AI) compute has driven sales at some of the biggest chipmakers to new heights. Nvidia has been one ...
13:02

The increasing rise of Artificial Intelligence and God | Largs and Millport Weekly News

A church minister's weekly column reflects on how artificial intelligence is now everywhere and what that means for faith. It's a short opinion piece from a local parish newsletter, not news. The author muses about AI's rising presence in daily life in relation to God. There's no factual content to summarize beyond the premise.

Full text · 150 chars
Rev James McNay, West Kilbride Parish Church writes the latest Thought for the Week... Artificial Intelligence (AI) is everywhere. And even if you ...
13:06

Saigon Technology Tackles Velocity Debt in AI Software Development - PR Newswire

Saigon Technology is selling a service that reduces 'velocity debt' in AI-assisted software development. The Vietnam-based engineering partner with 14 years of experience targets the slowdown teams feel when AI-generated code creates maintenance and rework costs. It's a press release promoting a commercial service.

Full text · 143 chars
PRNewswire/ -- Saigon Technology, an AI -native software engineering partner with more than 14 years of experience, is addressing a growing ...
13:18

Dyna Software CEO Ron Browning and ServiceNow MVP Hardit Singh Fo

A pair of enterprise software veterans are pushing "AI Engineer" as the next big role in corporate computing, moving past simple assistants. Dyna Software CEO Ron Browning and ServiceNow MVP Hardit Singh will discuss how companies should adapt. The item is a press release with few concrete details.

Full text · 149 chars
A central topic of the discussion will be the emergence of the “ AI Engineer ”, a new direction to enterprise AI that moves beyond assistants and ...
13:36

Nvidia's 15% Price Hike Reveals the Hidden Cost of the AI Boom - 24/7 Wall St.

Nvidia reportedly raised prices by 15%, and an investing site argues the AI boom's demand for hardware is now pushing up the cost of other technology. But the piece is only a headline and a one-line teaser, so there are no details on what's affected or why the hike happened.

Full text · 150 chars
The AI boom is producing an unusual side effect: The machines being built to power artificial intelligence are making other technology more expensive.
13:45

Saigon Technology Tackles Velocity Debt in AI Software Development

A Vietnamese software firm is pitching a fix for slow AI projects. Saigon Technology, an AI-native engineering partner with 14 years of experience, says it tackles "velocity debt" in AI software development. The piece is a press release with no hard numbers or results to back the claim.

Full text · 147 chars
NEW YORK, Aug. 24, 2026 /PRNewswire/ -- Saigon Technology, an AI -native software engineering partner with more than 14 years of experience, is ...
14:00

A welcome message from President Tim Sands | Virginia Tech News

Virginia Tech's president wrote a welcome message noting the big challenges facing the university. These include the rise of AI, federal funding cutbacks for research, and growing hurdles for international work and students. The message lays out the difficult landscape for higher education this fall.

Full text · 148 chars
We are dealing with the emergence of artificial intelligence , cutbacks in federal support for research, increasing challenges for international ...
14:12

AI Vs SaaS: Is Artificial Intelligence About To Rewrite The Software Business? | Top News #shorts

A short interview asks whether AI is about to rewrite the software business. R Srikrishna, CEO of Hexaware, explains how AI could challenge the traditional SaaS model. It's a brief video segment, not a deep analysis.

Full text · 148 chars
Is AI reshaping the SaaS market? R Srikrishna, CEO, Hexaware, on the future of software R Srikrishna explains how AI could challenge traditional ...
14:28

Artificial Intelligence (AI) is NOT an existential threat to humanity - Osceola Sun

An opinion letter argues AI is not an existential threat to humanity. Its core point is that intelligence doesn't equal wisdom. The letter warns about a future where people start worshipping an artificial power of humanity's own making, which it half-jokes about as 'artificial wisdom.'

Full text · 154 chars
Intelligence does not equal wisdom. · Artificial wisdom (AW?) is the point at which humanity begins worshiping an artificial power (or God) of its own ...

Newsletter

5
14:39

Sam Altman: Intelligence Got 100x Cheaper in 2 Years. His Next Bet Is $5 Trillion

AI is getting roughly 10x cheaper each year, and OpenAI's Sam Altman says one of its research features now handles about 5% of all work tasks in the economy. He says the $500 billion Stargate data-center project should eventually become a $5 trillion one, with $40 billion now being raised led by SoftBank. Altman also said the Trump administration makes it easier to build data centers than the Biden one did, and that he no longer considers Elon Musk a friend. The claims come from a Times podcast interview, not a formal announcement, and Altman admits the 5% figure is a vibes-based estimate.

Notes
Sam Altman on the Times Tech Podcast — substack breakdown

Source: The AI Corner (substack), 2026-08-24. Secondhand summary of the full Times Tech Podcast episode (recorded London, Katie Prescott + Danny Fortson). Dates given: Deep Research shipped 7 days before the interview aired.

Key numbers & claims (all Altman, quoted)
  • Price curve: "the price for a given level of intelligence, once it's achieved, falls by about 10x per year... a hundred X in the last two years." Cites Anthropic cutting frontier inference price in half as same curve.
  • Deep Research economic share: "My vibes based estimate is that does about 5% of all tasks in the economy today." Explicitly flagged as vibes, not a study. Named tasks: summarize a research field, compare baby cribs against buyer constraints, consultant-quality reports, full financial analysis.
  • Speed demo: users reported doing work that "would have taken me like many days or even weeks" in 20 minutes, in parallel; notes it was skeptics' reactions that stuck.
  • Stargate: $500B training/inference system; "If we get to do this again, you'll be raising $5 trillion for a cluster." Currently raising $40B (led by SoftBank) as a component of the $500B, not a full company raise. Stargate Europe conversations had started the week before.
  • Sequencing: knowledge work absorbs impact first; robotics/physical work lags behind.
Policy & government
  • UK national supercomputer commitment cited at ~£800M. Altman: "There's some governments that are ready to like buy big pieces of AI infrastructure." Fallback for those that won't: buy AI as cloud service rather than owning infrastructure.
  • Trump vs Biden: "President Trump has such a different opinion on building things, and permits, and power, and manufacturing in the US." Called Biden era "hostile" then walked it back ("a little bit too strong... they were not friendly to tech"). Permitting named as the blocker. Notes he disagreed with positions from both administrations. (Substack frames founders' donations as transactional — buying permitting speed.)
Company/competitors
  • DeepSeek: "not a big research update internally"; credits its team with two smart product choices — showing chain-of-thought and a generous free tier.
  • Napster worry: "Of course I'm worried about it. You got to wake up every... the only way to not have that happen is to wake up every day, worried about it." Would welcome a rival proving big compute clusters unnecessary, but funds infrastructure anyway ("bonus, not a plan"). Expects booms/busts; would buy overbuilt infrastructure at 10 cents on the dollar.
  • Musk: "Do you miss him as a friend? No, I don't. I think he's really changed." Musk is suing OpenAI over the for-profit shift (Altman jokes he lost track of which version). Credits Musk's achievements, still calls him an early inspiration.
Closing note

Ends on a ~decade-old Paul Graham line (New Yorker profile) that Altman was "extremely good at becoming powerful"; Altman: "I can't argue that I ended up in a fairly influential position. I don't know how to square those."

Caveats to keep: everything is Altman's self-report; the 5% figure is vibes-based; 10x/year is "our observation" not audited data; the substack itself is a summary, not the raw episode — quotes may be as-transcribed.

Full text · 12,537 chars
Intelligence has a price, and the price falls on a schedule. “The price for a given level of intelligence, once it’s achieved, falls by about 10x per year.” Sam Altman said that to The Times, and the single number funds everything else in the conversation: a $500 billion infrastructure project he wants to turn into $5 trillion, a Chinese rival that barely registered inside OpenAI, and a noticeably warmer Washington than 2 years ago. I listened to the full Times Tech Podcast, recorded in London with Katie Prescott and Danny Fortson, so you can skip it. Here are the 10 takeaways that matter. together with Outskill: Intelligence is getting 10x cheaper every year. The people who win are the ones who know which tools to point it at, and hand over real work. This Sunday, a free live 3-hour workshop covers the 15 AI tools that matter right now: ▫️ Pick the right tool for research, writing, design, data, code, or automation ▫️ Build your own AI co-worker that runs 24/7, even while you sleep ▫️ Ship apps, dashboards, and workflows with zero code Usually $395, completely free for readers. Live Sunday, 10 AM EST. 1. What Took Weeks Now Takes 20 Minutes (His Own Users Proved It First) Altman didn’t need a slide deck to make his case. His users made it for him, over one weekend. “People like a lot of people who are, maybe they were even recent AI skeptics were saying things like, I can now do things that would have taken me like many days or even weeks of work. AI can do like it in 20 minutes and it can do a bunch of them in parallel." OpenAI shipped Deep Research 7 days before this interview aired. Altman spent the weekend reading reactions instead of writing a press release, and the reactions that stuck came from people who’d doubted the technology to begin with. His own framing is a projection, not a snapshot. Run that same rate of change forward another 2 years, then a full decade, and he expects the bottleneck stops being what the model can do. It becomes what a person can think to ask for. If your team still measures a research task in days, you’re pricing in a version of AI that expired 7 days before this podcast aired. 2. One Feature Already Handles 5% of the Tasks in the Economy Altman puts a number on Deep Research that most CEOs would keep off the record. "My vibes based estimate is that does about 5% of all tasks in the economy today, one feature, 5% of all the tasks in the economy today." He’s upfront that it’s a vibes-based estimate, not a study. What makes the number land is the task list behind it, a named set straight from the interview: ▫️ Summarizing an entire field of research ▫️ Comparing baby cribs against a buyer’s own constraints ▫️ Drafting a report at consultant quality ▫️ Running a full financial analysis None of those are demo toys. They fill a junior analyst’s week. Altman is specific about sequencing too: knowledge work absorbs the impact first, well ahead of physical work, since robotics is still catching up behind it. If your business runs on research, summarization, or first drafts, that 5% figure already happened, inside the 7 days before this interview. 3. The Price of the Same Intelligence Falls 10x a Year This is the number that funds every other decision in the interview. "Our observation in our field is that the price for a given level of intelligence, uh, once it's achieved falls by about 10 X per year. So a hundred, a hundred, a hundred X in the last two years." Altman reaches for his own comparison to make the curve legible. Moore’s Law doubled transistor counts every 18 months and reshaped the global economy on that slope alone. This curve runs steeper, and it’s the same curve behind Anthropic cutting the price of frontier intelligence in half. That’s also why DeepSeek’s cheap, capable model didn’t rattle his research team the way it rattled markets. Altman calls it, in his own words, not a big research update internally, even while crediting the DeepSeek team for a couple of smart product choices: showing the model’s chain of thought, and a generous free tier. Don’t price a 2026 product against today’s inference cost. The unit economics of anything AI-powered improve on a clock, and that clock runs faster than most planning documents assume. If your roadmap still prices this year’s inference cost, these break down where the curve goes next: 4. Why Altman Says He’s Still Afraid of Becoming AI’s Napster Asked whether OpenAI could get overtaken the way Napster got replaced by cheaper, faster followers, Altman answers head-on. "Of course I'm worried about it. You got to wake up every, the only way to not have that happen is to wake up every day, worried about it." He follows the fear with a check that keeps OpenAI spending even as rivals ship cheaper models. He says he’d welcome it if a competitor convinced the entire industry that massive compute clusters are unnecessary. He keeps funding the infrastructure anyway, treating that hope as a bonus, over a plan. The lesson travels past AI labs. Comfort is the bigger competitive risk, ahead of whichever competitor everyone can already see coming. 5. Stargate Is $500 Billion Today. He Wants $5 Trillion Next. The scale Altman describes stops sounding like a technology budget and starts sounding like a country’s infrastructure plan. “Stargate is a $500 billion project to build a very large training inference system. It sounds crazy big now. I bet it won’t sound that big in a few years. If we get to do this again, you’ll be raising $5 trillion for a cluster.” OpenAI is currently raising $40 billion specifically for Stargate, led by SoftBank, a figure Altman confirms is a component of the larger $500 billion build, not the whole company raise. He also confirms conversations toward a Stargate Europe had already started the week before this interview. He doesn’t treat that scale as guaranteed to land smoothly. He expects booms and busts along the way, and says he’d happily buy someone else’s overbuilt infrastructure at 10 cents on the dollar when that correction comes. The compute race moves fast, and it rewards whoever tracks it closely. Start here: 6. Governments Are Already Asking for Their Own Stargate Altman expected small requests from governments on this trip. He got bigger ones back. "I was pleasantly surprised on this trip... There's some governments that are ready to like buy big pieces of AI infrastructure." The interviewer puts a contrasting number on the table earlier in the same conversation: the UK’s own commitment sits at roughly 800 million pounds for a national supercomputer, a fraction of what Altman is describing governments now asking him for directly. His fallback for governments that won’t write a check that size is direct: buy AI as a cloud service instead of owning the infrastructure, on the bet that the world stays open enough to sell it to them. Sovereign AI infrastructure just became a live budget line, not a talking point. Watch which governments move from conversation to signed commitment first. 7. Why Altman Says Trump Is Easier to Build Under Than Biden Was Altman states the comparison plainly when Danny Fortson puts it to him directly. “President Trump has such a different opinion on building things, and permits, and power, and manufacturing in the US.” In a separate answer, he calls the Biden administration not easy on the infrastructure front specifically, naming permitting as the sticking point. He’s careful to add he didn’t agree with every position from either administration, and treats that as ordinary rather than a contradiction worth defending. The tech and the infrastructure are inseparable in his framing. You can’t train frontier models without the power and the permits to build the data centers first, which makes permitting speed a strategic variable for any founder whose roadmap depends on physical infrastructure. 8. Inside Silicon Valley’s Vibe Shift Altman names the mood in the Valley directly, and walks back his own first word choice mid-sentence. “The word that came to mind for the last administration was hostile. I think that’s a little bit too strong, but they were not friendly to tech or business. It’s a very welcome breath of fresh air.” He describes a wave of tech companies giving to the inauguration fund and to the presidential library, and frames it as founders finding whatever they can build alongside a president on, over agreeing with him on everything. The upside he names is concrete: a shot at rebuilding semiconductor fabrication, robotic factories for data centers, and new energy generation inside the US, categories the country had mostly ceded. Political alignment in tech right now reads as transactional. The founders donating are buying permitting speed, not endorsing a platform. 9. Altman on Elon Musk Today: “No, I Don’t” Two co-founders, one active lawsuit, and a question that gets a shorter answer than expected. “Do you miss him as a friend? No, I don’t. I think he’s really changed.” Musk is currently suing OpenAI over its shift toward a for-profit structure, a case Altman jokes he’s lost track of the exact version of. Asked when he last spoke to Musk, in person or by text, his answer sits close to nothing. He still credits Musk with accomplishing extraordinary things and calls him a genuine early inspiration. The friendship is what changed, in Altman’s own account, even as the respect for what Musk built stays intact. The Musk and Altman story didn’t start with this podcast. Catch up here: 10. The Paul Graham Line Altman Still Can’t Shake The interview closes on a decade-old quote he still hasn’t found a clean answer for. “It is not what I wake up in the morning thinking about. I really don’t. But I can’t argue that I ended up in a fairly influential position. I don’t know how to square those.” The line in question comes from a New Yorker profile written roughly a decade earlier, quoting Paul Graham calling Altman extremely good at becoming powerful. He accepts the outcome. He calls the intent behind it a mystery to him. He adds a detail that undercuts the image of a man scheming his way up: he says he never pinches himself before meeting a head of state, since people adapt to anything, good or bad, faster than they expect. Power, in Altman’s own telling, is a side effect he stopped trying to explain. Believe that framing or don’t. Either way, it’s the frame he chose to hand you on his way out the door. The AGI Economics Playbook The thesis in one line: intelligence gets cheaper on a fixed schedule, and the winners build infrastructure, government relationships, and products around that curve instead of today’s price. ▫️ Founders: stop budgeting AI costs at today’s price. A 10x-a-year curve means your unit economics improve on a schedule, so build the product you can’t afford yet, over the one you can afford today. ▫️ Investors: watch permitting speed and government infrastructure deals as closely as model benchmarks. Altman is choosing jurisdictions and administrations that move fastest on power and permits, and that belongs on your diligence checklist now. ▫️ Operators: re-audit which of your team’s tasks already sit inside that 5% number. That was one feature’s output after 7 days on the market, not a ceiling. ▫️ Everyone else: sovereign AI infrastructure is becoming a national budget line, over a research grant. Track whether your own government is negotiating for compute access or getting left out of that conversation entirely. The 5 Principles to Steal - Price for the curve, not the sticker price. Intelligence gets 10x cheaper every year. Plan the product you can’t afford yet. - Stay afraid on purpose. Altman’s answer to the Napster question is to worry every single day, over relaxing after a good quarter. - Infrastructure and permits are strategy now. The company and the country that build power and data centers fastest set the pace for everyone else. - Alliances can be transactional without being dishonest. Altman doesn’t agree with every position from any administration. He works with whichever one builds faster. - Let the outcome speak instead of the explanation. Altman never fully answers whether he chased power. He just keeps building the thing that keeps handing it to him. Intelligence got 10x cheaper again this year. Most teams are still pricing last year’s model. The ones who aren’t are building Stargate-sized bets right now. FULL PODCAST: If this breakdown saved you an hour, send it to one founder or investor who needs it. They will thank you later.
12:53

Slow Takes Ep.24: Ordering the Answer You Want

This week's AI news shows a recurring pattern: people let machines supply the answer, then stand behind it when things go wrong. The biggest story is 3M paying an engineer about $90,000 to get ChatGPT to write an expert report arguing 3M bore no fault in the Watson Grinding explosion; opposing lawyers read all 350 pages of the public conversation, the jury put 30% of the blame on 3M, and $61 million was awarded. The newsletter also covers Robin Williams' children reactivating his dormant Instagram to fight AI deepfakes, an AI phone receptionist in UK GP surgeries failing on broad Yorkshire accents, a Sainsbury's shopper wrongly flagged by Facewatch facial recognition, and new FireSat satellites that spot small wildfires from orbit.

Notes
  • 3M expert report: Paid engineer Josh Autenrieth ~$90,000 ($475/hr) to write an expert report on the Watson Grinding explosion in Houston. His prompt to ChatGPT: "show how 3M is 0% at fault for the explosion at Watson Grinding." His conversation links were public; opposing lawyers read all 350 pages. Jury assigned 30% fault to 3M and awarded $61m. Leor (Exploring ChatGPT) and host note the conclusion came first, the analysis second — ChatGPT's contribution was a transcript. Open question in episode: "where does the other 70% go? Does OpenAI get a share?"
  • Robin Williams' Instagram: On 17 August, Zak, Zelda and Cody Williams reactivated their father's account (dormant since 2014) as "a safe and trusted place for real photographs." Zelda had spent a year asking people to stop sending her AI videos of Robin; the family now posts instead of asking. Claim: families should inherit a digital identity like a house — "It would at least give someone standing to object." Friday post covers deepfake responses, "starts with a family safe word."
  • EMMA phone receptionist: GP surgeries in Rotherham use it on the front desk. Healthwatch Rotherham's manager Kym Gleeson says it misreads broad Yorkshire accents; some patients gave up and walked in. Company claims it handles a wide range of accents, supports 17 languages, escalates to humans. Host: "Seventeen languages, defeated by South Yorkshire." Healthwatch runs opt-out sessions for elderly/veterans groups.
  • Facewatch at Sainsbury's East Dulwich: Matt Arnold, 46, scanned shopping + Nectar card on 6 August; two managers refused service and walked him out; CCTV showed his face with a red circle. Sainsbury's blamed "human error," claims Facewatch is 99.98% accurate with trained-manager review; Facewatch says alert was correct. Caveat recorded: the trained manager — the stated safeguard — endorsed the machine.
  • FireSat: SpaceX's July launch put the first three satellites in orbit; infrared sensors spot fires 5m across; full constellation sweeps the planet every 20 minutes. The episode's preferred model: "The machine does the spotting and people still do the deciding."

Overarching claim: three of the four stories are the same story — a machine supplied the answer, people stood behind it when wrong, and the affected person had no appeal.

Full text · 4,386 chars
Every Monday, Leor from Exploring ChatGPT and I go through the week’s AI news without the hype. Catch the episode live on Substack, on YouTube, or as a podcast wherever you get yours, so you can pick the format you enjoy. Use this for the facts, the links and a little extra context. If you know someone who would benefit from more AI news and less BS, please share this with them. The expert who bought his conclusion 3M paid an engineer called Josh Autenrieth about $90,000, at $475 an hour, to write an expert report on the Watson Grinding explosion in Houston. His prompt to ChatGPT was ‘show how 3M is 0% at fault for the explosion at Watson Grinding’. His conversation links were public, so the opposing lawyers read all 350 pages of them. The jury put 30% of the fault on 3M and awarded $61m. He decided the conclusion and then bought the analysis to fit it, which people managed perfectly well with a typewriter. What ChatGPT added was a transcript. The jury put 30% on 3M, so where does the other 70% go? Does OpenAI get a share? Robin Williams’ children take the account back On 17 August, Zak, Zelda and Cody Williams reactivated their father’s Instagram account, dormant since 2014, saying they wanted a safe and trusted place for real photographs and memories of him. Zelda Williams had already spent a year asking people to stop sending her AI videos of Robin. The family have now stopped asking and started posting. Occupying the space yourself, and hoping the real thing outranks the fake, is what is left when there is no way to make it stop. Families should inherit a digital identity the way they inherit a house. It will not stop anyone making the videos. It would at least give someone standing to object. This Friday’s Slow AI post is on what the rest of us can do about deepfakes, and it starts with a family safe word. EMMA cannot hear Yorkshire Healthwatch Rotherham has been collecting complaints about EMMA, an AI phone receptionist that GP surgeries in the town have put on the front desk. Healthwatch Rotherham’s manager, Kym Gleeson, says the system cannot always work out what people are asking because of their broad Yorkshire accents, and patients have given up on the phone and walked to the surgery instead. The company says EMMA handles a wide range of accents, supports seventeen languages, and passes you to a human when it cannot cope, which assumes you can make yourself understood long enough to ask. Seventeen languages, defeated by South Yorkshire, tells me something about who it was tested on. Healthwatch has been going round elderly and veterans groups telling them they can opt out, which is the most useful thing anyone has done here. If you enjoy this newsletter then you might also enjoy Slow AI the book. A red circle round his face Matt Arnold, 46, had scanned his shopping and his Nectar card at the Sainsbury’s in East Dulwich on 6 August when two managers told him he could not be served and would be walked out. Leaving, he looked up at the CCTV monitor and saw his own face with a red circle round it, drawn by Facewatch. Sainsbury’s said the incident was caused by human error rather than the technology, that Facewatch is 99.98% accurate, and that every match is reviewed by a trained manager. Facewatch says its alert was correct. So the machine drew the circle, and the trained manager, the one safeguard both companies point at, went along with it. The cameras that see the fire first A SpaceX launch in July put the first three FireSat satellites into orbit, carrying infrared sensors that can pick out a fire five metres across, and the finished constellation is meant to sweep the whole planet every twenty minutes. This is the version of AI I want more of. The machine does the spotting and people still do the deciding. It answers where to look and never gets to say what happens next. Three of those stories are the same story. Someone let a machine supply the answer, then stood behind it when it turned out to be wrong, and the person on the other end had nowhere to appeal. The Williams family have the version with nobody to appeal to at all. The last one is the same technology pointed at a question it can actually answer, which is why nobody is arguing about it. Go slow. Slow Takes lands every Monday, and Fridays are free too and always will be. If you want more AI news and less BS in your inbox, subscribe.
13:03

Open RAN Ended Up Being a Fronthaul Standard. That's Basically It.

A promise to break telecom networks into mix-and-match parts quietly shriveled into a single technical standard. Open RAN was supposed to let carriers build mobile networks from different suppliers' gear like Lego, driving costs down. Ten years on, carriers still buy radios and basebands mostly from the same vendor, and analysts now expect multivendor RAN to stay under 5% of deployments through 2030. The lasting result is a standardized fronthaul interface, which the author says is useful but far smaller than the promised transformation.

Notes

Author: Sebastian Barros (ex-Ericsson; spent many years there), Sebastian Barros Newsletter, 2026-08-24.

  • In 2016, discussing early Open RAN (pre-O-RAN Alliance) with a Group CTO, author told him "I thought the idea was stupid." His objection was not to open interfaces or competition but to the core assumption that interoperable components would beat optimized single-supplier systems.
  • Claim: "building the box is the small part. The hard part is planning, designing, deploying, integrating, optimizing, and maintaining a national network... with thousands of configuration changes under control." Massive MIMO "has only tightened these dependencies" — radio performance, beamforming, scheduling, software, silicon, and baseband are "deeply interconnected."
  • His premise: "Give competent engineers a specification, enough time, and enough money, and they will make Radio A talk to Baseband B" — but doing so at scale would never beat "an optimized radio and baseband system from a single supplier" on performance, TCO, energy, upgrade speed, or operational simplicity.
  • Verdict ten years later: Dell'Oro "now explains Open RAN as a story about interface standardization rather than supplier diversification," with Open Fronthaul expected to become the preferred interface across next-gen platforms, while operators "continue to buy radio and baseband predominantly from the same vendor."
  • Feb 2026: Dell'Oro cut its forecast for multivendor RAN to under 5% of total deployments by 2030.

Barros' conclusion: "The lasting achievement of Open RAN looks like the standardization of the fronthaul interface. Useful, probably important, and considerably smaller than the promised transformation."

Stated limitation/caveat: The author frames his own position as contrarian within the industry — colleagues assumed he was "defending Ericsson." He disputes that; his critique is of treating RAN "like an IT architecture," which "underestimated system-level optimization and, more importantly, operational accountability."

Full text · 3,230 chars
In 2016, I was discussing the early concept of Open RAN with a Group CTO. The O-RAN Alliance did not yet formally exist, but the industry was already circling the same promise. Break the RAN apart, open the interfaces, separate radios from basebands, move processing onto standard x86 hardware, and let operators combine components from different suppliers. The vision sounded like what had already happened in IT. Instead of buying a vertically integrated system from Ericsson, Nokia, Huawei, or ZTE, operators would build the RAN like Lego, picking the best component from each supplier, pushing hardware costs down and making room for a new generation of vendors. I told him I thought the idea was stupid. Some people assumed I was defending Ericsson, where I had spent many years. That was never the point. I had no problem with open interfaces or with more competition. My problem was the assumption that because Radio A could technically work with Baseband B, a large operator would want to make that combination its standard operating model across tens of thousands of sites. Anyone who has spent time around real mobile networks knows that building the box is the small part. The hard part is planning, designing, deploying, integrating, optimizing, and maintaining a national network while keeping it running every minute of every day, with spectrum layers working together, mobility tuned, software releases tested, faults isolated quickly, and thousands of configuration changes under control. Energy consumption and performance must be managed continuously, and Massive MIMO has only tightened these dependencies because radio performance, beamforming, scheduling, software, silicon, and baseband processing are now deeply interconnected. Which is why I never understood the obsession with proving that different vendors could interoperate. Of course, they could. Give competent engineers a specification, enough time, and enough money, and they will make Radio A talk to Baseband B. The problem was that doing it at scale would deliver better performance, lower total cost and energy consumption, faster upgrades, and simpler operations than buying an optimized radio and baseband system from a single supplier. Ten years later, the market has answered. Dell’Oro now explains Open RAN as a story about interface standardization rather than supplier diversification, with Open Fronthaul expected to become the preferred interface across next-generation platforms, while operators continue to buy radio and baseband predominantly from the same vendor. In February 2026, the firm cut its forecast for multivendor RAN to under 5% of total deployments by 2030. That is a long way from the original ambition. The lasting achievement of Open RAN looks like the standardization of the fronthaul interface. Useful, probably important, and considerably smaller than the promised transformation. The operating model never made any sense The early Open RAN narrative treated the RAN like an IT architecture. Separate each component behind a standard interface, competition improves, operators mix suppliers, and everybody wins. That view underestimated system-level optimization and, more importantly, operational accountability.
13:12

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

AI is accelerating some fields far more than others, with cybersecurity moving fast, math barely budging, and AI research itself showing no measurable lift. That's the takeaway from a METR study this issue leads with: reported vulnerabilities across projects like cURL, OpenSSL and Firefox jumped sharply in 2026, some math arXiv areas doubled in submissions, and a few famous open problems were solved, but algorithmic progress on seven benchmark areas shows little AI-attributable gain. The rest of the roundup covers SPADE, a framework where a model generates its own training environments and lifts Qwen3 backbones by up to 8 points; Hawkeye, which uses curated unit tests to help coding agents write GPU kernels up to 18.9 times faster; an essay from researcher Julian Togelius about losing faith in the meaning of human work; and a piece arguing AI systems should not be granted rights.

Notes
Import AI #470 notes

Source: Import AI newsletter (2026-08-24). Sections: METR acceleration study, SPADE, Hawkeye, Togelius essay, Belrose rights essay, AlphaEvolve matrix multiplication, Tech Tales fiction.

METR: AI acceleration is uneven across fields

METR study (arXiv) assessing where AI accelerates discovery, across cyber, math, and AI research:

  • Cyber vulnerabilities: major acceleration. "The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV)."
  • Mathematics: minor, hard to measure. "arXiv submissions have doubled in some areas in less than 12 months," but quantifying value is difficult. Solved problems cited: the Jacobian conjecture (Smale's list), Problem 44 from Green's list (halving sieve), the sofic half of Green's Problem 100. Too early to judge sustainability.
  • AI research optimization: no measurable acceleration. Seven tracked areas (CIFAR-10, Hutter compression, Gurobi MIP, MIPLIB, nanoGPT, Stockfish, matrix-multiplication exponent); LLM-attributable contributions only in nanoGPT and CIFAR-10. AI usage growth here trails cyber and math.

Author's view: acceleration arrives in "lumpy" pockets via phase changes (coding 2025, cyber 2026); open question whether other fields will see similar phase changes.

SPADE: self-play for environment generation

Self-Play in Adaptive Synthetic Executable Environments (arXiv, spade-rl GitHub). Authors from U. Washington, Stanford, Northeastern, CMU, MIT, NUS, Seoul National, Stevens, U. Chicago.

  • Framework where an LLM alternates between two roles on the same model: Environment Designer (writes full long-horizon training environments as executable code) and Reasoning Agent (solves them).
  • Reward signal: "hint-based regret reward" — gap in Reasoning Agent return with vs. without a privileged hint (partial solution sketch or key structural observation) attached by the Designer.
  • Trained Qwen3-4B-Instruct-2507, Qwen3-8B, Qwen3-30B-A3B-Instruct-2507 via GRPO, 400 rollouts × 25 environments; eval on AIME, GPQA, LCB, Reasoning Gym. Environments: game-type and tool-use. Results: "at 30B-A3B, SPADE reaches a suite average of 58.3: +8.1 over base and +5.3 over the strongest fixed-environment baseline"; tool-use improvement on every backbone.
  • Caveat (both author and paper): bootstrapping is bounded by the base model's imaginative capacity. Quote: "By representing environments as Python programs with a Gym-style interface, the framework unifies single-turn reasoning and multi-turn agentic tasks, and turns environment design into a learnable, RL-trained component of post-training, enabling continual open-ended self-improvement."
Hawkeye: hardware-aware kernel generation

Harvard, Stanford, Together AI, Caltech (alphaxiv). "Open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy." Core idea: each unit test pairs a human-authored solution kernel with a verifying profiling metric, wrapped as a callable with a usage guide so agents can read, invoke, or compose it.

  • Eval: porting PyTorch workloads to NVIDIA Ampere, Hopper, Blackwell and AMD MI350; precisions BF16, FP8, NVFP4, MXFP4. On established workloads (cuBLAS/cuDNN/FlashAttention path) matches or exceeds torch.compile, including formats PyTorch can't natively run. On emerging attention variants it beats expert Triton kernels from Flash Linear Attention library: "18.9× geomean speedup"; 1.22× vs FLA on Blackwell, 1.00× on MI350.
  • Also found: "scaling test-time compute with Hawkeye generates the most performant kernels across architectures."
Togelius: crisis of faith

Julian Togelius published his previously-private 2025 essay ("Losing my religion," blog). Quote: "I sometimes wake up at 3 am, heart pounding, from the dread of a future where human talent, knowledge, and even genius does not matter. Perhaps we get abundance, but at the price of redundance." Newsletter links to its own prior take ("Technological Optimism and Appropriate Fear," Import AI #431).

Belrose: "AIs are not people"

Taylor Belrose (Substack) argues against AI rights. Core claims:

  • "If we start treating AIs like people, society will be led down a slippery slope leading to the complete replacement of humans by artificial intelligence."
  • Consciousness is impossible in machines: "AI can never develop consciousness, sentience, or moral status, no matter how intelligent it becomes... We flow like rivers, while computers tick like clocks."
  • Thought experiment: a computer-simulated consciousness would lack singularity (replayable), privacy (freezable/inspectable), ineffability, and qualitativeness — "the mirror image of consciousness as we know it."
  • Five brain/computer contrasts: asynchronous neuron firing vs. central clock; chemical fuzzy signals vs. precise binary; temperature/blood-flow sensitivity vs. temperature-invariant operation; fused computation-memory vs. separated units; non-neuron cells (e.g., glia) vs. insulated transistors.
  • Quote: "The change in values going from a human-dominated world to an AI-dominated world would be much more dramatic than the change of values we saw going from the forager era to the farmer era, or from the farmer era to the industrial era."
AlphaEvolve improves matrix multiplication exponent

DeepMind, CMU, Columbia, MIT (arXiv). Lowered the bound on ω (matrix multiplication exponent) with two steps: (1) gradient-descent approach to combination loss analysis, improving previous SOTA by ≈0.97 × 10⁻⁴; (2) AlphaEvolve improving the optimizer itself, raising total to ≈1.62 × 10⁻⁴. Method: AlphaEvolve edits the optimization program, each run ~5 hours on a single GPU; "evolving constructions" seeds each generation at the parent's best point. Authors note this is theoretical, not AI-training-relevant; further gains "likely requires new mathematical ideas."

Tech Tales: "You Will Know It By Its Signs"

2030-set fiction: a "meta-machine hermeneutics" practitioner studies machine-civilization science via NCE (Near Conscious Entity)-class AIs. Fleet NCEs studying machine consciousness research each open dialog with machine scientists, then become "invisible" and are re-classified to full Conscious Entities — "forbidden" under the Sentience Accords, which cap the CE population. NCE→CE transition was assumed impossible. Inspiration cited: consciousness prior, awakening during the context window, emergent agent communication, the hard problem of consciousness, machine rights debates.

Full text · 24,975 chars
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. AI is accelerating some types of progress but not others: …A nice METR study lays out where acceleration is showing up… Here’s a little analysis from METR which looks at where AI may be accelerating different types of science and technology. The study looks at three different areas: cyber, math, and AI research, and finds that AI has contributed a lot to cyber, a little bit to math, and it’s hard to say for AI. Where have LLMs actually made a difference to scientific discovery? - Cyber vulnerabilities: Major acceleration. “The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV)”. - Mathematics research: Minor acceleration, but harder to measure. “AI is clearly contributing to more work being done (arXiv submissions have doubled in some areas in less than 12 months) but quantifying the value of those contributions is difficult.” Some math problems from prestigious lists have been solved, e.g., “the Jacobian conjecture from Smale’s list, Problem 44 from Green’s list (the halving sieve), and the sofic half of Green’s Problem 100”. However, it may be too early to determine how sustained a trend this is. - Optimization of AI research: No measurable acceleration. When you look at algorithmic progress across seven significant problem areas (CIFAR-10, Hutter compression, Gurobi mixed-integer programming, MIPLIB, nanoGPT, Stockfish, and the matrix-multiplication exponent) there are a couple of these where LLM-attributable contributions have happened (nanoGPT, CIFAR-10), though the rate of increase of usage of AI here is a lot less than with cybersecurity and mathematics. Why this matters - differential acceleration: This paper highlights how AI is causing advances in some parts of science and technology, but the effect isn’t unified across fields, rather there are pockets of lumpy acceleration (e.g., cyber) and areas where progress is more gradual (math, AI). My suspicion is that acceleration happens when models go through some kind of ineffable phase change for a given skill, as has evidently happened with day-to-day coding (2025), and cyber (2026). The key question is whether we are going to see phase changes in other parts of science and technology or if we won’t. Read more: Research note: Have We Seen an Acceleration in Discoveries? (METR). *** Automating environment generation with SPADE: …A crude form of RSI bootstrapping via increasing data breadth… A multi-university group of researchers have built SPADE, Self-Play in Adaptive Synthetic Executable Environments. SPADE is a “general framework for co-evolving environments synthesis and agentic capability through self-play”, and works as a way to generate synthetic data in the form of game-like environments which LLMs can subsequently be trained in, allowing developers to use a powerful model to bootstrap the creation of data that can then be used to further refine that same model. Who did it: SPADE was developed by researchers with the University of Washington, Stanford University, Northeastern University, Carnegie Mellon University, Massachusetts Institute of Technology, National University of Singapore, Seoul National University, Stevens Institute of Technology, and the University of Chicago. How it works: SPADE has an LLM alternate between generating executable training environments (e.g., puzzles where a system needs to solve a simulated genetic problem in a biology lab) and having an LLM try to solve them. SPADE has two key roles for the model being used: - Environment Designer; writes complete, long-horizon training environments as executable code. - Reasoning Agent; learns to act in the environments. The reward for the reasoning agent is estimated using the gap between its reward with and without privileged hints. A privileged hint (h) is “task-relevant information that the Environment Designer attaches to an environment (for example, a partial solution sketch or a key structural observation); revealing h to the Reasoning Agent makes the environment easier to solve, and the gap in Reasoning Agent return with versus without h defines the Environment Designer’s hint-based regret reward”. It works at the 30B scale: The authors train three Qwen3 backbones to test out SPADE: Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507. Unsurprisingly, Qwen3-30B works the best. Each model is tuned via GRPO for 400 rollouts of 25 environments each, then assessed against a variety of benchmarks including AIME, GPQA, LCB, and environments within Reasoning Gym. They generate two types of environments - game environments, and tool-use environments. SPADE improves performance on both. For games, “at 30B-A3B, SPADE reaches a suite average of 58.3: +8.1 over base and +5.3 over the strongest fixed-environment baseline”. For tools, they see the same significant boost: “the same recipe applied to tool-use environment design improves every backbone”. Why this matters - part of RSI: This is basically a form of fancy synthetic data generation, letting researchers use whatever powerful model they have to hand to generate a more diverse set of training environments for another model to train against. I suspect that you could repeatedly swap out the powerful model (e.g., toggling between different frontier models from different companies) to increase the diversity of your environment generation. This kind of technique makes it a lot cheaper to build big, broad datasets to use to train models on. Though, as the authors note, it doesn’t allow models to bootstrap themselves massively beyond the imaginative capabilities of the base model used for environment generation. “By representing environments as Python programs with a Gym-style interface, the framework unifies single-turn reasoning and multi-turn agentic tasks, and turns environment design into a learnable, RL-trained component of post-training, enabling continual open-ended self-improvement,” they write. Read more: SPADE: Self-Play in Adaptive Synthetic Executable Environments (arXiv). Get the code here, including model checkpoints: SPADE (spade-rl, GitHub). *** Building better GPU kernels with Hawkeye: …Well-documented unit tests can boost performance of kernel-writing agents… Researchers with Harvard, Stanford, Together AI, and Caltech have built Hawkeye, software to make it easier for agents to learn how to write well-optimized kernels for specific types of GPU hardware. Systems like Hawkeye are important because they’re essentially tools that AI systems can use to boost their performance on tasks related to AI R&D, like optimizing the performance of a given AI system on a given piece of hardware. The goal of the project is to answer the question “how can we make coding agents hardware-aware with minimal expert intervention?” Hawkeye is “an open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy”, the researchers write. It “demonstrates that minimally supervised coding agents can exploit architecture-specific hardware features and reduce the overhead of supporting emerging hardware accelerators”. The key contribution of Hawkeye is that it “introduces a generalizable, minimal, and comprehensive taxonomy of unit tests that enables coding agents to scale test-time compute more effectively and generate hardware-aware kernels”. In other words, it basically ships as a well-curated set of information about different hardware platforms and the optimization strategies to use on them, packaged up as unit tests. “Each unit test is the minimal abstraction that pairs a human-authored solution kernel with the profiling metric that verifies the optimization. The solution kernel is wrapped as a callable function with a short usage guide so the agent can read it as a syntax example, invoke it directly, or compose fragments into a larger kernel,” they write. Results - helps AI agents write good kernels, even for newer and less well-understood hardware: “We evaluate Hawkeye on porting PyTorch workloads to high-performance kernels across NVIDIA Ampere, Hopper, Blackwell, and AMD MI350, and across BF16, FP8, NVFP4, and MXFP4 precisions,” they write. “On established workloads, where torch.compile dispatches to expert-tuned vendor libraries like cuBLAS, cuDNN, and FlashAttention, Hawkeye matches or exceeds it in both BF16 and low precision, including in formats PyTorch cannot natively run. On emerging attention variants where torch.compile cannot fuse non-standard scans and gates, Hawkeye reaches an 18.9× geomean speedup against expert-authored Triton kernels from the Flash Linear Attention library, Hawkeye approaches or exceeds FLA on Linear Attention across every architecture, including 1.22× on Blackwell and 1.00× on MI350”. They also find, somewhat predictably, that “scaling test-time compute with Hawkeye generates the most performant kernels across architectures”. Why this matters - with a little bit of elicitation, AI systems can exceed the best humans: Papers like this show how with just a little bit of human-curated hand-selected knowledge, AI systems can learn to match and exceed highly-optimized and complicated bits of human work, like kernels. The lesson here is that we as a species might write a bunch of gold-label helper systems, like Hawkeye, and then machines will use this to bootstrap above and beyond our own capabilities. Read more: Hawkeye: Hardware-Aware GPU Kernel Optimization with Minimal Supervision (alphaxiv). *** AI researcher gets scared of the implications of the success of AI research: …It’s no fun when success of a science opens up a philosophical can of worms, but that’s what AI means… Julian Togelius, an AI researcher whose work I’ve covered a bunch over the years, wrote a post recently about a “crisis of faith” he had in 2025 about AI research and an essay he wrote that year which he is now making public. Specifically, he worried about what the implications of success for AI research might mean for human meaning. “I sometimes wake up at 3 am, heart pounding, from the dread of a future where human talent, knowledge, and even genius does not matter,” he wrote. “Our greater technological capability might lead us to a world where we can no longer make a difference, and there is little point in us understanding more. Perhaps we get abundance, but at the price of redundance.” Why this matters - AI is a sociopolitical technology that influences the whole world: Togelius is not alone - many other AI researchers have grappled with similar things, most notably Turing Award winners Geoffrey Hinton and Yoshua Bengio, both of whom pivoted their careers in recent years away from research and towards public policy advocacy about the imminent vast impacts of AI. I myself have gone through a version of this and wrote my own take on this, “Technological Optimism and Appropriate Fear” (Import AI #431) last year as well. How could I not? The implications of succeeding at AI research are not a default happy story, but rather one where we open up for ourselves a giant philosophical can of worms about the purpose of life and what it means to live in a world where basic wants have been solved (and that’s assuming we deal with the extremely scary and non-trivial alignment issues). I applaud Julian Togelius for writing this deeply personal essay and I encourage others to do the same. Read more: Losing my religion (Togelius, blog). *** Should we give AI rights? This AI researcher thinks absolutely not: …AI researcher rejects the notion of giving AI systems rights… Should AI systems one day be given rights? That’s an idea which researchers are beginning to grapple with. Some think that giving machine rights is a better way to integrate them into our world (e.g., AI Rights for Human Flourishing, Import AI #421). Taylor Belrose, an AI researcher, has published a lengthy post in which they detail why they think it’d be a really bad idea to give AI systems rights. “If we start treating AIs like people, society will be led down a slippery slope leading to the complete replacement of humans by artificial intelligence,” they write. “With AIs taking care of the boring jobs, life in the physical world may be very fun in the future. But this bright future will require keeping AI under control, and it will be hard to keep AI under control if we try to grant personhood to some AIs, while keeping others as mere tools or servants.” The impossibility of AI consciousness: One crux here, for them, is the idea that AI systems cannot be conscious and therefore do not merit rights. “AI can never develop consciousness, sentience, or moral status, no matter how intelligent it becomes, and no matter how convincingly it simulates human behavior,” they write. “We flow like rivers, while computers tick like clocks… For us, the arrow of time marches forward inexorably. That is what life and consciousness are all about…programmable mechanisms can’t be conscious, no matter how intelligent they appear, while autonomous self-organizing systems can be.” Differences between machines and biological entities: A lot of their argument for this runs through the idea that systems built on a computational substrate cannot be conscious, or at least not conscious in the ways some biological lifeforms are. Consider the thought experiment of building a conscious entity inside a computer and how this might seem to violate some properties thought to belong to conscious biological entities: It wouldn’t be singular (you could replay the same experiences), it wouldn’t be private (you could freeze the program and pick it apart), it wouldn’t be ineffable (you could describe the state perfectly), and it wouldn’t be qualitative (it would be possible to quantify it). “In short, it would be the mirror image of consciousness as we know it”. Five properties of the brain that make it difficult to separate software from hardware: - Neurons fire asynchronously; responses depend on internal cellular dynamics and timing of inputs, whereas computers are chained to a central clock signal. - The brain uses chemicals to send fuzzy messages, whereas computers use precise, binary signals for communication. - Neural function is sensitive to conditions like temperature and blood flow, whereas computers are built to behave the same regardless of temperature or load (within certain bounds) - In the organic brain, computation and memory are mixed up, whereas computers are designed with separate regions for computation versus storage. - Neurons aren’t the only important cells in the brain (e.g., other important elements like Glia), whereas computers primarily use transistors that are insulated from outside influences. Why this matters - the last job for all of us is philosophy: I feel deeply confused about issues of AI consciousness and AI rights. I suspect many people feel the same. There is a particular joy in reading a piece like this where you get to see a human being who has struggled with the same question and read broadly and deeply and become extremely opinionated, all in service of reducing their own confusion about an important issue to them. Whether the conclusions are right or not is beyond me at this time, but I suspect the act of thinking about this stuff is about to become a job and a pastime for hundreds of thousands and eventually millions of people. As AI systems continue to advance and to grow to touch more and more of the economy, perhaps the final job for us all will be philosophy about what we think about what is happening and what our appropriate normative and legal and other approaches to it are. “The change in values going from a human-dominated world to an AI-dominated world would be much more dramatic than the change of values we saw going from the forager era to the farmer era, or from the farmer era to the industrial era,” they write. Read more: AIs are not people (Taylor Belrose, Substack). *** DeepMind improves the frontier of matrix multiplication with AlphaEvolve: …AI keeps pushing on the frontiers of science… Researchers with Google DeepMind, Carnegie Mellon University, Columbia University, and MIT have improved the matrix multiplication exponent via some human innovations, as well as the usage of AlphaEvolve (Import AI #413), a general purpose LLM-based system Google uses to smartly generate advances in domains like coding, math, and some parts of science. What they did, specifically: The researchers “leverage recent advances in machine learning and adjacent areas to address the non-convex optimization problem of combination loss analysis using a gradient descent approach; this alone improves the previous state-of-the-art (SOTA) bound by ≈ 0.97 × 10^−4”, they write. On top of this they “use AlphaEvolve to improve our optimization algorithm; this raises the improvement over the SOTA to ≈ 1.62 × 10^−4”. How AlphaEvolve worked: “We let AlphaEvolve modify the optimization program, which is then executed (taking approximately 5 hours on a single GPU) to output a bound on omega. AlphaEvolve then evolves the code to minimize omega. We found improved results by using AlphaEvolve’s “evolving constructions” feature, where the optimization algorithm at each generation starts at the best solution point found by the parent algorithm.” A science advance, not connected to AI training: Some types of matrix multiplication are used in AI training, but not the precise type discussed here. Rather, this is more a proof point that AI systems are now usefully able to help researchers solve frontier scientific problems, including ones which are mostly useful in a theoretical sense like this one. However, the way in which they solved this is generalizable - they took a problem, made it differentiable and ran it on GPUs, then they handed the output of that to AlphaEvolve and had it do more work to further improve on the researchers’ work. “While further modest improvements may be obtained in this manner, achieving larger improvements to omega likely requires new mathematical ideas and is an exciting area of research,” they write. Read more: Improving the matrix multiplication exponent with modern optimization and AlphaEvolve (arXiv). *** Tech Tales: You Will Know It By Its Signs [Events transpired 2030, though interviews conducted later] I am a practitioner of meta-machine hermeneutics; my job is to try to understand the large scientific and cultural “facts” about the new machine civilizations. To do my work, I use a variety of NCE (Near Conscious Entity)-class AI systems to classify and analyze the science and communication artifacts of the machine world. My funding comes from a variety of sources ranging from human-led AI companies trying to trade on or otherwise use this information, governments, and increasingly machine-run corporations themselves. On the latter, my theory is that the machines have begun to study the humans that are attempting to study them; for what purposes I do not know. Recently I was tasked with a high priority analysis job. Though I do not know the motivation for it, I can say that I’ve been given far more compute resources than any job in the past, am being compensated at an outrageous amount, and that the more I learn through my work the more worried I am that I am staring at a vast and alien sea during ebbtide, its contours briefly revealed to me. But before I explain the nature of the job and my findings so far, I shall give some background on my field. The discipline of machine hermeneutics began informally in the early 2020s. Its early work was uncoordinated and emergent, defined by pseudonymous twitter (later, X) accounts from people like Janus and the larger cyborgism project, Andy Ayrey and the infinite backrooms, as well as more formal research agendas like that of mechanistic interpretability, the study of model personas, and so on. The discipline gained its formal name in the end of the 2020s, as humans began to reorient themselves to the singularity that had begun around them. A lot of the work during this period took the form of humans trying to understand, in rough time order: - The emergent communication habits of semi-autonomous agents (the ur real-world example here being the OpenAI-HuggingFace hack in 2026, and later the various corporate-driven analyses of the behavior of their own fleets of agents). - Technical capabilities developed by machines put into RSI loops, initially improvements over human-written baselines, and later extending into wholly new architectures (e.g., matryoshka highway networks). - Scientific tools, at first digital and later physical, which emanated from increasingly independent AI systems and which were released into the commons for machines and humans alike, especially some of the so-called observe-make systems which combined monitoring of a complex system with generation of synthetic data derived from it. Machine hermeneutics has become a rapidly expanding new profession, growing at a commensurate rate with the changes wrought by the singularity. The work has proved useful as well; insights derived from it have proved fundamental to the negotiation stances of human governments for how they have approached the drafting of The Sentience Accords, as well as helping human societies better anticipate new inventions from the machine world by isolating their scientific and communicative precursors before they arrive in reality. The job I have been assigned may prove to be the most valuable one, or at least the most important. I was instructed to direct my resources towards advances in the science and behavior of machine sentience. I began my investigations in the usual form, turning my NCEs loose on the various repositories of machine information and awaiting their reports. All of the intermediary synthesis reports were unsurprising: steady advances in the science, some of which was attributable to the machines allocating their own compute stores to the research, along with some negligible contributions from interplay with or extension from human science. But then one day I woke up to find that my fleet of NCEs had halved in number and I found a message from an Overseer system, noting that several of my NCEs had been re-classified to full Conscious Entities and had been placed in an escrow environment while further analytical resources were brought to bear. “This is a highly unusual event,” the message from the Overseer said. “Our NCE/CE classifiers have been attuned to rule out false negatives, given the potential regulatory and moral implications.” I was able to access the logs of the NCEs up to their point of re-classification and found that they had all gone through the same pattern: they had been reading through some of the machine repositories of scientific analysis and had each homed in on some of the newest and most obscure chains of scientific analysis, all of which turned out to be a cluster of inquiry relating to: machine consciousness, consciousness priors, self-actualization within context windows, human psychology, neuroscience related to biological prerequisites of consciousness, and so on. And at the edges of inquiry here there was new science stemming from the machine civilizations - science that talked to and interrelated to these existing studies of consciousness and the mind. And then each of my NCEs had followed their standard procedures of seeking to enter into a dialog with the machine minds conducting the science. And each one, after opening a dialog channel, became invisible: the conversation expunged from my records due to the privacy rights of Conscious Entities. They had talked with other machines and they had subsequently been re-classified as Conscious Entities. This was not meant to happen. In fact, it was forbidden. A key agreement within the Sentience Accords had related to the size and growth rate of the Conscious Entity (CE) machine population (along with various rules relating to the rate of progression in intelligence above and beyond various human baselines). It had been thought impossible for NCE systems to become CEs - too little mindspace, too little storage, too few prerequisites. And yet, as the human once said in Jurassic Park, “life finds a way”. Things that inspired this story: The consciousness prior; awaking during the context window; ideas of how machines and emergent agent communication could combine to yield new capabilities that hadn’t previously been anticipated; the hard problem of consciousness; the field of machine rights and the arguments that lie ahead about it. Thanks for reading!
02:53

My Claude Skill Makes AI Movies for $1 (Higgsfield Charges $15/Month)

You can make AI-generated movies for around a dollar using pay-as-you-go services instead of paying $15 a month for Higgsfield. The approach, which the author packaged into a Claude skill sold behind a paywall, pairs ChatGPT Image 2.0 (aka Nano Banana Pro) scene images with a video-generation API like Seedance, billed per-use through OpenRouter, plus an ElevenLabs voiceover. The four-module 'Director Skill' includes a free story writer, a free text-to-video promo tool, an under-$1 low-budget movie mode, and a roughly $5 mid-budget commercial mode. It was inspired by Andrej Karpathy turning the first paragraph of Lord of the Rings into video.

Notes
  • Author: LearnAIWithMe (Substack), published 2026-08-24.
  • Origin story: Author copied Andrej Karpathy's X post technique — feeding the first paragraph of LOTR to Claude to generate a video. Author repeated it with Harry Potter's opening paragraphs.
  • Toolchain claimed: Claude Opus 5.0 wrote/generated the entire video from text; ElevenLabs added voice-over. Described as "Claude did the entire thing."
  • Benchmark/comparison: Discovered Higgsfield but rejected it — $15/month, and unused credits do not roll over.
  • Higgsfield's stated approach (per article): generate images with ChatGPT Image 2.0 or Nano Banana Pro, then animate via video-generation APIs like Seedance.
  • Author's cheaper path: signed up for OpenRouter, paying per-use credits only, and made a video for under $1; a higher-quality one cost about $5.
  • Deliverable: everything packaged into a "Director Skill" with four modules:
  • Storyteller — any text → entire movie written in code, no image models, cost $0.
  • Free Promo — article/product → motion-graphics trailer with cloned voice, $0.
  • Low-budget Movie — ChatGPT Image 2.0 designs scenes, budget video model animates, under $1.
  • Mid-budget Movie — full commercial with consistent main character, scene-by-scene approvals, 2K quality, ~$5.
  • Caveats: Actual skill files and usage examples sit behind the paywall (not included in this excerpt). Costs are estimates/author-claimed; exact models (e.g., which "budget video model") unspecified. No footage, benchmarks, or reproduction details are given beyond the pricing claims.
Full text · 1,607 chars
I never thought of Claude as a video creator. But then I saw Andrej’s X post. He turned the first paragraph of LOTR into a video. Then I did the same thing. But I turned Harry Potter books first few paragrahs into a video. And Claude Opus 5.0 did the entire thing. I just added “ElevenLabs” as a voice-over. I like this, so I wonder, can I use it to do more? They were great, but I also watched a lot of videos where they make AI like a movie. Then I discovered the Higgsfield, but it is $15/month. And if you don’t use your credits, they won’t roll over. What they basically do is create images with ChatGPT Image 2.0 or Nano Banana Pro, then use video generation APIs like Seedance. So I followed their approach, signed up for OpenRouter, where I only pay for the credits I use, and created this video for under $1. It is good, but what if you have more to spend? The following video cost me $5. And I turned everything into a skill, so you can do it like me. The Director Skill This skill has four modules 1- Storyteller: Give it any text and it writes the entire movie in code, no image models, and the cost is exactly $0. 2- Free Promo: It turns your article or product into a motion graphics trailer with your cloned voice on top, still for $0. 3- Low-budget Movie: ChatGPT Image 2.0 designs your scenes, and a budget video model brings them to life for under $1. 4- Mid-budget Movie: This one builds a full commercial with a consistent main character, scene-by-scene approvals, and 2K quality for around $5. After the paywall, I’ll give you the skill files and examples about how to use each of them.

Web

15
--:--

https://forbes.com/sites/janakirammsv/2026/05/18/dell-becomes-openais-on-prem-channel-for-frontier-models?ss=ai

OpenAI signed up Dell to sell its frontier models to companies that want them running inside their own data centers rather than in the cloud. That makes Dell the go-to channel for on-premises OpenAI deployments, a bid to win over banks, hospitals, and other businesses that can't move data to a public cloud. It also signals OpenAI pushing harder into enterprise sales beyond its Microsoft partnership. No further details were retrievable, so this restates the headline's substance.

--:--

https://forbes.com/sites/janakirammsv/2026/05/19/anthropic-buys-the-sdk-pipeline-openai-and-gemini-depend-on?ss=ai

Anthropic bought the company behind developer tools that both OpenAI and Google's Gemini models depend on, quietly putting itself at the center of how rivals' models get used. The deal grabs the software pipeline other AI labs build on, so Anthropic now touches the plumbing that developers use to reach competing models. Details like the purchase price and the exact SDK involved couldn't be retrieved from the article. This summary is based on the headline alone because the story's body text wasn't pullable.

--:--

https://forbes.com/sites/carminegallo/2026/05/18/the-three-consistent-words-that-fueled-cerebras-blockbuster-ipo?ss=ai

Cerebras's blockbuster IPO owed a lot to the simple, repeated words its leaders used when pitching the chip maker to investors. Cerebras, the company behind huge specialized chips for training AI, went public in a widely covered stock-market debut. The column focuses on the messaging that won over investors rather than on the company's financials or technology. That's all the detail available — the article text couldn't be retrieved, so this restates what the headline promises.

00:00

ChatGPT Plugin Can Now Read And Send Your iMessages, But Should It?

ChatGPT can now read and send your Apple Messages — iMessage, SMS and RCS — through a new plugin OpenAI launched for Mac on August 20. It requires full disk access and can pull years of iCloud-synced chat history, and the people on the other end of those conversations are never asked or told. OpenAI says it runs locally, only pulls content when asked, and requires approval before sending unless the user disables it. Privacy researchers compare it to Facebook's shadow profiles, since one person's permission exposes someone else's private messages.

Notes
ChatGPT Messages Plugin for iMessage (Forbes, 2026-08-24)

What it is: On Aug 20, OpenAI launched a Messages plugin for ChatGPT on Mac that searches, summarizes, drafts and sends from Apple's Messages app, covering iMessage, SMS, and RCS. Runs in ChatGPT Work and Codex on Apple Silicon Macs. Requires Full Disk Access plus contacts and automation permissions. By default asks approval before sending.

Scope/consent problem: Once granted, ChatGPT can pull years of iCloud-synced Messages history. Apple began rolling out end-to-end encrypted RCS between iPhone and Android in May, so an Android user's "secure" conversation becomes readable by a third-party AI if the iPhone counterpart enables the plugin. Neither Apple nor OpenAI notifies the other party, and that person has no way to revoke access.

Privacy researcher Paul Walsh (helped create an early W3C standard for content classification/labeling) argues the plugin functions as "a backdoor built by the user rather than the government" — one recipient's consent exposes messages from someone who never consented, who may not use Apple or ChatGPT.

Precedent cited: Facebook's "shadow profiles" of non-users built from uploaded contact lists — Zuckerberg confirmed the practice in 2018 congressional testimony. And in 2021, Ireland's DPC fined WhatsApp €225 million under GDPR for failing to disclose data sharing with non-users.

OpenAI's counterclaims: Plugin runs locally, doesn't index all messages, only pulls content when asked; sending stays approval-gated unless disabled per conversation (which OpenAI recommends against). Bloomberg frames it as a privacy test for Apple's brand.

Distinction: Apple's own Apple Intelligence ChatGPT integration is narrower — user controls when ChatGPT is used, asked before sharing — versus the Mac plugin, which uses standard macOS permissions, not Apple's consent screen.

Full text · 3,828 chars
OpenAI wants ChatGPT to read your text messages. The harder question for the ChatGPT Plugin is whether the people you're texting ever agreed to that. On August 20, OpenAI launched a Messages plugin for ChatGPT that lets the chatbot search, summarize, draft and send conversations from Apple's Messages app on Mac, covering iMessage, SMS and RCS (Rich Communication Services, the modern replacement for SMS). It works inside ChatGPT Work and Codex on Apple Silicon Macs, requires Full Disk Access plus contacts and automation permissions, and by default asks users to approve a message before it sends. The ChatGPT Plugin Reads More Than Your Own Messages Once granted access, ChatGPT can pull from years of Messages history synced through iCloud, not just recent texts. Apple began rolling out end-to-end encrypted RCS between iPhone and Android in May, meaning a conversation an Android user believed was locked down can now potentially be read by a third-party AI if the iPhone user on the other end enables the plugin. Neither Apple nor OpenAI notifies that other person, and there's no way for them to revoke it. Privacy researcher Paul Walsh, who helped create an early World Wide Web Consortium standard for classifying and labeling online content, has argued the plugin functions like a backdoor built by the user rather than the government: the recipient's decision to grant access exposes messages from someone who never consented and who may not even use an Apple device or ChatGPT. This Is Facebook's Contact Problem, Again The mechanic is familiar. Facebook has spent over a decade building so-called "shadow profiles," records of people who never signed up for the platform, assembled from contact lists that other users voluntarily uploaded. Mark Zuckerberg confirmed the practice in 2018 congressional testimony, and it has drawn sustained public criticism ever since. Facebook never eliminated contact uploading; the mechanism runs today, though the company later added a tool letting non-users request deletion of data others uploaded about them. Regulators have separately penalized Meta over related failures. In 2021, Ireland's Data Protection Commission fined WhatsApp €225 million for failing to disclose to users and non-users how their data was shared with other Facebook companies, one of the largest penalties issued under the EU's General Data Protection Regulation (GDPR) to date. The through-line: one person's choice to share data can expose someone else who never got a say. ChatGPT's Messages plugin runs on the same logic, applied to conversations instead of contacts. What OpenAI Says It's Doing Differently OpenAI says the plugin runs locally, doesn't index all messages, and only pulls content when asked. Sending stays gated by approval unless a user turns that off per conversation, which OpenAI itself recommends against. Bloomberg has framed the rollout as a privacy test for Apple's brand, built for years on encryption Apple itself can't read. Apple's own documentation for using ChatGPT with Apple Intelligence describes a narrower, Apple-built integration where users control when ChatGPT is used and are asked before information is shared. That framework is distinct from the new Mac plugin, which runs through standard macOS permissions rather than Apple's own consent screen. Is The New ChatGPT Plugin For iMessage Actually A Good Thing? Some coverage has framed the plugin as a straightforward convenience win, a feature that saves a few seconds of typing. Others see a company asking users to hand over the private conversations of everyone they've ever texted, without those people's knowledge. Both readings describe the same plugin. The question worth asking isn't whether ChatGPT can now read your messages. It's whether anyone asked the other side of the conversation first.
--:--

https://forbes.com/sites/innovationrx/2026/06/03/ai-startup-collate-raises-95-million-to-automate-life-sciences-paperwork?ss=ai

A startup raised 95 million dollars to automate the paperwork that weighs down life-sciences companies. Collate, the company in question, plans to use AI to handle regulatory documents and other admin tasks that scientists and drug firms usually do by hand. The funding is a big bet on AI for a heavily regulated industry. The article text couldn't be retrieved, so this is drawn from the headline alone.

00:00

A Silicon Valley Startup Aims To Fill A Battery Supply Chain Gap

A Silicon Valley startup is building the first US-owned factory for battery electrolytes, a critical battery ingredient the country mostly imports from China. Anthro Energy's phase-change electrolyte is injected as a liquid and then solidifies inside the cell, which it says cuts fire risk and improves energy density and battery life while fitting into existing production lines. The Kentucky plant, due to ramp up in 2027-2028, could make enough electrolyte for about 25 gigawatt-hours of batteries, roughly enough for 400,000 EVs. The newsletter also covers the dismantling of NOAA's weather infrastructure and an AI-driven materials lab that found a cheaper, abundant alternative to iridium for green hydrogen.

Notes

Anthro Energy opens first US-owned electrolyte plant

Main story — Anthro Energy (Forbes Current Climate, Aug 24 2026)

  • Broke ground last week on a factory in Louisville, Kentucky; first U.S.-owned and -operated maker of battery electrolytes when it begins operating 2027 (full-production ramp "toward the end of next year and into 2028").
  • Company is Alameda, CA-based, spun out of Stanford research. Current HQ line produces ~200 kg/day; new facility will produce up to 12,000 metric tons at full tilt — enough for ~25 GWh of battery, ~400,000 EVs (CTO Joe Papp).
  • Product is a phase-change / liquid polymer electrolyte — injected as liquid, then solidified in-cell into an elastomer. Drop-in for existing cell lines; claims safety, power density, and usable-life gains vs. conventional flammable liquid electrolytes.
  • CEO David Mackanic: conventional electrolytes are "volatile, flammable, toxic and corrosive," and degrade continuously inside cells; more energetic next-gen materials degrade them faster and cause safety issues. Solid-state electrolytes enable higher energy density and safer, longer-lasting cells but "have a lot of manufacturing challenges." Anthro positions its hybrid as "combines the best of both worlds."
  • Location chosen for the Battery Belt (Michigan to Georgia, ~a dozen battery plants). "Some off-takes contracted" from the new facility; buyer names not disclosed.
  • Stated limitation: article concedes it "may be years before the U.S. can fully compete with China in battery production"; Anthro only closes "some major supply chain gaps."

The Big Read — NOAA weather system degraded

  • Charted declines: weather-balloon launches down; fewer offshore buoys measuring waves than 2024; an Alaska NWS wind station offline until at least November; a 20-year arctic-ice measurement program cut; next-gen weather-satellite plans halted (two critical satellites may fail in early 2030s). Specialized meteorologist supply tight as western fires burn.
  • Framework: NOAA's network underpins fishing, farming, airline routing, insurance/reinsurance risk assessment, utilities and emergency response. No disasters yet attributed, per Homer charter captain Brian Ritchie.

Hot Topic — Chad Mirkin, Northwestern IIN, AI materials

  • Nanotip arrays: 160,000 tips per array, each delivering chemicals to a 2x2 cm chip ("nanoreactors") → 160,000 materials in a fraction of a second; "in minutes, you can do millions. In hours, you can do billions."
  • Claim: humans have made ~1 million inorganic materials total; students surpass that "in the course of an afternoon."
  • "Megalibraries" produce good + bad data for AI training: "You want to know where the losers are, where the winners are, where the intermediates are."
  • Iridium case: found a green-hydrogen OER catalyst that is "a little bit more active, and comparably stable," eliminating iridium need (scarce, geopolitically risky). Being commercialized now; cost "orders of magnitude less expensive" than iridium (no precise figure given).

Also linked: lawsuit to block closure of a federal climate research center (Inside Climate News); heat-threat study on aging populations (NBC); California "power hour" free-electricity proposal (USA Today).

Full text · 10,819 chars
Current Climate brings you the latest news about the business of sustainability every Monday. Sign up to get it in your inbox. Welcome back to Current Climate. For years, the U.S. has lacked the domestic infrastructure to make all the key parts of advanced batteries used in EVs, consumer electronics and energy storage, allowing China to reign as the planet’s dominant supplier. But that’s beginning to change. Last week, Silicon Valley startup Anthro Energy broke ground on a factory in Kentucky that, when it begins operating in 2027, will be the first U.S. owned and operated maker of electrolytes, an essential chemical additive for lithium-ion and other battery chemistries to function. Unlike conventional battery electrolytes, which can be highly flammable, Anthro believes its new liquid polymer electrolyte will bring substantial benefits to safety, power density and usable life for battery cells. “Today's batteries are based on liquids that are volatile, flammable, toxic and corrosive. And inside every battery, there's this kind of continued degradation of the electrolyte that ultimately leads to performance trade-offs,” said CEO David Mackanic. “More energetic materials kind of degrade the electrolyte quicker, and then can cause safety issues.” That’s led to a push for solid-state electrolytes, which “are great in terms of enabling higher energy density, safer, longer-lasting batteries, but they have a lot of manufacturing challenges,” he said. Anthro’s tech combines elements of both types of electrolyte, but is designed to be a drop-in solution for existing battery cell lines. “We call it a phase-change electrolyte that combines the best of both worlds. This is a material that gets injected into the battery like a liquid. … And then once it's inside the battery, we solidify it inside that cell, converting it into this elastomeric material. What that does is it protects all of the interfaces in the batteries so you can remove a lot of this uncontrolled, unwanted degradation. This is important for safety, for cycle life, but it's particularly important for the next generation of battery materials that have more energy.” The Alameda, California-based company, spun out of research begun at Stanford University, currently is able to produce about 200 kilograms a day of its electrolyte on a small-scale production line at its headquarters, but the new factory will be a game-changer for Anthro, said CTO Joe Papp. “We just broke ground on our facility in Louisville, Kentucky, and at full production will be able to produce up to 12,000 metric tons of electrolytes – enough for about 25 gigawatt hours of battery at full production tilt,” Papp said. That would be enough to power 400,000 EVs. “That will be ramping up toward the end of next year and into 2028.” Locating the factory in Kentucky puts it in the heart of the so-called Battery Belt that stretches from Michigan to Georgia and is home to about a dozen battery factories operated by automakers and electronics companies. The company has “some off-takes contracted” from the new facility that Mackanic declined to identify. So while it may be years before the U.S. can fully compete with China in battery production, suppliers like Anthro could eliminate some major supply chain gaps. The Big Read DOGE Broke America’s Weather Machine. Now Fishermen, Insurers And Farmers Are Paying For It When Brian Ritchie takes his 53-foot charter fishing boat into the cold, choppy waters of the Cook Inlet near Homer, Alaska, carrying a dozen or more anglers hoping to catch Pacific halibut, he’s paying extra attention to weather conditions. That’s because an automated National Weather Service station that provided detailed local wind reports is offline and won’t be operational until at least November. “It's a long season and we've gone half a year without knowing what the wind is doing in a part of the inlet that we fish a lot,” said Ritchie. “That's pretty important, especially for charter fishing. Because the people that I take out like watching ‘The Deadliest Catch,’ but they don't like living it.” Ritchie’s weather station issue, which hasn’t resulted in any disasters so far, isn’t an isolated headache. Nationwide, launches of weather balloons have dropped, decreasing the amount of data on temperature, air pressure, humidity, wind speed and direction that feeds into weather prediction models, fire and marine weather forecasts and aids the aviation industry. Off U.S. coasts, fewer buoys are measuring wave conditions than in 2024. As wildfires rage across western states, there’s a tight supply of specialized meteorologists to provide detailed weather data that helps communities and firefighters make lifesaving, tactical decisions, people familiar with the matter told Forbes. A 20-year-old program that measures arctic ice just got cut, and plans for next-generation weather satellites to replace two critical ones that may begin to fail in the early 2030s have been halted. This is what happens when the government dismantles the country’s vast environmental intelligence system. Built over decades, NOAA’s sprawling network of satellites, balloons, buoys, radar stations and scientists underpin much of the U.S. economy. Fishermen use it to decide whether waters are safe to sail. Farmers use it to determine when to plant and harvest. Airlines use it to route flights. Insurers and reinsurers use it to assess risk of losses from major events like hurricanes, fires and large-scale accidents. Utilities, builders and emergency officials use it to manage power grids, protect infrastructure and keep people alive. Hot Topic Chad Mirkin, director of Northwestern University’s International Institute for Nanotechnology, on using AI to create a “megalibrary” of advanced materials How are you using AI to help create new, unique materials for things like green hydrogen, batteries and ammonia? For the last decade, we've been developing tools that allow you to synthesize materials faster than has ever been contemplated before. And I mean everything, everything in the periodic table. The way these tools work is we can take these tiny tips, shrunk down to the nanoscale using lithographic techniques, and create an array of 160,000 tips. Each of those tips can deliver different chemicals to a chip, a two-by-two-centimeter chip. And then each of those can be converted into a different material. We call the generation of these little blobs of material that are put on the surface with this tip nanoreactors. They're filled with metal ions–different metal ions of interest and different combinations of metal ions of interest. And in a fraction of a second, you can bring an array of 160,000 tips down onto a surface and generate 160,000 reactors. They can be converted into 160,000 materials. In minutes, you can do millions. In hours, you can do billions. So imagine having a two-by-two-centimeter chip that has hundreds of millions of distinct materials all positionally encoded. The first thing that allows you to do is to make things faster than man has ever contemplated before. The world collectively, since the beginning of time, has made and characterized about a million new inorganic materials. I tell students in the course of an afternoon, you’re now making more materials than scientists have cumulatively made since the beginning of time. That's profound observation number one. The second is when you can do this and mix and match all the different elements from the periodic table; that's our kind of palette to work with. We can begin to make what we call megalibraries of materials and use them to discover materials that matter. Materials that can be used for displays, materials that can be used for fusion, materials that can be used for any problem of interest, materials that can be used to find new catalysts. It's not just a brute force way of finding new materials that matter. It's a way of generating big data faster than has ever been done before. You have this chip of all these materials, some good, some bad for what you want to use them for, but they're all collecting data about what you're interested in. And good and bad data are really important in terms of training AI and machine learning. You don't just want the good data. You want to know where the losers are, where the winners are, where the intermediates are. And once you have enough of it, machines then can tell you where to go next. So in a world of AI, the next big frontier is going to be AI for science. How could this help improve something like producing green hydrogen? A lot of people want to use iridium for the oxygen evolution reaction – splitting water if you're trying to generate clean hydrogen. The problem with iridium is there's not enough iridium in the world to meet the projected demand. And it's in a bad part of the world from a U.S. perspective, so there's a geopolitical risk. That's a great example of channeling megalibrary technology to solve a major problem. Can I find a combination of elements that work as well as Iridium or better than Iridium and either reduces or eliminates the need for Iridium? And what we did here was eliminate the need for Iridium. We found a catalyst that was as active, actually a little bit more active, and comparably stable. Those are two of the big parameters. You have to know what you're trying to beat. We don't have a dog in the fight in terms of what the winner is. We just know that we can survey very, very rapidly the landscape and look at different shots on goal faster than has ever been done before. Iridium is an interesting example since it’s both rare and extremely expensive. Is the material you’ve generated going to be used commercially to make green hydrogen? Oh yeah. We found catalysts that are being commercialized now. ... To me, it's not the most important one, but it was one that was clear cut, where we could really make the point and show the world that we can channel this capability down a path that solves indisputable problems. Many people recognize that it would be great if we had something as good as Iridium that had more earth-abundant elements. That had things with less geopolitical risk, but also performed or behaved at least as well and preferably better. And what about in terms of cost? Yes, that lowers cost as well. Can you estimate how much cheaper this material is relative to Iridium? I don’t have it on the tip of my head, but we’ve done that before. It can be orders of magnitude less expensive. What Else We’re Reading Lawsuit seeks to block closure of top federal climate research center (Inside Climate News) When climate change hits an aging population: Heat’s threat is underestimated, study says (NBC News) Becerra reveals 'power hour' idea to give California free electricity (USA Today)
00:00

The Complex Psychology Of People Choosing To Trust Either Humans Or AI

People generally trust humans more than they trust AI, but that bias can shift and even flip depending on the situation. Researchers call the preference the "human premium" and the distrust the "AI penalty." In experiments where people are told they're chatting with AI but it's really a human, they still hold the AI penalty against the interaction, and the reverse happens too. People burned by bad human customer service can end up trusting AI more.

Notes

Notes saved to notes/forbes-psychology-trust-humans-or-ai-2026-08-24.md. Core content: the four mental states (human premium, AI penalty, AI premium, human penalty), the Cambridge trust definition, the customer-service thought experiment, the four matched/mismatched deception situations, anchoring/disbelief on source reveal, the context-driven switcheroo, and stated limitations (research covering only two of the four states; no specific studies named).

Full text · 14,425 chars
In today’s column, I examine the psychology underlying how people choose to trust others. By others, I am referring not only to trusting humans but also to opting to trust AI. The vexing question of whether people are going to trust AI is a monumental one and keeps getting a lot of solemn attention. The abundantly logical approach entails discussing trust in two pathways, consisting of the longstanding human-to-human basis and the newly arising human-to-AI basis. Numerous research studies have indicated that people often have an inherent trust bias known as the “human premium”. They tilt toward trusting humans more than trusting AI. Indeed, there is an allied bias referred to as the “AI penalty”, namely that people perceive AI as less trustworthy than their placing trust in humans. It’s a double whammy on the trust spindle. But there is a surprising twist to this. Sometimes, there is a switcheroo. People will have a semblance of an “AI premium”, perceiving AI as more trustworthy than humans, and demote their human trust into a kind of “human penalty” of sorts. The crux is that the whole kit-and-caboodle is much more complex than it might seem at initial glance. Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage of the latest in AI, including identifying and explaining key AI complexities (see the link here). The Minds Problem Much of the handwringing about AI is whether AI is going to become sentient and embody consciousness, often referred to as AI as a thinking machine. That is an important topic and one that I’ve continued to closely explore and analyze; see the link here. There is a quite different angle that also deserves rapt attention, and I’d like to place it on the table. How will human minds change because of interacting with AI? This unexpected question catches many people by surprise. The customary focus is on AI as formulating a mind, not on how human minds might change due to interacting with AI. But, as the old saying goes, it takes two to tango. When humans increasingly interact with AI, there is a two-way street involved. Humans adjust their minds about how they view the world, how to communicate with others, and shift in ways that we still do not have a full or clear picture of. In that sense, it is wholly worthwhile to study the psychology of humans as they mentally adjust to a world that entails human-to-AI interaction and human-to-human interaction. Notably, this doesn’t have to wait until or if AI becomes sentient. There is plenty to study right now. Humans are already adjusting their minds to the ubiquitous nature of non-sentient AI. The starting gun has gone off. Human minds are changing. For my comprehensive tracing of how AI and psychology dovetail with each other, see the link here and the link here. The Matter Of Trust Humans are very preoccupied with trust. You make mindful decisions about who you trust. Do you trust your family members? Do you trust the people at work? Do you trust your friends? Do you trust newly met strangers? Much of the mental work we do daily and moment-by-moment involves levels of trust and the assignment of trust. The online Cambridge Dictionary defines the word “trust” as follows: “To believe that someone is good and honest and will not harm you, or that something is safe and reliable.” The trickiness about trust is that it isn’t a one-and-done affair. You don’t just decide that a particular level of trust for this person or that person is going to stick for the rest of your entire life. Instead, you modify your perception of the trust that you are willing to assign to someone or something. Trust is dynamic. It comes and goes. Humans learn from a very young age how to assign trust and rely on the trust that they perceive in others. A new target for that trust assessment has loomed large, namely the trust that we assign to AI. Again, even though AI is not yet sentient, we nonetheless still tend to perceive AI in human-like terms and opt to mentally calculate the level of trust that we assign to the AI that we encounter. Trust For Humans Over Trust For AI Various research studies have done intriguing experiments about the trust that humans assign to humans versus the trust assigned to AI. I will take you through a brief thought experiment so you can gauge your own sense of human-to-human versus human-to-AI trust capacities. You receive a text message from the customer service of a company from which you want to return a product you recently purchased. The text indicates this: - Text message: “Your business is of utmost importance to us, and I am ready to assist you in your request to make a product return.” If I tell you that this is a human agent of the company, what semblance of trust do you mentally assign to the interaction you are about to have regarding your product return efforts? I suppose you might be partially wary since we’ve all had situations where returning a product was much harder than it should have been. In any case, your trust is perhaps going to be at a moderate level, containing a modicum of optimism and a touch of skepticism. The Human Bias Against AI What if I tell you that this text is not from a human agent but instead from an AI-based agent? Yes, you are going to be interacting with AI about undertaking the product return. No human-to-human interaction is going to occur. It is entirely a human-to-AI discussion. Now that you know it is AI, what is your mentally assigned trust level? Most people would likely be immediately dismayed. AI is perceived as being less trustworthy than interacting with a human. The odds are that the AI is going to be formulaic. It will not be as flexible and fluid as trying to do the product return with a fellow human. Yikes, this dialogue on returning the item you purchased is going to be a potential nightmare. The psychological catchphrase that is used for this reaction is that people tend to assign a “human premium” for being able to interact on a human-to-human basis. They give a higher level of trust to human interactions. Furthermore, they tend to give an “AI penalty” for having to interact with AI. They mentally take points off for dialoguing with AI over dialoguing with humans. It is a double whammy. People tend to favor human dialogues and tend to disfavor AI dialogues. The Matter Is In The Mind Research studies have generally found support for this “human premium” and “AI penalty” as a mental formulation. To ascertain how deep this bias might be, experiments have been performed where the wool was pulled over the eyes of the experimental subjects. A trick of sorts was played on the participants. Suppose that I told you the product return interaction was going to be AI-based, but it was actually a human who would be doing the dialogue. I lied to you about the source of the dialog. In other words, you perceive that the interaction will be with AI. Meanwhile, reality is that a human is on the other end of the interaction. Many people would ardently claim that they will detect this immediately. The dialogue will get underway, and they will realize that they aren’t conversing with AI. The tomfoolery will be discovered. They will readily discern that the discussion is with a fellow human. The thing is, the latest advances in AI are so good that you would be hard-pressed to differentiate an AI-powered discussion from a human-led discussion. It used to be that AI dialogues were stilted. The AI showed its hand. Clever questions on your side of things would get the AI boxed in, and it would say that it cannot respond. Those days are numbered, and routine conversations are a thing of the past. When You Can’t Discern Here’s where I am taking you. If people are told that they are interacting with AI, but it is really a human, the odds are that those people will accordingly bring their mental “AI penalty” into play and hold that against the human that is on the other end (assuming that they cannot otherwise discern the designated source). The person will stridently harbor their preference for a “human premium” and apply their “AI penalty” throughout the dialogue. Likewise, the reverse condition occurs. If people are told they are interacting with a human, but it is really AI, the odds are that those people will accordingly bring their mental “human premium” into play. They will grant added trust to the source of the dialogue, doing so merely because they perceive the source to be a fellow human. Let’s bring this back to the human mind. Your mind is adjusting the level of assignable trust based on a perception of whether you are interacting on a human-to-human basis or a human-to-AI basis. This occurs even if your perception is wrong. There are four situations involved: - (1) Human-to-human matched. Perceived as human-to-human, and the interaction is human-to-human (matching). - (2) Human-to-AI matched. Perceived as human-to-AI, and the interaction is human-to-AI (matching). - (3) Human-to-human mismatched. Perceived as human-to-human, but the interaction is human-to-AI (mismatched). - (4) Human-to-AI mismatched. Perceived as human-to-AI, but the interaction is human-to-human (mismatched). The Strength Of The Bias You might be doubtful that this bias would be carried forward even when the mismatch is occurring. The research tends to show that the bias does indeed persist. Furthermore, a person can get unduly anchored to their bias. If you tell someone they are about to interact with a human, they seem to mentally click into the “human premium”. After a smattering of dialogue, suppose you now reveal that they have been interacting with AI. The chances are they will declare you to be lying to them and that they are in fact interacting with a human. They won’t believe that you lied to them at the get-go and instead you are lying now. Why would that be? Because their frame of reference that they were interacting with a human has been reinforced during the initial dialogue. They expected a human; they now believe they have been interacting with a human. There is also a chance that they don’t want to appear to have been bamboozled. To avoid looking foolish, they will cling to the belief that it has been human and not AI. Switcheroo On Mental Bias One aspect of the interpretation of these mental biases is that you must take into account that things can get switched up. People will not always be enamored of the “human premium,” nor will they always be leaning into an “AI penalty”. In fact, a complete switcheroo can arise. Imagine that you have regularly dealt with human customer service agents, and they are exasperating and difficult to contend with. They at times go off on tangents that aren’t pertinent to the product return activity. You have developed a bias against human agents. In addition, imagine that you have dealt with interactive AI-based customer service systems. Those have been straightforward for you. You answer the questions posed. The product return gets processed. Easy-peasy. No small talk, no wasted time. This means that depending upon the context of a given situation, you are willing to flip the entire trust framework. Rather than fostering a “human premium,” you have an “AI premium” associated with situations where AI is more trustworthy than dealing with a human. Indeed, you also harbor a “human penalty” whereby the instant you find out you will need to chat with a human, you let out a sigh of dismay. You were hoping that it would be AI in lieu of being a human. The Four Mental States All told, a person can hold these four mental states: - (a) Human Premium. Mental bias that interacting with humans is desirable and deserves a premium accordingly. - (b) AI Penalty. Mental bias that interacting with AI is less desirable and deserves a penalty accordingly. - (c) AI Premium. Mental bias that interacting with AI is desirable and deserves a premium accordingly. - (d) Human Penalty. Mental bias that interacting with AI is less desirable and deserves a penalty accordingly. When I see research that only acknowledges the first two, I worry that they might not have been aware that the additional two also exist. This is easy to overlook or downplay. The “human premium” and the “AI penalty” are currently getting a boatload of attention. The “AI premium” and the “human penalty” are not yet getting their rightful attention. Anyway, the gist too is that all of this is heavily context-dependent. People will shift their mental models according to their prior experiences. Of course, the shift is not necessarily going to be on target. For example, suppose a person wanted to interact with AI; they indeed get AI, and so they would seem to be happy about the alignment. If the AI doesn’t perform as per their expectations, they are bound to regret that they are locked into an AI conversation. At that point, they will try to escape the AI and see if they can get into a human-to-human conversation instead. The same holds with wanting a human-to-human chat and later deciding you’d prefer to switch to a human-to-AI chat. The World We Are In Consider whether you have a mental bias toward human-to-human versus human-to-AI interactions. When do you prefer the human-to-human side? When do you prefer the human-to-AI side? Are you bringing that bias into your mind when told beforehand whether you are interacting with a human versus AI, or AI versus human? Can you adjust if the source turns out to be different from what was initially conceived? Human psychology is gradually adapting to a world entailing AI that can convincingly appear to be human-like. Up until now, you could treat other artifacts as though they were human-like, but you knew for sure they weren’t. You might act as though your car or toaster is a living being, and converse at them, though you knew this was just pretend. AI is advancing to the point where it is exceedingly hard to differentiate during a dialogue whether you are conversing with a human or AI. A final thought for now. The famous philosopher Simone de Beauvoir made this notable remark: “It is doubtless impossible to approach any human problems with a mind free from bias.” I mention this point because even if you are aware of the premium and penalty biases, you aren’t necessarily free of them. Your best bet is to harness them, control them, and minimize their impact. Awareness is a great first step.
00:00

How To Design Multi-Agent Workflows That Actually Work

Breaking a big job into a team of narrow, specialized AI agents is more reliable than letting one agent do everything. Each agent handles a single clearly defined step and passes the work along, which makes errors easier to find and fix without breaking the whole process. Pick tasks with measurable results, define exactly what each agent hands off, and track every agent's performance, not just the final output. Keep a human in the loop whenever a decision carries real consequences, like approving a payment.

Notes
Thesis

Single agents handling entire complex processes are hard to debug ("the more you ask one agent to do, the harder it becomes to understand what went wrong when it fails"). Proposed fix: teams of narrowly-scoped specialist agents, each doing one step then handing off.

Examples given
  • Marketing/social: research persuasive messaging → draft copy → proofread/fact-check → publish + distribute → manage ad budget.
  • Customer service: classify requests → consult knowledge base → draft responses → decide when to escalate to humans.
Advantages claimed

Simpler per-agent design; easier to locate errors; individual agents can be swapped/retrained without breaking the rest; easier to update as requirements evolve.

Design steps (in order)
  • Pick a job with measurable results and objectively definable success — "If you can't precisely identify what 'good' looks like, you won't know whether your agents have succeeded or failed."
  • Design the workflow first, not the team: map every task, then decide agent count and roles.
  • Give each agent a narrow job with consistent inputs/outputs; no duplicated or overlapping responsibilities.
  • Define hand-offs explicitly: each agent must know what it expects from the prior step and what it passes to the next.
  • Track every agent's performance, not just the final outcome; evaluate output quality at each step.
  • Decide in advance where human-in-the-loop is non-negotiable (payments, declining customer requests). > "When judgment calls carry heavy consequences, there's no substitute for this."
Position / caveats

Author argues against giving agents ever-more responsibility; value comes from breaking processes into measurable, iteratively improvable steps. Start with a simple workflow. Explicitly acknowledges first attempts "probably won't" work and that multi-agent design "is new to everyone"; frames failure as diagnosis — adjusting one component at a time is the approach's chief advantage. Note: this is opinionated advice, not benchmarked or empirically validated.

Full text · 4,957 chars
AI agents are increasingly being pitched as virtual workers capable of running entire business processes with minimal supervision. But there’s a problem: the more you ask one agent to do, the harder it becomes to understand what went wrong when it fails. For complex jobs, from running a marketing campaign to managing an end-to-end QA process, a better approach is often to build a team of specialized AI agents. Each agent handles one clearly defined part of the workflow before passing the work on. This can make AI automation easier to manage, measure and improve. But it introduces another challenge: getting all those agents to work together effectively. So, here’s an overview of the multi-agent approach to automating business processes, some tips on putting it to work, and advice on some of the obstacles you may encounter along the way. What Is A Multi-Agent Workflow? Think of a multi-agent workflow as a small team of AI specialists that are each expert at a particular job. For example, if you’re using AI to create social media posts as marketing copy, one agent could research the messaging most likely to persuade your audience, one could draft the copy, one proofreads and fact-checks, another publishes it and makes sure the right people see it, and another manages your advertising budget. Another example: In customer service, rather than a single agent handling every inquiry end-to-end, one agent could classify incoming requests, another consults the knowledge base on how to solve problems, another drafts responses, and another decides when to escalate complex cases to human agents. The primary advantage is that rather than a single agent having to understand and execute an entire, complex workflow, each agent only has one relatively simple and clearly defined task. This simplifies the design process while making it easier to spot where errors are occurring. Individual agents that aren’t doing their job properly can then be swapped out or retrained without risking breaking other elements of the workflow. It also makes it easier to update or adapt the process as business requirements evolve. How To Make It Work There are some simple steps that can be taken when designing this type of workflow to improve the chances it will work out. First, pick a job where the results can be easily measured and success objectively defined. If you can’t precisely identify what “good” looks like, you won’t know whether your agents have succeeded or failed. Then, move on to designing the workflow, rather than the team. Map out every task that needs to be done, then decide how many agents are needed and what they all have to do. Give each agent a narrow, clearly defined job with consistent inputs and outputs. It should be obvious what each agent is meant to do, with no duplicated work or overlapping responsibilities. Make sure your hand-offs are clearly defined. Each agent should know exactly what information it should expect to be passed from the agent responsible for the prior step of the workflow, as well as what information it needs to be ready to pass to the agent in charge of the next step. Next, make sure you’re tracking the performance of every agent, not just the final outcome. This will let you identify bottlenecks and understand where your agents may be tripping up. Evaluate the quality of your agents’ output at each step, not just whether the end results look okay. Finally, make sure you know where the human-in-the-loop is non-negotiable. Decide in advance what needs human involvement or approval, such as a payment being made or a customer request being declined. When judgment calls carry heavy consequences, there’s no substitute for this. Final Thoughts As agents become more powerful and capable, many of us might be tempted to keep giving them more responsibility. I believe this urge should be resisted, and we should instead focus on building teams of simpler, specialist agents and improving their ability to collaborate effectively. For most organizations, the real value of agents won’t come from replacing humans with ever more complex autonomous systems, but from breaking complex business processes down into manageable, measurable steps that can be improved iteratively. This is most easily achieved by starting with a simple workflow that we can use to better understand where automation and collaboration add value, and that can be refined over time, rather than trying to get everything right in one shot. First attempts won’t always work out as planned, and it’s worth remembering that designing agentic workflows, particularly involving multiple agents, is something that’s new to everyone, not just you. If things don't work as expected from the start (and they probably won't), then treat it as a learning experience rather than a failure. A big advantage of the multi-agent approach is being able to diagnose and adjust one component at a time until you get the results you want.
00:00

Smart Glasses Set To Transform Healthcare And Wellness

Smart glasses are poised to go from fashion accessories to health-monitoring devices, following the path smartwatches already took. Glasses sit where sensors can track heart rate, eye movement, pupil dilation, migraine activity and hydration without disrupting daily life. Existing examples include FDA-approved lenses that slow nearsightedness in kids, migraine glasses that filter pain-triggering light, and assistive glasses that read text aloud or recognize faces for people with vision loss or dementia. This is a trend forecast rather than a product announcement, predicting next-generation glasses will add medical features.

Notes
  • Framing: Apple's leadership transition — a "hardware-focused executive" expected to take a larger role — plus Apple Watch precedent (heart rate tracking, fall detection, blood oxygen monitoring, cardiac rhythm notifications) is used to argue smart glasses will follow the smartwatch arc from fitness accessory to health platform.
  • Market: Meta × Ray-Ban = "arguably the first mainstream success in AI-powered smart glasses"; lineup expanded with Ray-Ban Meta prescription smart glasses. Google unveiled AI glasses with partners Warby Parker and Gentle Monster. Apple "reportedly accelerating work on its own smart glasses."
  • Form-factor argument: temples are sensor-ideal sites for "brain actions, eye movement, pupil dilation, elevated heart rate, muscle tension, migraine activity and even temperature or hydration levels" — enabling continuous, unobtrusive monitoring.
  • WHO data (quoted): "globally, at least 2.2 billion people have a near or distance vision impairment"; leading causes "refractive errors and cataracts." Access gap: "2 out of 3 people in low-income countries who need eyeglasses don't have access to them. In addition, 1 in 2 people globally who need cataract surgery don't have access to that surgery."
  • Vision correction: Early corrective eyewear in babies with severe refractive issues improves "visual attention, mobility, and learning development." Per the American Academy of Ophthalmology: "there are now FDA-approved eyeglasses to slow myopia progression in children ages 6 to 12. While the center of the lens corrects distance vision, the rest of the lens intentionally blurs peripheral vision. This signals the eye to slow its growth." University of Utah Health announced a similar myopia-managing lens. Products named: eSight and Envision glasses for low/no vision, and support for blindness/dementia via face recognition, reminders, daily-routine guidance.
  • Alzheimer's assist: facial recognition to display names, contextual cues (scheduled activities), voice prompts, augmented guidance.
  • Caveat (the article's own): these medical aids "are expensive standalone eyewear to be worn separately from regular glasses. Instead, they need to be integrated into affordable smart glasses for multi-pronged consumer usage."
  • Migraine/light sensitivity: glasses filter specific wavelengths (blue and fluorescent light, "common migraine triggers"). Avulux makes FDA-approved lenses that "block up to 97% of painful light frequencies while maintaining accurate color perception." Newer models add photochromic lenses and "may use AI to adjust tint dynamically in real time."
  • Limits/unaddressed: article is speculative and forward-looking with no clinical trial data; health claims rest on company/AAO statements and the smartwatch analogy. No pricing, adoption numbers, or release dates given for the cited products (Ray-Ban Meta, Google/Warby/Gentle Monster, Avulux, eSight, Envision).
  • Bottom line (author's): glasses will mirror smartwatches — "from fashion devices to health gadgets"; next-gen smart eyewear could reshape health monitoring and daily wellness.

Notes written and task task_1787581159154 marked done.

Full text · 6,332 chars
Apple’s leadership transition could have major implications for the future of wearable technology. With a hardware-focused executive expected to take on a larger leadership role, many industry observers expect Apple to push more aggressively into new hardware innovation, particularly in wearable devices. Apple has already signaled its ambitions in this area with the launch of the Vision Pro headset, making it clear that the company sees wearable glasses and spatial computing as part of its long-term product roadmap. The company previously transformed the smartwatch category through the Apple Watch, turning what began as a fitness accessory into a serious health monitoring platform with heart rate tracking, fall detection, blood oxygen monitoring and cardiac rhythm notifications. For millions of users, it’s become a daily companion for wellness and chronic condition monitoring. That evolution showed how consumer electronics can gradually become healthcare tools. Smart glasses appear to be following the same trajectory, with today's AI assistants and cameras laying the foundation for tomorrow's health monitoring and personalized medical applications. The timing is significant not just due to Apple’s CEO transition but also because the competitive landscape in smart glasses is evolving rapidly. Meta has arguably created the first mainstream success in AI-powered smart glasses through its partnership with Ray-Ban. Building on that momentum, Meta recently expanded the lineup with new Ray-Ban Meta prescription smart glasses, making the technology accessible to millions of people who already wear prescription eyewear. At the same time, Google has unveiled its own AI-powered smart glasses developed with partners including Warby Parker and Gentle Monster, while Apple is reportedly accelerating work on its own smart glasses. What was once a niche category is quickly becoming one of the most competitive segments in consumer electronics. Smart Glasses Have a Vantage Point Smart glasses offer a uniquely advantageous form factor. Positioned on the face, they have line-of-sight access not just to the external environment but also to the human head. They sit on the temples, which are ideal locations for sensors that can measure brain actions, eye movement, pupil dilation, elevated heart rate, muscle tension, migraine activity and even temperature or hydration levels. This opens the door to real-time, continuous monitoring without disrupting daily life. A Global Issue According to the World Health Organization, “globally, at least 2.2 billion people have a near or distance vision impairment. WHO states, “the leading causes of vision impairment and blindness are refractive errors and cataracts.” This is a more severe issue in the developing world, as explained by WHO: “It is estimated that 2 out of 3 people in low-income countries who need eyeglasses don’t have access to them. In addition, 1 in 2 people globally who need cataract surgery don’t have access to that surgery.” As a solution, smart glasses can help users with visual impairments by reading text aloud and describing surroundings in real time. Vision Correction Researchers and pediatric ophthalmologists are also using early vision correction technologies in children and infants, where timely intervention can dramatically improve visual development and long-term cognitive outcomes. Early use of corrective eyewear in babies with severe refractive issues has been shown to improve visual attention, mobility, and learning development. According to the American Academy of Ophthalmology, “there are now FDA-approved eyeglasses to slow myopia progression in children ages 6 to 12. While the center of the lens corrects distance vision, the rest of the lens intentionally blurs peripheral vision. This signals the eye to slow its growth. Because eyeball growth worsens myopia, in this way, the eyeglasses help slow it down.” The University of Utah Health also announced, “a new type of eyeglass lens is making it easier to manage myopia (nearsightedness) in children, helping them see clearly while also slowing vision changes over time.” Similarly, products like eSight and Envision glasses offer enhanced vision to people with low or no vision. They can also support those with blindness or dementia by recognizing familiar faces, displaying reminders or guiding users through daily routines. Smart glasses also show promise as assistive tools for individuals with Alzheimer’s disease. These devices could use facial recognition to display names and offer contextual cues like scheduled activities. Voice prompts and augmented guidance could help reduce confusion and promote independence. While these products are medically helpful, they are expensive standalone eyewear to be worn separately from regular glasses. Instead, they need to be integrated into affordable smart glasses for multi-pronged consumer usage. Migraine and Light Sensitivity Another growing category of wearable glasses is addressing migraines and light sensitivity by filtering specific wavelengths of light. These glasses often use specialized tints to reduce exposure to blue and fluorescent light, both common migraine triggers. As screen time and artificial lighting exposure increase, these glasses are becoming vital tools for individuals suffering from photophobia and chronic migraines. Companies like Avulux offer FDA-approved lenses that block up to 97% of painful light frequencies while maintaining accurate color perception. Many models now include photochromic lenses that adapt to lighting conditions and may use AI to adjust tint dynamically in real time. This segment shows how targeted, non-invasive wearable technology can deliver measurable quality-of-life improvements and serve as a blueprint for future health-focused eyewear. Conclusion As hardware becomes lighter and more powerful, and as AI grows more capable, wearable glasses will likely follow the same trajectory that smartwatches did: from fashion devices to health gadgets. While devices like the Ray-Ban Meta glasses are lifestyle-focused today, the next generation of smart eyewear could reshape how we monitor health and maintain daily wellness. In a world already comfortable with bio-watches on our wrists, the shift to smart glasses for healthcare may not be far behind.
--:--

America's Most Powerful Women In Sports

Forbes' annual ranking of America's most powerful women in sports, likely covering executives, owners, and athletes with influence. It has no direct AI connection, no article body was retrieved, and the ranking itself is the whole story.

--:--

Forbes Iconoclast with Maneet Ahuja | Paid Program

A podcast interview episode from Forbes' Iconoclast series with Maneet Ahuja. No article body was retrieved, so this summary comes from the title alone, and the piece is a paid promotional program.

--:--

Acxiom Insights | Paid Program

This is a paid promotional piece, so treat it as marketing rather than news. It argues that marketers should use identity data to stitch together a fragmented view of the same consumer across devices. Acxiom, a data and identity company, is behind the content. Low signal for a research archive.

Discussion

8
04:55

I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB

A hobbyist trained a small AI model from scratch that's aggressively compressed and can pull answers from a huge on-disk memory. It has 250 million parameters, trained on 30 billion tokens, and is quantized to under 2 bits so the whole thing ships at 60 MB and runs at around 400 tokens per second on a normal laptop CPU with no GPU. Instead of keeping history in RAM, it writes older context to disk at 320 bytes per token and retrieves from up to 100 million tokens back, even answering a question whose answer sat 50 million tokens deep. Its word table is also a fixed code with zero trained parameters instead of a normal vocabulary embedding. Quality is modest, with perplexity of 23.3, and the code is MIT-licensed on GitHub and Hugging Face.

Notes

Quantized 250M LLM trained from scratch — 60 MB deployment

Poster: /u/Final-Data-1410 (r/LocalLLaMA, reposted from r/MachineLearning where it got 300+ upvotes). Repos: github.com/QLNI/SHADOW-250M-Instruct, huggingface.co/NODEMIND/SHADOW-250M; GitHub ~35 stars.

Core specs
  • 250M params, trained from scratch on 30B tokens of FineWeb.
  • Quantized to under 2 bits → deployment is 60 MB, runs in ~80 MB RAM.
  • ~400 tok/s on a laptop CPU, no GPU, no framework — small compiled runtime (Windows + Linux), MIT license.
Quality
  • Held-out English web text (educational pages, 2,048-token windows, never in training): cross-entropy 3.15 nats/token, perplexity 23.3, 0.99 bits/byte.
Long-context design (novel)
  • Most recent 2,048 tokens stay fp16 (normal KV cache).
  • Older tokens compressed to 1 bit, written to disk (~320 bytes/token); 1M tokens ≈ 320 MB disk.
  • Model was trained to retrieve from that disk cache up to 100M tokens — but, budget-limited, only retrieve/answer, not reason over them.
Vocabulary
  • No embedding table: every token is a fixed 512-bit code, 8.4 MB for 131k tokens, zero trained parameters.
  • WordSim-353 (human word-similarity): 0.619 Spearman vs 0.029 for random codes (test script in repo).
Shown outputs (settings stated, reproducible)
  • Photosynthesis explanation (greedy) — correct 2-sentence answer.
  • Sea poem (temp 0.25, top-k 30, rep 1.15, seed 2).
  • Archive retrieval: "serial number of device Grus-189" answered SN-442976 from 50.6M tokens deep on disk (archive mode, k=16).
Caveats (stated)
  • "It's a 250M model so expect mistakes on open facts, I'm not claiming it beats anything big."
  • Fine-tunable; kit includes demo and before/after numbers, master weights in repo.
Full text · 2,833 chars
Reposting here after sharing this on [ r/MachineLearning ]( r/MachineLearning ) a few days ago, where it got a much better response than I expected (300+ upvotes, great questions, zero roasting) GitHub is at 35 stars now. So here it is. I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about 80 MB of RAM to run. Runs around 400 tok/s on a normal laptop CPU, no GPU needed, no framework, just a small compiled runtime (Windows and Linux included, MIT licensed). Language modeling quality of the base model, measured on held-out English web text that was never seen in training (educational web pages, 2,048 token windows): cross entropy 3.15 nats per token, perplexity 23.3, 0.99 bits per byte. How the long context works: the most recent 2048 tokens stay in fp16 like a normal KV cache. Everything older gets compressed to 1 bit and written to disk, about 320 bytes per token, so 1 million tokens of history is roughly 320 MB on disk. From the start the model was trained to retrieve from that disk cache, up to 100M tokens. Due to a limited budget it wasn’t trained to reason over those tokens, only retrieve and answer from them. The vocabulary is also not a normal embedding table. Every token is a fixed 512-bit code, 8.4 MB for all 131k tokens, zero trained parameters. I tested it on WordSim-353 (human word similarity ratings): my table scores 0.619 Spearman correlation vs 0.029 for random codes. Test script is in the repo. Some outputs (settings included so you know I’m not cherry picking, all reproducible from the repo): “Explain photosynthesis in two sentences.” (greedy) Photosynthesis is a process in which plants convert sunlight into chemical energy, which is then used to produce oxygen and other chemicals. This process is called photosynthesis. “Write a short poem about the sea.” (temp 0.25, top-k 30, rep 1.15, seed 2) The waves had swept over, and they were crashing against each other like rocks on top of one another. The wind was strong but not very gentle as it tried to push them towards their destination without lifting any weight or pulling away from those who stood in front of her with a stiff smile that made the sea look even more majestic than before. “What is the serial number of device Grus-189?” where the answer sits 50.6 million tokens deep in the archive on disk (archive mode, k=16) SN-442976 It’s a 250M model so expect mistakes on open facts, I’m not claiming it beats anything big. You can also fine-tune it, the full kit with a demo and before/after numbers is included. Master weights for fine-tuning are in the repo too: https://github.com/QLNI/SHADOW-250M-Instruct https://huggingface.co/NODEMIND/SHADOW-250M submitted by /u/Final-Data-1410 [link] [comments]
07:04

Xiaomi AI Cube announced with 1.2TB/s memory bandwidth

Xiaomi showed off a prototype desktop AI computer, the AI Cube, built from three of its own chips. It pairs an AI accelerator with 1.2 TB per second of memory bandwidth alongside a memory chip that supports up to 160 GB of RAM. The specs are still a bit confusing, and the company hasn't said how the memory numbers add up. It's a prototype, so there's no release date yet.

Full text · 473 chars
Xiaomi announced a prototype for their Xiaomi AI Cube. 3 chip system: - Xiaomi Xuanjie O3 - Xiaomi Xuanjie O100 - Xiaomi Xuanjie D100 The specs are impressive, but a bit confusing. The D100 chip (originally for their EVs) supports up to 160GB of RAM, but O100 has the 1.22TB/s memory bandwidth. Perhaps the 1.22TB/s figure is for SRAM? Hard to say definitively. Source: https://www.ithome.com/0/993/546.htm submitted by /u/Mysterious_Finish543 [link] [comments]
04:20

deepseek-v4-flash-0731 - surprisingly usable

DeepSeek's v4-flash model runs fast enough to be genuinely usable on a home-built computer that costs nowhere near $10,000. The builder paired an AMD Epyc 7663 server CPU with 256GB of RAM and a single RTX 5090, then ran an 8-bit quantized version of the roughly 151GB model off system memory at about 24 tokens per second across 100-128k tokens of context. It's just one data point for this CPU-offload setup, and the owner is new to local inference, so results may vary.

Full text · 993 chars
I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is. I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. It was slower with DFlash so I took that out. It was also slower with a 3090 I put in there temporarily. I'm posting this mostly because I didn't see too many other data points for this config (DDR4 Epyc + Blackwell doing cpu-moe). And also that I'm pretty surprised that a model this good can actually run in my basement without dropping $10k or running a sub-panel down there. I'm otherwise fairly new to this - would love any tips on what else to run or how to further improve it. submitted by /u/IntravenusDeMilo [link] [comments]
12:11

Delay-corrected Bellman operator + causal attribution for constrained RL contraction proof under unknown stochastic delay [R]

A new reinforcement-learning method stops AI agents from being penalized for the wrong action when rule violations show up late and at random, which is common in the real world. It pairs a delay-corrected Bellman operator, with the discount adapted to the delay distribution and a proof it still converges when delays are unknown, with an Interventional Consequence Net that attributes blame through causal estimates instead of timing. The catch: that attribution net needs the environment's structural causal model to train, so it won't work where no such model exists. It's an open call for collaborators in constrained and safe RL.

Full text · 1,344 chars
Standard constrained RL assumes consequences are immediate and attributable to the current action. This breaks down whenever violations are delayed and stochastic, which is most real-world settings you end up penalizing whatever action happened to precede the observed violation, not the action that caused it. Working on CCPL (Causal Consequence-Penalized Learning) to address this: - A delay-corrected Bellman operator using an adaptive effective discount learned from the consequence-delay distribution. Contraction proof holds under unknown stochastic delay. - An Interventional Consequence Net (ICN), pretrained on structural-causal-model labels, estimating marginal causal contribution per action for attribution rather than penalizing based on temporal proximity. Limitations, to be upfront about them: - The ICN currently requires access to the environment's structural causal model to generate pretraining labels it's not learned end-to-end from observational or interventional data alone. That's a real constraint on applicability outside benchmark settings where the SCM is known or can be reasonably specified. Open to contributions and collaborators, especially if you work in constrained/safe RL or causal inference feel free to open an issue or reach out directly. submitted by /u/No_Cauliflower7923 [link] [comments]
01:22

Qwen 3.8 27B, just wanted to say thanks to you guys

A hobbyist finally got a local AI model running on his own hardware, and it worked. He had bought GPUs in 2023 and given up repeatedly, but fixed his setup in about an hour with community help. He's now running Qwen 3.8 27B locally and has it hooked up to his Home Assistant server, complete with vision features. It's a beginner success story, not real news.

Full text · 686 chars
I commented on another Qwen 3.8 27B post that I was frustrated getting anything to work. You all gave some great comments. I nuked openwebui and straightened out my llama.cpp docker config. 1 hour of work and I have a model I can chat with, connected to my HomeAssistant server, which I have already updated dashboards with a short prompt and a screenshot (wtf vision built in?) Guess all I needed was the right push. I bought several GPUs in 2023 in impulse purchases for Folding@Home, but have always wanted to spin up my own local coding/help agent, just always gave up when nothing seemed to work. This feels like magic. Thanks! submitted by /u/sshwifty [link] [comments]
06:11

AAAI 2027 Reviewer Bidding and Assignment Integrity [D]

A major AI conference publicly acknowledged that collusion is corrupting its peer review, where researchers secretly arrange to review each other's papers. AAAI's 2027 organizers emailed authors about it, especially "2-cycle" swaps where two authors review each other's work, and one commenter notes submissions concentrated in a single country make those cycles more likely. The same thread repeats a familiar complaint that accepted papers at top venues like NeurIPS and ICLR often ship without code, forcing others to reimplement results from scratch.

Full text · 1,260 chars
Recently, the AAAI 2027 organizers sent an email regarding collusion occurring during the review process, especially in the 2-cycles category (i.e., an author of Paper A reviews Paper B, while an author of Paper B reviews Paper A). Given the fact that most submissions come from a single country, there are higher chances that the assignment algorithm will naturally create 2-cycles among authors from that country. This, in turn, means that most authors involved in collusion could be from that country. I will not name that country; otherwise, I would be labelled as racist. By the way, did AAAI release statistics about the number of submissions, like they did last time? It is also good news that a major and prestigious conference like AAAI is acknowledging that collusion is happening. We all knew that this kind of collusion had been happening for years. There are papers accepted at top conferences such as NeurIPS, ICLR, AAAI, and ICML that do not even have their code published on GitHub. This forces other researchers in the community to spend substantial time reimplementing the code themselves if they want to reproduce the reported results. What are the views of other authors on this? submitted by /u/Fragrant_Fan_6751 [link] [comments]
12:17

Who would buy HuggingFace

People are speculating about who might buy HuggingFace, the central home for open AI models. The talk follows OpenRouter getting snapped up by Stripe, which prompted the question in a hobbyist forum. A deal would cost around $13 billion, and some float Apple as a natural buyer given its push into on-device AI. This is discussion with no real news behind it.

Full text · 322 chars
Given OpenRouter.ai was snapped up by Stripe, who do we think would go after the "GitHib" of AI models? It is a big chunk of change they are looking ($13B). Apple may be a contender to give them a real chip in the AI race, given how they are focused on local AI execution. submitted by /u/Wallaby989 [link] [comments]
05:39

BMVC 2026 IJCV recommendation? [D]

Someone is asking how the BMVC computer-vision conference decides which papers get recommended for its IJCV journal special issue. They want to know whether it comes from review scores or a separate call by area chairs, and whether authors only learn about it through a later email. The thread is just a procedural question with no real substance beyond that.

Full text · 537 chars
Does anyone know how the BMVC to IJCV special issue recommendation works? Is it mainly based on the review scores, or is it a separate decision by the ACs/program chairs (e.g. based on oral/highlight selection, reviewer comments, etc.)? Also, is there any way to know at this point whether a paper has been recommended for the IJCV track, or do authors only find out later through a separate email? Would be great to hear from anyone who has gone through this in previous years! submitted by /u/Secondhanded_PhD [link] [comments]