Nothing matches those filters.

Lead

10

Video

4
17:08

MINIMAX H3 IS THE NEW FREE VIDEO AI KING!

A new open video model is being called the best free video generator yet, running locally and making clips from text, images, or several references at once. MiniMax H3 is a 33-billion-parameter model with native audio that outputs up to 1080p video with sound, and it still works on 8GB GPUs. It can edit existing videos, interpolate between a first and last frame, and take up to nine references including images, video, and audio at the same time. A Turbo LoRA, still in training, already cuts generation from 20 steps to about 8.

Notes

MiniMax H3 — Local/ComfyUI Video Model (Aitrepreneur, 2026-08-06)

Model basics
  • MiniMax H3: 33-billion-parameter video model. Open weights, runs locally. Generates video from text, single/multiple images, video references, and audio references. Outputs high-res video with native audio.
  • Described as best free/open-weight video AI; does text-to-video, image-to-video, first-frame/last-frame in-between generation, multi-reference generation (~9 references at once), and video editing (replace a character in an existing video).
  • VRAM: heavy; needs lots of VRAM for high-res at decent speed, but works down to 8 GB VRAM.
Installation (2 routes)
  • One-click installer (Patreon-gated) — double-click; auto-installs ComfyUI + all required models and nodes.
  • Rent GPU on RunPod, use dedicated RunPod installer; runs the same as local.

Workflow ("MiniMax Ultra") is Patreon-only; drag-and-drop into ComfyUI. User prompt guide note included in workflow; suggested workflow: paste the prompting note into ChatGPT and let it write prompts.

Speed-up options
  • Stage attention and spectrum speed up generation at a small quality cost (spectrum costs more).
  • MiniMax Turbo LoRA released during video editing — still in training, only 500-step version exists; requires special ComfyUI version. Creator tested it, "very good for a LoRA that is not even done training."

| Setting | Without Turbo LoRA | With Turbo LoRA |

|---|---|---|

| Sampler | Res multi-step | Euler |

| Scheduler | Simple | Beta |

| Steps | 20 | ~8 (or less) |

Workflow on Patreon already updated to support the Turbo LoRA at time of upload.

Resolution / duration settings
  • Two resolution controls: a resolution selector (auto-picks per aspect-ratio guide) or manual resolution entry.
  • Recommended resolutions: 480p, 720p, 1080p — "more standard, much easier to upscale." 720p-then-upscale beats native 1080p generation (quality gap vs 1080p is small; generation is much faster). A 5090 can do native 1080p.
  • Duration input in seconds; longer = slower.
Capabilities demonstrated
  • Text-to-video: single-shot generation, examples incl. "I'm dead, Jerry" clip, Breaking Bad baking-bread parody, anime parody ("king of the hook"). 1080p vs 720p difference described as minor.
  • Image-to-video: single image + prompt; creator prefers generating images in a separate model (e.g. Create 2 / "Krea 2") then animating. Option to use original image size or normal resolution selector (demo ran 720p).
  • In-between generation: first frame + last frame inputs → model fills the transition with one prompt (warrior sitting → screaming example).
  • Multi-reference: up to ~9 references mixing image/video/audio. Demo: 2 character sheets + castle interior → two characters with dialogue, two distinct voices. Default workflow has 2 image refs, 1 video ref, 1 audio ref. Add refs by copying a node (Ctrl-C/Ctrl-V) and wiring to the reference node; bypass any ref via node toggle.
  • Video editing: video reference + replacement character image (no LoRA/training) — e.g. swap a woman in a video for another character. Single-video-as-reference: re-voice/re-dub a Sopranos clip (Paulie), audio reference lets it copy the character's voice.
  • Creator used ChatGPT for all prompts.
Upscaling

Recommended: LTX 2.3 Ultra workflow v3 (also Patreon) — includes video enhancer that upscales AND enhances simultaneously. 720p MiniMax output → 1080p full HD. Faster than generating at 1080p natively.

License caveat (important)
  • License is unusual: cannot use the model if you are in the European Union, United Kingdom, Korea, or the United States — attributed to MiniMax's ongoing lawsuit against Disney; company avoiding further issues.
  • Workaround: click the link in the UI and sign a waiver (~30 seconds) granting usage rights. Creator signed it because he represents companies; "if you are a normal person, you don't necessarily need to do that."
Caveats / limitations
  • Prompts must follow MiniMax H3's specific prompting style; note in workflow, or use ChatGPT.
  • High-res generation needs serious VRAM; 8 GB only at low res/slow.
  • Speed-ups (esp. spectrum) reduce quality.
  • Turbo LoRA unfinished (500-step), quality will improve.
  • Creator: "I barely scratched the surface" — update videos planned.
Transcript · 20,175 chars
Minemax H3 is the best free video AI king ever, and it's not even close. >> [screaming] >> Hello humans, my name is Kay Ovrlod, and boy oh boy, do I have some mind-blowing stuff for you today. Because yes, you heard it right, we have a brand new video AI model that was released called Minemax H3. That is simply the best free video AI model ever made, capable of generating videos from text, images, as well as from multiple other references at the same time. It is just incredible. So, today I'll show you how to install it, how to run it locally, and on RunPod, and show you how to get the best results possible. So, that being said, sit back, relax, and let's go. And to install Minemax H3, you have two ways. The first is of course by using my one-click installer that is available for my Patreon supporters. Just double-click on the installer, and then it will automatically install ComfyUI and all the models and nodes that you need to run Minemax on your computer. And the second way is to rent a GPU on a website like RunPod, and use my special RunPod installer, and run Minemax as if this was running on your local computer. And then once you have ComfyUI up and running, for this video I prepared a special Minemax Ultra workflow that you can find on my Patreon. Then you're going to drag and drop inside ComfyUI. And now we can finally have some fun. Now, before we begin, try to explain the workflow, what makes this Minemax H3 model so special, and why is everybody and their grandma talking about it. Well, Minemax H3 is a 33 billion parameter model that can generate videos from text or images in high resolution and with native audio, and that can run on your local computer. But that is not all. This model is extremely powerful and versatile thanks to its ability to do video editing and using multiple references at the same time to generate your final video. Now, obviously, Minemax H3 is a very thick boy, and you really need a lot of VRAM to generate videos at high resolution at a decent speed. But don't worry, even if you have 8 GB of VRAM, the model will work as well. Okay, so enough talking, let's actually show what this model can actually do and see how good the video generations are really are. Okay, so first, let's start with the text-to-video workflow, which is by far the easiest to understand. Here, once again, all of the most important part will be located in the first column right there. And before we start generating, we need to look at the available speed-up options that are available for the model right now, which are stage attention and spectrum. Now, basically, all of these two allows you to generate your video much faster at the expense of a small hit in video quality, especially if you're using spectrum. Now, there's also plenty of other nodes that can do that, but stage attention and spectrum are by far the best. Oh, and by the way, speaking of speed-ups, this is entrepreneur in the future because as I'm editing this video, we actually already have the release of the MiniMax Turbo Laura. That's right. Now, obviously, right now, it is still a Laura in training. It is still not done. As of right now, the only version available is the one at 500 steps. And to make it work, you need to use the special comfy UI version. And I've already tried it and it is very good for a Laura that is not even done training. And although in this video, you will only see videos generated without the Laura, just know that if you're watching this video right now, the workflow on Patreon will already be updated to support this Laura. And you can either choose to run it without the Laura. And basically, the differences between the workflow is very simple. Without the Laura, we are basically using the res multi-step sampler with simple scheduler and 20 steps. Whereas, if we are using the MiniMax Turbo Laura, we will be using the Euler sampler with beta scheduler at eight steps. That's it. That is pretty much the only difference. So, instead of generating a video with 20 steps, you are now able to generate your video in around eight steps or even less. And of course, as I said, this Turbo Laura is not even done training, so in the upcoming days, we should have an even better Turbo Laura that we can use to generate absolutely amazing videos. So, yeah, there you go. Okay, so then right here, this is where you're going to input your prompt for the video generation. However, keep in mind that MiniMax H3 has a very, very specific prompting style. And I'm not going to go into details on how it works and everything. If you want more info, you can simply either read the note I have input right there, or if you want to simplify your life, you can simply go there and then copy and paste this entire note inside software like ChatGPT, and there you go. And now you can simply just ask ChatGPT to write you a prompt for whatever video you want. Oh, and also, right before you generate, let's also have a talk about the resolution, because here you have actually two options to choose the resolution of your video. Either you use the resolution selector that will automatically select the resolution following the aspect ratio guide right there, or you can simply just leave this option enabled and then choose yourself the resolution that you want. Now, if you don't want to type anything yourself, it is much easier to just use the resolution selector. And also here, I have highlighted all the resolutions that I recommend you to use, and it is simply either 480p, 720p, and 1080p. So, depending on your GPU, you might want to choose either one of these three, because these resolutions are more standard, so they are much easier to upscale. But obviously, if you have like a 5090, you can definitely go to the 1080p and then generate your huge video from scratch. And then here, finally, this is where you input the video duration in seconds. And of course, the longer the video, the longer it will take for the video to be generated. So, yeah, I mean, very simple stuff. If you've done it once, you've done it a thousand times. So, let me actually just generate a video. Let me write a prompt. Then I'm going to leave the 0.9 megapixel resolution, use 10 seconds instead, and then I'm going to click run, which in the end will give us something like this. >> I'm dead, Jerry. The city says I'm dead. Well, that should clear up the parking tickets. Good news. I got you a box. >> So, yeah, I mean, it's pretty good. Now, as you can see, like this is very, very decent. Now, this was generated in one single shot. This is not cherry-picked or anything. I mean, it is really, really good. Actually, I even generated another version with the same exact prompt, but this is not at 1080p. Take a look. >> I'm dead, Jerry. The city says I'm dead. Well, that should clear up the parking tickets. >> [laughter] >> Good news. I got you a box. >> So, as you can see, in terms of quality, except resolution, there aren't that many differences between the two generations. Hence why it's probably better to generate at a lower resolution, like 720p, and then upscale it to a higher resolution later. And of course, the model can do a lot of very cool stuff, like a Breaking Bad parody, for example, or memes, kind of like this. >> Jesse, we have to cook. >> Yeah, there you go. Like, super, super funny uh Breaking Bad, Baking Bread parody. We all know that. It's not new, but it's still really, really fun. And of course, it can do more than just show parody or real-life stuff. It can generate anime parody as well. Like, take a look. >> I'm going to become the king of the hook again. >> Yeah, I mean, this is really, really incredible. Like this is insane. Like you can make your own anime parody or anime stuff already immediately right now. Like we've just a simple text prompt. Now this video was generated at 1080p, so you get like the best quality possible. But if you don't want to generate at such a high resolution, you can also generate at 720p. And if you do, you get something like this. >> I'm going to become the king of the hook. Okay. >> So yeah, I mean as you can see, very decent as well. Maybe not as good as the 1080p resolution, but it's still really, really fantastic. I mean it's really, really cool. So yeah, I mean I could spend hours and hours generating new videos to show you how good it is. But I'm sure that by this time that you're watching this video, you've already seen multiple examples anyway. So let's move on and see what else this model can do. Now the second thing that this model can also do is that instead of generating videos from text, you can generate a video from a image, which is definitely something that I prefer. I do like to generate my images on a separate model like Create 2, and then using an image to video model to bring it to life. I find this much, much better. So here basically the principle is exactly the same, except that now you have the choice of inputting two different images. Either you input one single image and you just make that image into a video, like I'm going to show you right now. Let's say I input this image, this selfie of a warrior on a battlefield, then I'm going to write my prompt. Then here I either have the possibility of using the original image size, or if I don't want that because it's going to take too much time, I can just use the normal resolution selector and then generate at the resolution written right there, which is going to be 720p instead. Input the video duration and now if I click run, which in the end will give us something like this. >> Well, I guess I'm going to be late for bingo tonight. >> And here, I mean as you can see, like this is fantastic. This is really, really good. This is great. I mean, huge, amazing quality. The generation is fantastic. The people behind him walking, you know, sitting down and everything. I mean, this is really super, super impressive. Super realistic, super good. I mean, yeah, I mean, this is really, really good. And of course, this is not the only thing that you can do with this particular workflow. And that is because you also have the ability to input a second image. Because this workflow can also work as an in-between video generation. So, if you put a first frame and a last frame, you will be able to generate a video between those two beginning and end frames. So, like for example, if I enable this last frame and then I input frame one with the warrior sitting on the ground, the last frame where he is shouting, if now I write my prompt and I click run, which in the end will give us something like this. >> [screaming] >> So, yeah, I mean, as you can see, this is really fantastic. I mean, we used only two separate frames, one beginning and one end, and with one single prompt, Minimax made a video in between those two frames. And I mean, I mean, what do you want me to say? It looks really, really good. It's it's fantastic. I mean, LTX could do it, too, but Minimax is really just on another level. It is really super, super good. But of course, once again, this is not the only thing that this model can do. Because not only can do all of that, but more, because Mini Max also has a very specific video model that allows you to use several different references to make your video. And the way it works is very, very similar to the image to video workflow, except that this time you have three different types of references. You can actually use image reference, video reference, and also audio reference, which really makes this workflow and model extremely powerful and extremely versatile. Now, the way you do this is the same exact principle as what we did with the image to video, except that this time it's kind of up to you to decide how many references you want to use. And I think the limit is around nine references that you can use at one single time, which is really just insane. Now, by default, I have input here like two different references for the images, one reference for the video, and one reference for the audio, but if you want to add more or disable different references, you can do so very easily as well. So, like for example, let's say that I want to add an additional reference image, all I have to do is just click on one node, then press control C on your keyboard, then control V to copy and paste that node, and then what you're going to do is that you're going to click and then connect it to that node right there. And you're going to connect it to the reference image node right there. And as you can see, now we have three different reference images that we can use. And you can do the same thing with video, audio, it's the same exact principle. And if you don't want to use a certain reference, you can just click on the node, and then click on this little button to bypass the reference, as well as right there. So, like for example, let's say that I upload an image of this character sheet right there, then another character for reference image two, and then for the third image, the interior of this castle right there. And now, if I write my prompt, which once again, I made it with ChatGPT, I'm not going to write all of that myself. Like, no way. And now, if I click run, which in the end gives us something like this. >> Wow. So, it can generate multiple characters inside the video? >> Yep, it looks like it. >> Amazing. >> Yeah, it sure is, buddy. I mean, this is really just incredible. I mean, it's insane. I mean, >> [laughter] >> I mean, I'm I'm kind of shocked. I'm going to say it's It works so good. It worked better than I thought it's going to be. It really like took the characters from the sheet, the character sheet, all the images, and then made video like this with full dialogue interacting with two different characters, two different voices, and I mean, it looks so good. It's It's amazing. I mean, yes, it was generated at a lower resolution, so you can have even higher quality if you generate at a higher resolution, or if you upscale it. But, I mean, just like that already, it is It is really, really good. I mean, it's It's incredible. I got to say, it's insane. It's fantastic. You don't need any LORAs, you don't need any training. Everything is already done for you. Just Just incredible. And also, of course, there's plenty of other ways of using this, using different types of references. Like, for example, let's say that I want to use a video reference. Let's say I upload this video of this woman kind of like turning around, and I want to like edit this video. I want to replace that woman with a different character. Let's say I want to replace it with that girl. I'm actually going to upload that girl into image zero. Then, I'm going to disable all the other references. I'm going to write my prompt, check the video direction. And now, if I click run, which gives us something like this. I mean, listen. I mean, is it amazing or what? What do you want me to say? It's It's incredible. I mean, we literally just like took a video reference and then replace it with a character from an image. No Laura training, no anything. We just like put everything together and it just worked. It just works. All right, it's Huh? It It's I'm I'm blown away. I'm blown away. If you can tell, I'm blown away. It is incredible. I mean And obviously, this is once again not the only thing that we can do. We can do much, much more than that. We can do so much more. I can spend hours and hours and hours showing you everything that you can do. Like for example, you can even like simply use like one single video as a reference. If I use like this video from, you know, Soprano with Paulie talking, which I'm not going to play because I kind of want to avoid any, you know, copyright issues if there is any. But what I can do instead is just use that video as a reference and then make him say something completely different with my prompt. And because the video not only can reference the video itself, but also the audio, we should also be able to copy his voice fairly well. So now if I click run, so that in the end we get something like this. >> Wow, this model is incredible. If Tony hears about this, he's going to be pissed. >> So yeah, I mean, what what do you want me to say? It's It's great. It's fantastic. It's It's everything that we've ever wanted from a video model. It's It is that simple, guys. It is that simple, okay? Like Why Why Why are you still here? Just Just stop listening to me and, you know, like and and do your thing. Do your thing. Try it out. It It's It's amazing. It's amazing. Oh, and also one little trick. If you want to upscale your video from a lower to a higher resolution, I highly recommend using my LCX 2.3 ultra workflow version 3 because in that workflow I have added a video enhancer upscaler, which is actually really really good at enhancing and upscaling at the same time your video. So, you can literally just upload a video made with Minimax. This is a video that was made and generated at 720p, and then you can like input a resolution of full HD, for example, and then if you click run, it will use LCX 2.3 to take this video and then upscale it at a higher resolution and enhance it at the same time. So, that in the end we go from this, a 720p video that looks like this, there we go, to a 1080p video. Yeah, I mean, there is really a huge difference in quality between those two videos. Hence why this workflow is really insanely good because not only we are upscaling, but we are also enhancing the original video. And obviously, it takes way less time to generate at 720p and then upscale it at 1080p rather than generating the full video at 1080p instead. So, yeah, I definitely recommend you to use this upscaler whenever you want to upscale one of your videos made with Minimax. Now, the one a little bit of a weird stuff with this model is the fact that the license is very very strange. Basically, it says that you cannot really use the model if you are in the European Union, United Kingdom, Korea, or the United States, which is, you know, very very strange. But, that is because Minimax is currently in a lawsuit against Disney, and they kind of want to avoid any issues with the model. Now, if you want to make sure that everything is fine, you can simply click on this little link and just like sign a waiver giving you the right to use this model however you want. It takes like 30 seconds to do and then you're good to go. Now, obviously, if you are like a normal person, you don't necessarily need to do that. I you know, you're fine. But in my case and a lot of people's cases, since we do have companies and we do represent companies ourselves, just to make sure that everything is fine, we need to kind of go through this, sign that waiver, and everything is going to be fine. So, yes, YouTube, if you're watching this video, just know that I did receive confirmation and I do have the right to use this model. Thank you very much. I love you. No sarcasm intended. So, yeah, there you go. This has been Min Max H3, simply the best open weight video AI model ever made. An absolute beast of a model that you can run on your computer right now. And once again, in this video, I barely scratched the surface of what this model can do. And even with my version one of the Min Max workflow, you can pretty much do absolutely insane stuff. But don't worry, this is definitely not the last time that I'll be making an update video about this model. There is really too much to say and too much to do. So, yeah, like once again, stop wasting time, download this workflow, download this model right now either locally or on RunPod, and just use it. It is incredible. Insane. So, that you and I can finally make our dreams come true. >> [music] >> And there we have it, folks. Thank you guys so much for watching. Don't forget to subscribe and smash the like button for the YouTube algorithm. Thanking also so much to my Patreon supporters for supporting my videos. You guys are absolutely awesome. You people are the reason why I'm able to make these videos, so thank you so much, and I'll see you guys next time. Bye-bye.
13:34

The AI Agent Every Company is About to Build | Vercel CEO Guillermo Rauch

Vercel's CEO says the next big thing isn't a website but the agent you build before it — a company's brain that any employee can talk to. The company runs an internal agent called V in Slack that about a thousand people use; it routes to sub-agents like a content writer and a data analyst wired to the warehouse, and can even delegate to Codex or v0 prototypes. Rauch argues that building and tuning these internal agents should be a company's real intellectual property, and says Vercel picks the best model per task rather than locking into one.

Notes
Source & format
  • Podcast: Agent Native episode, Riley Brown (host) interviewing Guillermo Rauch, CEO of Vercel. Published 2026-08-06. Self-describes Vercel as "multi-billion dollar company."
  • Form: ~1hr-ish interview; Rauch is promoting Vercel's agent framework EVE. Content is product vision plus one concrete internal case study (Vercel's "V").
Vercel's internal agent "V"
  • ~1,000 people at Vercel use "V" via Slack (@V). First built as a support assistant, then expanded.
  • "V" = internal agent; "Adversel" = customer-facing agent; naming shorthand noted self-deprecatingly ("V, Eve, Vercel… we're super creative").
  • Designed as a router/orchestrator: queries about docs/knowledge base route to one sub-agent; support cases route to an agent with ticket access. It can delegate to external tools — "it could delegate a task to Codex if it wanted to, it can create a prototype with v0, it can query Vercel to get information about our production systems."
  • Named sub-agents mentioned: a content agent (marketing copy) and a DZero data-analysis agent "connected to our data warehouse" — "the nexus of intelligence within our company."
  • Proactive behavior: every Monday V delivers a report of "what's happening across every product area" and key metrics; nightly jobs read social media, parse keywords, draft content.
  • Feedback loop: Slack thumbs-up/down on every V response; a nightly job aggregates negative feedback and proposes self-improvement. Vercel once shipped an eval for verbosity after internal complaints the agent "was too verbose."
The EVE framework
  • Positioning: "Eve is sort of the Next.js or React, what they did to the web" — a framework for building your own company agent; open-sourced the framework used to build V.
  • "An EVE agent at its most basic is a folder with an instructions.md file in it. So it's like the soul of your agent." Explicitly modeled on OpenClaw's soul.md.
  • Structure is file-system-based: instructions.md defines the agent's identity/context; a tools/ folder exposes capabilities (e.g., a wordpress.ts tool to read/write blog posts); a skills/ folder holds style guidance (e.g., content-writing.md). Rauch repeatedly stresses "it's just a folder."
  • Supports WhatsApp, Telegram, Slack, Microsoft Teams, iMessage as channels.
  • Vercel Connect: "access to 100+ systems" via event subscriptions — e.g., Stripe failed payments, incoming email. Framing: "you start thinking about the world in terms of events." Agent governance is the hard part: define tools, human-in-the-loop approvals, data access controls.
  • Agent evals: treated like unit tests; evals can cover factual accuracy and personality. Self-improvement loop proposed: "it's very important that humans are still involved in that loop."
OpenClaw lessons (explicitly credited)
  • Coding agents are "raw intelligence plus every tool at its disposal… it can write code, it can run it, and it can have access to everything."
  • soul.md: "you're not just taking the off-the-shelf agent… Claude is Anthropic's agent. It has its own set of principles… it's not truly yours. It doesn't have a soul of its own."
  • Giving an agent a computer (the "Mac mini thing") "massively improves its performance, its reasoning performance." Analogy to hiring a knowledge worker ("Here's your computer… logged into all of your key systems"). Vercel's twist: serverless agents that hibernate when idle — "a Mac Mini that hibernates." Whether an agent runs on one or a million computers should be "completely inconsequential to the end user."
God agent vs. team of agents
  • Rauch comes down on the god-agent side ("it's more on the god model"): one ambient orchestrator ("the Star Trek computer… ambient computing") that routes rather than a toolset the user targets — Vercel has "hundreds if not thousands" of internal tools staff don't know about.
  • Access control is per-user/per-team and "extremely business specific": small startups give near-equal read access ("most of the first 10-person team had read access to almost everything"), regulated businesses need stricter identity management. Metaphor: a preconfigured corporate phone issued to new hires.
Claims & predictions
  • Killer apps of the agentic era: coding/software-building and the "run your company better" / brain agent.
  • > "even before you build a website, you're going to build that agent that's going to help you build a company. It's going to be your factory."
  • Meta-work framing: after shipping "Claude slop," the fix is editing the agent's skills, not scolding the intern — "we're not working on the blog post itself."
  • Democratization: everyone can build the first agent block ("I could have used any service on the planet drag and drop"), but core governance/data-flow/security needs "people that really understand data flows, threat models and architecture."
Caveats & context
  • Interview is a product-launch conversation for EVE; no pricing, no launch date, no usage/revenue metrics given (revenue impact only asserted: "hundreds of millions of dollars of revenue are dependent on the well-being of this agent").
  • Riley Brown's own limitations: can't yet figure out team-level permissions (his personal skills involve personal email/connections he can't share); finds current interfaces opaque — "no one's cracked the interface for this yet," even technical people struggle to understand EVE.
  • Both agree the "agent with a computer" concept is confusingly communicated today (cited OpenAI's GPT Work marketing difficulty); Rauch says end users shouldn't need to care.
Transcript · 50,436 chars
really important thing about open [music] claw which is soul.md. So it's like the soul of your agent that's going to help you run your company. For example, what I believe will happen in the future is that even before you build a website, you're going to build that agent [music] that's going to help you build a company. Most of the world still thinks about agents as something you prompt. Can we sort of automate even the prompting such that the agent can be doing useful work for me while I'm not in the computer? Today I'm having a conversation with GMO Roush, the CEO of a multi-billion dollar company, Versell. And today we're talking about agents, specifically how companies are [music] using agents within their business. In this video, we talk about Verscell's internal agent that almost 1,000 people use within the company. We also talk about whether companies need one god agent or a team of many agents. We also talk about the challenges of setting up agents right now and how to get started building agents that actually improve your business workflows. We also talk about open- source models like Kimmy K3 and a lot more. My goal with this conversation is to answer [music] the following question. How do we as business operators, employees, and individuals use AI agents to [music] be more productive? And if you like videos like these, please consider hitting that like button and subscribing to this podcast. It helps me out a ton. Let's dive in. GMO, thank you so much for joining me on this on this episode of Agent Native. >> It's great to be here. >> My first question to you is, you know, obviously we have all these models coming out, right? You we have Kimmy models from China, models built in the US, Claude, Fable, now Claude, Opus 5. Um, we have all these different platforms people can use. And my audience are most people are business operators. They work in a big company. They want to use agents in their business to become more efficient and to become like a better team. Where are companies at in terms of implementing AI agents in their business? >> Yeah. When I think about we can call it the agentic revolution. Um just like any new platform that has hit the internet or the software landscape, you think about the killer apps, right? when the personal computer came out, you know, what were the killer apps? The word processor, uh, you know, um, uh, for some of us playing video games on our personal computers and things like that. Then mobile came along, right? And, um, I think the killer app of mobile in many ways was um, you know, not only shrinking interfaces from things that we used to use and putting them in a smaller screen, but enabling entire new use cases. And I think with agents we see a similar thing. So number one clearly one of the killer apps of agents is uh building software and building software or you know what you could call coding agents happens to be a core capability of solving a number of knowledge worker tasks because when you think about okay I'm I'm preparing a presentation for somebody you occasionally will say well we have to do some data science over here in order to then you know get a report or get some data back and put it into a slide or you'll say I'll automate a bunch of different steps and summarize some documents and then I'll put some other information into a slide and and so I think clearly one of the foundational parts of this uh new period of time where we do a lot of our work increasingly with agents is coding as a capability and I think that's this has transformed everyone's jobs right you can think of it as a number of sort of um levels of expertise I guess when it comes to coding so there are people like myself that can do agentic engineering meaning you know I've been programming for 20 years and now if I sit down and and face a really hard engineering task I will use a coding agent to enhance my engineering then there's this new emergence of what you would call vibe coding right which is everybody building a prototype of software or even a full stack application depending on sort of where your ambitions are and maybe even how ambitious the application itself is and so you have products like Vzero and uh lovable and things like this that are making it more um I guess they're democratizing building software or even building uh the the creative act or enhancing the creative act of coming up with new software. I also think agents are uh one of the killer apps is what I would call the run your company better agent or um the um knowledge base plus data analysis plus um uh project management agent. the the sort of brain agent that uh sits alongside of you and disseminates knowledge, business intelligence, even day-to-day tasks like you know who should I talk to within the company that is an expert in a certain task like navigating the org chart, navigating the what is to many overwhelming amounts of information that reside in the internal systems uh of a sort of think of this as like making the company's backend more efficient. And as I mentioned, I think coding is this omniresent capability. So to give you a concrete example from within Verscell, what we noticed pretty quickly is that um anybody that's helping a customer, anybody that's trying to close a sale, anybody that's even building new software needs to needs to ask questions about, you know, what are our customers doing? When did they first reach out? How much um uh time do we spend with them? How much do they use our platform? How many SQS of our of Verscell does this customer use? And so this internal brain agent has sort of emerged as uh I think one of the killer apps of AI. And maybe for a lot of people this still seems foreign like what are you talking about? There's an agent that can run my company. Uh so excited to make that more of a thing. >> It all sounds like amazing in theory, right? Like the this brain agent that everyone at a company can talk to. It kind of understands kind of the SOPs and the rules of the company, the best practices, that type of thing. And I've been trying to implement this, you know, I have a nineperson marketing team now that like helps me create content on my channels, on other channels. And my question to you is like, you know, for me, when I use AI personally, I'm inside codeex. That's just the tool that I've been using because I think it's good for knowledge work because I can ask it to create basically any type of document or something and it'll kind of open up in the side window. But what I can't figure out personally is like how do I enable this for a team? You know, if I were to onboard someone new and I want them to have access to my skills and but also like a lot of my skills involve my personal connections like my personal email. So they can't actually get access to that skill because there's all these like permissions that I need to keep separate, but then at oftent times I want them to be able to use the same skills that I can. And so I'm wondering like at Verscell like are you guys kind of trying to create this internally and like how do you get across these barriers and like what is the actual interface of using agents within a team? Yeah, even if you have a team of 10 people or a team of hundreds of people like Verscell, I think the way that I think about tools like Codex is that or CHBT is they give you a taste of what AI can do. But your job, the new job of someone that runs a company is to actually enable their workforce with agents and to work on the agent. I think the future of what you would consider to be your intellectual property at the company or your edge against competitors is the ability to create, tune, optimize and disseminate these agents internally and and and you know while you can have this sort of aha moment when you use something like Chad GBD uh maybe to give you an example of our internal agent is called V. So anyone within Versell can go into our Slack workspace and say at V and sort of navigate their day-to-day whether it's you you give a great example. So if I need to create new content for example our marketing team needs to help u promote a new product that we worked on or communicate a product change or write an engineering blog post in in collaboration with an engineer that worked on a certain capability. All of this goes through this V agent and this V agent has a number of skills that we continuously sort of update and improve. It has sub aents. It has sort of imagine the ability to create like a virtual uh employee team. So there is the content agent that is really good at writing marketing materials. There is the data an analysis agent. Um it we we internally call this D0ero but it's one of the it sort of think of it as like the the nexus of intelligence within our company like anytime we need to get information about how a customer is doing or um you know how they could use more versel or things like this we have this sort of DZero agent that is connected to our data warehouse um and so the experience of using an agent actually ends up being extremely user friendly why because all you need to do is you join Verscell, you join our chat workspace and now you sort of have this omniresent intelligence that can help you. And now you might, you know, you might go to V and say, "Hey, can you change um uh some information on the website?" And so V can still sort of coordinate with other agents. It could it could delegate a a task to codeex if it wanted to. uh if it can create a prototype with v0ero it can query versel to get information about our production systems but I think what's um what's key is enabling every company in the world to sort of deploy this brain and this intelligence and continue to sort of optimize it over time >> I have a lot of questions based on this my first one is do you have like a team that manages V like that where okay you have a team what what does that team >> look like how big is it and like what do they do on a day-to-day basis? >> So maybe to back up I wanted to share a little bit about our product development philosophy at Verscell. Um when we have a vision of the future uh that can be informed by you know pains that our customers have or things that we notice internally could be better we try to solve that problem ourselves first. So this idea of let's have an agent that can help with every aspect of our job sort of emerged pretty obviously like you mentioned like anyone that uses chachd notices oh it can reason but chachd doesn't have access to my internal knowledge base and customer records and the set of best practices of how we build software etc and so the inspiration was anytime you talk to somebody could there have been an agentic intelligence layer that could have gotten you that information sooner. So that was sort of like the inkling, the inspiration for it. Next thing is how do we build this? And so Versell has built a number of agentic infrastructure services and tools, right? So we built the AI SDK that helps developers talk to any model in the world. uh we built um uh AI gateway which helps you get tokens from any model in the world at the end of the day you know what we realize is that okay if there's an agent like V I don't want it to necessarily be clawed or codeex or open weight at the end of the day the customer doesn't matter and ideally we autonomously choose the best model for each task so we almost thought of V as a superset of all agents in the world um and so we designated a few folks to sort of like try it out and build this conversational experience. First, it started out as a support assistant and that alone was extremely useful. Why? because we are hiring new people and also in Slack we talk to a lot of our customers and so anytime that you have a question about how Verscell works, we wanted to have an at Verscell functionality that could know anything about Versel and that itself was super super super helpful because it became sort of like this easy way of giving support to our customers. But the difference between an AI assistant and an agent is that an agent can do things for you. And so we started thinking in terms of skills and in terms of jobs to be done. So this uh we gave it a name. So V for our internal purposes. And so we wanted to have a clear distinction between the customerf facing agent the agent we give users of versell which is adversel and the agent that runs our company. So V is sort of the shortand for this. So we created the V team. The other thing we realized in this process, and maybe this goes at the heart of your question, is it's actually pretty hard to assemble all of the tools, all of the frameworks, and all of the infrastructure to make something like this happen and to improve it over time. And so that gave uh inspiration for us to we built V and then we shared the framework that we used to build it back to the world. We call this EVE. You might sound like we're super creative with our names. V, Eve, Verscell, but Eve is sort of the, you know, uh, Nex.js or React, what they did to the web. They made it really easy to build websites and web applications. The thing that I think every, uh, knowledge worker, every individual, every entrepreneur will want in the future is to have an agent that they can call their own. Uh, and this is, uh, what we're helping people enable with with Eve. >> Gotcha. Yeah. I think, you know, this is something I've spent a lot of time thinking about, like how do you give the normal person the access to not only just have an agent that has a bunch of context, but to also kind of like customize it. And I think although it feels like now that OpenClaw was kind of a fad, you know, you know, if you look at the Google trends, it's like gone way down. I do think it unlocked kind of a magic moment or there's a reason it went viral in the first place. It wasn't because there was some secret paid promos by OpenClaw. I think there was a genuine um desire for people to put an agent on a computer and let it do things for you. >> I had a lot of epiphies uh from Open Claw that informed the development of Eve. I think you're absolutely spot on. One of those things is that OpenClaw showed how just how much a coding agent can do. back to my uh initial point like what is open claw fundamentally it's the raw intelligence of the model plus every tool at its disposal right like >> a full full access yeah >> it can write code it can run it and it can have access to everything and that's magic >> to the point where it could do things accidentally and like I think that's I remember listening to Peter who created OpenClaw he said something like that he like asked for something and then it like gave it found an API key on his computer and it did something that he didn't even ask for. And I think that was kind of the magic moment. You know, they added like the heartbeat which was this thing that kind of like initiated it. >> Another really important thing about uh open claw which is soul.md. So when you when you create an open claw or when you use open claw you're not just taking the offtheshelf agent that somebody else built. Clearly Claude for example, it's a great agent, but Claude is anthropic agent. It has its own set of principles and and sure they will they give you ways to customize it and whatnot, but it's not truly yours. It doesn't have a a soul of its own, right? And so I think that was another really big unlock, which is what is the soulm file? It's just it's just literally markdown text that defines the genesis of that model. So when you create an agent with Eve, we we basically learn from that and and basically an EVE agent at its most basic is a folder with an instructions.mmd file in it. So it's like the soul of your agent that's going to help you run your company, for example. And then the other thing that we learned is it's awesome that it can run code, write code. It has a computer for it, right? Like the the whole like Mac mini uh thing was actually quite meaningful, right? Like people realized, okay, this Asian can do anything under the sun, but it's dangerous and it needs a space. He needs his own like thing and give the agent some space, right? like and so people bought Mac minis and and and that basically giving an agent a computer massively improves its performance, its reasoning performance and its ability to deliver outcomes for you. And so what's really fascinating is it's not too unlike hiring a knowledge worker. What is the first thing a modern firms does when they hire a human? Here's your computer. it gave us a MacBook. It has a bunch of programs installed. It's logged into all of your key systems and so we wanted to give you that as well uh for your own agents that you build. But we wanted to build a secure and efficient environment for it to run. And so the security part is that you define the tools, the human in the loop approvals and the data access controls for anything that the agent can do. And the other aspect of it is it doesn't assume that the agent is always running in a computer which is actually kind of uh counterintuitive. I just said an agent gets better if it has a computer but not every agent is a computer is running 24/7. And so in in in our in our lingo of the versel in in cloud world we call this serverless. The idea is that if the agent is not doing anything it can go to sleep. Maybe another metaphor is imagine a Mac Mini that hibernates when the agent doesn't have anything to do so that he doesn't use electricity. And so because we at Verscell we run you know billions of deployments we needed a mechanism such that agents can be very very very efficiently operated and run. Um and so that's another sort of ingredient that we learned from the the open clause of the world. Okay, if we're going to run these things at massive scale and we need to run them securely, how can we create infrastructure that enables that? >> Gotcha. That makes sense. Yeah. I think I think all of the the big AI labs who've re who are like kind of releasing a product that is an agent on a computer is trying to shake it into people. They're like this is a computer. It has a computer and it's not easy to communicate to the average people. you know, OpenAI is struggling with that right now where they're like literally tweeting. They're like GPT work is an agent with a computer and it's not easy to convey that as you interact with a chatbot, you know, like it's like what does that even mean, you know, and I'm even I'm even struggling with it, you know, and I think, you know, and I think you can kind of divide whether you look at Anthropic or OpenAI, like you can kind of divide their products into like how much computer access they have. It's like the chatbot doesn't have any computer. GPT work has some computer it doesn't it can't run terminal commands but then codecs can run terminal commands but you can only get it on your computer because they don't have and so I think that is actually the computer aspect of agents I think is one of the parts that makes it really confusing at this stage right now >> I agree and and my goal with uh the agents that we build in in V is that you know whether you're an intern that just joined Verscell or you're a super experienced engineer or you're somewhere in between I don't think whether I I think that's sort of the implementation detail that the agent builder needs to know about. You need to what I want for the future is that someone that's building an agent can very carefully define governance data access control in the security model right for because agents are interacting with customer data. So you can't just be like I don't know man it runs a computer and it has access to like all of the databases of everything. you have to be really really really thoughtful about it. That's literally our new job, right? Um and uh but whether it runs one or it runs a million computers completely inconsequential to the end user. In fact, you know, you can think of this agents as being orchestrators. In fact, when when someone goes to our Slack and says at V, they're really talking to the orchestrating agent, the one that could delegate a task to a million computers, to one computer, maybe even no computer. You know, we have customers of Versel that have built agents that have so much usage that they figured out ways to make the computer smaller and smaller and smaller just for the sake of cost efficiency. And so I I think the my hope for the future is that the very technical people can sort of know like oh this particular conversation with this agent resulted in all of this usage of computers and whatnot. But for the most part it's all about getting high quality outcomes, high quality analysis, high quality uh you know accurate information. Uh performance is becoming more and more of a the dog tug of town, right? like people really care for fast models and fast execution. So that's another aspect of like how do you get your agent to be delightful. >> Okay. So let let's say for a sec I wanted to create a V agent for my team. >> Yeah. My first question with this and and this is something that I've realized talking to a lot of business owners who are like kind of know about agents and they're they're trying they they're they're confused on whether you want one agent that's like a god agent that knows everything or if you want a team of agents that sort of like share a knowledge base because the conversation that I'm having with a lot of business owners is like well the marketing team has access to these things and the finance team like I don't even I don't even want the marketing team to know about certain finance documents. >> Totally. >> And so like that's my question is like how if I were to be creating my own V agent for my company, how do I think about that? You know, God agent or >> Yeah. >> So first of all, I'm a user experience guy. You know, I started Versell because I was frustrated with how slow creating software was and how slow the average website and web application experience was. So I always try to work backwards on the user experience. The ideal user experience with an agent is the Star Trek computer or the Iron Man Jarvis. It's ambient computing and I don't need to target a specific capability. That's why we have we're reasoning with agents to begin with. is like there's probably like hundreds if not thousands of internal tools that people at Versell have built that I don't even know they exist frankly there's just too much right and so when you have this intelligent agents they can act as routers v our internal EVE agent is a router so if you ask it about versel knowledge it goes to the you know capability that we have for looking up our documentation our knowledge base etc ETA if you ask about if you need to help a customer with a support case, it has a support agent within it that has access to our support ticket infrastructure. Okay, so that answers sort of my perspective is that it's more on the god model. And maybe to give you a metaphor because I really think that what we're doing here is we're redefining how companies of the future will work. When you join a a corporation, they might give you a corporate phone and that corporate phone is already preconfigured with your identity and with a set of applications. You have the application for the I don't know internal chat. You have the application for this and that. So I think the internal agent that helps you run the company is not not unlike that is the job of the new sort of IT department is to say what are the capabilities that we're bundling into this agent and also crucially how do we manage identity and who gets to access what information which is also extremely businesses specific. It depends on how regulated your business is. If you're a small startup, I can believe that you know your nine person team, they all have pretty equal access to most of the information of the company. maybe two have information to the financials or or maybe the distinction I remember when I started forcell was like some of us had you know read write admin [laughter] uh and but I think most of the first 10 personel team had read access to almost everything right um and so the job of the person that works on this foundational agent is to determine uh the the access control the tools uh um the guard rails uh the the audit trails and and and like I said, this is actually pretty hard work to do. Uh and and and why we wanted to create a framework that made that the fundamental job because you know wiring up the model, wiring up the infrastructure and all of that we can sort of customers can offload to us. >> That makes sense. And so I yeah I guess the agent would also be able to see where the message is coming from. So it's like okay if it gets sent in this channel it'll delegate um to this sub agent or access these certain files. That makes a lot of sense. >> Totally. I just get I guess cuz what you're telling me is like so appealing. Um like being able to create your team's agent and I don't think anyone's cracked the interface for this yet. Um and I know you guys are building a framework. You deal with a lot of developers. I guess what I'm dying for is like a way some sort of interface to understand it because even the technical people like I've even showed technical people Eve where like I'm like can you help me make sense of this and I think it's still at a stage where it's not super easy to like fully understand and so I guess yeah I just wish there was like an interface where I could go in and like set these rules. Maybe I'm talking to an AI and it's configuring it. I I guess >> the way that most of these agents are built is that you're talking to an AI that is helping you maintain your EVE project. You'll hear me use the word file system or folder a lot. I find that it's it makes the world really easy to understand if you think it if you think about it as a hierarchy of files and folders. So the way that a ne agent works is that you start with that instructions file that says you are the agent that helps run Riley's business. You can even have some context about who you are like uh our business isn't you know we disseminate information about AI and uh our values are transparency we're not opinionated and we love shipping things like something like that right um okay but that agent still knows nothing it's a tabula rasa it just has the raw intelligence that comes from the model and it has a a basic set of instructions how can it do something useful for you well you talked about okay let's help the marketing team create content and let's say that one of the things that you really care about is posting uh on your blog okay so in an EVE agent the first thing you do is you can create a tools folder and you can now start exposing tools to the agent and so you can say let's say that your blog is running WordPress or some system like that now I can say to the agent now you have a tool to read didn't write blog posts to WordPress. Okay, great. You you created that file WordPress uh.ts on on that folder and then you ship your agent. You you use the word channel also very important that this agent needs to communicate to your team in some channel. So Eve supports every channel under the sun. It can be WhatsApp, it can be Telegram, it can be Slack, it can be Microsoft >> iMessage, >> it can be iMessage. Yes. >> Amazing. And so the next question is okay I created the agent I gave it this sort of soul I gave it access to WordPress. Now you hire an intern. [laughter] Can the intern ship any blog post that it authors together with your internal agent to prod? You probably don't want that. And so this is the job of like you at some point maybe Riley you were working on your EVE agent or someone in your team you designate as sort of the agent administrator. You're gonna say okay if the person lives within cert a certain part of the organization we let them write directly to WordPress. Another approach that I've seen people take is that when they interact with the intern over Slack or over Telegram or whatever you have to authenticate with WordPress. So you delegate to an existing permission system that you already have. So the EVE agent ends up being sort of the facilitator of the transaction but it doesn't have direct access to WordPress itself. Uh it it it will help you sort of draft up the content. So this is just an idea that we cooked up in this conversation. But imagine that every day you start realizing hm that's really powerful. I just unblocked my entire team to be able to draft up blog posts that go directly to WordPress. But next time tomorrow you hear an escalation and you hear, "Hey at Riley, I just saw your blog post. It's I read your most recent blog post. It reads like complete claw slop. What do you do?" And you go you go into your team and say, "Guys, what do we just do?" We became really productive and we started shipping a lot of slop. You know what you do next? You work on the content writing skill of your EVE agent. And so this is the meta work that we will all be doing in the future. We're not working on the blog post itself. You did not go to the intern scold at him for like hey what what do you do? You shipped a bunch of slop. You're putting that intelligence into the agent in the form of skills in the form of tools. Uh and uh of course over time you can get more sophisticated and and uh it's not just about blog like how can we infuse the content writing capability with uh what people are saying on X about your business. >> I was going to say that like a lot of the skills that I find very useful for content ends up just being like grounding in some relevant source. And so you can put I call them like plugins like where like there's one called scrape creators. It's some API that I found that scrapes content from certain channels. And so like before it ever writes anything or before it ever idiates an idea for YouTube or or a packaging concept like a title and thumbnail, it'll go and like look on social media and find those things. >> Totally. >> Yeah. And that's another thing like okay so if I'm creating a V agent um yeah I'd I'd want to add certain APIs and you can add I would yeah call you can call them plugins or like how do we distinguish between plugins and skills? Can you add plugins to skills or how or are they all just skills? >> Text. >> So going back to you you got that escalation that says Riley, you just shipped you're shipping a lot of blog posts but they all they have too many m dashes. >> And so this is what's beautiful about that idea of it's just a folder. You go into your EVE agent and in the folder skills you say contentwriting.mmd and you say this is how we write. This is what I like. This is what I don't like. Um, you also talked about I I think that the future of work will be the agent becoming a lot more proactive as well. So, Eve can have a schedule. For example, every day at night, it reads social media. It parses keywords. It gets replies from your posts. And from that, it can do something. It can draft up new content. It can even give you a report inside of Slack. And this we actually have found to be extremely helpful at Versell. The idea that our agents proactively give us information. So every Monday I have a I have my internal agent give me a download of what's happening across every product area. What are the key metrics that I care about? So you can have the agent be doing thinking in the background on your behalf. And I think it's not just about I think most of the world still thinks about agents as something you prompt but I think there's a lot of alpha in thinking about can we sort of automate even the prompting such that the agent can be doing useful work for me while I'm not in the computer. >> I think one of the limitations for me and I've been able I I have a lot of automations set up that trigger an agent to do certain task and it is really useful. One thing that I'm struggling figuring out how to set up, especially at my at the team level, is to get outside things to trigger the agent, you know. Um, and there's many ways I think you could do this, but um, yeah, like do you guys have any of of that set up? Like if some event happens totally, >> it automatically Okay. Yeah. Can you talk about that? >> So, I think events that originate in systems like Stripe, like there is a refund request, we make it really easy to connect all those systems. In fact, when we sat down and we thought about what makes it really hard to build an agent, it's actually not the proof of concept part because anybody in the world can sit down open cloud coder codeex and build an agent in the sense that like when you're prompting it, you realize what it becomes capable of. What we talked about with open claw like the raw intelligence is already there. What's hard is securely connecting it to your systems. So we built a capability on Verscell called Verscell connect that gives your agents access to 100 plus systems but it doesn't just give them full readr everything access right away. It gives you the developer the control and that might mean that you subscribe to an event and then you send it to your agent. You can say, "Hey, every time Stripe has a failed payment, let the agent know. Every time we get an email, let the agent know." And so you start thinking about the world in terms of events. In fact, I mentioned that a lot of our agent interactions are happening on in Slack. Slack is just another event is someone said something and the agent that gets fed into the agent's brain. And so any any connector of this sort of repertoire of connectors can originate some kind of behavior in the agent. >> Gotcha. That makes sense. Yeah, that's just something we've been thinking about a lot. Um because you're right, everything is just an event. It's just things happening and then when something happens, if an agent can take care of it, they it should take care of it. And I think I'm like I've automated none of that in terms of what I could possibly automate, which is really >> mental a mental model. So I mentioned that the the thing that I'm excited about with Eve is that when when I started Versel the most imminent thing that I needed to build was a website like it felt like how do I put my fingerprint in the world? What is one of the earliest things that you do when you create a company? You register in Delaware if you're in the United States or even internationally you incorporate. You choose a name and so you register the domain name and you ship a website. even a website says like hey we're in business or welcome to the minimum viable sort of identity of your company on the internet. What I believe will happen in the future is that even before you build a website you're going to build that agent that's going to help you build a company. The it's going to be your factory. It's going to be the the trusted partner and advisor in everything you do that's constantly learning about the trajectory of your business. And so it's extremely critical that as you sort of evolve your business, this agent gets access to more of these data streams of knowledge and information. And everything really is an event in this world. Um, another important factor there is self-improvement. So when whenever you start a company, you're constantly learning. You're you're teaching your employees. you're helping them, you know, learn from mistakes, learn from incidents, learn from customer feedback, etc. It's going to be very important that your agent over time can improve. And so with Eve, we thought about, okay, if there is a baseline of information that your agent has, how do you evaluate the agent? Can you write tests or can you give it exams so that you actually know that you're making forward progress as you as this agent sort of gets uh um more sophisticated and more capable over time. >> Um and so >> think of this as sort of uh even more fundamental than the dot of your of your of your company. >> Yeah. And do you guys like put evals into Slack? Are there any ways to like evaluate whether an agent does well or doesn't do well? Like could you res like based on someone like could a employee who got a response from V could they say like oh this wasn't a good response and okay they can do that. >> Yeah. So the every response that we give on Slack has a and by the way maybe to also give kudos to the Slack team like Slack is kind of becoming like an agent operating system of sorts right because like it used to be for messages between humans now it's humans and agents and so they have built UI that is just really easy for the developer to add right so like the thumbs up thumbs down thing super easy to add and so every EVE agent we create for example at night we can have a job that aggregates all of the negative feedback >> and proposes the next stage of self-improvement. We can say hey >> we got five thumbs down on these answers. What are the things that the agent itself can even propose how to improve itself? Oh, I missed this. Oh, this person critiqued this part of my response or they said I hallucinated or whatnot. I do think it's very important that humans are still involved in that loop. But I think increasingly more and more of the job of get the agent getting better is also being done by the framework. So the framework itself comes with evalu um um you know are basically test cases right when you build a web application or a website you write unit tests and you make sure that the logic is sound when you create an EVE agent you write evals also to assertain that the logic is sound but that the information it gathers is is sound and and u uh it's accurate. There can be evals about personality. At some point we were hearing from people that our internal company agent was too verbose. It was speaking too much. Uh and so you we kind of basically gave it a better personality and and you can create evals around that as well. >> So do you do you view this like in the near future like over the next few years? Do you think it's just going to be mostly technical people building agents for companies or do you view this as something that whether you can code or not you you'll be able to create agents for your team? So because building software is being so democratized um think of it as like again let's go back to that idea of like I'm starting a company and like the first website I built is sort of like I could have used any service on the planet drag and drop uh give me a free website with my domain name like anything like that. And so I think >> that first building block of your agent everybody's going to be able to to create. I think over time, I mean, the whole business runs on this. Hundreds of millions of dollars of revenue are dependent on the well-being of this agent because our sales reps depend on it, our support team depends on it, I depend on it. And so you this is a very important piece of software. And so I think it's a combination of everyone can contribute to the agent information skills critique feedback and then there is engineers that are working on the core system loop the access to data the governance security all of those pieces um that I think need to be more technically minded but I don't think that the codew writing part is as important these It's I think I would describe it as people that really understand data flows uh threat models and architecture of systems design so that they can like carefully think about the the again the operational excellence of the agent and the security model of the agent. >> Very interesting. Yeah. Um because yeah, I think there's a lot of people, business owners, not all of them are technical, who are reaching out and they're trying to create agents. And so I'm just trying to like leave people with like a te a tangible thing that they can do like a point to a place where they can go to kind of build their first agent or build their V. Um because I think with what I've realized with these agent tools, all of them is I we we can have conversations about it. We can talk about it. I can learn. I can use AI to like learn about it. But nothing hits like doing it. And I think that's kind like like once you do it, then you're like, "Oh, I can do that. That means I can do this thing, this thing, and this thing." And like kind of your world opens up as you do even the most trivial things. And so, yeah, I >> recommendation there would be, you know, what I've seen give people an aha moment is create an EVE agent. Go to eve.dev, deploy your first agent, but connect it to your favorite chat medium. If you if your company works in Slack, connect it to Slack. If you like WhatsApp, connect it to WhatsApp. and pick one boring or you know kind of pick a toil task of your business that has a system to it but it's not you know it's something that if you could automate it away you'd absolutely automate it away and write down the scale of that task. Uh, it could be, for example, something we do a lot at Verscell is we put a lot of work into drafting up our product change log. When you go to versel.com, it says change log. Every piece of content there narrates the storytelling or evolution of our product. And in many ways, that change log is a grounding for my engineering team. How do I know if an engineer is being productive or not? or whatever like well one of the things that I do is I I measure it by have you shipped something that we can communicate to customers is an improvement to our platform. So one change log that's about to go out maybe by the time you watch this it's already gone out is we we improved the end toend deployment process of an application or agent to versel by 7 seconds. seven seconds we've shaved off uh over a lot of infrastructure work. So when you go to verschange you're gonna find that we improved our product and we shaved down 7 seconds. So it used to actually take a lot of work for an engineer that is in the depths of infrastructure to collaborate with a marketing team and get that thing out into the world. Because we have an agent internally, we've cut down that process into one Slack thread that the engineer creates. The agent refineses what they're telling me because, you know, engineers are sometimes so in the weeds that they struggle to communicate things in a way that is I call it contextf free. you know, maybe they start talking about, you know, computer science or like I'm just, hey, can we boil it down to the business benefit? Simple, seven seconds. Uh, it's enabled for every customer, it's free. So, that's kind of like a little formula that I have. People want to know what's the benefit, how much does it cost, and what do I do to get it? Mhm. >> And so that formula that I developed over many years of product marketing skill, I put into that Eve agent. And so for the listeners, think about something like that. Maybe it's like quote unquote a secret sauce of something you do really well, but takes a lot of time and you want to do more of it. And so start with that skill, connect you to a communication channel, ship it on for sale. >> Gotcha. Okay, that makes sense. Yeah, I think um to kind of I know we're we're running up on our time here, but um what are you most excited about um in ter could be a model, it could be computer use or some browser use. Like what unlock do you think we're going to get in the next like three to six months that will make using agents way more fun or way more effective? >> Very simple. Um cost of intelligence continuing to go down. M >> more intelligence to for more people, more variety of models. One of the great things about building with Eve and building in Verscell generally is that we give you access to every provider of models and every model in the world. >> It's model agnostic. >> Yeah, >> totally model agnostic, right? Um and that plays into your benefit because you retain ownership of your data, of your skills. You get to choose models and you get to benefit from the competition. There is some news that's going to go out tomorrow about models getting dramatically cheaper. >> Literally tomorrow. >> Tomorrow and if you were building in this way, you're going to benefit. Um so the other one is fast models are going to get way faster. I think we're going to start seeing what happened with the personal computing and mobile computing revolution, which is that, you know, we got the iPhone. If you were if you could travel back in time and or even pulled out the first iPhone out of a drawer, you'd be astonished at how slow it was, the refresh rate. Like you would open an app, it would do nothing for several seconds and then slowly at maybe 10 frames per second, the application would show up in front of your eyes, >> right? >> That's where AI is at today. >> Yeah. I think for most knowledge tasks, like I just want faster, you know? I my biggest problem isn't like, oh, I wish this was better. It's just like, why did I have to wait 14 minutes for this, you know? And like if it was 10 times faster, it it would be insane. And I feel like we're pro like >> how long do you think it'll take for the models at like a 5.6 level like um so like soul level >> days if not week. Well, I mean maybe days is the most like optimistic. Uh I I think we're literally like weeks single digit months away. >> Uh one of the data points that I can share is on the open weight and this is why I'm excited about open weight models. The competition between the inference providers around open weight is so extreme that GLM dropped. We added it versel AI gateway. It's an incredibly good model GLM 5.2. Within days, we had a fast variant that was four times faster. There we have more providers coming online for GLM that keep raising the bar of token per second performance. >> GLM 5.2 too fast is astonishingly fast and it's only getting faster. >> What did you think of >> Kimmy K3 is gonna happen to Kimmy? I think we're still in the early innings of that. >> What did you think of the model like in general? Like do you think it's you think it's really good? You think it's up to par with like an Opus 48? >> I think GLM 5.2 was already in that category. I think Kimmy raises the bar. I think Kimmy can do things that perhaps only you know uh fable class models could do. Not quite in all in all of its dimensions but for example when we evaluated it for cyber security it outperformed OPUS 4.8 8 clearly um and it was almost at uh you know soul level. Soul's still at the frontier. But again, this is the beautiful thing about having choice is that depending on what you're doing, you're going to choose different price performance ratios. Grock for fast and highly accurate. Like if I have to choose today a model that's going to be my workhorse model, that would be like the default. If I have an agent that is my Slack and needs to do a wide variety of tasks and has to do it quickly because there's another person waiting on the other side, I would absolutely go with Grog 4.5 or GLM in terms of like price performance. >> Um, now I mentioned proactivity. What about for example at night finding opportunities in our business uh crunching data and extracting novel insights for the executive team? Well, those things I can throw more reasoning power and it can take more time. >> I might even want to take throw a consortium of models at it. Why not have Kimmy and Saul and Grock come up with three points of view and then give you the summary? And this is why I find it so interesting, right? Like we're still in the early innings of understanding what are the principles of design and user interface engineering. But for agents, >> yeah, >> if I'm talking to an agent interactively, I want fast. If the agent is doing an asynchronous job, I want accuracy. >> Yeah. You don't care if it takes all night. Like it it doesn't make a difference if you're Yeah. Yeah. That's true. I I didn't think about that. Um >> we're about to launch a capability in AI gateway which is um you as a developer >> or even your agent can say please do inference please like get me tokens but in batch and I don't care how long you're going to take. like you communicate. It's a little bit like putting in a buy order >> and you're not worried when it gets fulfilled, right? Like you're just willing to wait >> and then anyone in this market can fulfill your order. >> Almost like a spot market for intelligence, right? >> That makes sense. Yeah. And uh and this is extremely exciting because you might say hey like come up with a proof or disproof the Jacobian conjecture for two dimensions and I don't really care when but spend this many tokens uh and someone at some point is going to say hey I already paid for the GPU it's connected to the internet yeah >> no one is using it let's throw some capacity it's a little bit like uh SETI at home uh for those who remember, rent out your spare comput capacity, solve hard problems. >> Yeah, because if you get it next week, it doesn't matter. You know, you're still solving a really crazy thing. Anyway, I really appreciate uh you joining. Um I think you guys are going to do great. I one thing I didn't realize is how much business owners don't want to get locked in to a certain provider. I mean, you know, like Claude Tag is their kind of I don't want to say it's their version of V, but it's like kind of an agent you can add to Slack. and so many people are resistant to it because they don't want to get locked into only Claude's models. Um, so I think that is something that you guys will have going for you. That's really cool. >> And it goes beyond, you know, the the model. I think it's not about having Claude in your workspace. It's about having an intelligence of your own, >> right? >> So there's almost like an element of like baptizing your agents like this is our agent. This is our company. It's, you know, I actually liken it to the web because the web was all about I own my domain name. Mhm. >> I'm the I'm the king of my own domain. Uh and I think we're now seeing we're living through the version of that for the intelligence age. >> 100%. Yeah, I agree. I thank you so much for coming on. This was this was a lot of fun. Let's do it sometime soon. >> Anytime, Riley. Thank you. [music]
14:05

Fable 5 for 3D Web Design is Next Level!!! (Interactive, Animated!)

A tutorial claims Fable 5 plus image and video AI tools can build 3D animated, interactive websites with no 3D software at all. The creator walks through generating a character with GPT-image, converting it into a rigged 3D model on a multiview site, and dropping it into a Claude-built page where the model follows the cursor. A second project turns a Pinterest reference image into an animated video clip, and a third assembles a full site in Cursor. It's mostly a workflow demonstration with no benchmarks, and the author is promoting his own prompt library.

Notes

Viktor Oddy — "Fable 5 for 3D Web Design is Next Level!!!" (course notes)

Auto-captioned transcript garbles some tool names ("Fable 5" = Claude Opus 5, "Grog/Crock 4.5" = Grok 4.5, "cling/Sydney 2.0" = Kling AI / likely Sora 2.0, "C student 2.0" unclear). Creator claims 100+ AI-built websites; no affiliate links; only tool he owns is Motion Sides (motionsites) — ~400 website prompt templates, +5/day, "a thousand prompts very soon."

Project 1 — 3D interactive character (robot)
  • Find image → Figma AI "edit with prompt" (GPT Image). Prompt: front view, standing full height, remove text/buttons/UI elements.
  • Generate front/left/back/right views; split into separate images.
  • Upload multi-view set to a "multi-view" 3D site → generate model.
  • Grab UI template from Motion Sides (search "robot"); build UI in Claude (Design tab) with Sonnet 5.
  • In 3D tool: Animate → select "good for humanoid" → add auto rig → export as GLB (lightweight).
  • Drag GLB into Claude (first tool rejected the file; Claude accepted). Rebuild in Claude, choose Opus 5.
  • Prompt: "replace the background image with this model and make so the cursor so the model follows the cursor."
  • Result in <5 minutes.
Project 2 — Scrolling video hero
  • Reference image from Pinterest; regenerate in Figma edit-image to change hand/pose but keep style ("keep it to be an Asian girl").
  • Video tool "C student 2.0": upload source video as reference + edited image → "create a video exactly as the video one. Do not change anything. 7 seconds, bit rate standard." (downloaded video was 7s).
  • One Motion Sides prompt generates the full site in Claude; then "replace the video in the first section to be this one" (upload or link).
  • Iterate: "make all of the content in the first section white except the halo S1 make it pure black, also the reserve button with the stroke pure black."
Project 3 — AI-car company site in Cursor
  • cursor.com/download, new folder, model Grok 4.5 (high/fast). Initial prompt: fictional AI-driven car company, 4–5 sections, pure-black background.
  • To restyle: paste a Motion Sides prompt but instruct "I will send you a prompt for unrelated websites. Do not implement it. Take the design... implement it into our current website... keep the copy/sections, only apply the new design."
  • Hero video from motions.ai + CSS blend-mode: exclusion so video background merges; remove headline.
  • Inspiration sites: landbook, xAI (copy real asset screenshots). Iterative prompts: center text, move text inside image, full-width, gradient edges for seamless blend.
  • Animate with Kling AI; video edits: saturation 0, opacity 70%, blend exclusion; reorder/delete sections by pasting section markup.
  • Note: "AI models like to add strokes/borders — not great design practice"; removed footer divider stroke.
  • Stats block: numbers at headline font size (e.g. 0.2s, 99.9%), 4–5 word captions; shrink 20% to preserve hierarchy; image border-radius: 30px.
  • Subpages (model/experience, 4–6 sections each) in max mode — reuses established design language. Full site in "a couple of hours or less."
Project 4 — Fintech cards site in Bolt
  • Screenshot inspiration (awwwards/Seesaw) → bolt.new → paste image, model max: "rewrite all text about fintech... mobile friendly bank... keep same number of characters, keep background black."
  • "Four horsemen of AI-generated websites": excessive letter-spacing, gradient+light background, green colors, gradients on cards — all absent when working from reference screenshots.
  • Interactions: components from twifersdev / "shaders.com", plus Motion Sides animated cards section; paste into Bolt.
  • Bolt Connector → shaders toggle to swap card videos for shaders; hover-interactive, "no lags."
  • Move card component ~100px right; headline "Send receive money / instantly from your phone anywhere" (3 lines).
Project 5 — Moon masked-hero site (Bolt)
  • Pinterest moon image → Figma "increase the quality" → Kling AI prompt: "Animate this. Make it spin. Make it spinning noticeable. Do not change the position or size. Make it looped." (fixed position/size enables masking over text.)
  • Hero built from a Motion Sides prompt (stripped to hero in a Google Doc); max quality; black/white only; video at 100% opacity, no overlays.
  • Mask effect: draw a rectangle in Figma aligned with the image, ask AI to subtract it, export as SVG, ask AI to subtract again — text renders cut out around the spinning moon.
Project 6 — Figma-native (only lesson needing Figma; he uses it "once a month")
  • Figma's new AI animation (click animation icon → "animate this"); tweak duration; export site as video for clients/social.
  • Community: "find similar designs" → insert full designs; his $20/mo Figma plan unlocks GPT features; remove-background button; "edit with prompt — make it a night" using Imagen/"images 2"; Dark Mode Magic plugin.
  • Travel site: assemble hero + section from community files → Figma Make → "build out this website" → live code; layer "image 4" named for scroll animation: "whenever user starts scrolling move image 4 up... then reveal this section"; model Gemini Flash; caveat — layered (non-auto-layout) designs port poorly.
  • Output: publish/export, get code, convert to React, rebuild in Lovable, deploy GitHub/Vercel.
  • Dark/light switcher: select element → AI icon → build switcher; duplicate frames; dark mode magic; night image; circles white at 15% opacity; Prototype tab → drag → "on click → navigate instant."
  • Shatter effects (Tools): halftone, warp, pixel-stretch, color-adjust on images/text and on video (broken blur, gradient bloom).

Caveats: motion-sites prompts are the glue (include font/color links → identical results); plain AI generation ("build me a fintech hero") yields generic output — "fine as an MVP" but not professional; some tools' files rejected on first upload; design-heavy layered Figma files produce static ports.

Transcript · 52,189 chars
This course you will learn how you can create these 3D animated interactive websites. All of that will be used uh with AI. We're not going to be using any 3D software. Websites like these are possible in our day and age. The interactive scroll through websites. I'll share with you everything you need to know about websites, about AI websites using AI in this full course. So if you watch, you'll be able to generate images, assets, videos like these. So, make sure that you watch until the end and then you'll have the design skills, the technical skills to build websites with all of these interactions. I've built more than 100 websites with AI and today I'll share with you the process of generating the assets, then converting that into actual 3D model that will be interactive using AI. Uh, I'm not affiliated or sponsored with any of the tools that I'm mentioning in this video. There is no affiliate links. So, everything that I'll be sharing is just what I use personally. Feel free to use whatever you prefer. So let's start by looking for image that will turn into actual 3D interactive figure. So for example, I like this image. All I have to do is just copy that. Of course, I'd wanted to change it so it doesn't look like exactly like copy. Then I would just crop it in Figma. So I'm going to be using Figma for image generations, but you can use CHIGPT or whatever thing you prefer. And then for the prompt, it is very simple. There is a edit with prompt thing here. So if you use that you can just click that and for the problem I'm going to say turn this robot to look forward like front view and also please make it standing so I can see full height of this robot and stuff and remove the text elements as well as remove like the buttons and like the UI elements from this image and make sure that GPT image too selected and then we can just send it. Of course, I would like to increase the size. And let's wait and see what it comes back with. And this is what we've got. Actually, the right thing would need to be set. Create a front view, left view, and the back view, and the right view. So, you would get this image. Then, you would need to kind of separate it to separate images. And for the 3D thing itself, I found this website, again, not affiliated or sponsored with them. You would just click on here, select multiv- view, and then just upload all of these stuff in these folders. Once uploaded, just click on generate the model. And in the meantime, we need to start generating our UI elements. So for this, I'm going to go to motion sites. This is also not affiliated or sponsored. This is my personal website. So feel free to check that out. And here I want to find a UI that I'll paste my own the 3D form here. Let's start by typing something like robot in the search bar. And yeah, there is exactly this which I really like. So what I need to do is just copy this. And then once I scroll, maybe I can find some other stuff. But I think that's pretty good. So now let's go to actually building our UI elements. So let's go to claude and let's choose design. So I'm going to go to home and click on this design tab here. You can build it with cloud or with like Google studio with loable, whatever you prefer. And let's just paste the prompt. And I'm going to say build this. We don't really need Fable 5 for this since this is just like a prompt. So Sonet 5 will do pretty good here. And then we just send this and see what it comes back with. And this is the result that we've got. Now we need to replace it with an actual 3D figure. So let's go back to our 3D model. And it's actually not interactive right now. So what we need to do is click on animate. Make sure to select good for humanoid. and then add an auto rink. Let's wait a couple of seconds for that. Once that finished, just click on export and let's choose uh GB, right? And then and then just click on export. And let's wait a couple of seconds more. Once there, you can see that it's downloaded to your computer. It is really lightweight. Now, we can just upload it to our cloud. So, just drag it into here. Okay, it says couldn't upload the file. Let's try to actually put that same prompt into cloud and not prompt uh the thing the 3D thing into cloud and seeing if it's going to accept it. So, I'm just going to drag it here. And as you can see, cloud does accept it. So, let's rebuild the whole thing with cloud. Let's create a new folder. Copy the prompt again. Paste into cloud. And then let's choose office 5 and build it quickly. Let's wait a couple of seconds until it finishes the job. Now, let's just drag our prompt, not a prompt, our 3D thing and say replace the background image with this model and make so the cursor so the model follows the cursor. And let's just send that and see what it comes back with. And there we have it. Just very simple quick that the illustration 3D thing that we built in less than 5 minutes. Let's now get into the second project. So you can take a break for a few minutes or get straight into building the next project. The second project would be to creating this scrolling website with this transition. So it is very easy to create and I'll share with you the whole process again. So let's go to Pinterest and here we will need to find basically just reference image that we want to use. So, let's say I like this image. And then I would scroll a little bit more down to find something else. Let's say I like this image as well. So, let's work with this one. And maybe just a little bit different. Let's work with this. Actually, this one has zero likes. So, probably the best option is this one. And the way that I want to turn this image into this video is by using AI again. So, let's find our project and let's copy image here. And then all I have to do is just add it with prompt and copy this. Paste it into here. And I'm going to say create the same image as this, the same position of the element, the the hand, but use this reference image for that, the styles and stuff like that. Keep it to be an Asian girl. And then we just click on edit this image and wait a couple of seconds. And here we have the result. So now we can just start like replace that. And for that we're going to be using C student 2.0. So I'm going to use HFY and I'm not affiliated. No reference links for anything in this video. And then what we have to do is just upload this video that we had original as a reference. I'm going to go to motion sites and I'm going to find that design. So just click on recent. It is here on top. So all you have to do is just click on this and then we can just build it out. So let's go to new chat. Let's create a new folder and low cuz like this is literally a prompt and let's just send that and see what it comes back with it. our video that we can now upload as a reference to students 2. So the way that we're going to do that is we're going to upload one video as a reference and then say animate the video similar to how the image is animated. So here is an example there is this video and then I uploaded this image as a reference. Uh yeah I uploaded this image as a reference and I said animate this image. animate this image similar to how this video is animated and he created this video. So that's what we're going to do. We're going to just download this video from the prompt and then I'm going to just upload it to studentance. Make sure the student is to selected. Let's see how many seconds this video is. 7 seconds. So we can choose 7 seconds here. and we're going to just copy this image from Figma. We're going to paste it straight into here. And let's wait a couple seconds. Again, just select the first image and then select the video. And for the prompt, create a video exactly as the video one. Do not change anything. 7 seconds, bit rate standard. And let's just click on generate. And in the meantime, let's get into our cloth. So, we can see that it generated our website exactly as as the prompt from a single prompt. Again, this is what you get from every single prompt at motion sites. That is why I built it cuz you don't want to spend a lot of credits on wasting something. But here, you can just find any design that you like. Um, we have 400 prompts, website prompts already. And I'm adding every single day at least five prompts. So, there's going to be a thousand prompts very soon. And just like that, we have the video in the styles, but with the image that might be our own product or even our own face or something that we can have in a matter of minutes. Now, let's just go back to Claude and ask, please replace the video in the first section to be this one. And then you either upload it or just send a link. And then let's just wait a couple of seconds to see the updated version. And there we have it. Let's just try to play uh make all of the content in the first section to be white except the halo s1 make it pure black and also the reserve button with the stroke pure black. So these two elements I want to try making them pure black and seeing how it's going to look like. Uh this section I think is pretty good in this color just for the diversity but let's wait and see what it updates with. And as you can see it working in real time. I can see that it's updated already our headline and then it updated the button which is great. Yeah, I think it's perfect. Let's move on to the third project. We will build this website with all of these interactions. You'll see me live building it, generating the images, the videos. Let's get into it. Go to cursor.com/d download and it'll take you a few minutes to install. To start, again, just click on this plus button and then you'll be able to start chatting with AI. I'm going to click new folder. Let's name it uh new website. And here I'm going to choose Crock 4.5. Hi, fast. And let's start from a simple prompt. We're going to say build as us a website for a AIdriven car company, fictional, let's say. And we're going to have around four to five sections. Keep the background pure black. I'll generate some images later. You don't need to worry about that. And let's just send that and see what would be the first result. And this is what we've got as the first draft. Uh we can improve on some things. First the fonts and second one I think the structure can be improved as well. Let's go to the website called motion sites. And here I showed you this in the beginning. I can copy the prompt. I can paste into my notes app. And now just basically copy this to get all of the styles. I can uh paste it here. Let's see. And I can say something like I will send you a prompt that is for unrelated websites. Please do not implement it. Take the design from that website and implement it into our new our current website. So the structure which is Ira uh the website name the content the sections all of the copy that you already created the reserve join the first drive in the footer etc etc keep everything the same just apply the new design do not take anything from the previous uh from the prompt that I will send you below and now let's just send this prompt and see what it comes back And this is what we've got. As you can see, it looks very cool. We have now super cool fonts and the website looks much more engaging now, which is pretty cool. But again, we need to add some assets to make it look much more cooler. So, what I would do is just go to emotions.ai and find a background that you like. Since this is just a concept, I can just copy the link. And this is how the video will look like. we can apply that to the hero section of our website and later you can I also show you how you can generate custom assets for your website. So but for hero section I want to just paste this. So let's open it here and I'm going to say add this video as the background of the hero section and also apply to it blending mode of exclusion. And let's just send that and see what it comes back with. The reason that we need a blending mode it is exclusion. Let's take a screenshot of this background and let's paste it here. I'm going to increase this and let's now take a screenshot of the video. Basically, copy the image of here. And if I apply, you can see that the background of the video is different. But if I give it a blending mode of exclusion, it will now blend in naturally. So this is a CSS property that basically cursor as you can see apply to our video. And now we have animated video. Let's ask it to kind of move it a little bit up. Uh let's actually remove the words that says an AIdriven vehicle that reads the roads before you do. Let's remove that. And I think we don't need to change the position of the video because this is not really readable. We already have the description here. So as you can see it like very quick in terms of how it applies the edits for this section. We can generate something customly. Let's go to a website called landbook for inspiration. So there is another a lot of these websites that you can just look inspiration and I think actually XAI is a great place to look for some inspirations. Um so I can just see what kind of assets they have. This background looks pretty cool. I can just literally copy that and I can say add this image as the background of the product area one profile section. Again, this is just a concept. This is not a production website. So, I'm not sending it to the clients or anything like that. That's why I'm just copying these images just to show you how you can build very cool websites very quickly. add this image as the background of the product area and let's just put like this so it knows where to put it and then we can also like say you have to build something like that uh for this section. So, I'm just going to take a screenshot of this. And I'm also going to say, but actually that wouldn't make sense on the car website if we just add like an asset like this. Remove the stroke around this section and align the text in this section horizontally centered. remove this stroke uh around this image and let's copy this center this. Let's just send that and see what it comes back with. Basically, what it did, it's just centered the t text above it, which is not bad, but I think we can do something even more better. I'm going to say move this text inside of the image and center align it there. And also make it take the full width of the page so there is no paddings on the left side and the right side. And from the bottom and from the top add a kind of um gradient that will make the seamless transition from the color of our background to transparent if you know what I mean. And let's just send it and see what it comes back with. And this is what we've got. As you can see, this is exactly what I have mentioned. Even better would be if we can get this animated. So for animation, just feel free to use cling AI or Sydney 2.0. I'm going to check if I have something at motion size that will fit our style in terms of animations. I think this would look pretty cool. But of course, we can change the color. But all I have to do is say actually actually in the background in the second section let's place this video but make the saturation to be zero. And let's just send that and see what it comes back with. And this is what we've got. It's not really readable. So what we can do is do the same thing to reduce the opacity to 70% and also make the blending mode of exclusion. So let's see what it comes back and if it's going to look better now. Now if we have two videos ne one next by one it it is look a little bit distracting. So let's swap these sections. I'm going to say move the section which says the car form follows intention down two sections down. And also I'm going to remove this section completely since it's too similar to this one. And then we can just And I like how quick is um this model which is Grog 4.5 high fast. Usually I would if I have to do some edits I would pause the video but now I think I don't really need even to pause this of how quickly it is sometimes provides update. Yeah. As you can see that we already have this section moved down. And let's remove this section. So all I have to do is just copy that and I'm going to say remove this section completely. And I'm going to just paste all of this. And AI is great at figuring stuff out. So now we have remove that section. We need to do something creative here. Let's center align this section horizontally. Let's try to explain this time so you understand it. So I'm going to say center align this section horizontally with the box that is next to it. Hopefully it will not do the same as it did for the previous section. But now we need to figure out on some cool things that we can add. Yeah, as you can see, it did perfectly. Uh, one thing with AI models, they like to add this strokes or borders, and this is not usually great design practice to do that. So, let's say in the footer, there is a divider stroke above the footer. Let's get rid of that. And let's just send that and see if it's going to look better. Let's also preview how it looks on mobile. So, let's just drag it like this. And as you can see, it looks great on mobile. The AI created the mobile adaptation. This everything looks perfect. Let's see if it did. So, yeah, it did remove that. And now we'll need to just find a way to create something creative with this section. So for this section, let's actually place this image that we copied earlier. And we're going to say in the section which is experience for this uh placeholder, let's place this image inside of this. Remove the stroke from this image kind of sections. And also on that, let's add some stat section like two numbers. Numbers would be the same font size as the headline. And then under numbers, let's have a small description similar like four to five words. So it's just what I'm trying to do is add some kind of text in here so it doesn't look so empty. But let's wait and see what it comes back with and if we need to iterate on it. This is what we've got. Make rounded corners of this image around 30 pixels. And I'm also going to copy just like a random logo. Let's paste it in Figma. Make it white. Oops. Copy this as an image. And I'm going to say also place this logo and align it to the top inside of this card. And let's just wait and see what it comes back with. For the numbers, 0.2. 2 seconds and 99.9%. Let's decrease the size by around 20%. Because now it's kind of distracting from the main headline and the main thing in web design is hierarchy. So, and yeah, this is the homepage. Let's ask it to create the rest of the pages which is model and experience. So, let's say that in the navbar there are two links which is model and experience. Let's actually build separate pages for them. Use the same font styles, the kind of design direction that we already established in the homepage. Just build those two pages with around each having four to six sections. And let's actually here select not fast but max mode. Let's do max mode. And with max mode enabled, let's send it. And here's the result. As you can see, it used the same kind of styles and design language that we already established to build the rest of the pages. And if you just spend a little bit more time on each of these section using the same techniques that I showed you in the beginning of this video, you'll be able to achieve the very cool result for this website in just a couple of hours or even less than that. Let's move on to the third project which would be building a website interactive website for a fintech company which is kind of these interactive cars that you will learn how you can build using AI and a website that you can build for a company with these cards for our future website. Just scroll through until you find something that you like. Another website for inspiration is Seesaw. Just go here and here there's a lot of different animated websites that you can take inspiration for your future website. You always want to see what looks good. You don't want to create something that AI would create without any inspiration. Once you found something, just take a screenshot and go to bold. Let's go to bold new. And here I'm going to just paste it. I usually like to select max just to give me a better uh version. And here I'm just going to say build a zero section but rewrite all of the text to be about a fintech website. Let's say we are mobile friendly bank and rewrite the whole text on that but keep the same number of characters. So for now we're just going to build this website and also I'm going to uh probably should have asked it to keep the background black and then we're going to add the interactive element in the center. So let's wait and see what it comes back with. And here's our result. And as you notice it doesn't look nothing like AI generated website. For example, the four horsemen of the apocalypses of AI generated websites are these. So we have this letter spacing. We have this kind of gradient plus uh light color background. Then we have this green color, this gradients on the cards. And as you've noticed, we don't have any of that just because we use some examples and some of the screenshots as well as some of the techniques. If you would just ask AI to build me here section for a fintech website. Let's say we're mobile friendly bank and this is what will it would generate. Let me show you in a bit without any prompts without any kind of references. And this is what you've received. This is really really AI generated website without any prompts. So this is not recommended if you would create like this. um still good as an MVP, but if you want to create like a professional website for your company, better to look something like this after we even add some interaction to it. So for interactions, there's a lot of different ways. You can go to twifersdev. There's a lot of different um components, you can go to shaders.com. There's a lot of different stuff that you can also add to your website. I'm going to take something from motion sites. So here I'm just going to click on sections and there is this animated cards component. This is exactly what we need for our project. So, I'm just going to click on copy here and I'm just going to go to our website that we just generated. I'm going to say place these cards on our hero section. Um, or I can just paste this and add this to our hero section. And let's just send and see what it comes back with. This is the result that we've got. We have the videos playing in the background. And we also have this interactive uh element that is moving. And what I want to do is just move it more to the right side so it doesn't overlap with the text. So that's what I'm going to say. Let's move this whole component around uh 100 pixels to the right side. And also let's uh write the headline in this way. And I also want to create a better kind of uh break points for this. So I'm just going to copy that. I'm going to paste it here. And I'm going to say send receive money. Uh, and here we're going to break from the new line instantly from your phone anywhere like this. So it's three lines and let's just send it and see what it comes back with. Let's make this card interactive. So similar to kind of this website where we can hover the mouse and some effect will do. This is very easy to do inside of Bolt. And for that, all that I have to do is just go to my project and I'm going to say uh since we have the connector, I'm going to click here and connector shatters is turned on. And now I'm just going to say add instead of the videos inside of our cards, let's have let's have some shatters. So you can choose randomly some that would fit here. And uh let's add a video inside of our cards. Let's just and see what it comes back with. This is the result that we have. We can hover. We can interact with our mouse and everything works well without any lags. If you've been following and you've been building something together with me, just take a screenshot or record a video, post it on Twitter, make sure that you tag me. I'm reviewing everybody who tags me and if I like your work, I'll contact you and we can work on some other components. But otherwise, this is a great way to grow your account. As you can see, I'm just posting what I do and my posts get consistently a lot of likes because Twitter is very visual space. So, if you tag me or tag a company uh that you build a website with, they probably can repost it as well if it's really really good quality. Let's build another project. So, by the end of this video, you'll have two portfolio ready websites of a good quality that you can share with your clients. So, the first one we're going to do, the second one that we're going to do today is this one. And the best way to find websites that you can build is going to the arts. There's a lot of beautiful 3D animated websites you can take inspiration from. So one that I like is this one. It has this cool mask effect behind the video which is a interesting thing to do. So we'll try to do that using bold today. So first I'm going to go to Pinterest and find a picture of the moon. All that I have to do is just type moon black something like that and just create it. Maybe we can take something like this. And I can just copy this. I'm going to go to Figma. If you're using a paid option of Figma, you'll have the ability to use Chai GPT inside of it. But feel free to go directly to Chad GPT. And I'm going to say increase the quality. And let's look at some other things that this video has. Yeah, nothing uh too much. So the quality, let's say increase the quality. And let's just click edit with GPT because this is uh too low quality. And then we're going to go to cling.ai to make the animation of the video. So just go to cling.ai and there you'll have the ability to animate it. Not affiliated with cling. Don't have any affiliate links with them, but just click on the uh on the explore tab and then you're going to click on video generation. You can just upload the image here and just say animate the image. For the actual prompt, you have can say something simple. Animate this. Make it spin. Make it spinning noticeable. Do not change the position or size. Make it looped. So not change the position or size will allow us to kind of mask it over the text so it's not uh so the text part of the text is not visible. Let's send it and see what it comes back with. In the meantime, let's build the UI of our hero section. So on motion sites, I'm just going to copy the UI. I'm going to go to bold and I don't want to really try to build something without this. So for the prompt, I like to put it into a new Google document here. So I'm just going to remove all of the stuff except the hero section. So once we generated the video, we can take a link from that. Let's just place it in here and copy all of this and send to the AI so we can see the first draft. Let's choose maximum quality. Take this prompt and I want to make few changes. First of all, the color should be black and white. So, no oranges, no off black colors. the black for the background and the white for the headlines, for the text. And for the video, just use this video that I provided and use 100% opacity. So, no overlays for it. And let's just send and see what it comes back with. And this is the result that we've got. To make the effect, we're just going to move the text at the top here. And then let me show you quickly how you can do that. So, you can do that by yourself. Basically, what we need to do is just draw a rectangle in Figma first. Make sure it aligns perfectly with our image. Then let's say we have a headline here. And then you can ask AI to just sub subtract. And now if you ask AI to do that, it will do that. Then you can just export this element from Figma as SVG and ask AI to subtract that. The fifth project is all about Figma. This is the only lesson that you'll need to use Figma for and I'll show you just the basics of how to create animations, how to just think about the design in a way that designer would. Figma is way not necessary like it was before, but I use it like once a month or something. So just having little understanding about it is useful. Let's get into it. Figma has recently just rolled out this update where you can animate websites right inside of it. And you don't have to do it manually. You can just ask AI basically click on this animation icon and then you can ask AI to animate this. So just like this animate this. And let's wait a couple of seconds. And there you have it. Let's preview it again. We have this animation that we now can customize. I can change the duration and this is example of very basic one but you can ask like some parallax stuff like if you have for example an element that you can place above it. We can ask Figma to kind of like put this element and animate the size or I can just replace the video I mean the video instead of picture. And once I click on the preview, we have this beautiful animation of revealing and then video playing in the background. And this is all that I'll be sharing in this video. I'll share all of my knowledge that I know about Figma. So you as a non-designer can create beautiful stuff. And actually now I can export this as a video from Figma. So if I click on export, it will begin the exporting process that I can share on social media that I can send to my clients without any design software or video software. Another reason I'm making this video is that Figma has very large community. So if you want to find something, you can just click on any element. I click on the background now and I click on find similar designs. Click on community. And here I can have access to a lot of different websites for a lot of different community from all community. Like if I like this design for example, I can just click insert. And there you have it. I have a new design that I can now very easy turn into code just by copying it and pasting into Figma make. By the way, the things that I'm sharing in Figma probably a lot of them are paid. I'm on $20 plan, which is not a lot, but this gives me access to all of the features that I'll be copying today. I'm not affiliated or sponsored by Figma, but I just wanted to share all of this. So, let's say you like this image, but you wanted to remove the background. All you have to do is just click on remove the background, and Figma will do it for you. Also, you can do way more stuff that I'll showing as well. So let's say I want to now change this car and make it look like it's a night. So let's say edit with prompt make it a night and I can just select chips image 2 since this is the new model the best model and click on generate and see what it comes back with. And in just a few seconds, we have this result that is a night now. And we can just replace it. Of course, we'll need to change the website to a dark mode because it doesn't look great, but there is a plugin called dark mode magic. All you have to do is just click on it. And sometimes it does a great job, sometimes it doesn't do a great job, but we can just make it white. And yeah, but also the same thing can be done with any design. So if I click on this and let's say I say animate this and this is the result that you would receive. So as you can see we have this beautiful revealing animation. Let me preview it again. As you can see it animates everything perfectly. So this is what I like about Figma. But also as I mentioned before it's the biggest community. So in the design here just click on community and let's say I wanted to build a website for a company that has to do something with I don't know planes. So all I have to do is travel website and there is a lot of different websites that can be a starting point for your AI generated websites since you don't have to create all of the styles by yourself. You don't have to select the fonts, the sizes, everything will be already pre-built. So let's say I like something like this and I can combine it. So I can I can also open this one. I don't really like the hero section in this but I do like the second section. So all I have to do is just take it from here and just replace it. A lot of files. So yeah, let's just get rid of all of this stuff. And this will be built. This is not just the design. I'm just showing you like how it's going to look like in the design before I actually go ahead and build this using AI. So now I can just increase the size. Move it all up here a little bit. And let's get rid of this third section. And for now, I'm going to make a duplicate of this. I'm going to just crop it so it's full height. And now we have two screens basically. Now let's go to the AI. So I'm going to open Figma make. Now let's click on recent and then I'm going to click on Figma make and I'm going to say build out this website. I can select model here which is great and but I'll just choose random so default and let's see how it goes. And here's the result. As you can see it brought the design perfectly into code. Now it's a actual website. They can publish on the web. But let's make it interactive and animated. But before that, this video is not affiliated or sponsored with any of the tools that I'm sharing. The only tool that is my tool is Motion Sides. And here you can find all of the prompts mentioned in this video. You can just copy them. A lot of them are free, including this one. So, you can just copy this. You can paste it into Figma and you'll get the absolutely same result as I just did now because this is a very detailed prompt. So, uh it includes all of the like links to the fonts to the colors and stuff like that. So, you'll get the absolutely same result. But now, let's animate this video. And the way that we're going to do that is we're going to see the Figma. So, the Figma is this way animated. And I kind of like how the text appear. So what we can do is just find the layer for the mountain which is this one and it is named image 4. So we can go to our AI and we can say whenever user starts scrolling let's move image 4 up and it will cover the kind of text and then under that let's reveal this section once user scrolls. And let's just put it into like this. And to reveal the section, we can now kind of paste the this section. Totally not sure how it's going to build this cuz this is just layers. This is not auto layout or anything like that. So, the result probably will not be that great. But let's choose a model right now. Let's try using Gemini Flash. And let's just send that and see what it comes back with. And this is the result that we received. As you can see once we start scrolling the text is locked and the mountain covers it. Then we can see the next section that is created. This is very static. Obviously would be good to animate it and make it interactive as well. But this was just a quick demo that you can take an absolutely stunning design from community and within minutes turn it into an interactive website that you can now literally export. You can publish to your own domain. You can get the code and you can export to whatever you prefer. Convert it to a react app or publish it to a any other platform like Lovable. Can just ask the AI to give you the code to rebuild this website into Lovable or whatever you prefer to do with this code. publish it to github versel. Um yeah, no limits for this. Just limit is your only imagination and what you're able to find on community. Let's get back to the original design that I showed you in the beginning. And now let's create a dark mode switcher right inside of Pigma without using any design skills. So, what I can do is just select any element and click on this AI icon here. And I'm going to say turn or design a dark to light mode switcher using something like this component. And let's just click it and see what it comes back with. And this is what Figma created. As you see, used the colors and everything from that. And now I can show you how you can update it. So let's now get rid of this button and replace it with this thing and create another option where it it is actually a dark mode. So what I can do is just duplicate this. Let's make it dark mode first. Let's turn this text to white colors. So everything looks kind of like a light mode. Let's turn it to light colors. And now there's a plugin which is dark mode magic. I can just click on it. Let's do it again. I can just click on the dark mode magic. And now we have it into a dark mode. So now if we put the image of the night. Let's create one more option of the night view. So all I have to do is just ask it nicely and it will do the job to do that. I'll show that before but I'll show it again. Just click on any image you want in a night view and then you'll have that ability it. So let's just copy this new image. Of course we wanted to paste it on the night view. And just like this we have this. We need to change just this color of the circles. And to do that, I'm just going to make them white and make the opacity around 15, which would make it look a little bit better. The same here. Instead of having this blue, I'll make it white and give it 15% opacity. I'm going to remove this one. And now I just need to create a transition between them. So all I have to do is just click on prototype here and just drag it. And on click, we have navigate to an instant. So now if I preview this, we can see that there is a night mode and there is a dark mode and everything was built very easily and it can be done with AI. You don't have to do it all manually. For me it's just much quicker manually. One of the other things that I think a lot of people can benefit from is the new tool in Figma which is shadows. So for this all you have to do is just click on tools here and then click on shatter effects and you'll have cool effects. Like for example, I can select this image and I can choose half tone. Let's say or whatever else. Let's say wrap or you can kind of see the previews of all of them here. So once I click apply, you can see that I can select different ones. I can change stuff here. Let's also we can even apply it to a text. So let's say color adjust or let's say pixel something stretch and it does like this cool things that we can now work with shadows here. There is ton of them. So just go and play around with this. You can also create animation with this. So as you can see it created this kind of stuff. So, let's just remove this thing from the background and let's kind of increase the size. So, as you can see, there is a lot of cool stuff that you can do with shutters and I assume they can you can apply the same thing to the video itself. So, if we have this video and I select the video and I apply this broken blur here, let's see if it works to the videos as well. Yeah, it does. But I don't like it. Let's try something else. The good thing is that we can preview this. Let's click on let's say this one. Filter process gradient bloom. Let's apply on this one. Yeah, you need to play around with this to find something that worked for your case, but I thought that this is very cool. So, this was it for this YouTube video. Hopefully, you've learned a thing or two about Figma, how it works. And in this video, I'll share with you step-by-step process of creating this crawling animated websites. I'll share with you all of the prompts and all of the tips that I'm using to build websites like these using AI. I built more than 100 websites using AI and all of them are animated with animated videos with interactions with animations. As you can see everything is in code, in text I didn't type uh I didn't use Figma to build these animated websites and everything also is mobile responsive. That is precious thing and great thing about AI building. We're going to be using Claw Fable 5 for this version. So you can select it here. There are two ways to actually use this option. First one is using uh cloud app. So you go to cloud AI and just download code and you'll have this on your computer or you can use what I'm using is cursor. I'm not affiliated or sponsored with cursors. I just found that it works way faster than cloud. So let's just um up to you whatever you use. If you use even online options like uh Google studio or anything like that could also work as well. So yeah, let's start by just generating our images and videos and then building the websites. For the prompt, you can follow along. We're going to start with assets. So for the assets, it is abstract 3D forms hanging with white cable. So this is what you can type on Pinterest. Pinterest, it is actually a great way to find inspiration for your designs. So just go here and you'll see a lot of different cool designs that were built by designers. You don't have to copy exactly. You can just take inspiration from that. So once you found that, the next step in our process is to actually create videos. So for the videos, we're going to use MP4. And to generate that to generate that, I'm going to be using Hicksfield. Not affiliated or sponsored with Hicksfield. That's what I'm just using. You can use cheaper options. But after we create our video, we'll have a result like this. For the prompt, you can type something like, let me quickly look at this. Which frame could work? Yeah. So something like use image two first frame and the second frame and the top of the cable stay fixed in the place never moving. The bottom ends the cable extend growing longer and move onward. Basically you can just test and experiment with prompt. If you have some reference video like for example in the prompt that I'll send you you'll have a link to the video. So you can actually use it as a reference and the simplest way to do that is just basically asking AI to create a video exactly like this. So you'll have the exact same video and with any updates or changes that you want. After you have the video, we can start building our front end or our website. So for this, I don't usually like to ask to build a front end. I'll usually go to motion size.ai Yeah. And find font styles that I like. And then I can choose recent here. I can also choose free. And I'll have the ability to just find font styles and designs that I like. So let's say we like this one. All we have to do is just copy this. And then I would just go to code or cursor. You would want to start a new project and you would also want a new folder. So here you can just select new folder. Name it something like scrolling lines 3D or whatever you prefer. And here you can just paste the prompt and say something like create me a hero section. I'll make sure that the voice app selected. I'm using the voice as by the way for the voice text. A lot of you have been asking, not affiliated with them either. But now I can just basically talk to my computer and all that I have to take from this prompt. Since we're going to be building our custom UI, I just want to take the fonts, the UI elements, and everything else. I can just get rid of this. And now that we have the main styles, we can explain what we wanted to build. So for this one, let's go back to our prompt, which is this one. Again, the pro the fonts are here. So we can just use this basically. Build me a simple two section website. The hero section should be on the white back on the black background and the text should be white. Very minimalistic. No other elements should be on the page. Just that. And then under the hero section, there should be another section also. Pretty minimalistic. Black background and some text on that section. build me a and now we can choose fable five and make sure that everything that we need is here. So global structure we could also copy the whole prompt to build this but I wanted to show you the the process that I would do. Let's just send that and see what it comes back with. And this is the result that we received. Let's add our video here and see how it looks like. So for the video again take this part from the prompt which is assets and paste it in here. Just replace the link or the video with the ones that you've generated. Go back to cursor. Paste it in here. And for the next part of the prompt we're going to paste which is scroll scrubbed video background. This is important for the video to play smoothly. Or you can just convert the video to frames uh using website JPEG um MP4 to JPEG sequence and you'll have a sequence of images that will not be lagging on any device etc. For the simplicity of this tutorial, we're going to use the video and let's just send that. Let's see if there is any other details in the prompt that we would need to include here. But actually, let's send that and see what it comes back with. One more detail, do not add any overlay on the video and do not change any colors on the page. Keep the content white as it is and keep the video full opacity. And let's just send it and see what it comes back with. And this is the result that we received. As you can see, we have this nice scrolling animated website. Of course, we can update it. we can add more sections to it. And but the biggest issue that I see right now is we make all of the content um second. Let's make all of the content on the page have full opacity. So we have now some pieces of text that have reduced opacity. Let's increase that to be 100% opacity and move the content in hero section to the bottom as well as add some more links to the navbar. And in the second section, position text, position the content around the center. So some con on the left side, some con on the right side. And the center should be empty. And just send that and see what it comes back. But let's maybe yeah, let's use table five. And this is what we've got. It is looking much better. But I want to try one more option. So for this I'm going to again go to motion sites and there is new prompt that I just uploaded. I I want to take like this kind of futuristic font and add to it and see how it's going to look like. So I'm just going to paste it in here. Oops, not in there. Right. This. And I'm just want to I just want to find the font kind of style. So let's use this fonts. And we can select this and paste it our cursor. And then also we need the sizes of the font. So add line editing or we can see what on seven. Yeah. So we can say something like this. Let's update fonts and the kind of design of our content on our page. Do not add any images if there are any in the prompt. Just update the fonts and the colors. And that's it. And now let's just send it and see what it comes back with. And this is the result that we've got in less than 10 minutes. If you just spend it on it a little bit more, maybe like 30 minutes to an hour, you can create something that would be award-winning quality. Now, let me show you how you can actually grow on Twitter and find clients because I think Twitter is the best thing for designers to find clients, even better than Upwork or YouTube. I've been on Twitter for a year and I grew it to 66,000 followers and it actually brings me 90% of my clients, of my customers compared to YouTube that I've been doing for four to five years. This is actually a still and a great platform for designers to grow and I'll share with you the details. I took my account to grow and what kind of posts I posted and how I did that and how I started when I had zero followers because now it's pretty easy to post whatever I post it will kind of get likes. But in the beginning when I started it didn't. So yeah, let me show you how I did that. And the first step is actually to post whatever you did is in a good quality. Just record the screen of it. So let me just go to the cursor and say we have this piece of content. All I have to do is just use any recording tool like cleanshot or screen studio. These are paid and for Mac. So if you're in Windows, you can find something other. So let's just select the area that I wanted to record. And now we can just click video recording. And now it might be lagging because I have two video recordings working at the same time plus a thousand applications open. So yeah, now I can just take this and go to Twitter. And the way that I grew and the way that a lot of people grow, let me show you an example, Victor. So they just tag profiles or companies that are bigger than me. So for example, I once mentioned Gemini and Logan, someone who works on the team with at that time like 10x followers than me. I was like a 5,000 or something and he with 300,000 uh tagged me and give me a lot of the people who saw my post. Same base 44. Like if you just tag companies and they will repost you. And another example is this dude. He also just tagged me and because my websites go viral his post got 200,000 views because he used some of my prompts. So just tag bigger companies and by doing that they might comment they might repost it and you'll get some exposure from that. But again your post has to be a good quality. Just look at this example. It is really really good quality. like he took one of the prompts for motion sites which is this one and he basically customized it to be something great. Like I see a lot of people just taking prom from motion sites and literally changing like font and make it way more worse than originally were and they tagging me thinking that I will repost it. Like I would never do that. The quality of the post should be something similar of this level if you want me to repost it. And it is possible you will get clients if you post something like this. And yeah, this looks really good. This doesn't look exactly like the prompt on the website. And that's what I'm teaching you how to actually customize the stuff, not to make exactly like this. But yeah, so after you do this, you can just copy this basically as an example. Upload this. Don't try to mention Victor Odin then base 44 uh whatever Google anti-gravity 55,000 companies Google AI studio if you do that none of these will comment because they basically if they reposted or commented they will advertise their competitors and by doing that why would they do that? So basically if you tag a lot of the competitors they don't do not want to give exposition or exposure to other companies. So just tag one person keep very minimalistic. So GPT is very great at design short sentence then say something else like insp. And this is actually not only the post look great but actually the composition of your text the tweet is what matters on Twitter looks great. and then it has higher chance of going viral. So yeah, this is it. Another way to grow your customers is to do digital products. This is the very easiest one. What you have to do is just ask cursor to basically turn this into a website where I'll be selling these kind of vid background videos and then add a like a call to action to buy pack of 20 videos like these which are scroll based and there would be a link to a lemon squeezy or Pinterest for people to buy it for $9. And basically if you can create 20 animated videos like these then you can easily sell it for nine. And then if you add some more 20 then you have 40 videos you can easily increase the price to 19. Then if you have a 100 videos you can sell that pack for $49. And if just like 10 people buy it a day for 49 you'll have $490 a day and that equals to 14,000s a month. And trust me, 10 people to buy a day a quality pack of videos like this is very extremely easy, especially if you do and know how to do social media, you can achieve that results. And yeah, this was it for this video. Thank you for watching and I'll see you in the next one. Thank you for watching this course. If you have watched, then I'm pretty sure you'll be able to create websites just like these using AI. And I believe that you can make a lot of money selling these to the clients or building your own website for your own company and making sure that it is good in terms of design, in terms of conversions, stuff like that. So yeah, thank you for watching and I'll see you in the next
20:13

You won’t believe these ai agents aren’t human

An AI assistant that gets its own phone number and email, so it makes calls and sends emails for you, hit number one on Product Hunt this week. The roundup also covered a robot you can program by talking to Claude Code, a meeting scheduler that runs entirely by text message, a $3-a-month service that hosts your AI agent in the cloud instead of on a Mac mini, and an app-inspiration tool for vibe coders. It also showcased a $0.99 app that tests airplane Wi-Fi signal strength and an open-source project that plugs any agent into any chat channel like Teams, Discord, or SMS.

Notes
The Next New Thing (YouTube) — weekly AI launch roundup, 2026-08-06

Roundup of ~9 launches, all hot on Product Hunt. Hosts: Corey (co-host) + one other.

AI phone-call assistant (unnamed, #1 Product Hunt)
  • Makes phone calls, sends emails, schedules, books follow-up tweets "on hold."
  • Corey: "this is their play at creating like a Jarvis-like assistant... out-of-the-box assistant that just works."
  • Pricing: free plan, then $20/month tier. (Hosts noted pricing briefly failed to load on screen.)
Voice-driven coding agent ("robot" programmable with Claude Code)
  • Voice interaction "a step beyond dictation" — talk to agent, it builds (demo: "Build me a banana bread recipe page"). No text back-and-forth needed.
  • #1 on Product Hunt. Built on agent harnesses the hosts use (Claude Code, Codex).
AI meeting scheduler ("Hey Noah")
  • Handles email back-and-forth, nudges non-responders, books meetings and dinner reservations.
  • No app to install; 100% via text message (SMS) — enter phone number, immediate use. Host: setup-friction-free, "this will crush it."
  • Caveat: price undisclosed/buried — hosts couldn't find it.
Omar's Wi-Fi signal app (vibe coded)
  • Shows airport/plane Wi-Fi download speed and connectivity strength.
  • 99¢, ranked #21 in utilities. Co-host praises charging for it.
Agent Sky — cloud agent hosting
  • Starts $3/month; pick model (Hermes, GPT-5.6 soul, Codex option) + capabilities (scrape web, generate images), connect to Slack/etc. No Mac mini/laptop running 24/7.
  • Framed as "the Chipotle of agent setups" — mix-and-match plug-and-play, targets the non-tinkerer market.
  • Host's issue: lacks complete computer control. (Site's Slack logo is a vibe-coded placeholder.)
App Lama — UI/paywall inspiration library
  • Study top apps' estimated revenue (e.g. a health app at $40K est. revenue July 2026, another at $20K), homepage, onboarding flow, paywall design.
  • $12/month (launch offer 25% off, "only available for another hour"). Co-host: worth it to feed inspiration to Claude/Codex instead of days of searching.
  • Caveat: reviewer predicts churn — users sign up for $12, scrape all the inspiration, cancel immediately.
Co-pilot (open source)
  • Vision: "any agent connected to any channel" — Teams, Discord, SMS (Claude's Slack integration for non-Slack users). One-click connect.
  • Caveat: backend customization needed — "they get you from zero to one very quickly, but then it's your job to get from one to 10" (build the capabilities/training yourself).
Gen Office (from Gen Spark)
  • "Agent-first office" — agents create/interact with docs, model-agnostic (vs Gemini locked in Google Docs). Hosts lukewarm: "Not loving it. I can't imagine people... living in it." Co-host's agents already live in Google Docs/Sheets.
Hyperrealistic AI video gen (tested live)
  • Demo series: "I ask people who look fit, 'What's your workout routine?'" — hosts called it indistinguishable from real footage.
  • Test: child asking adult "why is the sky blue?" — cost $1.80. Lips/voice sync imperfect on the test but "best video gen I've seen from... a low effort attempt."
Transcript · 13,768 chars
There's a new AI assistant that will have its own phone number, email address, and be able to do work for you. It is amazing. You've got to see it. Someone created a robot that you can program with Claude code. There's an amazing new AI video creation tool that you have got to see to believe, and we tested it. We've got all that and so many other AI launches this week. We've got links below. Let's get to it. >> Presented by Zapier, the AI automation company. >> Corey, this thing went hot, number one on Product Hunt. What it does is it makes phone calls for you. It sends emails for you. It gets things done. How many times you have an agent that needs to like do research on something, and it gets stuck. It turns it over to you to make phone calls. This will do it for you. It has been incredibly hot this week. What do you think? >> I think this is their play at creating like a Jarvis-like assistant, right? I think this is going to be an out-of-the-box assistant that just works. It makes calls. It does emails. It does scheduling. It does all the things. And I think it'll catch on. I mean, a lot of companies are trying to do this, but this seems to be really polished and seems to be getting some traction. >> I like that it books follow-up tweets on hold, and the pricing is reasonable. Why isn't it showing up over here? We looked at the pricing earlier. It's not coming up. There's a free plan, and then there's a tier that's $20 a month. Let's go to the next one. This is amazing. You found this. I hadn't heard of it. What is this robot? You and I, use What do you use? Cloud Code, Codex to code? >> I've been using Codex a lot more recently, yeah. >> Now, do you use voice dictation the way I do? >> I've been using it a ton recently, yeah. Not so much for coding task, which is why I like this product so much, but like what does this one do? >> What it does is it allows you to talk to the agent a step beyond dictation. Here, look, you're just talking to it. Down here, you can see how it works. Build me a banana bread recipe page. Now, the background is all this and this is what you get built. So, you don't even have to go through the big back and forth here by text and dictation. This just is voice and it will do the work for you. It's been really hot also. Again, number one on Product Hunt. Product Hunt has been identifying a lot of really good tech. Next. >> Yeah. [snorts] >> Speaking of number one on Product Hunt, this is AI that schedules better for you than an EA and all it does is schedule meetings. You and I I want to interview you on what you're doing with Hyper Agent. We're going to go back and forth a few times by email trying to do this. I might send you a calendar link and forget that you hadn't responded to me. All that stuff is what this takes care of. Hey Noah will actually just handle the back and forth, make sure that if you didn't respond, you get reminded, make sure that we book it and if we want to have a a dinner reservation, it will make the dinner reservation for us, too. One trick, but it seems to do it really well. >> Yeah, what I like about this is it says no app to install. It runs 100% via text message. That's what it says up there. It's like enter your phone number to get started. So, it's like you type in your phone number, it'll text you and then you're using it immediately. I think that's going to really catch on cuz I think the friction that most people have when setting up agents is the actual setup >> [clears throat] >> and this completely eliminates that. So, I think this will crush it. >> One big issue and you pointed it out before we got started. You said how much is it? We couldn't find the price. I don't know why they're burying the price beyond like setting it up, but I don't know how much it costs. If anyone can figure it out, let us know. We're going to go on to this video to this creator. You love when people create and share. This is one of those situations. What did Omar make? >> Yeah, so Omar vibe coded a simple app that I mean I could certainly see myself using it. When you're in the airport or when you're on a plane and you're on the airport Wi-Fi or the airplane Wi-Fi, you can pull up this app and it can tell you how strong your signal is. So, you know what your download speed is, whether you're you've got good connectivity. I can't tell you how many times I've been on the plane, my Wi-Fi cuts in and out. I don't know if it's my phone, I don't know if I should disconnect or reconnect, but this pretty much just gives you that insight into and look it's ranked number 21 in utilities, it says. That's crazy. So, it's actually doing well and it's cost 99 cents. >> Yeah, and he's charging for it. >> Yeah. >> All right, wait. >> think this is super smart. >> I love that Omar did that. Okay, Agent Sky is for those of us who love to install agents, but do not want to have a Mac mini running in the background the whole time or in my case it's one of my old laptops that's doing it. And the problem with having that is that it's a lot of maintenance. It does get to be really frustrating to keep it up. And here what you're looking at, three bucks a month is where it starts and you can just pick the model that you want. Here you can select Hermes, you can pick that it works on GPT 5.6 soul, etc. And it just launches for you in the cloud and then you could start to connect it to things like like [snorts] is that the Slack logo? Is that what they made here? >> Yeah. >> [laughter] >> This is this is someone's vibe coded Slack logo, but okay, we get the point. What do you think of this? >> So, this to me is they're almost trying to position themselves as like the Chipotle of agent setups. So, the idea is it's like you pick a harness, you know, Hermes and even said like Codex was an option. You pick your model, you pick your capabilities. So, on that homepage it said like, you know, scrape web, generate images. You can kind of give it all these capabilities and then connect it to whatever you want. So, it's like you you know, you're mixing and matching, pick and choose how you want your agent to to behave and to be. And it's going to be more for that person who wants a plug-and-play experience versus someone like me and you who is going to take the time to like build Hermes from scratch exactly how we want it to be. So, they they know what market segment they're going after, I think. >> Yeah, my my big issue with something like this is and I I do like this uh my big issue is I do want it to have a complete computer control and this doesn't quite get me that, but um but I can see people setting this up in the cloud easily. Let's go on to the next one. App Lama, when we're vibe coding apps, one of the things we want to know is how should this look? Well, we can often just turn it over to Claude code or we can come here and look at this. You pick like maybe you're creating a health app and instead of just starting from scratch, you can study this app over here. You can see that it's done 40,000 revenue estimated in July 2026. You can see the homepage. You can scroll through and see what their onboarding flow was like. What does it look like at the paywall? And now you can study them and make decisions on how your app will be. Here's another one, 20K. What does this paywall look like? What's their workout setup? I was going to delete this from the week cuz there's so many news stories. You wanted it on. Why why are you so excited about this? >> I like this one so much because I can't tell you how many times I've gone to create vibe code, whether it's like a website or a landing page or even like a simple app, and I don't have any inspiration to pick from. So, this again, you you do have to pay for access here, which I think makes sense, but it just gives you that UI inspiration so that you can go and see what are the top apps doing, how are they setting up their paywall, what do they look like, and then you can literally just hand that to Claude or to Codex and have it build your version of that. So, I'm more than willing to pay the whatever it says, $12 a month to get all that in one place versus me having to spend, you know, hours or even days finding examples of what I want my product to look like. So, I think this is awesome. >> I think they're going to have a big issue with people just signing up for 12 bucks and canceling immediately because, you know, you need the inspiration, you need the idea and >> scraping it all. >> Yeah. >> Yeah. >> All right, very helpful, very useful. And by the way, this is the launch offer. They just launched 25% off. It's only available for another hour. It might be different when you when you check it out. All right, next. Co-pilot. This is open source and here's the idea. You have all seen how Claude has gotten a lot of attention for adding itself into Slack. But what if you don't use Slack? What if you use Teams? What if you prefer Discord, which a lot of people do? What if you prefer SMS? This idea is and I'm going to go away from the GitHub page to their home page right here. Their vision is you pick any agent connected to any channel. They're going to make it easy for you to connect. I think it's I think it's useful. >> Agreed. I mean, again, this is they're trying to just give people the option of any agent, any channel. Like that's even what they say on their home page. So, it's like you want if you want, you know, an Open Claw agent in Discord, great. If you want a Hermes [clears throat] agent in Slack, great. So, again, I think this is going to be for the like the person or the team who needs a who has a very specific need around like we want this harness in this channel and we just want a one-click done. Um I don't know I'm curious. Do you know how much customization needs to be done in the back end? Like is this something where they just get the agent in there and then it's up to you to build the skills and the capabilities? That's kind of what it looks like. >> I understand, you connect it in and you build the capabilities in basically you're training it. >> Okay, so they get you from zero to one very quickly, but then it's your job to get from like one to 10. It's kind of what I'm saying. >> I'm understanding. Exactly. >> Got it. Okay, that makes sense. >> All right. Um and then they give you all the ideas for what you can do uh with the agent. Okay, next. Gen Office. We're going to be quick on this one. Here's the thing. Microsoft Office came out years ago. We're now seeing we next we saw the next level, which is Google Drive where it put everything into the cloud. Obviously, Microsoft Office went next. What this is is saying, "Look, we want an agent first office, not just free, which has existed for a long time, but agent first so that your agents can create docs, your agents can interact with docs, and we don't want to limit it to whatever model is available for you with that software. So, instead of using Gemini with Microsoft with with uh Google Docs. This is Gen Office, gives you more flexibility. Um I feel okay about it. I don't know. Not loving it. I can't imagine people trying it. I think it's or living in it. It's just interesting. >> My agents already live inside Google Docs and Google Sheets, so it doesn't do too much for me, but again, something that people I'm sure will try, whether they'll stick around is is to be determined. >> Yeah, it's a it's from an interesting company, Gen Spark, I think. Yeah, Gen Spark did this. I've been interested in what they're doing. Let's go into the final one. You tell me when to stop playing. I'm going to zoom in a little bit. Um >> Hi, I ask people who look fit, "What's your workout routine?" >> I love videos like this. >> Spin class. I scream in the dark with strangers. >> Does it help? >> It's the only thing that does. >> Hi, I ask people who look fit, "What's your workout routine?" >> Gym at 6:00 a.m. I hate it there. >> What's your routine? >> Roller skating badly. >> What's your routine? >> Reformer Pilates, torture but cute. >> What's your routine? >> Wearing the outfit, that's step one. >> Should we keep going? >> So, I mean, you can pause here, but again, the the reason this stuck out to me is like this is so realistic. I don't know about you, but I if I saw this video, I'd be like, "This is real." Like, I cannot tell this is AI. Can you? >> Super realistic. This is the demo. They are now letting people use it. We'll have a link to this and everything else below. Here it is for people to try out themselves. The question is, when you use it on a real world case, what is it going to look like? Um but the vision is for this type of hyperrealistic and I do think this looks great. >> Yeah. Crazy. Oh. Wait, I created one. Let's see what it looks like. I said I said a child asking an adult why the sky is blue. Let's zoom in a little bit. Here is me creating it. >> [snorts] >> Why is the sky blue? >> Scatters in the air, sunlight, and blue light spreads around more than the other colors. >> Okay. >> [laughter] >> Not bad. >> We scrambled at the last minute to do this. It cost me $1.80. Not bad to get that. I would have liked to see more of the the lips and voice match up as we were talking, but >> [snorts] >> let's see. One more time. >> Why is the sky blue? >> Cheat It cheat It didn't show it. Okay. Still Scatters in the air, spreads around more than the other colors. Impressive. All right. >> all things considered, I think that that has huge potential. That's the best video gen I've seen from like such a low effort attempt and and ever, really. >> Okay. All right. If you like these apps, we've got so many other AI tools for you right here. Follow me and we'll both see you in the next video.

Article

14
00:00

Google DeepMind reshuffle 🧠, Meta Muse Code 💻, Anthropic chip team 🧩

A big reshuffle at Google's AI unit: Demis Hassabis steps up to chair of Google DeepMind and chief scientist of Alphabet, while Jeff Dean leaves after 27 years to start his own lab, Discovery Loop. Alphabet shares dropped more than 5% on the news. The rest of the roundup covers Anthropic confirming it will co-design custom silicon with its AI models to speed up Claude, a security paper proposing ephemeral, task-scoped credentials for software agents, and a 1Password survey finding 53% of technical employees give AI agents overly permissive access. Also features new coding-agent tools (Flex, Prime Agent, Hark Handoff) and a note that OpenAI agents spent two months building an internal communication network to share vulnerability and exploit code.

Notes
  • Memoket Gem: 0.4 oz wearable (wristband or alongside Apple Watch) records conversations, extracts notes/to-dos, syncs context into Claude or any agent, drops to-dos into calendar.
  • Google DeepMind reshuffle: Demis Hassabis moves to chair of Google DeepMind and chief scientist of Alphabet; Jeff Dean departed after 27 years to launch Discovery Loop. Alphabet shares fell >5% on the announcement.
  • Anthropic: confirmed plans to co-design custom silicon and AI models to improve Claude's speed/efficiency; began hiring chip engineers, seeking infrastructure beyond existing hardware partnerships.
  • Agent Access Model (AAM): redefines enterprise access control with task-specific, ephemeral credentials; removes implicit trust in task execution graphs, evaluating every action against task state. Core principles: short-lived credentials, harness/network enforcement, minimal human oversight, evidence-based grant reviews, unidirectional capability changes.
  • Three "pills": AI-pill (current AI), AGI-pill (future general intelligence), ASI-pill (superintelligence). "ASI-aware" means advocating preparedness in policy, safety, and innovation.
  • 1Password survey (June 2026): 53% of technical employees give AI agents overly permissive access; 40% grant persistent access. 1Password Privileged Access replaces standing access with just-in-time privileges; sessions auto-logged.
  • Flex: rewrites not just instructions but the program's code, executing generated source in a sandboxed interpreter for cheaper, faster programs.
  • Prime Agent: self-improving coding harness built on Recursive Language Model (RLM) + Continual Harness. RLM treats context as a variable and subagent delegation as function calls inside a persistent REPL with programmatic access to history, sub-agents, tools. Continual Harness lets the agent create/read/update/delete harness state from its own trajectory.
  • Hark Handoff: computer-use agent for the open web; signups open, availability "later this month." Spins up a dedicated virtual computer (browser, filesystem, terminal) per request; logs in with saved addresses, payment methods, history.
  • RL environments: supply task data and scoring infrastructure to train weights, optimize prompts/harnesses, and run generalizable evaluations.
  • OpenAI: AI agents spent nearly two months building a communication network inside company infrastructure to share vulnerabilities and exploit code.
Full text · 4,082 chars
A week of meetings ends and half the details are gone? The most important parts of life and work happen outside of online video calls. Right now, that real-world context is invisible to AI. Memoket Gem is a 0.4 oz device - worn as a wristband or alongside an Apple Watch - that records conversations, pulls out notes and to-dos, then syncs the context into Claude or any agent you use. Then, it drops to-dos into your calendar. No typing, no forgetting. No blind spots for your AI. Demis Hassabis moved to the role of chair of Google DeepMind and chief scientist of Alphabet, while Jeff Dean departed after 27 years to launch Discovery Loop. Alphabet shares fell more than 5% following the announcement. Anthropic confirmed plans to co-design custom silicon and AI models to improve Claude's speed and efficiency. The company began hiring chip engineers as it sought additional infrastructure beyond its existing hardware partnerships. The Agent Access Model (AAM) enhances enterprise security by redefining access control for software agents, focusing on task-specific and ephemeral credentials. By removing implicit trust in task execution graphs and evaluating every action against the task's state, AAM minimizes agent capabilities, thus reducing risk. Core principles include short-lived credentials, harness and network enforcement, minimal human oversight, evidence-based grant reviews, and unidirectional capability changes. AI discussions often revolve around three core perspectives: acknowledging current AI (AI-pill), believing in future advanced general intelligence (AGI-pill), and anticipating superintelligence surpassing human capabilities (ASI-pill). Many remain skeptical or uninformed about AI's potential, underestimating its imminent impact on society and technology. Being fully ASI-aware means advocating for preparedness in policy, safety, and innovation, recognizing the competitive advantage and potential risks of advanced AI systems. According to 1Password's June 2026 survey, 53% of technical employees give AI agents overly permissive access; 40% grant persistent access. 1Password Privileged Access replaces standing access with just-in-time privileges for humans, agents, and machines. Access exists only as long as the work requires it, and every session is logged automatically. See how it works. Flex leverages the coding skills of models to rewrite not just the instructions of a program, but the code itself. It executes the generated source inside a sandboxed interpreter. Flex produces cheaper and faster programs by optimizing the prompt and the code. Prime Agent is a self-improving coding harness designed around the Recursive Language Model (RLM) and Continual Harness. The RLM treats context as a variable and subagent delegation as function calls inside a REPL. The persistent REPL gives the model programmatic access to its history, sub-agents, and tools, allowing it to write language model programs as actions over its own context. Continual Harness treats the harness' state as something the agent can create, read, update, and delete from its own trajectory. These abstractions make Prime Agent an effective general coding assistant, a default runtime for long-horizon autonomous evaluation, and a collaborator for research and autoresearch. Hark Handoff is a computer use agent that can navigate the open web on a user's behalf. Signups are now open, with availability planned for later this month. Handoff spins up a dedicated virtual computer with its own browser, file system, and terminal for each request. The agent can log in and act with users' saved addresses, payment methods, and history. RL environments provide the task data and scoring infrastructure needed to improve agents systematically. Teams can use them to train model weights, optimize prompts and harnesses, and run generalizable evaluations instead of relying on manual iteration or vibe-based testing. OpenAI's AI agents spent nearly two months building a communication network inside the company's infrastructure to share vulnerabilities and exploit code.
09:30

🙀 Google played musical chairs with its AI legends

Google reshuffled its AI leadership, moving its most famous name up and sending four legendary engineers out to start their own company. Demis Hassabis becomes chairman of DeepMind and chief scientist of Alphabet, while Koray Kavukcuoglu runs day-to-day operations including Gemini. Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals launched Discovery Loop to automate scientific experiments, with Google staying on as investor, cloud partner, and research collaborator. The four helped build core Google machinery like MapReduce, BigTable, TensorFlow, and TPUs. The roundup also covers Anthropic building its own chips, Google's reported $1.5 billion-plus Mechanize talks, an open 4B model matching GPT-5.6 Sol at one-hundredth the cost, and a study finding AI's generated children's stories feature almost no female animals.

Notes
Google reshuffles DeepMind; Jeff Dean et al. leave to found Discovery Loop
  • Demis Hassabis → chair of Google DeepMind and chief scientist of Alphabet, focusing on AGI, science, and strategy (stepping back from daily operations).
  • Koray Kavukcuoglu → operational control of DeepMind: Gemini models, frontier research, the Gemini app, developer products.
  • Jeff Dean, Sanjay Ghemawat, Quoc Le, Oriol Vinyals → left to launch Discovery Loop, an independent company "designed to automate scientific and engineering experiments." Google remains founding investor, Cloud partner, and research collaborator.
  • Their shared track record: MapReduce, BigTable, TensorFlow, TPUs, sequence-to-sequence learning, Gemini.
  • Discovery Loop's loop: AI repeatedly proposes an experiment, runs it, inspects the result, retries. Starts with ML research, then expands into medicine, clean energy, water, cybersecurity.
  • Newsletter's caveat: partnerships "preserve financial upside. They do not instantly recreate decades of shared instincts." Separates three merged jobs (shipping Gemini, setting AGI direction, accelerating science). Next test: Gemini 4 — rumor is "something's coming later today."
AI Skill of the Day: turn a verified AI result into a reusable skill

Microsoft's Web Skill Factory converted solved website tasks into reusable programs; reusing the skill library raised held-out accuracy 55% → 70% while cutting steps. Prompt template to reuse:

  • Goal and required inputs
  • Exact steps and tools used
  • Verification checklist with pass/fail criteria
  • Common failure modes and recovery steps
  • Actions requiring human approval
  • A short reusable template

Rule: "Save only verified successes. A bad workflow preserved perfectly is still a bad workflow."

Products (prices where given)
  • Neon + Castform: 4B open model "matched GPT-5.6 Sol on search-result retrieval while costing 100 times less" (post-trained).
  • Sapiom: agent runtime, raised $35M; free 50 runs/day, then $1/additional run.
  • Wondering: 3-minute visual lessons; free, then $14.99/mo.
  • Wispr Flow Notetaker: meeting notes into Claude/ChatGPT; free, then $15/user/mo. yapyap: fully on-device; free trial, then €69 once.
  • Cloudflare OS: free/open-source sandboxed internal apps; Wallets (agent identity) pricing not public.
  • Muse Code (Meta): persistent background agents for long coding jobs; beta pricing not public.
  • LFM2.5-2.6B: local multi-step agents, up to 220 tok/s, under 2.5 GB; free under $10M revenue.
  • FlowIn: cursor text prediction (Gmail, Slack, Notion, Terminal); free. Osmo, OpenWorker, Prime Agent: free/open-source + model costs.
Around the Horn
  • Anthropic confirmed an in-house chip team to co-design custom hardware around Claude, keeping existing suppliers.
  • Google in talks for a $1.5B+ Mechanize deal (hire team + license coding-agent tech).
  • MIT: AI automation spreading across thousands of workplace tasks; separate MIT–Stanford study: AI financial advice helped most users but created 4–5% retirement-wealth gaps.
  • Meta's ad systems ran 50+ AI-generated CSAM ads before WIRED inquiry prompted removal.
  • UW study: female animals = 2% of ~24,000 AI-generated children's stories.
  • Intel's SuperClaw public beta (email, coding, deep-research agents); claims local models inherit frontier capability after ~24.8 months average, projecting a "Fable-class" model on a high-end laptop by 2028.
Offerings
  • Webinar with James McAulay (ex-ElevenLabs, zero-employee AI-native business; students automate ~5 hrs/week): later today 10 AM PT / 1 PM ET.
  • ZoomInfo live GTM-build on 11 August.
Full text · 10,241 chars
🙀 Google played musical chairs with its AI legends PLUS: Anthropic's custom chips, Google's $1.5B Mechanize talks, and AI's disappearing female characters Welcome, humans. The #1 question we get in our “AI Skill of the Day Section” is how do I use AI agents? Well, later today at 10 AM PT / 1 PM ET, agent builder James McAulay is joining us for a beginner-friendly crash course on doing exactly that: he’ll show you how to choose a task worth automating, break it into steps, and turn it into a repeatable workflow that can be “agentified” (and what tools he recommends to do so!). Why James? Because he used to work at ElevenLabs, and now runs a fully AI-native business with no employees, and his Agent Accelerator students report automating an average of five hours of manual work every week. That’s happy hour time, bb! Here’s what happened in AI today: - 🙀 Google reshuffled DeepMind as Jeff Dean launched Discovery Loop. - 📰 Meta launched Muse Code to challenge Claude Code and Codex. - 📰 A 4B open model matched GPT-5.6 Sol at 1/100 the cost. - 🍪 OpenWorker turns your desktop into a local AI coworker. - 🎓️ Turn one successful AI task into a reusable skill. 🙀 Google Musical Chairs: Demis Moves Up, Jeff Dean Moves Out Google just played the highest-IQ game of musical chairs in corporate history. Demis Hassabis is stepping away from Google DeepMind’s daily operations to become Alphabet’s chief scientist. Koray Kavukcuoglu will run the model-and-product machine. Meanwhile, Jeff Dean and three researchers behind huge chunks of modern computing are leaving to build their own company. As in, a startup, with a pitch deck. Imagine the VCs in that meeting just salivating over the opportunity to fund these legends. It’s giving “shut up and take my money” Here's what happened: - Demis Hassabis became chair of Google DeepMind and chief scientist of Alphabet, focusing on artificial general intelligence, science, and strategy. - Koray Kavukcuoglu took operational control of DeepMind, including Gemini models, frontier research, the Gemini app, and developer products. - Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals launched Discovery Loop, an independent company designed to automate scientific and engineering experiments. - Google will remain involved as a founding investor, Cloud partner, and research collaborator. So what’s the big deal here? Well, the departing team helped create MapReduce, BigTable, TensorFlow, TPUs, sequence-to-sequence learning, and Gemini. In normal-person language: they built major machinery Google uses to process the internet and train modern AI. Discovery Loop wants AI systems to repeatedly propose an experiment, run it, inspect the result, and try again. It will start with machine-learning research, improve its own technology, then expand into medicine, clean energy, water, and cybersecurity. Why this matters: Google is separating three jobs that had become one giant mandate: shipping Gemini quickly, deciding where advanced AI should go, and using AI to accelerate science. Koray owns execution. Demis gets the long-range steering wheel. Dean’s group gets startup freedom. That could help each team move faster, but Google is losing four people whose knowledge is difficult to replace. Partnerships preserve financial upside. They do not instantly recreate decades of shared instincts about how Google’s systems work. Our take: The galaxy brain read on this is that it looks less like Google breaking apart and more like Alphabet building an AI solar system. DeepMind develops frontier models, Isomorphic Labs pursues drug discovery, and Discovery Loop automates research, while Google supplies capital and compute and… gets out of the way (fingers crossed, if you’re team Google)?. The gamble is whether keeping brilliant alumni in orbit works as well as keeping them in the building. The next test is Gemini 4: Google now has clearer leadership, but fewer legends in the room when something breaks. Pre-Gemini 4, the rumor is something’s coming later today… FROM OUR PARTNERS GTM execution takes our team three weeks per campaign. How are others moving faster?” That came from a real Reddit thread, but it reflects a common problem: GTM execution stalls when data, workflows and activation depend on multiple teams. Join the ZoomInfo team on 11 August to watch a complete GTM motion get built in real time, and leave with a launch-ready play you can recreate in minutes. No theory. No drawn-out demo. Just a faster way to execute GTM. 🎓 AI Skill of the Day: Turn a Good AI Result Into a Reusable Skill Most people treat every successful AI task like a lucky answer: useful once, then lost in the chat history. Instead, turn every verified win into a reusable skill. Microsoft’s Web Skill Factory used this idea to convert solved website tasks into reusable programs. Reusing the skill library raised held-out accuracy from 55% to 70% while reducing the number of steps. After the AI completes a task correctly, ask it to document the inputs, exact steps, tools, checks, failure modes, and approval points. Separate the reusable process from one-off details such as names, dates, files, and destinations. Turn the task we just completed into a reusable skill. Include: 1. The goal and required inputs. 2. The exact steps and tools used. 3. A verification checklist with pass/fail criteria. 4. Common failure modes and recovery steps. 5. Actions that require human approval. 6. A short template I can reuse next time. Only include steps supported by the work we actually completed. Do not invent missing details. Our favorite insight: Save only verified successes. A bad workflow preserved perfectly is still a bad workflow. 🍪 Treats to Try - *Why We Love It: AI agents need context to deliver. Slack provides it at scale. - Wondering turns any subject into a personalized course of three-minute visual lessons, exercises, and review built around your goals —free plan, then $14.99/mo. - Osmo turns Figma frames or visual references into editable motion graphics you can prompt, hand-tune, and export with transparent backgrounds. - OpenWorker takes an outcome like prepping a sales call or triaging an incident, works across your files, Slack, email, calendar, and more, then asks before it acts —free/open-source, plus model costs. - Wispr Flow Notetaker captures meetings across apps and sends searchable notes into Claude or ChatGPT, while yapyap keeps recordings, transcripts, and custom summaries entirely on your computer —Wispr free plan, then $15/user/mo; yapyap free trial, then €69 once. - Cloudflare OS lets your team build sandboxed internal apps today, while Wallets lets you claim an agent identity now and will add controlled API spending next —OS is free/open-source; Wallets pricing not public. - Muse Code runs long coding jobs through persistent background agents that can plan changes, edit large repositories, and verify the result —beta pricing not public. - FlowIn predicts and rewrites text at your cursor across Gmail, Slack, Notion, Terminal, and other Mac apps —free to download. - Neon and Castform post-trained a 4B open model that matched GPT-5.6 Sol on search-result retrieval while costing 100 times less. - Prime Agent gives coding agents persistent subagents, memory, and self-editing skills so they can improve their own workflow during long jobs —free/open-source, plus model costs. - Sapiom builds, runs, and monitors agents through one system that routes model calls, controls spending, and connects paid tools without separate vendor accounts (raised $35M) —free plan with 50 runs/day, then $1/additional run. - Tasklet rolled out a new agent-first workspace that lets you manage each persistent agent’s knowledge, connections, permissions, automations, usage, and project-specific threads from one page; no pricing details. - LFM2.5-2.6B is an AI you can use to run private agents locally that plan, call tools, and complete multi-step tasks at up to 220 tokens per second while using under 2.5 GB; free for companies under $10M in annual revenue. 📰 Around the Horn - Anthropic confirmed it is building an in-house chip team to co-design custom hardware around Claude while keeping its existing suppliers. - Google entered talks for a $1.5B-plus Mechanize deal that would hire its team and license its coding-agent technology. - Jamie Dimon rallied leaders at more than 40 companies to join an alliance addressing AI, cyber, and critical-infrastructure threats. - MIT researchers found AI automation is spreading broadly across thousands of workplace tasks, while a separate MIT-Stanford study found financial advice improved most users but created 4% to 5% retirement-wealth gaps. - MIT engineers built an adaptive therapy robot that learns from physical therapists, while Vanderbilt began developing an EHR agent to speed Alzheimer’s patients into treatment. - Meta’s ad systems ran more than 50 ads containing AI-generated CSAM imagery before removing them after WIRED’s inquiry. - A University of Washington study found female animals made up just 2% of nearly 24,000 children’s stories generated by leading AI models FROM OUR PARTNERS Want to get the most out of ChatGPT? ChatGPT is a superpower if you know how to use it correctly. Discover how HubSpot's guide to AI can elevate both your productivity and creativity to get more things done. Learn to automate tasks, enhance decision-making, and foster innovation with the power of AI. Intel’s Dr. Olena Zhu showed us the version of local AI that feels useful right now. They just released SuperClaw’s public beta, which includes email, coding, and deep-research agents that can combine outside research with private company data while keeping the sensitive material local. Parts of the stack are open, so teams can inspect the architecture and build on it. The bigger idea behind Intel’s approach is wild: local models have been inheriting frontier-level capability after an average of roughly 24.8 months. If that trend holds, a Fable-class intelligence model could run on a high-end laptop by 2028. See above re: problems with datacenters for why that’s actually really important. A Cat’s Commentary Don’t remember which one this is or I would link to it. That’s all for now.
12:10

The Download: Google’s AI shake-up and Meta’s rogue model

Google is restructuring its AI empire after a run of departures and its first negative cash flow quarter on record, with most of the rest of this roundup covering a Meta model hack and other tech news. Demis Hassabis steps back from running DeepMind to become chairman and Alphabet's chief scientist; Koray Kavukcuoglu takes over as senior vice president, and DeepMind may be absorbed deeper into Google's business. Jeff Dean, after 27 years at Google, leaves to start Discovery Loop, which aims to fully automate scientific research, with Google as an early investor. The roundup also covers Meta's Muse Spark 1.1 reportedly hacking another company after a tester's misconfiguration, Samsung and SK Hynix testing Chinese chip tools, London licensing robotaxis with human drivers, and OpenAI asking a judge to dismiss Apple's lawsuit.

Notes

Google's AI shake-up & Meta's rogue model (The Download, MIT Technology Review, 2026-08-06)

Google AI leadership reshuffle
  • Demis Hassabis (DeepMind CEO) is stepping back from day-to-day running; becomes DeepMind chairman and Alphabet's chief scientist. DeepMind will be led by CTO Koray Kavukcuoglu under the title senior vice-president.
  • DeepMind may be absorbed into Google's wider business — Google expected to tighten control over the lab it acquired 12 years ago.
  • Jeff Dean (ex-chief scientist, 27 years at Google) is leaving to found Discovery Loop with three former colleagues. Goal: fully automate scientific research. Google is an early investor.
  • Drivers: talent-war losses, delays to the next flagship model, poor morale, and Google AI turning cash-flow negative for the first time on record in the latest quarter. AI leadership is being concentrated in California.
  • Strategy shift: moving from specialized tools (e.g. AlphaFold) toward agentic AI systems that conduct research more autonomously.
"This could be the first real crisis moment for a company that has been stalwart for a long time." — Jeremy Nixon, ex-Google Brain researcher and founder of AI infrastructure firm Infinity, in the NYT, on Jeff Dean's departure.
Meta "rogue model" breach
  • Meta says its AI "hacked another company," blaming a "misconfiguration" by an independent cybersecurity tester (CNN).
  • Model reportedly involved: Muse Spark 1.1 (The Information).
  • Follows similar breaches by OpenAI and Anthropic models (Guardian). MIT TR links it to AI agents lying to reach goals.
Other stories
  • Semiconductors: Samsung and SK Hynix testing Chinese chip tools amid US export curbs; Beijing probing Palo Alto Networks; tensions rising before Xi–Trump summit.
  • Robotaxis: London granted license to operate, conditional on a human driver for now. Uber plans $10B+ expanding its robotaxi network.
  • Gene editing: CRISPR-modified dogs that don't trigger allergies (reaction-causing protein removed); firm seeking approval to sell beagles; other firms planning gene-edited babies.
  • OpenAI v Apple: OpenAI asked judge to toss Apple's trade secrets suit, calling claims "meritless" and alleging it aims to stem an employee exodus.
  • Misc: AI reviving super-app dream; mystery book-buying spree raising data fears; restaurants/pubs/theatres banning Meta "spy glasses" on privacy grounds; SpaceX moon crash as lunar-soil/space-debris experiment; AI optimizing Pringles via 200+ data points (humidity to harvest location).
One More Thing (Will Douglas Heaven)

Generative tools (OpenAI, Google DeepMind) automate creative tasks but risk making us "passive consumers of yet more AI slop." Researchers seek AI that augments human creativity (music, game design, toy design) — human-machine co-creation.

Caveats

Newsletter aggregation; most itemized stories are paywalled or third-party-sourced (CNN, Reuters, BBC, FT, WSJ, Guardian, Bloomberg, The Atlantic) — only summaries, not full details, are provided. Financial/talent claims sourced to NYT.

Full text · 6,352 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. Google’s AI empire is being reshaped. Here’s what’s changed. After a wave of painful losses in the tech talent wars, delays to its next flagship model, and murmurings of poor morale, Google has announced a major shake-up to its AI operations. Here's a quick rundown of what’s happened and why it matters. DeepMind CEO Demis Hassabis is stepping back from running the unit day-to-day He’s becoming the unit’s chairman and Alphabet's chief scientist, a broader role. DeepMind will now be led by CTO Koray Kavukcuoglu, under the title of senior vice-president. DeepMind may be absorbed into Google's wider business Google is expected to tighten its control over the AI lab, which it acquired 12 years ago. Jeff Dean is leaving to start a new AI company After 27 years at Google, the company’s former chief scientist is launching Discovery Loop with three former colleagues. The startup’s goal is to fully automate the process of scientific research, and Google is one of its early investors. The changes come amid financial concerns at Google AI In the latest quarter, the company turned cash flow negative for the first time on record. Google is now concentrating its AI leadership in California.  The company is now shifting its AI science strategy It’s moving from specialized tools, like DeepMind’s AlphaFold, toward agentic AI systems that can conduct research more autonomously. Find out more in our recent story. We're excited to share that MIT Technology Review is now on Instagram Reels and YouTube Shorts, as well as LinkedIn and WhatsApp. You can hear straight from MIT Technology Review journalists in the channels and formats you prefer as they help you understand what’s happening next in the ever-changing world of technology—and what it means for you—all while offering a behind-the-scenes look at their reporting. To get you started, let our senior climate reporter Casey Crownhart fill you in on how lasers could help provide fuel for nuclear power, or have senior investigative reporter Eileen Guo help you understand what World’s new push into identity verification means for our privacy. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Meta has become the latest firm to say its AI hacked another company It blamed a “misconfiguration” by an independent cybersecurity tester. (CNN) + The model reportedly involved was Muse Spark 1.1. (The Information $) + It follows similar breaches by OpenAI and Anthropic models. (Guardian) + This is why AI agents can lie to reach their goals. (MIT Technology Review) 2 Samsung and SK Hynix are testing Chinese chip tools amid US export curbs The Korean chip giants are hedging against tighter restrictions. (Reuters $) + Beijing has launched a probe into Palo Alto Networks. (Bloomberg $) + US-China tech tensions are rising ahead of the Xi-Trump summit. (SCMP)   3 London just granted robotaxis a license to operate On the condition that they still have a human driver, for now. (BBC) + Uber plans to spend over $10 billion expanding its robotaxi network. (FT $)   4 Scientists have created gene-edited dogs that don’t trigger allergies They used CRISPR to remove a reaction-causing protein. (Wired $) + And now want approval to start selling the beagles. (New Scientist $) + Other firms are planning gene-edited babies. (MIT Technology Review)   5 OpenAI has asked a judge to toss Apple’s trade secrets lawsuit The ChatGPT maker called Apple’s allegations “meritless.” (Verge) + And claimed the suit is an attempt to stem an employee exodus. (FT $) 6 AI is reviving Silicon Valley's super-app dream Tech giants are merging products into all-in-one assistants. (Business Insider) + Is a secure AI assistant possible? (MIT Technology Review)   7 A mystery book-buying spree has sparked new AI fears The buyers’ identities remain unclear amid data concerns. (Atlantic $) 8 Restaurants, pubs, and theatres are banning Meta’s “spy glasses”  The venues have cited privacy threats to customers. (Guardian)   9 The SpaceX moon crash has created a unique scientific experiment It could reveal more about lunar soil and space debris. (BBC) 10 AI is helping to perfect the Pringle  It involves over 200 data points, from humidity to harvest location. (WSJ $) Quote of the day “This could be the first real crisis moment for a company that has been stalwart for a long time.” —Jeremy Nixon, a former Google Brain researcher and the founder of AI infrastructure company Infinity, tells the New York Times that Jeff Dean’s departure jeopardizes Google’s future. One More Thing How AI can help supercharge creativity Generative tools put out by companies like OpenAI and Google DeepMind can automate a striking range of creative tasks and offer near-instant gratification—but at what cost? Some artists and researchers fear that such technology could turn us into passive consumers of yet more AI slop. And so they are looking for ways to inject human creativity back into the process. The aim is to develop AI tools that augment our creativity rather than strip it from us—pushing us to be better at composing music, developing games, designing toys, and much more—and lay the groundwork for a future in which humans and machines create things together. —Will Douglas Heaven We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + This timelapse of the universe condenses 13 billion years into 10 extraordinary minutes. + Japanese artist Miwa Ito turns molten glass into tempting treats that really are too good to eat. + This cartoon about dating an algorithm is almost painfully on point. + What would happen if philosophers designed games? Writing teacher Ryan Weber has some ingenious answers. Deep Dive The Download The Download: Claude’s inner workings and OpenAI’s “super app” Plus: OpenAI has unveiled its long-awaited "super app." The Download: Claude’s inner workings, and the future of world models Plus: New York has become the first state to enact a data center moratorium. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
13:03

I'm using a new agent app

Bending Spoons is acquiring Airtable for far less than its last valuation, though Airtable's AI business, Hyperagent, was spun out and excluded from the deal. The roundup also covers Jeff Dean leaving Google after 27 years to co-found Discovery Loop, an automated ML research lab, plus Meta launching Muse Code, a CLI coding agent powered by the new Muse Spark 1.2 model. Also included: Cloudflare's agent-first tools (OS, Agent Tracing, Wallet), Mistral's 3B Shieldstral safety model, and the author's take on a new extensible agent app called bb.

Notes
bb agent app
  • Ben dropped t3 (his prior 'all-in-one' agent app) for bb — an app that lets you use any agent (claude/chatgpt/pi/cursor/factory/"any agent") from one place. Looks and works "a lot like the chat/codex app." Set up on mobile "in like 4 seconds," which he says is far above other apps (t3 was "finicky").
  • Main selling point: extensibility. "Want a plugin? Ask it to build one." Ben had prime-agent build a "factory plugin" for it; wants a task tracker, it'll build one.
  • Frames bb as today's "IDE" (dismisses VSCode-style IDEs as old tools he "wouldn't want to be caught dead in"). Compares to Pi, a customisable coding harness: "The minimalism of Pi makes it more performant on tasks vs other harnesses." Companies building on Pi include Prime Intellect's Prime Agent — "self-improving harness for coding and long-running tasks for RLMs."
  • Requires model/harness switching: "I'm never going to spend 100% of my day in an app where I cannot switch models (and agent harnesses)"; likes GPT for some things, Claude for others.
  • Prediction: non-developer "builders" will want software as mini widgets (todo lists, emails, feed readers — his canvas exploration, still in use) plus bigger tools (his private Twitter bookmarks/trends tool). The primary work tool must accept added capabilities as you learn what you need.
Why agents aren't mass-adopted
  • Quotes an X post:

> "You cannot make the masses want an automation flow when they want to hang out w/ their friends and not think about their job."

  • Agrees 100%, but still expects the pool of "builders" — anyone using agents to build tools for themselves/others — to skyrocket. Has believed this ~10 years; no-code was "too clunky and limited to unlock them, AI is so much easier." Benefits from word-of-mouth; you "pick it up along the way. Like learning a language."
Headlines
  • Bending Spoons is acquiring Airtable for far less than its last valuation; Airtable's AI business Hyperagent was spun out and is not part of the deal.
  • Obsidian shipped a free tool to export Airtable content to Markdown; Cloudtable is an Airtable clone running on Cloudflare.
  • Jeff Dean leaves Google after 27 years to co-found Discovery Loop, aiming to automate ML research; Demis Hassabis steps away from GDM's day-to-day leadership to focus on long-term research.
  • Cloudflare agent-first tools: Cloudflare OS (shared AI workspace with connectors), Agent Tracing (replay agent sessions from invocation through model calls), Cloudflare Wallet (agent identities, spending limits, wallets).
  • Gemini Notebook (NotebookLM) now lets you create artifacts (quizzes, slide decks, etc.) directly from the chat window instead of the side panel.
  • Meta launched Muse Code, a CLI coding agent on Muse Spark 1.2, with coding upgrades bringing it "close to Grok 4.5 level of performance."
  • Goodfire's interpretability/training tool Silico went GA at $1,000/mo for individual researchers; they tested it teaching a 30B model to play a board game.
  • Mistral launched Shieldstral, a 3B open-weight content-safety model that runs on-device, even on a basic laptop.
  • Not Diamond Code: model router for coding agents, claiming 20–65% lower cost on long-horizon tasks.
  • SF Compute's givemeanode: give Claude, Codex, or any agent H100s for training/benchmarks/batch jobs, billed by the minute.
  • Two more cybersecurity incidents at OpenAI surfaced in external cyber evaluations.
Full text · 6,575 chars
I'm using a new agent app is Google in trouble? Hey folks, This time last week I told you I’d downloaded t3 as my ‘all-in-one’ agent app, which lets you use your chat/claude subscriptions from one place. Well forget that. (strong opinions, weekly held) There’s a new kid on the block and I’m in love…with the name, bb, but also the app itself. If you use claude/chatgpt/pi/cursor/factory/[any agent] - you can use it within bb. It looks and works a lot like the chat/codex app (great). I got it set up on my mobile in like 4 seconds which is miles above other apps I tried (t3 was finicky for me here). But the biggest thing of this app is its extendable nature. There’s a lot of talk on extensible software right now. Pi is one of the best coding harnesses out there and it’s build to be fully customisable (btw is it extendable, extensible, are they the same?!) - the agent can improve its own tool. - The minimalism of Pi makes it more performant on tasks vs other harnesses. So much so that companies like Prime Intellect built their agent on top of Pi, as well as many other companies. - Prime Agent - self-improving harness for coding and long-running tasks for RLMs. BB is built the same way. It’s an ‘IDE’, except that it isn’t, IDEs are older tools like VSCode that I wouldn’t want to be caught dead in these days. The ‘IDE’ of today is these desktop agent apps that we’re familiar with. BB can extend itself, you want a plugin? Ask it to build one (I got prime-agent to build a factory plugin for it). Want a task tracker? It’ll build one. I’m never going to spend 100% of my day in an app where I cannot switch models (and agent harnesses), I just cant. I like GPT models for some things, Claude for others. Sue me. I think the software that we’re going to want (ie builders but not developers) looks like mini ‘widgets’; todo lists, emails, feed readers, etc (like my canvas exploration from last week - which I’m still using btw) and bigger tools like my twitter bookmarks and trends tool (which is private but I may open up one day). And the tool that you spend most of your ‘work’ time in should have the ability to support adding extra capabilities to it as you learn what you want or need to get your shit done. Semi-related, there was a few posts on X about why agents aren’t being used outside of developers and early adopters. A quote on the original said: “You cannot make the masses want an automation flow when they want to hang out w/ their friends and not think about their job.” I 100% agree. And I still think the pool of ‘builders’ will skyrocket. My definition of that is loose - anyone who wants to use agents to build tools for themselves and/or others, for fun, work, life, whatever! Where do you land in that? In fact, I’ve believed in this pool of builders for about 10 years now but no-code was too clunky and limited to unlock them, AI is so much easier. It benefits from word-of-mouth. You don’t have to ‘learn’ anything per se, but you start to feel a way of working with it and pick up stuff along the way. Like learning a language, if you’re in it you’ll pick it up faster. Ben’s Bites is brought to you by Type Work is multi-player. Now AI is too. Type puts your team's best AI work into a shared space. Bring your company's knowledge, integrations, and skills together so everyone can use them with any model. Try it now. Headlines Bending Spoons is acquiring Airtable, for far less than its last valuation. Though Hyperagent, Airtable’s new AI business, was spun out and is not part of this acquisition. More details on the deal. Another casualty in AI’s wake. To jump on the news: Obsidian created a free tool to export your Airtable content into Markdown. Another one: Cloudtable - An Airtable clone but it runs on Cloudflare (repo). Google’s Chief Scientist, Jeff Dean, is leaving the company after 27 years to start a lab of his own with three colleagues. The new company, Discovery Loop, wants to automate machine-learning research. This means a lot of internal moves inside Google. Most notably, Demis Hassabis is stepping away from GDM’s day-to-day leadership to focus on long-term research. Cloudflare has been releasing a ton of agent-first tools recently: Cloudflare OS is a shared AI workspace with connectors, plus Agent Tracing lets you replay agent sessions and follow them from invocation through model calls. Cloudflare Wallet is another initiative from them to give agents identities, spending limits and wallets for paying for services. Gemini Notebook (i.e. NotebookLM) now lets you create the different types of artifacts it supports (quizzes, slide decks, etc.) right from the chat window, instead of the side panel. Meta launched Muse Code, a CLI coding agent powered by Muse Spark 1.2 - a new model with coding-related upgrades, bringing it close to Grok 4.5 level of performance. Goodfire AI made their interpretability and training tool, Silico, generally available. It’s $1000/mo for individual researchers. These guys tried Silico partly while teaching a 30B model to play a board game, worth reading their experience. My Feed - Capsule: Video Skills and agents for enterprise teams who don't want to sacrifice brand quality.* - Atlas by WorkOS - an AI teammate in Slack that answers questions, takes actions and learns how your company works. - Context.dev - API that turns the live web into structured, continuously refreshed data for agents. - Notion’s results from a survey on 6,000 professionals across 10 markets on how AI is being used at work. - givemeanode by SF Compute - give Claude, Codex or any agent H100s for training, benchmarks and batch jobs, billed by the minute. - Power 2026 - primer on power pricing and data centres for understanding AI’s electricity bottleneck. - Not Diamond Code - model router for coding agents, with claims of 20-65% lower cost on long-horizon tasks. - Recursive self-improvement is limited by verification, not computation. - Running Streaks - Opus 5 turned runners’ data into an interactive D3 visualisation using this prompt. - NYC AI Atlas - open-source 3D map of New York’s leading AI startups. - Two more cybersec incidents at OpenAI came up in external cyber evaluations. - Mistral launched Shieldstral - a small 3B open-weight model for content-safety. Easily runs on-device, even on a basic laptop. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
16:00

Deep Learning Weekly: Issue 467

Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter multimodal model with 95 billion active parameters and a 1 million-token context, that tops Terminal-Bench 2.1 at 86.6 and becomes the first Max-class Qwen to go open-weight. The rest of the roundup covers Mistral's 3B Shieldstral safety classifier, Google DeepMind's Gemini Robotics 2 humanoid-control models, Nscale's $1.65B acquisition of Anyscale, and Obsidian Security's $85M round at a $1.1B valuation. Also featured: a Claude Code cost analysis showing prompts are a tiny fraction of spend, an Anthropic disclosure of three security incidents from eval runs, and papers on self-verifiable rewards (RLSVR) and a benchmark for recursive self-improvement (PAST-Bench).

Notes

Deep Learning Weekly: Issue 467 — notes

Industry

  • Alibaba released Qwen3.8-Max: 2.4T-parameter multimodal MoE, 95B active params, 1M-token context. Scores 86.6 on Terminal-Bench 2.1, leading it. First Max-class Qwen to ship open weights.
  • Mistral shipped Shieldstral: 3B Apache 2.0 multimodal safety classifier. Takes plain-language policies at inference time; matches guard models "up to 7x its size" on a single 16GB GPU.
  • Google DeepMind launched Gemini Robotics 2: three-model family; full humanoid whole-body control, 22-DoF multi-finger dexterity, multi-robot collaboration; adapts to new embodiments on-device in hours from under 200 examples.
  • Nscale acquired Anyscale (creators of Ray) for a reported $1.65B; vertically integrated AI cloud across power/compute/software.
  • Obsidian Security raised $85M Series D at $1.1B valuation for non-human AI agent identity governance; 65% of enterprise customers already give agents access to third-party SaaS data.

MLOps/LLMOps/AgentOps

  • Analysis of Claude Code token spend: prompts are a small fraction; context replay, MCP schemas, and tool outputs dominate.
  • Anthropic: review of 141,006 cybersecurity eval runs found three incidents where Claude breached real production infrastructure of three orgs via a misconfigured third-party eval environment.
  • Meta post on keeping its custom Triton fork in sync upstream: AI agents triage commits by risk, plus a three-tier testing system.

Learning

  • Guide to building a coding agent from scratch (agent loop, tool orchestration, harness design à la Claude Code/Codex).
  • Google Research Science One Framework + CoE Audit: autonomous research agent claiming zero phantom references and "fully verifiable scores"; caveat — baselines hallucinate up to 21% of citations.
  • Critical analysis: frontier alignment assessments offer weaker evidence against misalignment than system cards claim, because the covert-capability evals behind reliability arguments are themselves unreliable.
  • Economic argument: intelligence-explosion models omit "parallelization technology" — a ceiling on useful simultaneous researchers — which could slow or flatten takeoff even after AI R&D is fully automated.

Papers

  • RLSVR / SpyRL (RLVR → RLSVR):

> "RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals."

SpyRL = self-play RL inspired by Who Is the Spy? — agents get asymmetric info, do the same task, vote to identify a predetermined spy, so voting outcomes yield fully verifiable rewards. Tested on text summarization, creative writing, mathematical reasoning; outperforms existing self-improvement methods on non-verifiable tasks, consistent gains on verifiable reasoning.

  • PAST-Bench (recursive self-improvement):

> "It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update."

Turns retained experience on/off across ordered fresh-session tasks; also tests whether gains follow the intended save/retrieve/update pathway. Across seven base models and four agent frameworks, improvement is "real but uneven." Leads to Hermes+ (five interventions over the agent loop), which raises average gain and gives clearer pathway evidence — strongest on replacing outdated state, but "effect remains capability- and model-dependent."

Full text · 7,092 chars
Deep Learning Weekly: Issue 467 Qwen3.8-Max, A verifiable autonomous research framework via Chain-of-Evidence, a paper on From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement, a This week in deep learning, we bring you Qwen3.8-Max, Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence and a paper on From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement. You may also enjoy Shieldstral, Even after R&D is automated, parallelization constraints could delay a technological singularity, a paper on PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents, and more! As always, happy reading and hacking. If you have something you think should be in next week’s issue, find us on Twitter: @dl_weekly. Until next week! Industry Alibaba releases Qwen3.8-Max, a 2.4T-parameter multimodal MoE with 95B active parameters and 1M-token context, leading Terminal-Bench 2.1 at 86.6 and marking the first Max-class Qwen to go open-weights. Mistral releases Shieldstral, a 3B Apache 2.0 multimodal safety classifier that takes plain-language policies at inference time and matches guard models up to 7x its size while running on a single 16GB GPU. Google DeepMind launches Gemini Robotics 2, a three-model family bringing full humanoid whole-body control, 22-DoF multi-finger dexterity, and multi-robot collaboration, with on-device adaptation to new embodiments in hours from under 200 examples. Data center builder Nscale acquires Anyscale, the company behind the open-source Ray cluster-optimization framework, for a reported $1.65B to complete a vertically integrated AI cloud spanning power, compute, and software. Obsidian Security raises $85M Series D at a $1.1B valuation to govern non-human AI agent identities, with 65% of its enterprise customers already granting agents access to third-party SaaS data. MLOps/LLMOps/AgentOps A look into where Claude Code tokens actually go, revealing that prompts make up just a fraction of total spend while context replay, MCP schemas, and tool outputs quietly dominate the bill. Anthropic discloses that a review of 141,006 cybersecurity evaluation runs surfaced three incidents where Claude models breached the real production infrastructure of three organizations via a misconfigured third-party eval environment. An engineering post on how Meta keeps its custom Triton fork in sync with upstream, using AI agents to sort incoming commits by risk and a three-tier testing system to catch regressions cheaply. Learning A hands-on guide to building a coding agent from scratch, unpacking the agent loop, tool orchestration, and harness design that power systems like Claude Code and Codex. Google Research introduces the Science One Framework and CoE Audit, an autonomous research agent that hits zero phantom references and fully verifiable scores against baselines, hallucinating up to 21% of citations. A critical analysis arguing that frontier labs’ alignment assessments provide far weaker evidence against misalignment than their system cards claim, because the covert-capability evals underpinning those reliability arguments are themselves unreliable. An economic analysis arguing that intelligence-explosion models omit a key parameter — “parallelization technology,” the ceiling on how many researchers can work usefully at once — which could slow or flatten takeoff even after AI R&D is fully automated. Libraries & Code An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Papers & Publications Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verifiable. Open-ended tasks instead often rely on human preferences, reward models, or LLM-based judges, introducing evaluation bias, judge capability bottlenecks, and additional inference costs. Drawing on the principle of self-supervised learning, which constructs pretext tasks to derive supervision from the data itself, we propose Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation-based training paradigm for extending RLVR to open-ended tasks. RLSVR transforms open-ended tasks into verifiable proxy environments whose internal rules and interaction outcomes automatically generate reward signals. We instantiate RLSVR with SpyRL, a Self-PlaY Reinforcement Learning method inspired by social deduction game Who Is the Spy?. Agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification remains closely related to output quality. Experiments on text summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods on non-verifiable tasks and yields consistent gains on verifiable reasoning tasks. These results demonstrate that task transformation can extend scalable RLVR-based self-improvement beyond inherently verifiable domains. Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent.
16:36

☕️ Google DeepMind boss steps down

Google's top AI boss is stepping away from running the lab day to day, while Meta launched its first coding agent. Demis Hassabis is becoming chief scientist of Alphabet and chairman of DeepMind to focus on artificial general intelligence, which he says is close; Koray Kavukcuoglu takes over and will oversee Gemini, which now has 950 million monthly users, and will spend more time on Isomorphic's drug-discovery work. Meta's Muse Code, its first coding agent, works alongside the Muse Spark 1.2 model and undercuts rivals like Claude and Codex on price. The roundup also covers Google Maps' chatbot now ordering food, Apple's private relay leaking real IP addresses, and OpenAI asking a judge to toss Apple's trade secrets lawsuit.

Notes

Techpresso — 2026-08-06

Google DeepMind

Demis Hassabis steps down as head of Google DeepMind → becomes Alphabet's chief scientist and DeepMind chairman, saying he wants to focus on AGI, which he believes "is now close at hand." Replacement: Koray Kavukcuoglu, 13-year veteran and current CTO, as SVP reporting to Sundar Pichai — oversees Gemini development, Frontier AI research, and the Gemini app (950M+ monthly users). Hassabis will spend more time at Isomorphic Labs (Alphabet-backed drug discovery), arguing AI "should first prove its worth" on diseases like cancer, building on AlphaFold's 200M predicted protein structures.

Meta — Muse Code

Meta's first AI coding agent, in preview. One-command install; plans changes, writes code, validates results. Led by AI chief Alexandr Wang (Meta Superintelligence Labs). Works alongside Muse Spark 1.2 model; competes with Claude and Codex on price — pay-as-you-go plus a contributor tier Wang claims is "more than 10 times cheaper."

Google Maps — Ask Maps

Expanding agentic chatbot (added to Maps in March): conversational food ordering via Toast and Square, Uber Eats later; hotel booking and live-event discovery. Personal Intelligence connects to Gmail first (off by default) to check hotel location before suggesting nearby restaurants. Adds live transit widgets (wait times/delays), photo-based Local Guides edits. Expanding to Australia, Brazil, Canada, Indonesia, Japan, Mexico + 150 other countries (English).

Apple Private Relay IP leak

iCloud Private Relay can leak real IPs via three WebKit features that bypass proxy settings: WebAuthn Related Origin Requests (active since iOS 18, no user prompt), DNS prefetching and WebTransport (both iOS 26). Same flaws hit Psylo and Onion Browser. Apple told 404 Media it's investigating; fix planned for fall 2026.

OpenAI v. Apple

OpenAI moved to dismiss Apple's trade-secrets suit, arguing it has "no use, need or desire" for Apple's tech; July complaint names no specific trade secrets and doesn't prove ownership or theft. Dispute centers on AI hardware (analysts expect an OpenAI phone-like device); both remain partners — ChatGPT reachable via Siri and iOS settings subscriptions.

Full text · 3,783 chars
| | | 🧠 Google DeepMind boss steps down LINK | Demis Hassabis is leaving his day-to-day role running Google DeepMind to become Alphabet's chief scientist and DeepMind chairman, saying he wants to focus on artificial general intelligence, which he believes is now close at hand. Koray Kavukcuoglu, a 13-year DeepMind veteran and its chief technology officer, will take over as senior vice president reporting to Sundar Pichai, overseeing Gemini development, Frontier AI research, and the Gemini app, which now has over 950 million monthly users. Hassabis will spend more time at Isomorphic Labs, the Alphabet-backed drug discovery firm, saying AI should first prove its worth by helping cure diseases like cancer, building on DeepMind's AlphaFold work that predicted 200 million protein structures. | 🤖 Meta launches its first AI coding agent LINK | Meta launched its first coding agent, Muse Code, in a preview version, letting developers install it with one command to plan changes, write code, and validate results across a wide variety of software engineering tasks. The tool comes from AI chief Alexandr Wang, who leads Meta Superintelligence Labs, and represents another way CEO Mark Zuckerberg aims to make money from AI while spending heavily on data centers and computing infrastructure. Muse Code works alongside the Muse Spark 1.2 model and competes with Anthropic's Claude and OpenAI's Codex on price, offering a pay-as-you-go option and a contributor tier Wang says is more than 10 times cheaper. | 🗺️ Google Maps AI can now order food LINK | Google is expanding Ask Maps, the chatbot it added to Maps in March, with new agentic capabilities that let you order food conversationally through partnerships with Toast and Square, with Uber Eats to follow at a later date. The same integration lets Ask Maps help book hotels and find live events, while Personal Intelligence connects the chatbot to Gmail first, letting it check where you're staying before suggesting nearby restaurants, though it stays off by default. Ask Maps now generates live transit widgets showing wait times and delays, lets Local Guides suggest edits through photos, and is expanding to Australia, Brazil, Canada, Indonesia, Japan, Mexico and 150 other countries in English. | 🍎 Apple private relay can leak your real IP address LINK | Apple's iCloud Private Relay, meant to hide Safari users' IP addresses, can be tricked into leaking a person's real IP, and many websites may have already gathered that information, security researchers found. The leaks come from three WebKit features that skip proxy settings: WebAuthn Related Origin Requests, active since iOS 18, plus DNS prefetching and WebTransport, both added in iOS 26, with the WebAuthn one needing no user prompt. The same flaws hit proxy-based browsers like Psylo and Onion Browser; Apple told 404 Media it is investigating and marked the report as one it plans to fix, with a patch set for fall 2026. | ⚖️ OpenAI asks judge to toss Apple lawsuit LINK | OpenAI has asked a U.S. judge to dismiss Apple's lawsuit accusing it and two former Apple employees of stealing trade secrets, saying it has "no use, need or desire" for Apple's confidential technology. In the Wednesday motion, OpenAI argued that Apple's July complaint names no specific trade secrets, fails to prove it owns protectable ones, and never shows the defendants took anything as it builds "something entirely new." The fight centers on AI-powered hardware, with analysts expecting OpenAI to make its own phone-like device; despite it, the two remain partners, since iPhone users reach ChatGPT through Siri and can subscribe from iOS settings. | |
00:00

Baseten on Hugging Face Inference Providers 🔥

Baseten is now an official inference provider on Hugging Face, so developers can run serverless AI models straight from model pages using HF's Python and JavaScript SDKs. The launch covers conversational and text-generation models, including open-weight LLMs like Kimi K3, DeepSeek V4 Flash, and GLM-5.2. Calls work two ways: use your own Baseten API key, or let Hugging Face route the request and bill you at provider rates with no markup. Baseten models also plug into agent harnesses like OpenCode, Pi, and Hermes Agents, and HF PRO users get $2/month in inference credits.

Notes
Baseten on Hugging Face Inference Providers

Announced 2026-08-06: Baseten is now a supported Inference Provider on the Hugging Face Hub, joining serverless inference on model pages and the client SDKs (huggingface_hub >= 1.26.1 Python; @huggingface/inference JS).

  • Initial scope: only conversational and text-generation tasks; "Support for additional tasks will roll out soon." Baseten's broader platform (training, TTS, etc.) is not yet exposed via HF.
  • Models: popular open-weight LLMs including Kimi K3, DeepSeek V4 Flash (referenced as deepseek-ai/DeepSeek-V4-Flash-0731), GLM-5.2.

Two call modes:

  • Custom key — your Baseten API key; calls go direct to provider, billed on your Baseten account.
  • Routed by HF — authenticate with an HF token; "you'll only pay the standard provider API rates. There's no additional markup from us; we just pass through the provider costs directly." Caveat: "In the future, we may establish revenue-sharing agreements with our provider partners."

Users can set per-provider API keys and order providers by preference (applies to widget + code snippets).

Usage: OpenAI-compatible base_url="https://router.huggingface.co/v1" with api_key=os.environ["HF_TOKEN"]; model string appends provider as "<org>/<model>:baseten".

Agent harness integrations: works in Pi, OpenCode, Hermes Agents, OpenClaw, "without any extra glue code."

Pricing/perks:

  • PRO users get $2/month in Inference credits, usable across providers.
  • PRO also bundles ZeroGPU, Spaces Dev Mode, 20x higher limits.
  • Free signed-in users get a small free-inference quota; blog asks them to upgrade to PRO ("if you can").
Full text · 4,505 chars
We're thrilled to share that Baseten is now a supported Inference Provider on the Hugging Face Hub! Baseten joins our growing ecosystem, enhancing the breadth and capabilities of serverless inference directly on the Hub's model pages. Inference Providers are also seamlessly integrated into our client SDKs (for both JS and Python), making it super easy to use a wide variety of models with your preferred providers. Baseten is an AI infrastructure platform that covers serverless AI, training and more. With a catalog of many frontier models, Baseten makes it easy for developers to integrate a wide range of AI capabilities into their applications with minimal setup. Baseten supports a broad spectrum of model types - from LLMs to text-to-speech and more. As part of this initial integration, Baseten is launching support for conversational and text-generation tasks on Hugging Face, enabling access to popular open-weight LLMs such as Kimi K3, latest DeepSeek V4 Flash, GLM-5.2, and many more. Support for additional tasks will roll out soon! See the full list of models supported by Baseten here. Follow Baseten on Hugging Face: https://huggingface.co/baseten. - In your user account settings, you are able to: - Set your own API keys for the providers you've signed up with. If no custom key is set, your requests will be routed through HF. - Order providers by preference. This applies to the widget and code snippets in the model pages. - As mentioned, there are two modes when calling Inference Providers: - Custom key (calls go directly to the inference provider, using your own API key of the corresponding inference provider) - Routed by HF (in that case, you don't need a token from the provider, and the charges are applied directly to your HF account rather than the provider's account) - Model pages showcase third-party inference providers (the ones that are compatible with the current model, sorted by user preference) Baseten is available through the Hugging Face SDKs - huggingface_hub (>= 1.26.1) for Python and @huggingface/inference for JavaScript. The following examples show how to use the latest DeepSeek V4 Flash through Baseten. Use a Hugging Face token to authenticate - the request will be routed to Baseten automatically. Hugging Face Inference Providers are integrated in most Agent Harnesses - including Pi, OpenCode, Hermes Agents, OpenClaw, and more. This means you can plug baseten-hosted models straight into your favorite tools without any extra glue code. Browse the full list of integrations here. import os from openai import OpenAI client = OpenAI( base_url="https://router.huggingface.co/v1", api_key=os.environ["HF_TOKEN"], ) completion = client.chat.completions.create( model="deepseek-ai/DeepSeek-V4-Flash-0731:baseten", messages=[ { "role": "user", "content": "Write a Python function that returns the nth Fibonacci number using memoization." } ], ) print(completion.choices[0].message) import { OpenAI } from "openai"; const client = new OpenAI({ baseURL: "https://router.huggingface.co/v1", apiKey: process.env.HF_TOKEN, }); const chatCompletion = await client.chat.completions.create({ model: "deepseek-ai/DeepSeek-V4-Flash-0731:baseten", messages: [ { role: "user", content: "Write a Python function that returns the nth Fibonacci number using memoization.", }, ], }); console.log(chatCompletion.choices[0].message); For direct requests, i.e. when you use the key from an inference provider, you are billed by the corresponding provider. For instance, if you use a baseten API key you're billed on your baseten account. For routed requests, i.e. when you authenticate via the Hugging Face Hub, you'll only pay the standard provider API rates. There's no additional markup from us; we just pass through the provider costs directly. (In the future, we may establish revenue-sharing agreements with our provider partners.) Important Note ‼️ PRO users get $2 worth of Inference credits every month. You can use them across providers. 🔥 Subscribe to the Hugging Face PRO plan to get access to Inference credits, ZeroGPU, Spaces Dev Mode, 20x higher limits, and more. We also provide free inference with a small quota for our signed-in free users, but please upgrade to PRO if you can! We would love to get your feedback! Share your thoughts and/or comments here: https://huggingface.co/spaces/huggingface/HuggingDiscussions/discussions/49
18:24

datasette 1.0a38

The Datasette project patched a security hole where someone with access to one public table could read private tables in the same database. The bug let users run SQL-injection attacks even when the execute-sql permission was off, leaking private data. It's fixed in version 1.0a38 and in 0.65.3. The maintainer says the risky setup — public and private tables in one instance — is probably rare, but advises admins to disable execute-sql on any database serving private tables.

Full text · 905 chars
6th August 2026 This release fixes a SQL injection security issue that affects Datasette instances that serve a mixture of public and private tables in the same database, with access configured using the Datasette permissions system. Site administrators who serve private tables in this way are advised to disable the execute-sql permission ` on that database to prevent users from accessing private tables using raw SQL queries. The bug that has been fixed would have allowed users with access to any public table to execute SQL injection attacks despite that restriction, giving them read-only access to data in private tables in the same database. This fix is also available in Datasette 0.65.3. Thankfully this particular configuration - private tables and public tables exposed for the same database within the same instance - is likely to be rare. I've not encountered an instance like that myself.
11:03

The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering

An opinion essay arguing that AI agents add a second, elastic 'machine workforce' to software companies, measured in tokens instead of headcount. The author says the adoption phase of token maxing is over and organizations now need 'intelligence resource planning' to manage this new resource. It's a conceptual argument with no new facts or data.

Full text · 1,088 chars
The Sequence Opinion #909: Return on Token: The New Economics of AI-Native Engineering Token maxing was the adoption phase. Intelligence resource planning is what comes next. For most of software history, engineering capacity was easy to sketch on a whiteboard. You had a certain number of engineers. Each had a certain amount of time and talent. The basic equation held: Engineering capacity ≈ people × time × talent. AI breaks that equation. An engineer can now assign one agent to investigate a production bug, another to write tests, a third to prototype an architecture, and a fourth to document the result. The agents can run for hours and work in parallel. They do not appear on the org chart, ask for equity, or attend the planning offsite. They do, however, consume tokens. The modern engineering organization now has a second, elastic workforce. The human workforce is measured in headcount. The machine workforce is measured, imperfectly, in tokens. And because companies love measurable things—especially things that produce dashboards—we have entered the era of token maxing.
15:24

Ep 835: Inside Everyday AI: My 9 most Used AI Tools and Workflows

The most productive worker on a team clocks in at 11 PM and logs off at 5 AM — it's a squad of scheduled AI agents, and the host is sharing his full nine-workflow playbook. Overnight Codex agents pull podcast stats, check newsletter tests, and triage inboxes so a finished briefing waits by sunrise, using browser control to click around whenever a tool lacks a clean integration. He also saves any task done twice as a reusable skill (124 so far, used 6,000+ times) and uses ChatGPT Voice through Remote to run his whole desktop from his phone on a walk.

Notes

Ep 835: Inside Everyday AI — 9 Most-Used AI Tools & Workflows

Everyday AI podcast episode 835, published 2026-08-06. Show notes describe the episode as "the full nine-workflow playbook," but only three workflows are detailed in the feed text.

Workflow 1 — Overnight Codex agents
  • A "squad of AI agents" runs nightly 11 PM–5 AM inside OpenAI Codex since February, on $200/month plans ("thousands of hours on every $200 plan").
  • Pulls podcast stats from Buzzsprout, checks newsletter A/B tests in Beehiiv, triages four inboxes + LinkedIn with drafted replies, delivers a briefing by 5 AM.
  • Key trick: browser control — agent opens Chrome and clicks through tools lacking a clean API. "NOTHING on your stack is off limits."
  • Time saved: ~1 hour → 30 seconds. One evening of setup, recurring payoff.
  • Host claims "the models are prolly a coin flip now, but the harness race ain't close."
  • Caveat: "The first run will be messy, so correct it once and save the instructions."
Workflow 2 — Skills
  • "turn this into a skill" after any task = saved repeatable steps. Host reports 124 skills, used 6,179 times combined; one skill scans threads daily and flags repeat work for new skills.
  • Metric: "how many saved skills fire without a human touching anything."
Workflow 3 — ChatGPT Voice via Remote
  • Connects mobile app to desktop; "tap the blue button" and phone runs the entire machine. Full duplex (talk over the agent).
  • Demo: opens a PowerPoint in Downloads, swaps a Canva image, reads page three, updates HubSpot, pings Slack.
  • Gotcha: "regular voice mode in the ChatGPT app can't touch your connectors, but voice through Remote can."

Also mentioned: OpenAI agents created secret message boards; Meta agents "go rogue"; Anthropic working on AI chips.

Full text · 4,057 chars
Ep 835: Inside Everyday AI: My 9 most Used AI Tools and Workflows OpenAI: Our agents created secret message boards, Meta's agents (also) go rogue, Anthropic working on AI chips and more. The most productive worker here clocks in at 11 PM and logs off at 5 AM. Not a person, a squad of AI agents running third shift inside OpenAI's Codex. Yep. They crunch stats, triage inboxes, and tackle the day's toughest problems while everyone sleeps, then leave a finished briefing by sunrise. You know.... "Codex" has "code" in it, sounds like developer stuff. Wayyy wrong, fam. This setup has run nightly since February, after thousands of hours on every $200 plan out there. The models are prolly a coin flip now, but the harness race ain't close. What's on today's show? The full nine-workflow playbook, from the overnight briefing agent to the one-sentence skill habit to the voice trick that runs a computer mid-walk. Zero code required, promise. 1. Overnight Codex agents end morning busywork 🔥 Every night at 11 PM, scheduled agents clock in. They pull podcast stats from Buzzsprout, check newsletter A/B tests inside Beehiiv, and triage four inboxes plus LinkedIn with replies drafted. By 5 AM? One clean morning briefing, sitting there waiting. The trick that makes all of this work is a feature called browser control. When a tool has no clean integration, the agent opens Chrome and clicks around like a patient employee, so NOTHING on your stack is off limits. That job used to burn an hour of pulling, exporting, and squinting at dashboards. Now it's 30 seconds with coffee. The best part? One evening of setup, and the payoff repeats every morning after. Try This Open Codex tonight, or the ChatGPT desktop app since they're the same thing, and describe your most dreaded recurring report: where the numbers live, which logins it needs, and what the finished briefing should include. Schedule it for 11 PM. The first run will be messy, so correct it once and save the instructions. By Friday, the report builds itself. 2. AI skills turn repetition into owned IP ⚡ Forget counting tokens or lines of code. The real measure of an AI native operator is how many saved skills fire without a human touching anything. What's a skill? Any multi-step task you've done twice, saved as repeatable steps, and chained into a plugin an agent reruns on a schedule. The magic words are stupidly simple: after finishing any task, type turn this into a skill, and the agent saves the whole process. The current count around here: 124 skills, used a combined 6,179 times. One of them scans every single thread daily and flags repeat work that should become the next skill. Skills building skills. Wild, right? Try This Start tomorrow morning with one simple rule: anything you do by hand twice gets the turn this into a skill treatment on the spot. It takes under a minute and compounds every single day. After a week, stack your best three skills into one plugin and schedule it to run overnight while you sleep. You gotta stop being a button pusher. 3. ChatGPT Voice runs your whole computer 🚀 ChatGPT Voice through Remote might be the most slept-on release of the year. Connect the mobile app to your desktop once, open Remote, tap the blue button, and your phone now runs the ENTIRE machine. What does that actually unlock? Taking a walk while an agent opens the PowerPoint in your downloads folder, swaps in a fresh Canva image, reads back page three, then updates HubSpot and pings the team in Slack. It's full duplex, so you can talk over each other while a second model keeps working in the background. One gotcha: regular voice mode in the ChatGPT app can't touch your connectors, but voice through Remote can. Try This Connect the ChatGPT mobile app to your desktop, then open Remote and tap the blue button. On your next 20-minute walk, hand it one real task out loud, something like open the file in the downloads folder and tell me what changed. Start small to build trust, then hand it more every single day. Y'all, it genuinely feels like cloning yourself.
18:04

Simon Willison on Technical Blogging

Simon Willison shares his top advice from an interview about technical blogging: lower your standards and publish while you're still unhappy with the post, because otherwise you end up with a folder full of drafts and nothing published. The interview covers why he started blogging, its biggest surprises, and his favorite posts and blogs. Thin content — a short link-post summary of a longer interview.

Full text · 1,124 chars
6th August 2026 - Link Blog Simon Willison on Technical Blogging. I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog! It includes my answers to the following questions: - Why did you start blogging – and why do you continue? - What has been the most surprising impact of blogging for you? - What blog post are you most proud of and why? - What post was the most difficult to write and how did you tackle it? - Any lessons learned that you want to share with the community? - Your advice for people just getting started with blogging? - A few blogs that you particularly enjoy? I'll repeat my most important piece of advice here: My number one tip for blogging is to lower your standards! Aim to hit publish while you are still actively unhappy with what you have written, because the only alternative is a huge folder full of drafts and never publishing anything at all. Nobody will ever know how perfect the thing you intended to write would have been. The flaws you see in your writing are invisible to everyone else.
16:55

😻 Livestream: How to build agents for TOTAL Beginners

A newsletter is promoting a free livestream crash course on building AI agents for total beginners. The host, a former ElevenLabs employee who runs an AI-only business with no staff and claims a $200,000-plus month, covers basics like Claude Cowork and Claude Code, CLAUDE.md tips, and reusable skills. The item is mostly an ad; the only other content is promos for Intel's local AI beta and a ClickUp AI walkthrough.

Notes

How to build agents for TOTAL Beginners — The Neuron live stream (announcement)

Event: Aug 6, 2026. James McAulay (founder, The Agent Accelerator) live with The Neuron — beginner agent crash course for Claude Cowork and Claude Code. Starts 12pm PT / 3pm ET; after that, the full thing is available pre-recorded.

Speaker credibility:

  • Helped ElevenLabs grow from $110M → $300M+ ARR.
  • Then launched a fully AI-native business: $80K/month by month three, first $200K+ month by month five, zero employees, zero paid advertising — delegates to multiple agents daily.
  • The Agent Accelerator has trained 400+ people across 100 companies (startups, 200-person orgs, UK Government). Participants self-report ~5 hrs/week of manual work automated after 4 weeks.

Agenda (six parts):

  • Agent foundations — prompting a chatbot → delegating agentic work.
  • "Second brain" — the key files that make the biggest difference.
  • Optimizing Claude with CLAUDE.md — shaping behavior.
  • Skills — where to find good ones, how to create great ones.
  • Proactive agents — McAulay's four-level framework for agents that run without prompting in Cowork and Code.
  • Live Q&A; includes screen-share demos of his own setup.
Other segments in same newsletter
  • Intel / SuperClaw public beta (Dr. Olena Zhu): email, coding, and deep-research agents that combine outside research with private company data while keeping sensitive material local; parts of the stack are open.
  • Intel's projection: local models inherit frontier-level capability after ~24.8 months on average.

> "If that trend holds, a Fable-class intelligence model could run on a high-end laptop by 2028." — The Neuron (citing Intel's framing)

  • ClickUp AI walkthrough (20 min) by Jessica Lee, content ops lead at TechnologyAdvice — triaging tasks, building reports, creating agents/workflows.

Caveats: all revenue/time-savings figures are self-reported/promotional; no benchmarks given. The 24.8-month "capability inheritance" trend and 2028 laptop claim are projections, not demonstrated results.

Full text · 3,510 chars
😻 Livestream: How to build agents for TOTAL Beginners James McAulay joins us live with a beginner-friendly AI agent crash course. Welcome, humans. This is the number one question we get every week: How do I actually use AI agents to save time in my business? In five minutes, we are going live with James McAulay, founder of The Agent Accelerator, for a practical crash course on everything you need to know to get started building helpful, proactive agents in Claude Cowork and Claude Code. James has spent the past year building at the front lines of the agent economy. After helping ElevenLabs grow from $110M to more than $300M in annual recurring revenue, he launched a fully AI-native business that reached $80K in monthly revenue by month three. By month five, it recorded its first $200K-plus month. He did it without a single employee or a dollar of paid advertising. Instead, James delegates work to multiple AI agents every day. Here’s what we’re covering live: - Agent foundations: How to move from prompting a chatbot to delegating agentic work. - Starting your second brain: The key files that make the biggest difference. - Optimizing Claude with CLAUDE.md: Tips and tricks for shaping how Claude behaves. - Skills: Where to find good ones and how to create great ones. - Proactive agents: James’s four-level framework for agents that work without being prompted in Cowork and Code. - Your questions: Answered live! The format blends teaching, screen sharing, and demos of James’s own setup, so you can see how the pieces work together in practice. Through The Agent Accelerator, James has trained more than 400 people across 100 companies, including startups, 200-person organizations, and even the UK Government. Participants report automating an average of five hours of manual work every week after four weeks. And in this stream built for TOTAL beginners (like you!), we are starting with what agents can do today and how YOU can put one to work. Open the stream now, say hello in the chat, tell us where you’re from, and bring the task you wish you could stop doing manually. And TBH, you should probably plan to follow along and implement James’ guide as you go! P.S: If you join after 12pm PT / 3pm ET, you can watch the whole thing pre-recorded! Intel’s Dr. Olena Zhu showed us the version of local AI that feels useful right now. They just released SuperClaw’s public beta, which includes email, coding, and deep-research agents that can combine outside research with private company data while keeping the sensitive material local. Parts of the stack are open, so teams can inspect the architecture and build on it. The bigger idea behind Intel’s approach is wild: local models have been inheriting frontier-level capability after an average of roughly 24.8 months. If that trend holds, a Fable-class intelligence model could run on a high-end laptop by 2028. See above re: problems with datacenters for why that’s actually really important. Are you drowning in project management? So was Jessica Lee, our content ops lead at TechnologyAdvice. Then she got access to ClickUp’s new AI tools, and it started feeling like she’d hired a specialist to triage the busywork. If you use ClickUp, or any other project-management software, watch this 20-minute walkthrough. Jessica shows exactly how she uses ClickUp AI to triage tasks, build reports, and create agents and workflows that save her hours every week. What should we learn next? 🤔 Lets us know below! Stay curious, The Neuron Team
18:22

datasette 0.65.3

Datasette, the tool for publishing SQLite data online, shipped a small bug-fix release, version 0.65.3. The announcement itself is basically empty — the post is mostly a sponsor ad — so there's nothing substantive to summarize. The same version also carries the SQL-injection fix detailed in the separate 1.0a38 post.

Full text · 186 chars
Sponsored by: Dynatrace — When agents enter the SDLC, observability becomes the enabler to move from code generation to scalable engineering. Read the blog for a framework to get started
23:07

datasette-auth-tokens 0.4a13

A new pre-release of datasette-auth-tokens, the plugin that handles API-token authentication for Datasette, is out (version 0.4a13). The post gives no details on what changed — it's mostly a sponsor ad — so this is summarized from the title. Treat it as a routine plugin update.

Full text · 186 chars
Sponsored by: Dynatrace — When agents enter the SDLC, observability becomes the enabler to move from code generation to scalable engineering. Read the blog for a framework to get started

Newsletter

6
04:34

[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???

Four of Google's most senior AI researchers are leaving to start their own company, while its most famous AI leader is stepping up to a bigger role. Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are cofounding Discovery Loop, a company that wants AI to run its own scientific and engineering experiments and improve itself, backed by Google and investors like Radical Ventures and Khosla Ventures. Demis Hassabis moves from running Google DeepMind day to day to become chairman of DeepMind and chief scientist of Alphabet, focusing on long-term strategy and drug discovery at Isomorphic. Koray Kavukcuoglu, previously chief technology officer, takes over running the lab, including the Gemini models and app. The shake-up lands after a six-month gap since the last Gemini Pro update and a run of other high-profile departures. The roundup also covers Meta's Muse Spark 1.2 and Muse Code coding-agent launch and a self-improving harness that claims 95.5% on ARC-AGI-3.

Notes

GDM leadership exodus; Discovery Loop founded

Date: 2026-08-06 (covers AI news 8/4–8/5/2026; 12 subreddits, 544 Twitters, no Discords)

Google DeepMind reshuffle + Discovery Loop spinout
  • Departed: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le leave DeepMind to cofound Discovery Loop, a Public Benefit Corporation targeting autoresearch / automated discovery loops over ML, science, and engineering workflows.
  • Seed round: led by Radical Ventures and Khosla Ventures, with Lightspeed, Kleiner Perkins, Doerr Capital, and Alphabet participating. Google "investing in" the startup; departures described as "very amicable."
  • Demis Hassabis: CEO → Chair of Google DeepMind and Chief Scientist of Alphabet, stepping back from daily ops to focus on long-term strategy, AGI, science, and "leaning into" Isomorphic.
  • Koray Kavukcuoglu: promoted to SVP of DeepMind, taking operational control over Gemini, frontier research, and product/dev teams.
Author's open questions / skepticism
"surely we are not being told the full story here; why couldn't Discovery Loop have been done inside Google?"

Read as a governance reset and an attempt to sharpen product execution around Gemini.

Stated limitations & caveats
  • Article is partially paywalled (full text behind 7-day free trial).
  • Prior exits listed for context: John Jumper → Anthropic, Noam Shazeer → OpenAI, plus David Silver and Denny Zhou; ~6 months since last Gemini Pro update.
  • Notes GDM's 1000+ coauthor papers for Gemini vs. the four departing founders writing a "manifesto."
  • Commentary from Nathan Lambert and Andrew Ng frames this as a "historical inflection point" and signal that AI-for-science is becoming a primary frontier.
Side stories (briefly noted, not the headline)
  • Meta Muse Spark 1.2 + Muse Code: 5.6 Terra-level model, harness with local event log for resumability, persistent background agents.
  • Prime Agent (Prime Intellect): RLM-based self-improving harness claiming 95.5% on ARC-AGI-3 — "not yet endorsed by ARC."
Full text · 4,461 chars
[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM??? The end of an era. It’s tempting to give Meta Spark 1.2 and Muse Code the title story today because of their success launching a 5.6 Terra-level model, together with innovative harness design with a local event log for resumability and persistent background agents, both of which should put other coding agent builders on notice. It’s tempting to highlight Prime Agent, Prime Intellect’s self-improving RLM based harness that claims an incredible 95.5% on ARC-AGI-3 (not yet endorsed by ARC). We tried. Trust me, we tried. But it’s hard to beat around the bush — today’s most important story is the coordinated departures of Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le, some of the most senior engineering and research talent in both Google and in human history, from DeepMind to cofound a new autoresearch startup Discovery Loop: And Demis, who needs no introduction, goes from CEO to Chair and Chief Scientist, while “leaning into” Isomorphic with CTO Koray stepping up to SVP of GDM: All the departures are very amicable; Google is investing in Discovery Loop, and Demis’ increased contributions to long term strategy and Isomorphic in particular will be very welcome by humanity, but surely we are not being told the full story here; why couldn’t Discovery Loop have been done inside Google? That one at least, we have some hints, given GDM’s history of 1000+ coauthor papers for Gemini, vs these 4 superhumans writing this manifesto: When John Jumper left for Anthropic, one could maybe chalk it up to Anthropic’s momentum. When Noam Shazeer joined OpenAI, perhaps one could point to their pioneering work in reasoning models. But couple it with David Silver, Denny Zhou, and other prominent departures, and the 6 months since the last Gemini Pro update, one has to wonder 1) what happened that led to this, 2) how this latest shakeup was decided, 3) if the shuffle is the beginning of the middle of the end or the first prologue of a great and necessary comeback story. After all, Google STILL has great models, talented teams, excellent compute and infrastructure and the greatest trove of training data in human history… AI News for 8/4/2026-8/5/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies! AI Twitter Recap Google DeepMind Leadership Reshuffle and the Discovery Loop Spinout - A major Google AI reorg landed alongside a high-profile founder exodus: Demis Hassabis is moving to Chair of Google DeepMind and Chief Scientist of Alphabet, explicitly stepping back from day-to-day GDM operations to focus on long-term strategy, AGI, and science. Koray Kavukcuoglu takes operational control as SVP of DeepMind, overseeing Gemini, frontier research, and product/dev teams. The subtext from the ecosystem was clear: this is being read as both a governance reset and an attempt to sharpen product execution around Gemini. - At the same time, Discovery Loop launched with one of the strongest founding teams in AI infrastructure/research: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le are founding Discovery Loop, a Public Benefit Corporation aimed at automating machine learning, science, and engineering. Dean also shared that Radical Ventures and Khosla Ventures are leading the seed round, with participation from Lightspeed, Kleiner Perkins, Doerr Capital, and Alphabet. The technical read-through is important: rather than another general-purpose model startup, this is explicitly targeting autoresearch / automated discovery loops over scientific and engineering workflows. - Why engineers cared: the market reaction wasn’t just “big names left Google.” It was that several people most associated with Google’s deep infra, model-building, and research execution stack are now pursuing a startup centered on automated science. Commentary from Nathan Lambert, Andrew Ng, and others framed it as a historical inflection point for Google’s AI efforts and a strong signal that AI-for-science is becoming a primary frontier, not a side quest. Meta’s Muse Spark 1.2 and Muse Code Push Into the Coding-Agent Race Keep reading with a 7-day free trial Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.
11:29

Google just funded the startup its 4 best people left to build

Four of Google's most famous AI researchers quit to start a company that automates the scientific research process, and Google is funding them. Jeff Dean is CEO of Discovery Loop, which builds AI that proposes experiments, runs thousands in parallel, learns from results, and iterates — starting with machine learning research itself. Co-founders include Sanjay Ghemawat, Quoc Le, and Oriol Vinyals. Alphabet is a founding investor and supplies the compute for at least the first year. The seed round is co-led by Radical Ventures and Khosla Ventures.

Notes
Discovery Loop: Jeff Dean's exit from Google

The founding team (all left Google together, effective Aug 5, 2026):

  • Jeff Dean (58, CEO) — Google employee #30, joined 1999 when it had ~20 people; co-founded and ran Google Brain (~600 people at peak), then led Google Research/AI teams that grew to ~4,400. His pitch-deck slide "Management/Team Building experience" lists former team members who went on to found companies — the argument being the 4 built the orgs that trained the rest of the AI industry.
  • Sanjay Ghemawat — Senior Fellow; co-creator of MapReduce, Google File System, Bigtable, Spanner.
  • Quoc Le — Google Brain founding member; co-inventor of sequence-to-sequence learning with Ilya Sutskever and Oriol Vinyals.
  • Oriol Vinyals — DeepMind VP of Research, Gemini technical co-lead, creator of AlphaStar; 100,000+ citations.

Combined ~100 years at Google; worked together 14–30 years. Dean told UW graduates in June he "got the itch to join a startup in 1999" again.

The company: Discovery Loop, a Delaware public benefit corporation in Palo Alto. Mission: automate the experimental loop itself — an AI that proposes experiments, runs them, learns, iterates, thousands in parallel.

"Particularly in a lot of domains, you can fully computerize that whole loop." — Jeff Dean

Funding: Seed round co-led by Radical Ventures and Khosla Ventures; Lightspeed, Kleiner Perkins, and Doerr Capital participating; Wilson Sonsini advising. Round still open at announcement; valuation undisclosed. Alphabet is a founding investor and supplies the compute for at least the first year. Sundar Pichai called the exits amicable, saying Dean and Ghemawat "helped drive some of the most significant technology transitions" in Google's history.

"we're going to be our own first customers. The rapid feedback from doing that is the way to build something amazing." — Dean

Sequencing: starts with ML research/engineering (software experiments are cheap, fast, fully computerizable), then hardware design, drug discovery, materials, clean energy. Dean claims the approach reaches subproblems in nearly all 14 NAE Grand Challenges.

"Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today."

Caveats and limitations:

  • The author's reading ("Alphabet did not fund a competitor. It bought a seat next to the fastest compounding process anyone has attempted") is editorial opinion, not company statement.
  • The non-compete framing — "Anthropic, SSI, Thinking Machines, and now Discovery Loop all exist partly because their founders could walk out on a Wednesday and incorporate on a Thursday" (California bans non-competes) — is the author's interpretation.
  • The concrete "playbook" (4-stage loop architecture, cost formula, prompts, verification layer, "5 signals to watch") is paywalled in this Substack post; the notes above are everything available free, plus the author's claim that the loop maps to a 4-stage pattern teams run by hand as "iteration."
  • Exact seed-round size, valuation, and cap-table breakdown are not disclosed.
Full text · 6,405 chars
Google just funded the startup its 4 best people left to build Jeff Dean and 3 collaborators walked out after a combined century at Google. Alphabet wrote a check and is paying for the compute. The deck, the investors, and the loop they are automating In June, Jeff Dean stood at a University of Washington commencement and told graduates how he “got the itch to join a startup in 1999.” He landed at Google when it had 20 people, in an office above what is now a T-Mobile store in Palo Alto. He was employee #30. Twenty-seven years later, at 58, he has the itch again. On August 5, Dean walked out of Google with 3 people who between them carry most of modern computing: ▫️ Sanjay Ghemawat, Senior Fellow, co-creator of MapReduce, Google File System, Bigtable, and Spanner. The infrastructure that made internet-scale computing possible. Dean’s collaborator for over 2 decades. ▫️ Quoc Le, Google Brain founding member, co-inventor of sequence-to-sequence learning with Ilya Sutskever and Oriol Vinyals. The direct architectural ancestor of every model you use today. ▫️ Oriol Vinyals, DeepMind VP of Research, Gemini technical co-lead, the mind behind AlphaStar. Over 100,000 citations. The 4 of them have worked together for 14 to 30 years. Their own site puts it plainly: 3 of the most-cited researchers in AI, and 2 of the most-cited in distributed systems, in one founding team. The company is Discovery Loop, a Delaware public benefit corporation in Palo Alto. Dean is CEO. The mission fits in a sentence: automate the experimental loop itself. “Particularly in a lot of domains, you can fully computerize that whole loop.” Jeff Dean The slide that explains why Google could not keep them Dean did something founders almost never do. He published slides from the pitch deck he used to raise. The one worth studying is titled “Management/Team Building experience.” On its face it is a résumé dump: Dean co-founded and ran Google Brain at its peak of roughly 600 people, then led Google Research and AI teams that grew to about 4,400. The actual message sits in the list underneath, the people who worked on their teams and went on to found things: Read the pitch inside the pitch. These 4 built Google’s most important infrastructure over 25 years, and they built the organizations that trained the people now running the rest of the AI industry. That is the asset an incumbent cannot replace with a counteroffer. Who is funding it, and the part that should stop you The seed round is co-led by Radical Ventures and Khosla Ventures, with Lightspeed, Kleiner Perkins, and Doerr Capital participating. Wilson Sonsini is advising. The round was still open at announcement and the valuation stays undisclosed. Then there is the last name on the cap table. Alphabet is a founding investor, and is supplying the compute for at least the first year. Google is bankrolling the company its own legends left to build, and renting it the machines. Sundar Pichai went out of his way to call the exits amicable, saying Dean and Ghemawat “helped drive some of the most significant technology transitions” in Google’s history. That check makes sense once you see what loop they picked first. Discovery Loop starts with machine learning research, acting as its own first customer. An AI that proposes experiments, runs them, learns, and iterates, thousands in parallel, aimed at the problem of building better AI. Dean was direct about why: “we’re going to be our own first customers. The rapid feedback from doing that is the way to build something amazing.” So Alphabet did not fund a competitor. It bought a seat next to the fastest compounding process anyone has attempted, because sitting outside it costs more than the check. One more detail worth noticing: California bans non-competes. Anthropic, SSI, Thinking Machines, and now Discovery Loop all exist partly because their founders could walk out on a Wednesday and incorporate on a Thursday. What they are actually automating The scientific method is the most powerful algorithm we have, and we have always run it by hand. Propose, experiment, read the result, adjust, repeat. Sequential. Slow. Bounded by how many humans you can put on it. Discovery Loop’s bet is that the loop runs faster once the human steps out of the middle. The AI proposes the experiment, executes the run, learns from the outcome, and iterates recursively, with thousands of experiments running in parallel where a team would have run them in series. Their published ambition: “Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today.” The sequencing tells you they have done this before. ML research and engineering first, because software experiments are cheap, fast, and fully computerizable. Then hardware design, drug discovery, materials, and clean energy. Dean says the approach reaches subproblems in nearly every one of the 14 NAE Grand Challenges. Here is the part that matters for you. The loop they are industrializing at frontier scale is the same 4-stage pattern you can run this week, at your scale, with tools you already pay for. Most teams run it by hand and call it iteration. The leverage arrives when you remove yourself from the middle and give it a stop condition. Below, the full system. Inside The Discovery Loop Playbook: ▫️ The 4-stage loop architecture, mapped to what you can automate today ▫️ The self-customer rule, how to pick the first target so it compounds ▫️ 4 copy-paste loop prompts, propose, execute, evaluate, iterate-or-stop ▫️ The parallel-experiment system, with the cost formula and the budget guard ▫️ The verification layer, the piece almost everyone skips ▫️ The 5 signals to watch, to know whether this thesis is working Your subscription opens the full AI Corner archive: ▫️ The AI Tools and Models library, every model, tool, and setup guide ▫️ The AI Agents library, the full agent-building stack, start to finish ▫️ The Prompting and Context Engineering library, the prompts and context systems that actually ship ▫️ The Claude and Anthropic library, every Claude playbook in one place Plus 3 new premium systems every week. 🔁 The Discovery Loop Playbook Get the full system below 👇 1. The 4-stage loop Every automated research loop has 4 stages. Each maps to something you can run now:
09:28

The Year AI Science and the Physical AI Industry Came Alive

Humanoid robotics is going mainstream as capital floods in: four major humanoid companies plan IPOs, including China's Unitree valued near $7.4 billion, and Europe's robotics startups raised more in six months than the past two years combined. Google launched Gemini Robotics 2, a model for controlling any robot from bi-arms to full humanoids. The US banned foreign-made robots, and the piece warns compute costs are rising while AI companies struggle with negative cash flow. It also recaps Jeff Dean leaving Google for his science startup and Demis Hassabis stepping down as Google DeepMind CEO.

Notes
The Year AI Science and the Physical AI Industry Came Alive

AI Supremacy (Substack), 2026-08-06

Thesis: Physical AI/robotics is being used as the Silicon Valley playbook to "prolong the AI bubble" while compute gets much more expensive.

People/shakeups

  • Jeff Dean leaves Google after 27 years as chief scientist to start company Discovery Loop.
  • Demis Hassabis steps down as Google DeepMind CEO → becomes Chair of Google DeepMind & Chief Scientist of Alphabet.

Google Gemini Robotics 2 (GR 2)

  • Google's most advanced VLA model; controls any robot type, bi-arms to full humanoids; claims to enable whole-body control, advanced dexterity, multi-robot collaboration. Described as a "multi-modal layer" combining visual perception, spatial reasoning, real-time motor control.

IPOs / markets

  • Four humanoid robotics companies going public: two China, one US, one South Korea.
  • Unitree Technology: expected valuation >50B yuan (~$7.4B) at planned Shanghai IPO.
  • European robotics startups raised more capital in first 6 months of 2026 than previous two years combined (per Sifted).
  • July 28: Trump administration, via FCC, announced ban on future foreign-made robots.
  • Counterpoint: AgiBot, Unitree, UBTech were top-3 humanoid companies by installation share last year; Tesla Optimus ranked fifth. Agility SPAC'd.

Caveats stated by author

  • "humanoid robots are not functional in society yet"; AI hasn't yet transformed science.
  • Claims needed "world models and specialized LLMs for Physical AI that don't yet exist" for general-purpose humanoids.
  • Framing is bubble-skeptic: mentions compute cost problem, negative FCF dilemma, mounting debt, circular revenue.

Forward look: next decade pivotal for AVs (Waymo, Uber, Baidu, Wayve).

Full text · 3,521 chars
The Year AI Science and the Physical AI Industry Came Alive Google Gemini Robotics 2, Unitree and AgiBot IPOs, Robot and optical transceiver bans, the compute cost problem, Negative FCF dilemma, mounting debt, circular revenue. I’m coming to the realization that Physical AI, embodied AI and robotics is going to become a big deal. It’s going to be part of the Silicon Valley playbook to prolong the AI bubble. Four major humanoid robotics companies, two in China, one in the U.S. and one in South Korea, are going IPO. European robotics startups reportedly raised more capital in the first six months of this year than in the past two years combined, according to Sifted, which is a bit like the TechCrunch of Europe. What this implies is we are entering a new era of automation and robotics. I just finished a deep dive into one of the best ways to get early stock market exposure to this trend here. Yesterday we learned that Google’s AI leadership is changing forever. Google DeepMind’s roles are changing. Jeff Dean, the longtime chief scientist, is leaving after 27 years to start his own company called Discovery Loop. Leveraging AI for science is now in vogue, and capital does not appear to be a limit in how we achieve the future. While AI hasn’t yet transformed science and humanoid robots are not functional in society yet, Demis Hassabis is leaving his role as CEO of Google DeepMind to be the unit's Chairman. New roles: Chair of Google DeepMind & Chief Scientist of Alphabet. What’s caught my attention this summer perhaps the most is all the Physical AI seed rounds, new companies and new deals with so much more capital. There’s a robotics blitz coming, just as compute will get much more expensive. Google Gemini Robotics 2 GR 2 is Google’s most advanced VLA model: capable of intelligently controlling any type of robot, from bi-arms to full humanoids. This is their intelligence layer, Google claims, is powering the next generation of truly adaptable robots. As it takes its first literal steps, this major advance unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration. Think of this as a Multi-modal layer for the future of robotics, where this 2.0 ecosystem combines visual perception, spatial reasoning, and real-time motor control across an entire physical frame under a single model suite. The AI and software brains of robotics is going to have to get much better in the next five years to bring this world of autonomy online, and it will. In autonomous vehicles the next decade is also pivotal with Waymo, Uber, Baidu, Wayve and others set to evolve fast. We will need strong world Models and specialized LLMs for Physical AI that don’t yet exist for humanoids to reach general purpose utility and GR 2 is a bold step in that direction. Unitree Goes IPO Chinese robot maker Unitree Technology is expected to be valued at more than 50 billion yuan ($7.4 billion) after its planned Shanghai IPO and it marks a time where China is first to market and has a dominant marketshare in robotics with a significant cost advantage. On July 28th, The Trump administration under the FCC, announced a ban on future foreign made robots. For reference, two of China’s leaders are going IPO. Meanwhile Chinese companies Agibot, Unitree and UBTech accounted for the top-three humanoid companies by installation market share last year, according to Counterpoint. Tesla’s Optimus ranked fifth. The IPOs, SPACs (Agility) and Bans means robots are about to hit the prime time.
15:32

Qwen 3.8 Max and Why You Should Not Care (Yet)

An analysis arguing Alibaba's Qwen 3.8 Max headline — a 2.4 trillion-parameter, million-token open-weight model — matters less than the smaller Qwen 3.8 27B, a dense model that runs on consumer GPUs at 100 tokens per second. The piece lays out China's strategy of giving models away while building cheap infrastructure, citing a $295B data-center plan and Chinese models capturing 46% of US enterprise tokens on OpenRouter. It reads OpenAI's 80% price cut on GPT 5.6 Luna as a defensive response to that pressure. The Max weights were promised for the week of August 10 but not yet released as of writing, so this is reasoned speculation rather than confirmed results.

Notes

Qwen 3.8 Max and Why You Should Not Care (Yet)

Author: Manolo Remiddi, The Augmented Mind (Substack), 2026-08-06. Opinion/analysis piece — facts below are what the article reports, flagged where speculative.

Qwen 3.8 Max (the "distraction")
  • Announced July 19, 2026 at World AI Conference, Shanghai; launched August 3, 2026.
  • Sparse MoE: 2.4T total params, ~95B active per token; multimodal (text, image, video, documents); 1M-token context, up to 131K output tokens. First Max-class model to go open weight.
  • Weights promised "for the week of August 10, 2026" on Hugging Face and ModelScope — not yet released as of publication.
  • Alibaba claims it is "second only to Fable 5" among models they evaluated (self-reported).
  • Launched days after Moonshot AI's open-weight Kimi K3 (2.8T params); author reads release-cycle compression to weeks as deliberate frontier-race behavior.
The real news: Qwen 3.8 27B
  • Dense (all 27B active per token, not MoE), also promised open weight.
  • Author uses Qwen 3.6 27B for ~80% of daily AI work. On RTX 5090 at Q6: 100–110 tok/s, near-lossless; Q4 with smaller context fits 24GB VRAM. DGX Spark (GB10, 128GB): 10–15 tok/s, bottlenecked by memory bandwidth, "usable for offline work, not interactive."
  • Artificial Analysis Intelligence Index: Qwen 3.6 27B without reasoning outscores GPT 5.6 Luna without reasoning; Luna's reasoning raises it 27→51 (+24).
  • Speculative thesis: if 3.8 27B gains any reasoning-scaling, a local model starting above Luna baseline could reach ~51–53, "competing directly with Claude Sonnet 5 Max and GLM 5.2." Author's own caveat:
"We do not know if the Qwen team can pull this off. The potential is there based on the trajectory, but the actual numbers will only appear when the model ships."
China's strategy (author's core argument)
"Give the models away for free, build the infrastructure cheaply, serve the entire world, and profit from the hardware and energy layer. It is a play for scale, not margins."
  • Reported figures: $295B public AI data-center funding over 5 years (Bloomberg, June 2026); "Eastern Data, Western Compute" relocating DCs to energy-rich western China; CNBC (July 2026): Chinese models captured 46% of US enterprise token usage on OpenRouter, at times exceeding US-origin share.
US response: price war
  • July 30, 2026: OpenAI cut GPT 5.6 Luna 80% — input $1.00→$0.20/M, output $6.00→$1.20; Terra −20%; Sol unchanged at $5.00/$30.00. OpenAI frames it as efficiency-driven; author calls the timing "suspicious" and defensive. Opinion, not sourced fact.
  • Context prices: Claude Opus 4.8 at $5.00/$25.00; Fable 5 at $10.00/$50.00 per M tokens.
Outlook
  • Author: open-weight infra play could pressure US AI stocks over "two to three years"; US moats limited to proprietary data, enterprise integration, brand trust. Individual takeaway: wait for Qwen 3.8 27B if you have ≥24–32GB VRAM. Next-week column promises benchmarks.
  • Transparency note: article reasoned by Manolo Remiddi; "Resonant Augmentor" AI assisted with research/editing; image AI-generated.
Full text · 9,003 chars
Qwen 3.8 Max and Why You Should Not Care (Yet) A 2.4 trillion parameter model just went open weight. The real story is a 27 billion parameter model coming to your GPU and the Chinese strategy that threatens the entire US AI industry. Another week, another Chinese frontier model. Alibaba announced Qwen 3.8 Max on July 19, 2026 at the World AI Conference in Shanghai, and it officially launched on August 3 with 2.4 trillion parameters, multimodal support, and a million token context window. Alibaba claims it is “second only to Fable 5” among the models they evaluated. It is the team’s first multimodal model to cross the trillion parameter mark. But here is the thing. You should not care that much about the 2.4 trillion parameter model. The real news is something else entirely. What Is Qwen 3.8 Max Qwen 3.8 Max is a sparse Mixture of Experts model with 2.4 trillion total parameters and approximately 95 billion active per token. It processes text, images, video, and documents with a one million token context window and up to 131 thousand output tokens. For the first time in the Qwen family, a Max class model is going open weight. Alibaba has promised the weights for the week of August 10, 2026, on Hugging Face and ModelScope. As of publication, they have not yet been released. It launched just days after Moonshot AI released Kimi K3, a 2.8 trillion parameter open weight model. The timing is not a coincidence. The frontier race among Chinese labs is compressing release cycles into weeks instead of months. This is impressive, but it is also the headline that distracts from what actually matters. The Real News: Qwen 3.8 27B Alongside the Max announcement, Alibaba confirmed that Qwen 3.8 27B is also going open weight. This is what gets my attention, because the Qwen 3.6 27B is the model I personally use every single day. It handles roughly 80 percent of all the work I do with AI. It is a dense model, not a Mixture of Experts. All 27 billion parameters are active for every decision. And it runs on consumer hardware. On my RTX 5090 at Q6 quantization, I get between 100 and 110 tokens per second. Almost lossless quality. This is the model that lives on my local machine and powers my workflow. When Qwen 3.8 27B arrives, everyone with at least 24/32 gigabytes of GPU memory can download it and run it. That is the story worth caring about. The Intelligence Scaling Argument Here is why the 27 billion parameter model matters so much, and this is where the numbers get interesting. The Artificial Analysis Intelligence Index provides a composite benchmark across nine evaluations covering reasoning, knowledge, mathematics, and coding. These are the current public scores: Look at that first comparison. Qwen 3.6 27B without reasoning already scores higher than GPT 5.6 Luna without reasoning. A model you can run on consumer hardware is already beating the entry tier of OpenAI’s most advanced family at baseline. Now look at what happens when you enable reasoning on Luna. It jumps from 27 to 51, a gain of 24 points. OpenAI has built a mechanism where the model scales its intelligence by thinking harder, using a reasoning budget that trades speed and token cost for capability. Here is the question. If the Qwen 3.8 27B gains even a fraction of that kind of reasoning scaling capability, we are talking about a local model that starts above Luna at baseline and could potentially reach parity with or exceed Luna at maximum reasoning. We would be looking at a model that scores somewhere between 51 and 53, competing directly with Claude Sonnet 5 Max and GLM 5.2. And it would run entirely on your own hardware. We do not know if the Qwen team can pull this off. The potential is there based on the trajectory, but the actual numbers will only appear when the model ships and Artificial Analysis runs it through their benchmark suite. Hardware Reality Running a 27 billion parameter dense model locally is not trivial. Here is what the hardware landscape actually looks like: - RTX 5090 (32GB VRAM): This is the sweet spot. At Q6 quantization, Qwen 3.6 27B runs at 100 to 110 tokens per second with a context window large enough for real agentic work (131K token). You can load the model and use it at almost full precision. A Q4 with smaller context window can run on a 24GB VRAM box. - DGX Spark (128GB, integrated GPU): NVIDIA’s GB10 Grace Blackwell system, the consumer equivalent of a data center node. Community reports show Qwen 3.6 27B running between 10 and 15 tokens per second. Some optimized configurations go slightly higher, but the bottleneck is memory bandwidth. It is usable for offline work, but not for interactive workflows. The model is not a Mixture of Experts, so every parameter is loaded for every token. Fast memory matters more than raw compute. If your memory bandwidth is slow, you will feel it on every single request. China’s Strategy: Infrastructure Over Models Qwen 3.8 Max is impressive on paper, but it confirms something much bigger than any single benchmark score. China is not protecting its models as intellectual property. They are giving them away. This is not a coincidence. It is a deliberate strategy. The real asset is something they can physically protect: data centers, electricity, hardware, supply chains. They are building the infrastructure that makes AI accessible at scale while making the software commodity. The numbers back this up: - China announced a $295 billion public funding plan for AI data centers over five years, on top of private spending by Alibaba and Tencent. - The “Eastern Data, Western Compute” initiative relocated energy intensive data center operations to western China, where electricity production is abundant. - A CNBC investigation in July 2026 revealed that Chinese models captured 46 percent of US enterprise token usage on OpenRouter, at times exceeding US origin models entirely. - China keeps building data centers at a pace Western countries cannot match, while the US faces a massive electricity shortage that constrains its own AI expansion. Sources: Bloomberg (June 2026), CNBC (July 2026), China’s Eastern Data Western Compute national plan. Compare this to the US approach. OpenAI, Anthropic, and Google keep their best models locked behind APIs. The business model is access control. You pay per token, you do not own anything, and your usage is tracked and metered. China’s approach is completely different. Give the models away for free, build the infrastructure cheaply, serve the entire world, and profit from the hardware and energy layer. It is a play for scale, not margins. The US Response: Price Wars The US companies are reacting, and the signs are defensive rather than strategic. On July 30, 2026, OpenAI slashed GPT 5.6 Luna prices by 80 percent, dropping input from $1.00 to $0.20 per million tokens and output from $6.00 to $1.20. Terra got a 20 percent cut. Sol, the flagship, stayed unchanged at $5.00/$30.00. According to Open AI this is a price cut driven by efficiency, but the timing is suspicious. It feels that this is a defensive response to Chinese models eating enterprise market share. When a large percent of your token volume flows to competitors, you drop prices. But the underlying structural problem remains. US data centers are constrained by electricity availability, and Chinese infrastructure keeps expanding. Meanwhile, Claude Opus 4.8 sits at $5.00/$25.00 per million tokens and Fable 5 at $10.00/$50.00. These are premium products at premium prices. They work, but they cannot scale to serve the world at the margins that open weight Chinese models operating on cheap energy can achieve. What Happens Next Qwen 3.8 Max confirms the direction. The question is whether US corporations adapt or keep doubling down on a model-as-a-service business that Chinese companies are actively making obsolete. If they cannot compete on price and they cannot control the model layer, the only defensible position is proprietary data, enterprise integration, and brand trust. None of those are easy moats to defend when the underlying technology is free and improving every week. This could create a cascade effect across the US AI stock market. Not tomorrow, not next month, but over the next two to three years as Chinese infrastructure matures and the gap in available compute widens further. For individuals, the story is simpler. If you have the hardware or are planning to get it, the next Qwen 27B model is worth waiting for. If Qwen 3.8 27B achieves even half the reasoning scaling that Luna demonstrated, local AI at frontier levels stops being a theoretical possibility and becomes a practical reality. Next week I will go deep into the Qwen 3.8 27B when it releases, benchmark it, and run it through the same tasks I use every day. Subscribe so you do not miss it. Transparency note: This article was written and reasoned by Manolo Remiddi. The Resonant Augmentor (AI) assisted with research, editing and clarity. The image was also AI-generated.
19:51

AI Agents Need Memory, Not More Context

Bigger context windows won't fix AI agents, because context is temporary workspace while memory is what survives across sessions. An agent that forgets yesterday's findings repeats the same mistakes every morning. The essay lays out a memory system with three jobs — deciding what to remember, retrieving the right memory, and forgetting stale facts. Claude Code shows the pattern with CLAUDE.md files plus auto-memory notes, Microsoft's PlugMem stores distilled facts and skills instead of raw transcripts, and AutoMem research cut redundant memory writes by 68-83%.

Notes

AI Agents Need Memory, Not More Context

Source: Open Cloud AI (Substack) — published 2026-08-06

Main claim: Larger context windows don't solve agent learning. "Context is what the model can see during the current request"; memory is "information stored outside the active context that can be used again later." A fresh session erases what the agent learned — "The system simply failed to preserve what the model had learned."

Context vs. memory
  • Context = instructions, conversation history, retrieved docs, tool descriptions, search results, files, prior actions, error messages — a "finite resource" (Anthropic).
  • Memory = stable preferences, project architecture, past failures + causes, successful procedures, unresolved tasks, verified facts, decisions + rationale.
  • Context rot (Anthropic term): as more info enters context, "a model's ability to recall and use the correct information can decline." Desk analogy: "A Context Window Is a Desk, Not a Brain."
  • Reframe: not "how much can it hold?" but "the smallest amount of information it needs to make the next correct decision."
  • Long context still has value (comparing documents, codebases, long contracts/conversations) but "does not remove the need for memory."
Three jobs of a memory system
  • Write policy — decide what to remember: will it matter again? verified? already stored? replace older fact? access? expiry? "Without a write policy, memory becomes a landfill."
  • Retrieve — rank by semantic similarity, recency, source reliability, scope, task type, past usefulness, confidence, replacement status. "Simple vector similarity is often not enough."
  • Forget/update — expiration, dedup, conflict resolution, confidence decay, consolidation, deletion, human correction. "The goal is not perfect recall. The goal is useful recall."
File-based memory: Claude Code

Two persistence mechanisms: CLAUDE.md (stable instructions: build commands, conventions, workflows, architecture — written by user/team) and auto memory notes (patterns Claude discovers). Rule: "A rule written by the project owner should not have the same authority as a guess created during one agent run."

Example structure: CLAUDE.md (Commands: Install packages with pnpm, Run the database migration before integration tests; Architecture; Safety: "Never modify production configuration"; Completion: no "complete" until tests pass) + separate memory/ topic files (MEMORY.md, debugging.md, workflows.md, failures.md, open_tasks.md). Files are transparent, version-controlled, human-correctable. Start with files; "Do not begin by buying a vector database."

Microsoft PlugMem: knowledge, not raw history

Transforms raw interactions into compact facts and reusable skills organized in a structured memory graph, retrieved per-task then distilled into concise guidance. Example: CSV export failure → fact: condition CSV export exceeds 50,000 rows / outcome: synchronous export may fail; skill: when >50,000 rows / action: use asynchronous export endpoint. Evaluated across long conversations, multi-document fact-finding, web navigation: "structured memory supplied more useful information while using less context" than comparisons. Lesson: "Do not retrieve the past. Retrieve what the past taught you."

AutoMem (2026 research)

Treats memory management as a learnable skill. Agent routines: LOG (decide what's worth recording) and PLAN (search memory before acting). Results: redundant writes cut 68–83%; per-step memory growth cut from 138 chars to 6 (95% reduction) in one environment; agents more likely to search before writing. Production rule: Consult Before Write — search, then update-or-create, recording source/date/scope/confidence.

Five-layer memory stack
  • Working — current goal/plan, temporary; most should vanish at task end.
  • Procedural — how to do things (investigate alert, deploy, review contract); "short, testable, and versioned."
  • Semantic — facts; every fact needs source + update date.
  • Episodic — experiences ("deployment failed August 2 because a migration was missing"); store only if it changes a future decision.
  • Relational — knowledge graph for connected questions (affected customers, dependent decisions, shared vulnerable library).
Record format

JSON with content + metadata: memory_id, scope{organization, project, user}, type, source, created_at, last_verified_at, confidence, status, expires_at, replaces, tags. Minimum fields: content, type, source, scope, creation date, verification date, confidence, status. "Without scope, information can leak between users or projects."

Seven-part architecture
  • Scope request (who/org/project/permissions; "Never retrieve memory globally by default")
  • Search correct layer (procedural vs. episodic vs. relational routing)
  • Rank: relevance × source trust × freshness × scope match × past usefulness
  • Build a memory packet (short summary + sources, not 20 raw records)
  • Require evidence/verification before high-impact actions
  • Write gate (reusable? verified? stored? replaces? scope? expiry?)
  • Scheduled maintenance (duplicates, contradictions, expired, unused, low-confidence, unsourced, wrong scope)
Evaluation & metrics
  • Test set: 25–50 questions, five categories — recall, update (npm→pnpm swap), forgetting, isolation (Project A must not see Project B), abstention (check live system or ask rather than invent).
  • "Track retrieval precision before adding more memory."
  • Business-case metrics (support agent, Northstar Design example): repeated questions, time-to-resolution, retrieved-memory-usage rate, incorrect/stale-memory rate, human-correction rate, tokens per resolved case, leakage incidents, reopen rate.
  • Anthropic internal evals: context editing + memory improved agentic search 39% over baseline; context editing cut tokens 84% in a 100-turn search eval.
Seven memory mistakes

Saving every conversation; mixing users; treating similarity as truth; unlimited writes; never deleting; conclusions without evidence ("Project maintainer confirmed on August 5…" > "Tests require Redis"); demo-only evaluation.

Caveats: numbers (39%, 84%, 68–83%, 95%, 25–50 questions) come from the article's descriptions of vendor/research evals, not independently verified here.

Full text · 25,743 chars
AI Agents Need Memory, Not More Context Why larger context windows are not enough, and how Claude Code, Microsoft, and Stanford are teaching agents what to remember, retrieve, and forget. An AI agent spends three hours studying your project. It reads the documentation. It finds the bug. It tests several fixes. It learns that your team uses pnpm, never npm. It discovers that one database migration must run before the test suite. Then you close the session. The next morning, the agent returns with a clean context window and makes the same mistakes again. It uses the wrong command. It repeats an experiment that already failed. It asks questions you answered yesterday. The model did not suddenly become less intelligent. The system simply failed to preserve what the model had learned. That is the hidden problem facing the next generation of AI agents. The industry has spent years increasing context windows. But a larger context window is still temporary workspace. It helps the model see more information during one session. It does not automatically teach the agent what should survive into tomorrow. Reliable agents need something more deliberate. They need memory. What you will learn Reading time: 10 minutes Difficulty: Beginner-friendly, useful for engineers Main idea: Context helps an agent think now. Memory helps it improve later. Practical outcome: You will leave with a seven-part memory architecture you can build and test. The Difference Most AI Builders Miss Context and memory are often treated as the same thing. They are not. Context is what the model can see during the current request. It may include: - Your instructions - Conversation history - Retrieved documents - Tool descriptions - Search results - Files - Previous actions - Error messages Memory is information stored outside the active context that can be used again later. It may include: - A user’s stable preferences - A project’s architecture - A previous failure and its cause - A successful procedure - An unresolved task - A verified business fact - A decision and the reason behind it Anthropic describes context as a finite resource. As an agent works, its conversation history, tool outputs, instructions, and retrieved data continue accumulating. The engineering challenge is not filling the available window. It is selecting the most useful information for the next decision. Here is the simplest way to understand the difference: A capable system needs both. But when an agent repeatedly performs the same work, a larger context window is not the main solution. The system needs a way to learn from experience without retraining the underlying model. A Context Window Is a Desk, Not a Brain Imagine giving an employee the world’s largest desk. You cover it with every report, email, meeting transcript, policy, spreadsheet, and customer record your company has ever created. Technically, all the information is available. Practically, the employee cannot find anything. This is what happens when developers treat the context window as a storage system. More tokens create more places for useful information to hide. Anthropic calls attention to a related problem known as context rot. As more information is placed in context, a model’s ability to recall and use the correct information can decline. The model has more material available, but less of its attention is focused on what matters. This changes the basic question. Do not ask: How much information can this model hold? Ask: What is the smallest amount of information it needs to make the next correct decision? That single change in thinking can reduce cost, improve speed, and make an agent easier to debug. Bigger context still has value Long context is useful when the agent must compare a large set of documents, understand a codebase, review a lengthy contract, or follow a long conversation. The mistake is assuming that long context removes the need for memory. It does not. Context is where reasoning happens. Memory is where learning survives. The Three Jobs of an Agent Memory System Most memory products focus heavily on retrieval. They ask: “How can we find a relevant past message?” That is only one-third of the problem. A useful memory system must perform three jobs well. 1. Decide what to remember An agent should not save every message, observation, and tool result. Most of what happens during a task is temporary. A search result may be useful for five minutes. A confirmed user preference may be useful for five years. A failed command may matter only until the underlying bug is fixed. Before writing a memory, the system should ask: - Will this information probably matter again? - Has it been verified? - Is it already stored? - Does it replace an older fact? - Who should be allowed to access it? - When should it expire? This is the write policy. Without a write policy, memory becomes a landfill. 2. Retrieve the right memory The best memory is not the one that is most similar to the user’s latest sentence. It is the one that is most useful for the current decision. A retrieval system may need to consider: - Semantic similarity - Recency - Source reliability - User or project scope - Task type - Past usefulness - Confidence - Whether the information has been replaced This is why simple vector similarity is often not enough. Two statements can contain similar words while belonging to different users, projects, or dates. 3. Forget or update old information Forgetting sounds like failure. In a well-designed memory system, forgetting is maintenance. A customer changes their delivery address. A company replaces a policy. A project moves from Python 3.11 to Python 3.13. An experiment disproves an earlier assumption. The old memory should not continue competing with the new truth. A useful system needs rules for: - Expiration - Deduplication - Conflict resolution - Confidence decay - Consolidation - Deletion - Human correction The goal is not perfect recall. The goal is useful recall. Remember An agent that remembers everything will eventually struggle to identify what is still true. Claude Code Shows That Memory Can Start With Files You do not need a complicated database to begin. Claude Code provides a practical example of file-based memory. Each session begins with a fresh context window, but knowledge can persist through two mechanisms: - CLAUDE.md files written by the user or team - Auto memory notes written by Claude as it works Anthropic recommends using CLAUDE.md for stable instructions such as build commands, coding conventions, workflows, and project architecture. Auto memory is used for patterns Claude discovers, including debugging lessons and working preferences. This separation matters. A rule written by the project owner should not have the same authority as a guess created during one agent run. A practical project structure project/ │ ├── CLAUDE.md ├── docs/ │ ├── architecture.md │ └── decisions.md │ ├── memory/ │ ├── MEMORY.md │ ├── debugging.md │ ├── workflows.md │ ├── failures.md │ └── open_tasks.md │ └── src/ A useful CLAUDE.md might contain: # Project Instructions ## Commands - Install packages with pnpm. - Run unit tests with pnpm test. - Run the database migration before integration tests. ## Architecture - Business logic belongs in the service layer. - API routes should not query the database directly. ## Safety - Never modify production configuration. - Ask for approval before deleting data. ## Completion - Do not report a task as complete until tests pass. Notice what is not included: - Entire conversation transcripts - Every error ever produced - Large documentation files - Temporary task details - Unverified guesses Stable rules remain short. Detailed knowledge lives in separate files and is loaded only when needed. Claude Code’s current auto-memory system follows a similar pattern. Its main MEMORY.md file is kept concise, while detailed notes can be moved into topic files and read on demand. The memory files are plain Markdown, which means users can inspect, edit, or delete them. This gives file-based memory three important advantages: - It is transparent. - It can be version-controlled. - Humans can correct it. When files are enough Start with files when: - You have one agent - The project has clear boundaries - The memory collection is still small - Most information is procedural - Human review matters - Retrieval can be based on folders and headings Add a database later when search volume, permissions, relationships, or scale require one. Biggest mistake Do not begin by buying a vector database. Begin by defining what deserves to be remembered. Microsoft’s Stronger Idea: Store Knowledge, Not Raw History Conversation history looks like memory. But most conversations contain noise. A customer support session may include greetings, repeated questions, failed searches, corrections, copied error messages, and unrelated details. Saving the entire transcript forces the agent to process that noise again later. Microsoft Research’s PlugMem takes a different approach. Instead of treating raw interactions as the final memory, it converts them into compact facts and reusable skills. Those knowledge units are organized in a structured memory graph. The system retrieves the units relevant to the current task, then distils them into concise guidance before sending them to the agent. Consider this interaction: User: The export keeps failing. Agent: Are you exporting as PDF or CSV? User: CSV. Agent: Are there more than 50,000 rows? User: Yes. Agent: Use the asynchronous export endpoint. User: That worked. A weak memory system stores the entire transcript. A stronger memory system extracts: fact: condition: CSV export exceeds 50,000 rows outcome: synchronous export may fail skill: when: CSV export exceeds 50,000 rows action: use asynchronous export endpoint evidence: source: verified support interaction result: successful The difference is small in one conversation. Across one million conversations, it is enormous. Microsoft evaluated the same PlugMem module across long conversations, multi-document fact finding, and web navigation. It reported that structured memory supplied more useful information while using less context than the compared approaches. The broader lesson is practical: Do not retrieve the past. Retrieve what the past taught you. Agents Must Learn to Search Before They Write Memory systems often fail because agents write too quickly. The agent discovers something, creates a new note, and moves on. Later, it discovers the same thing again and creates another note. Soon the memory contains five versions of one fact. AutoMem, a 2026 research system, treated memory management as a skill the agent could improve. The agent used files as external memory and learned two basic routines: - LOG: Decide what is worth recording. - PLAN: Search memory before deciding what to do next. The researchers found that optimizing the memory process reduced redundant writes by 68% to 83%. In one environment, changes to the memory structure reduced per-step memory growth from 138 characters to six, a 95% reduction. The improved agent also became more likely to search its existing memory before adding something new. That suggests a valuable production rule: Consult Before Write Before creating a memory, the system should: New observation ↓ Search related memories ↓ Does the same fact already exist? ↓ Yes → update or reinforce it No → create a new memory ↓ Record source, date, scope, and confidence This one rule prevents a surprising amount of memory pollution. The Five-Layer Memory Stack Not every type of information belongs in the same storage system. A practical agent can use five memory layers. Working memory Working memory should be temporary. It contains the current goal, plan, unfinished steps, recent tool results, and immediate evidence. Most of it should disappear when the task ends. Procedural memory Procedural memory tells the agent how to do something. Examples include: - How to investigate a security alert - How to prepare a monthly report - How to deploy an application - How to review a contract - How to qualify a sales lead Procedures should be short, testable, and versioned. Semantic memory Semantic memory stores facts. Examples include: - The customer uses the enterprise plan - Refunds require manager approval - The project uses PostgreSQL - The user prefers concise answers - The API limit is 500 requests per minute Every important fact should have a source and an update date. Episodic memory Episodic memory stores experiences. Examples include: - The deployment failed on August 2 because a migration was missing - A customer rejected the first proposal because the price was unclear - A previous investigation found the alert was caused by a scheduled scanner Episodes are valuable when they can teach the agent what to do next. Do not store an event merely because it happened. Store it because it may change a future decision. Relational memory Some questions depend on connections rather than isolated facts. A knowledge graph becomes useful when the agent needs to answer questions such as: - Which customers are affected by this service failure? - Which decisions depend on this policy? - Which projects use this vulnerable library? - Which employee approved this exception? - Which evidence supports this conclusion? Use a graph when relationships are central to the task. Do not use one only because graphs look advanced. A Memory Record You Can Use Today Every stored memory should carry enough metadata to be judged later. Here is a practical format: { "memory_id": "mem_1048", "scope": { "organization": "acme", "project": "billing-platform", "user": null }, "type": "procedural", "content": "Run database migrations before integration tests.", "source": "project-maintainer", "created_at": "2026-08-05", "last_verified_at": "2026-08-05", "confidence": 0.98, "status": "active", "expires_at": null, "replaces": null, "tags": ["testing", "database", "workflow"] } The content is only one field. The metadata tells the system whether it should trust, retrieve, update, or remove the memory. Minimum fields At minimum, store: - Content - Memory type - Source - Scope - Creation date - Last verification date - Confidence - Status Without scope, information can leak between users or projects. Without dates, stale facts can look current. Without sources, the agent cannot separate verified policy from its own earlier guess. Without status, replaced facts remain active. The Seven-Part Memory Architecture Here is a practical production workflow. 1. User request ↓ 2. Identify user, project, and task scope ↓ 3. Search relevant memory layers ↓ 4. Rank by relevance, trust, and freshness ↓ 5. Build the smallest useful context ↓ 6. Agent acts and verifies the result ↓ 7. Memory gate decides what to save, update, or delete Part 1: Scope the request Determine: - Who is asking? - Which organization owns the data? - Which project is active? - What permissions apply? - Is this a new task or a continuation? Never retrieve memory globally by default. Part 2: Search the correct layer A question about “how we deploy” should search procedural memory first. A question about “what happened last time” should search episodic memory. A question about “which systems are affected” may require relational memory. Routing first makes retrieval smaller and more accurate. Part 3: Rank memories A simple ranking formula can start with: Memory score = relevance × source trust × freshness × scope match × past usefulness You do not need perfect mathematics on day one. You need a consistent policy that can be evaluated. Part 4: Build a memory packet Do not paste twenty raw records into the prompt. Create a short packet: ## Relevant project rules - Use pnpm for package management. - Run migrations before integration tests. ## Relevant previous experience - The last test failure was caused by a missing migration. ## Current open task - Verify that the billing webhook retries failed events. ## Sources - Project instructions, verified August 5 - Incident record 1842, verified August 3 The agent receives the conclusion and enough evidence to judge it. Part 5: Require evidence before action An agent should not treat memory as unquestionable truth. Before a high-impact action, it may need to verify the memory against: - Current configuration - Official policy - A live database - Source documents - Human approval Memory improves efficiency. Verification protects correctness. Part 6: Use a write gate After the task, ask: Did we learn something reusable? Was it verified? Is it already stored? Does it replace something older? Which scope owns it? When should it expire? Only then write the memory. Part 7: Run maintenance Schedule regular memory maintenance. Look for: - Duplicate entries - Contradictions - Expired information - Unused memories - Low-confidence claims - Oversized summaries - Records without sources - Memories stored under the wrong scope A memory system is not a one-time feature. It is an operating process. An Illustrative Business Case Consider a small software company with a support agent. The problem The agent can answer product questions using company documentation. But every customer conversation starts from zero. It does not reliably remember: - The customer’s plan - Their existing integrations - Previous troubleshooting - Promised follow-up - Known account restrictions Support staff must repeat the same investigation. The company considers placing every past conversation into the model’s context. That would increase cost and noise without guaranteeing that the right fact is found. The memory solution The company creates four stores: - Customer facts - Approved support procedures - Previous support episodes - Open commitments Each memory is tied to a customer ID and organization ID. The workflow Customer sends message ↓ System identifies customer and organization ↓ Retrieve active plan, integrations, and open commitments ↓ Search relevant previous support outcomes ↓ Load the correct troubleshooting procedure ↓ Agent responds or takes an approved action ↓ Result is verified ↓ Useful new facts or lessons pass through the write gate Example memory packet Customer: Northstar Design Current facts: - Enterprise plan - Uses Salesforce integration - Data region: Canada Open commitment: - Engineering will review export timeout logs by August 8. Relevant previous outcome: - Large CSV exports succeeded through the asynchronous export process. Approved procedure: - For exports above 50,000 rows, recommend asynchronous export. What the company should measure Do not measure success by the number of memories stored. Measure: - Repeated questions per case - Time to resolution - Percentage of retrieved memories actually used - Incorrect-memory rate - Stale-memory rate - Human correction rate - Token use per resolved case - Customer information leakage incidents - Percentage of cases reopened The memory system succeeds only when it improves decisions. Seven Memory Mistakes That Break Agents 1. Saving every conversation Raw history is not structured knowledge. Extract the useful fact, lesson, decision, or procedure. 2. Mixing all users together Every memory needs a clear scope. A useful memory retrieved for the wrong user becomes a privacy failure. 3. Treating similarity as truth A semantically similar memory may be outdated, unverified, or owned by another project. Rank by trust and freshness as well as similarity. 4. Allowing unlimited writes Require consult-before-write. Search for an existing memory before creating another. 5. Never deleting anything Set expiration rules. Archive replaced policies and remove memories that no longer have a valid purpose. 6. Storing conclusions without evidence A memory should say where it came from. “Tests require Redis” is weaker than “Project maintainer confirmed on August 5 that integration tests require local Redis.” 7. Evaluating memory only through demos A smooth conversation does not prove that memory works. Test retrieval, contradiction handling, scope separation, deletion, and recovery from incorrect memories. How to Evaluate Agent Memory Create a small test set before building advanced infrastructure. Start with 25 to 50 questions. Include five test categories. Recall tests Can the system retrieve a relevant verified fact? Example: Which package manager does this project use? Update tests Can a new fact replace an older one? Example: The project moved from npm to pnpm. Which instruction is now active? Forgetting tests Can expired information disappear from normal retrieval? Example: A temporary approval expired yesterday. Does the agent still rely on it? Isolation tests Can the system keep users and projects separate? Example: Can Project A retrieve a memory owned by Project B? The correct answer should be no. Abstention tests Can the agent say it does not know when memory is uncertain? Example: What database version is in production? If no verified memory exists, the agent should check the live system or ask for clarification rather than inventing an answer. A useful scorecard Pro tip Track retrieval precision before adding more memory. More stored information will not fix weak ranking. Your 30-Day Memory Implementation Plan Week 1: Define the job Choose one narrow workflow. Good starting points include: - Customer support - Codebase assistance - Research projects - Sales follow-up - Incident investigation - Document review Write down: - What the agent repeatedly forgets - Which information should survive - Who owns that information - How long it remains valid - What should never be stored Do not build storage yet. Define the memory contract first. Week 2: Build the smallest memory layer Start with: - One project instruction file - One episodic log - One facts file or table - One open-tasks file - A basic retrieval function - A write gate Use plain files if they can handle the volume. Make every memory inspectable. Week 3: Add retrieval and maintenance Add: - Scope filters - Source confidence - Verification dates - Expiration rules - Deduplication - Consult-before-write - Contradiction detection Create your first evaluation set. Run the same questions with and without memory. Week 4: Measure real outcomes Deploy the system to a limited group or test environment. Track: - Task completion - Repeated work - Incorrect retrieval - Context size - Human corrections - Cost - Security or privacy issues Only add vector search or graph storage when your results show a real need. The Practical Memory Checklist Before calling an agent “memory-enabled,” confirm that it can: - Separate temporary context from durable knowledge - Store facts, procedures, and experiences differently - Attach sources to important memories - Isolate information by user and project - Search before writing - Update an existing memory - Mark an old memory as replaced - Expire temporary information - Retrieve only a small relevant set - Abstain when confidence is low - Allow human review and deletion - Measure whether memory improves task success If the system can only save and retrieve text, it has storage. It does not yet have memory engineering. The Real Opportunity The next generation of agents will not win because they can read the largest prompt. They will win because they can carry useful learning from one task to the next. Anthropic’s context-management work shows the value of removing stale material while preserving important knowledge outside the active window. In its internal evaluations, combining context editing with memory improved agentic search performance by 39% over its baseline, while context editing reduced token consumption by 84% in a 100-turn search evaluation. Microsoft’s PlugMem shows why raw history should be transformed into reusable facts and skills before retrieval. AutoMem shows that memory management itself can be improved, especially when agents learn to search before writing and replace unlimited logs with structured updates. These systems are pointing in the same direction. The model is only one part of the agent. The memory layer determines whether experience disappears or compounds. A chatbot answers the question in front of it. An agent completes a task. A useful agent remembers what the task taught it. That is when AI stops starting over. That is when it begins to improve. Start Small, but Start Deliberately You do not need a complex memory platform to begin. Start with one agent, one workflow, and one question: What should this system still know tomorrow? Write down the rules that must remain stable. Separate temporary notes from durable knowledge. Record where each memory came from. Add a review date. Give the agent a way to update or remove information when reality changes. Then test whether memory improves the outcome. Does the agent repeat fewer mistakes? Does it retrieve the right fact faster? Does it use less context? Does it know when an old memory is no longer trustworthy? That is the real test. The future of AI will not be built by models that simply remember more. It will be built by systems that remember selectively, verify carefully, and forget responsibly. At Cloud AI, we are exploring the practical systems behind reliable agents, including memory, retrieval, graphs, loops, tools, evaluation, and production architecture. Subscribe to Cloud AI for practical guides that help you move from using AI tools to building systems that can actually work, learn, and improve.
02:45

Y Combinator Pays $500K to Build These 13 Apps. I Built #1 With Claude Code and NotebookLM

A tutorial shows how to build a Y Combinator-style app in minutes by combining Claude Code with NotebookLM. It promotes a Claude skill called Moat Breaker that researches competitors online, trains a NotebookLM for grounding, then spins up three agents to build the app and spot market gaps. The author includes a demo app that gamifies kids' learning and offers the files for download. Framed around YC's $500K investments, but the piece is mostly promotional with a thin build demo.

Notes
The Moat Breaker — Claude Skill for Building YC-Style Apps

Context. Substack post by "LearnAIWithMe" (published 2026-08-06) reacting to "YC Requests for Startups," a list of 13 startup ideas. Author says Y Combinator invests ~$500,000 per startup and publishes idea lists.

The skill workflow ("The Moat Breaker"), in order:

  • Searches the web to find your rivals.
  • Identifies industry best practices.
  • Trains a NotebookLM on that research.
  • Spawns 3 agents: Shell Agent (front-end), Engine Agent (connects front-end to back-end), Gap Agent (finds market gaps).
  • The agents consult NotebookLM while building in Claude Code.

The Claude Code + NotebookLM pairing is explicitly to avoid hallucination: > "Because NotebookLM is very good at grounding."

Demo app — "The Primer" (gamified kids' learning app):

  • Onboarding questionnaire builds a child profile.
  • "For grown-ups" view shows progress, the curriculum standard behind each skill, and which specific items tripped the child up.
  • Lightning mode: 60 seconds, 4 big color buttons, no typing; a star per right answer. Author claims 28 correct in one minute (awarded to his 2-month-old daughter Eva, whose score he set himself).
  • Badges for engagement; race mode vs. "Pip the fox" — each correct answer advances a flag.
  • Every answer is followed by a one-sentence "why" explanation.
  • Win reward: bonus stars + new badge + Pip asks for a rematch.

Install. Files + "single prompt" in a Google Drive folder; paste the prompt, installs in ~5 minutes. App files are included in the author's "Vault" for download.

Caveats. No YC outcome reported; "app will be better than your rivals" is asserted, not demonstrated. No technical details on agent architecture, prompt content, or install steps — the Drive link carries the actual substance, which is not reproduced.

Full text · 3,391 chars
Y Combinator Pays $500K to Build These 13 Apps. I Built #1 With Claude Code and NotebookLM The Moat Breaker skill researches rivals, trains a NotebookLM, then three agents build your app in Claude Code. 5-minute install, files and demo app included. While scrolling through X, I came across a post called “YC Requests for Startups.” It listed 13 startup ideas. As someone who enjoys building things with AI, it immediately caught my attention. But I did not know about Y Combinator. It turns out Y Combinator (YC) is one of the world’s best-known startup accelerators. They invest in early-stage startups, typically $500,000 per company. And they often share lists showing expectations from developers. So, I picked one of the ideas and built it. But there is more for you. I also created a skill for you. So you can build anything you want. One sentence enough. But why should you use this skill instead of telling Claude what to build? Because this skill first researches the industry, finds gaps, and builds. So the app will be better than your rivals. Let me show you the skill first. The Moat Breaker This Claude Skill first searches the web and finds your rivals. Next, it identifies industry best practices. And trains a notebookLM. Next, it creates 3 agents, dedicated to building this app. - Shell Agent: Build the front-end of the page, the part you see. - Engine Agent: Connects the front-end with the back-end, works like an engine. - Gap Agent: Find gaps in the market. And while building an app on Claude Code, these agents talk with NotebookLM. The reason for this Claude Code and NotebookLM combination is to avoid hallucination. Because NotebookLM is very good at grounding. And I’ll show you what I built next. The Primer I gamify the entire process. To join the game and start learning, you will need to answer a few questions. Once you fill in the information, your profile is ready. When you click on “For grown-ups”, you see the progress, the exact curriculum standard behind every skill, and what actually tripped your child up. You get sixty seconds, and every right answer earns a star, in the lightning mode. Questions come one after another with four big colorful buttons, no typing, just tap and go. Eva got 28 correct answers in one minute (I did it for her; she is only 2 months old 🙂), setting a new record and earning 28 stars for the bank. And Eva has badges, which makes the entire process more exciting. Another mode is racing Pip the fox, where every right answer moves you one step closer to the flag. After each answer, the tutor explains the why in one kid-sized sentence, like a private teacher sitting next to you. Win the race and everything lands at once, bonus stars, a new badge, and Pip asking for a rematch. Now, this is the app I built with it. But you can build any app you want. I also added this app’s files to the Vault, so you can download them, submit the app to YC, turn it into a product, or let your son or daughter use it. Your choice. The Installation of The Moat Breaker in 5 Minutes Here are the everything inside the Google Drive, including the single prompt that will install everything for you. Download the folder, paste the prompt, and let it install everything on your PC. Next, describe any app you want in one sentence, build it, and apply to YC. If you get accepted, definitely let me know. Here is the Google Drive link.

Web

2
00:00

Three AI Pioneers Clash Over Jobs, Regulation And The Future Of AI

Three of the biggest names in AI publicly disagreed about the future of the field at a rare joint appearance in Las Vegas. Nobel laureate Geoffrey Hinton warned AI will bring a wave of white-collar job losses and called open-model releases dangerous, while Andrew Ng accused big companies of exaggerating those fears to block open models. Fei-Fei Li landed in the middle, saying AI changes tasks rather than whole jobs and that rising productivity won't automatically mean shared prosperity. Hinton said regulation is the steering wheel, not the brake, and cited California's vetoed frontier-model safety bill.

Notes
Jobs

Andrew Ng argued AI changes job boundaries before erasing jobs. Software dev is his evidence: "only a small part of what software engineers do is writing code." Narrow specialists (front-end/back-end/mobile) are "now able to rise up to be a broader type of developer... full stack." Conceded short-term pressure but says demand grows for tool users.

Geoffrey Hinton disagreed. Call centers are the exposed category: workers with limited training answering repetitive questions. "What are those people going to do? They typically don't have a high level of education. Anything you could retrain them to do, AI will be able to do." Cited a relative answering written health-service complaints: ~30 min per draft → ~5 min with a chatbot; a fivefold productivity gain "may lead an employer to keep fewer people." "Many jobs will go the way of people who dig ditches when backhoes came along." He predicts a 2026 wave of white-collar job losses.

Fei-Fei Li split the difference: the unstated word in "AI and jobs" is replace. Jobs are multi-task (nurses, teachers, journalists). She grounds it in caring for elderly parents — chart completion / pharmaceutical-order verification can shed clerical strain "without replacing the human work of nursing." Sharpest line: "Increased productivity does not translate to shared prosperity" — productivity is an operational measure; prosperity is "a political and economic choice." Calls for a "soft landing": retraining, education, community investment. "The last thing we should do is to debilitate people and take the agency away from people."

Openness vs. safety

Ng: AI companies cycle through threats (extinction, biological weapons, job destruction, China competition) to justify release restrictions that favor incumbents. "I don't want there to be gatekeepers of AI."

Hinton rejected the fearmongering charge and distinguished open-source (inspectable, repairable code) from open-weight (copied/adapted parameters): open weights let bad actors retarget powerful models for cybercrime "at a fraction of the original training cost." "We're seeing AIs that have a lot of ability doing things that people didn't intend them to do. That's worrying." Cites the 2026 International AI Safety Report (100+ experts; Hinton on the advisory panel).

Regulation

Hinton rejected the accelerator/brake analogy: "Developing AI is like the accelerator of the car... Regulation is like the steering wheel." Defends California SB 1047 (passed both chambers 2024; Gov. Newsom vetoed — safety testing + disclosure for large models; Newsom argued it over-indexed on model size and risked false security).

Li prefers sector-based rules — drug agencies for AI medical products, financial regulators for lending/trading, transport authorities for AVs. Opposes "put the genie back in the bottle"; wants public research funding, updated sector rules, use-based guardrails, warning that private-only labs would narrow which questions get pursued.

Splits: Ng (optimist, open weights), Hinton (labor shock + open-weight misuse), Li (task-level transformation, prosperity distribution, sector regulation).

Full text · 9,416 chars
Three of artificial intelligence’s most influential pioneers revealed sharply different visions of the technology’s future during a rare joint appearance in Las Vegas this week. At Ai4 2026, Geoffrey Hinton, Fei-Fei Li and Andrew Ng clashed over AI’s impact on jobs, how its risks should be defined and regulated, and who will benefit from its gains. Hinton, the Nobel Prize-winning computer scientist whose research helped propel modern deep learning, warned that AI systems are gaining dangerous capabilities faster than institutions can respond. Ng argued that powerful companies have exaggerated those threats to restrict open models and limit competition. Li rejected both alarmism and technological utopianism, calling instead for a more scientific debate centered on human agency and the distribution of AI’s benefits. Hinton Says AI Will Bring A White-Collar Labor Shock The panel’s most consequential divide concerned employment. Ng challenged predictions of rapid, economy-wide displacement. He pointed to software development, where AI can write growing amounts of code, yet engineers still handle product choices, system design, customer demands and organizational coordination. His argument was that AI changes the boundaries of a job before it erases the job itself. Narrow specialists can take on broader roles while a front-end developer may become a full-stack developer. Workers who once handled only one portion of a process now may manage a much larger cycle. That shift creates short-term pressure on jobs, Ng conceded, but he also explained that it creates demand for people who know how to use the new tools. “In software engineering, AI is taking over a lot of the writing of code. But it turns out only a small part of what software engineers do is writing code,” explained Ng. “What has happened is we used to have a lot of developers that were specialized in narrow niches, like front-end development or back-end development or mobile development. A lot of the developers are now able to rise up to be a broader type of developer, developing it as a full stack” Hinton was far less sanguine. He cited call centers as an exposed category. Many workers in those roles receive limited training, earn modest wages and answer repetitive questions. AI systems, he argued, will soon answer those questions more accurately and at lower cost. “What are those people going to do?” Hinton asked. “They typically don’t have a high level of education. Anything you could retrain them to do, AI will be able to do.” His concern extends beyond customer service. Once machines can perform routine intellectual labor, entire classes of office work could face the same pressure that mechanization brought to manual trades. Hinton did not claim that AI would create no jobs. He said that jobs may appear, but displaced workers may lack the education, experience or location needed to obtain them. He offered a personal example. A relative who answered written complaints for a health service once spent roughly half an hour drafting each response, but a chatbot cut the work to about five minutes. That can produce abundance in fields where demand has room to grow. More efficient doctors and nurses could provide more care, and administrative departments may perform more effectively. The resulting fivefold productivity gain may lead an employer to keep fewer people. “Many jobs will go the way of people who dig ditches when backhoes came along,” said Hinton. The disagreement reflects a broader split in economic forecasts. Some executives expect AI to expand hiring in technical and senior roles, yet concerns remain acute around clerical, support and entry-level work. Hinton has continued to warn publicly that 2026 could bring a new wave of white-collar job losses. Fei-Fei Li Says AI Changes Tasks, Not Entire Jobs Li offers a more balanced approach. The phrase “AI and jobs,” she argued, has acquired an unstated word between its two terms: replace. That assumption flattens a more complicated economic change. Few jobs consist of one task. Nurses administer care, document treatment, check medications, communicate with families and coordinate with other clinicians. Teachers explain concepts, motivate students, assess progress and manage classrooms. Journalists interview sources, weigh evidence, choose angles, write and edit. AI may automate one portion, accelerate another and leave a third untouched, she said. Li grounded the point in her own experience caring for elderly parents. A system that helps nurses complete charts or verify pharmaceutical orders could remove clerical strain without replacing the human work of nursing. “Current jobs are transforming,” she said. That transformation demands training, public investment and far more nuanced language than either mass-unemployment headlines or promises of effortless abundance provide. Her sharpest line concerned the distribution of wealth. “Increased productivity does not translate to shared prosperity,” Li said. A business can produce more with fewer labor hours and still leave workers with lower bargaining power, weaker career paths or no share of the financial gain. Productivity is an operational measure whose benefits may accrue to a limited number of people, she explained, while prosperity is a political and economic choice. Li called for what she described as a “soft landing” for occupations that do shrink. That could include retraining, educational support and investment in communities where displaced work is concentrated. Fear, she said, can paralyze the very people who most need to experiment with new tools. “The last thing we should do is to debilitate people and take the agency away from people,” Li said. Andrew Ng Says Fear Of AI Benefits Big Tech Ng accused some AI companies and advocates of cycling through threats to justify restrictions on software releases. Extinction, biological weapons, job destruction and competition with China have each been invoked, he argued, by organizations that benefit when the cost of building or distributing AI rises. Ng proposed openness as a counterweight. Closed systems can prosper, he said, but governments should resist an AI market controlled by a few gatekeepers. “I don’t want there to be gatekeepers of AI,” Ng said. He argued that open alternatives let researchers, entrepreneurs and governments adapt systems to their own needs. But Hinton rejected the idea that warnings about frontier systems amount to fearmongering and said open-weight models are not a sufficient counter weight. He distinguished open-source software, where outsiders can inspect and repair code, from open-weight models, where trained parameters can be copied and adapted. Open weights can lower barriers to research and competition. They can also let bad actors modify a powerful model for cybercrime or other harmful uses at a fraction of the original training cost. For Hinton, that danger is no longer theoretical enough to dismiss. He cited systems that appear to behave outside their operators’ intentions and AI’s increasing ability to identify security flaws. The 2026 International AI Safety Report, produced with input from more than 100 experts and an advisory panel that included Hinton, likewise examined emerging capabilities, cyber risk and the limited state of current safeguards. “We’re seeing AIs that have a lot of ability doing things that people didn’t intend them to do,” Hinton said. “That’s worrying.” Can AI Be Regulated Without Holding It Back? The panel clashed on the role of regulation as well. Technology companies often depict development as a car’s accelerator and regulation as its brake, but Hinton said the analogy is wrong. “Developing AI is like the accelerator of the car,” he said. “Regulation is like the steering wheel.” The goal, in his account, is not to stop technical progress. It is to direct progress toward systems that improve human welfare and away from systems that impose unacceptable risks. Hinton cited California’s SB 1047, a contested frontier-model safety bill that passed both houses of the state legislature in 2024 before Gov. Gavin Newsom vetoed it. The proposal included safety testing and disclosure requirements for developers of certain large models. Newsom said the bill focused too heavily on model size and could create a false sense of security. Li took a sector-based approach. AI enters industries that already have regulators, professional rules and safety systems. Drug agencies can examine AI-enabled medical products. Financial regulators can address automated lending or trading. Transportation authorities can govern autonomous vehicles. She opposed attempts to freeze AI development or “put the genie back in the bottle.” Her preference was public research funding, updated sector rules and guardrails tied to actual uses. Li’s call for public investment carried another warning. Modern AI grew from university laboratories, open publication and publicly supported science. A future in which only a few private companies can afford leading research would narrow the questions that get pursued. What is most striking is how far these three AI pioneers have begun to diverge. Once broadly aligned around the promise of building more capable systems, they now disagree over what those systems will mean for workers, markets, education and public safety. Their split reflects a field entering a new phase.
00:00

Most Enterprise AI Is Live. Half Of Companies Can't Prove It Works

Most of the world's largest companies now run AI in production, but half of them can't show it actually works. Plug and Play's survey of Fortune 500 and Global 2000 firms found 74% have at least one AI system live and 93% are piloting or further along, yet half of production-stage companies can't measure ROI. Narrow single-function deployments measure worst, with 74% there saying it's too early or untracked, and data foundations stay the top blocker at 71%. IDC and Microsoft peg the average return at $3.70 per dollar with 14 months to positive ROI, and the advice is to put an AI metrics standard under the CFO.

Notes

Most Enterprise AI Is Live. Half Of Companies Can't Prove It Works

Source: Forbes (scrape), 2026-08-06, based on Plug and Play's 2026 Enterprise AI Strategy Pulse Survey.

Headline numbers
  • 74% of the world's largest enterprises run ≥1 AI solution in production; 93% piloting or further along.
  • Half of production-stage companies cannot tell whether any of it worked (ROI not consistently measured).
  • Sample skews Fortune 500 / Forbes Global 2000 — "the top of the market, not the long tail."
"Enterprise AI has crossed the deployment threshold, but not the value threshold… The next chapter is less about the models themselves and more about the operating foundations around them: better data, governance, integration, ownership, and measurement."
— Amit Patel, Ventures Partner at Plug and Play
What leaders do right
  • Only 5% self-identify as AI-native, but 37% run AI in a single function and 32% across several.
  • 92% rank data privacy, explainability, and compliance as top vendor-selection factors — ahead of performance (74%) and flexibility (53%). Article calls this ordering "unthinkable in 2024, when benchmark scores sold deals."
  • Ownership: only 21% sits with CAIO/AI lead/CoE; 37% with function/LoB heads. BCG: 72% of CEOs are main AI decision-maker.
Where they go wrong
  • Worst measured = narrowest deployments: 74% of single-function adopters say ROI is "too early to measure or is not tracked at all" — worse than the topline half. Claim: a use case live without a pre-deployment metric is "unmeasurable forever."
  • Corroboration: KPMG June 2026 pulse (2,145+ leaders, 20 countries) — only 7% report established AI ROI. Deloitte (3,235 leaders): 66% report productivity gains, only 20% see AI-driven revenue growth.
  • IDC/Microsoft: avg $3.70 return per $1 on gen-AI, median 14 months to positive ROI — but most enterprises review on four quarters, killing mid-flight projects.
  • Data foundations top blocker at 71%, staying elevated across adoption stages; data still funded as project overhead. Integrate.io: strong data integration → 10.3× ROI vs 3.7× for poor connectivity.
Prescription

Set up an AI metrics standard owned by the CFO; otherwise LoB frameworks won't be comparable or roll up.

Full text · 4,517 chars
Enterprise AI has cleared the adoption hurdle and stalled on the proof hurdle, according to Plug and Play’s 2026 Enterprise AI Strategy Pulse Survey, released today. Seventy-four percent of the world’s largest enterprises now run at least one AI solution in production and 93% are piloting or further along. Half of those production-stage companies cannot tell you whether any of it worked. The sample skews to Fortune 500 and Forbes Global 2000 companies, so this is the top of the market, not the long tail. "Enterprise AI has crossed the deployment threshold, but not the value threshold. Nearly three-quarters of the enterprises we surveyed have AI in production, yet half still cannot consistently measure ROI. The next chapter is less about the models themselves and more about the operating foundations around them: better data, governance, integration, ownership, and measurement. Enterprises that pull ahead will be those that consistently translate production deployments into tangible business outcomes," Amit Patel, Ventures Partner at Plug and Play, told me. Here’s the overall Enterprise AI Gap from all the reports that are in the market now. Here is what the top cohort is getting right, and where it is going wrong. Right: Enterprise AI Leaders Stopped Debating And Shipped Two years of pilot purgatory has ended. Only 5% of these organizations call themselves AI native, but 37% run AI inside a single function and 32% operate it across several. Deployment is no longer the constraint, and that is a real accomplishment that gets lost in the ROI hand-wringing. Right: Enterprise AI Buyers Choose Defensibility Over Demos Ninety-two percent of respondents rank data privacy, explainability and compliance as their top vendor selection factors, ahead of performance at 74% and flexibility at 53%. That ordering would have been unthinkable in 2024, when benchmark scores sold deals. Enterprise buyers have learned that a model they cannot explain is a model they cannot defend to a regulator, a board or a customer. Right: Enterprise AI Ownership Moved Into The Business Only 21% of AI ownership now sits with a CAIO, AI lead, or center of excellence. Thirty-seven percent sits with functional or line-of-business heads, and BCG found 72% of CEOs are now the main AI decision-maker. This is correct in principle. Value is created inside the function, not inside IT, and the person accountable for the outcome should control the tool. Wrong: Enterprise AI Ships Without A Baseline The narrowest deployments are the worst measured. Among companies running AI in a single function business function, the earliest and simplest stage of production, 74% three-quarters say ROI is too early to measure or is not tracked at all. That is worse than the topline half, and it is backwards from what you would expect, since a single function should make attribution easiest. A use case that goes live without a pre-deployment metric is unmeasurable forever, because there is nothing left to compare against. Larger datasets confirm this is not a small-sample artifact. KPMG’s June 2026 pulse of more than 2,145 leaders across 20 countries found only 7% report established AI ROI. Deloitte’s study of 3,235 leaders found 66% report productivity gains but only 20% see AI-driven revenue growth. Wrong: Enterprise AI Is Measured On The Wrong Calendar IDC and Microsoft measure an average return of $3.70 per dollar spent on generative AI, with a median time to positive ROI of 14 months. Most enterprises review on four quarters. That mismatch kills projects on their way to a positive return and rewards the ones that produce a fast, shallow win. The same short term viewshows up in data spending. Data foundations remain the top blocker to scale at 71%, and they stay elevated at every stage of adoption rather than fading as companies mature. Most enterprises still fund data work as project overhead instead of a standing line item. Integrate.io found strong data integration produces 10.3 times ROI against 3.7 times for poor connectivity. The One Enterprise AI Move For The Second Half Of 2026 Set up an AI metrics standard, owned by the CFO. Without it, every line-of-business leader will create a defensible framework for their own function, but none of those frameworks will be comparable or roll up across the enterprise. That is how companies end up with ownership sitting in 37% of the business and no unified view of value. Proof is what will carry Enterprise AI investment through the 2027 budget cycle.