Nothing matches those filters.

Lead

3

Video

2
03:26

New Deepseek, Seedance 2.5, Minimax H3, Gemini Robotics, AMD models: AI NEWS

DeepSeek shipped a new small model that rivals the best open-source AI while being about 100 times cheaper. V4 Flash scores close to GLM 5.2 and Claude Opus on coding and security benchmarks despite being 70% smaller, costs around 3 cents per million tokens, and its weights are already out on Hugging Face at about 167 GB. The roundup also covers ByteDance's Seed Dance 2.5 and MiniMax H3 video generators, Netflix's style-transfer tool, a fast transcription model, Moonshot releasing Kimi K3's open weights, and AMD's first open model trained entirely on its own chips.

Notes
AI news roundup (week of ~Aug 2, 2026) — AI Search (YouTube)
ByteDance Seedance 2.5 (video)
  • Succeeds Seedance 2 (already billed as best video model). Strong on high-action fight scenes and character consistency. Multimodal: accepts reference videos, 3D scenes, green-screen footage, storyboards as inputs.
  • Generates up to 30-second clips (competitors cap at 15–20s); currently 720p only, with 1080p/4K "coming soon." Accepts up to 50 reference inputs (audio, image, video).
  • Availability: live on ByteDance platforms (e.g., Dreamina) in select countries, not yet in the US. API not out; rollout planned "next week."
  • Pricing: a 10s clip ≈ 460 credits; 1,000 credits = $10 → ≈ $4.60/clip. Presenter calls it "one of the most expensive video models to use," far pricier than MiniMax H3.
MiniMax H3 (video)
  • From the company behind Hailuo (apparently rebranded as the MiniMax H series). Multimodal input; 2K resolution output, up to 15s.
  • 10s clip = 120 credits ≈ $1.20 (1,000 credits = $10), ~3x cheaper than Seedance 2.5.
  • Will be open sourced, likely next week. Presenter predicts it becomes the best open video model.
DeepSeek V4 Flash (new version)
  • Reported to match MiniMax M3 and sit "just one point below GLM 5.2"; ~10 points above previous V4 Flash, well above V4 Pro.
  • Shares the DSpark architecture (same structure as V4 Flash DSpark) — claimed efficiency/throughput breakthrough. Clobbers V4 Pro on agentic coding, software engineering, cybersecurity; occasionally matches/beats GLM 5.2 or Opus 4.8.
  • Claims: ~70% smaller than GLM 5.2 and 100x cheaper than Claude Opus; ~$0.03 per million tokens, "by far the cheapest model at frontier intelligence."
  • Open weights on Hugging Face at 167 GB (fits one DGX Spark); Unsloth GGUF versions, smallest 1-bit ≈ 82.5 GB.
Kimiko 3 — Moonshot AI (as named in video)
  • Promised July 27 release, delivered on that date on Hugging Face. 2.8T-parameter MoE, 104B active. Uses Kimi delta attention and attention residuals; native vision via MoonVIC encoder.
  • 1.56 TB weights; Unsloth 1-bit GGUF ≈ 594 GB. Presenter: "currently the most powerful open-source model."
AMD Instella MoE
  • Trained from scratch on AMD Instinct + ROCm (no CUDA). 16B total / 2.8B active params. Reportedly beats Gemma 3 4B and a smaller Qwen 3.5.
  • Publishes pre-training/mid-training checkpoints plus recipes and code. Uses multi-head latent attention and a "far skip collective" for communication/compute overlap. RL-finished version, "Think", ≈ 32 GB.
Gemini Robotics 2 — Google DeepMind
  • Family spanning feet-to-fingertips control; predecessor was upper-body/tabletop only. Demoed on the Apollo 2 humanoid (walk → locate → pick up → carry → place).
  • Three variants: Gemini Robotics 2 (VLA: language + camera vision → motor commands); Gemini Robotics ER 2 (environment/task reasoning, failure correction, multi-robot coordination); Gemini Robotics on-device 2 (offline, local). Improved dexterity: unscrewing light bulbs, tying trash bags, sealing Ziplocs. Tryable in AI Studio; trusted-tester signup.
Other releases
  • Netflix IDV2V (open source): restyle an existing video via one edited keyframe; changes background/lighting/clothing/style while preserving identity and motion. Up to 720p; ~80 GB model — high-end hardware only.
  • Crisper Whisper 2 (transcription): verbatim mode (keeps stutters, hesitations, laughter, repeats) vs intended mode (polished), with word-level timing. 4 model sizes, smallest 0.2B (<500 MB, no GPU), largest 2B (3 GB). Self-benchmarked above ElevenLabs; lowest timestamp error. Free HF Space demo.
  • Redesign (image→layers): splits a flat image into editable layers (recolor/reposition/resize); stacks PaddleOCR, Qwen Image Layered, DINO + SAM 2. Claims to beat Qwen Image Layered on its benchmark. <4 GB, but requires an OpenAI API key (swappable for local model).
  • Inkling Small (Thinking Machines, ex-OpenAI CTO's lab): omnimodal quarter-size of Inkling — 276B total / 12B active, 532 GB. 1 pt below full Inkling on Artificial Analysis; sometimes beats it per-compute. Advantage is audio understanding; lags DeepSeek V4 Flash on intelligence.
  • Prism (robotics control): fuses force, velocity, contact, friction, joint-angle signals → higher manipulation success rates than alternatives. Code released.
  • Ideogram Object Remover: brush-select removal including shadows and reflections; lowest error on Removal Bench vs Nano Banana 2 and GPT Image 2 Medium. Free with daily credits.
  • Wonder (Adobe, video world model): real-time key-navigable scenes, also from video input (walk around in 3D). "Far from perfect — noise and artifacts." Code "coming soon."
  • Phi-0 (world model): "physical language" — reasons about physics before rendering frames; targets interactive worlds, autonomous-driving and robot-training video. Outperforms similar world models on physical coherence. Code "coming soon."
  • Gemini voice dictation (macOS only): hold Fn, speak into any app; strips fillers/mistakes/repeats, handles mid-sentence corrections, inserts at cursor; optional Gemini reasoning for tasks like summarizing highlighted docs. Similar to TypeLess/Whisper Flow.
Caveats
  • Seedance 2.5 unavailable in US; API pending. Wonder and Phi-0 are demos, models unreleased — the presenter explicitly doubts Adobe will ship.
  • MiniMax H3 open-source timing is a promise ("likely next week"), not confirmed.
  • Whisper 2's superiority claims come from the developers' own benchmarks.
  • Names as pronounced in video ("Kimiko", "Kwen"/"Seedance") may differ from official spelling.
Transcript · 29,510 chars
AI never sleeps and this week has been absolutely insane. ByteDance releases the best video model out there, Seed Dance 2.5. MiniMax also releases their best video model, MiniMax H3. And the awesome thing is this will be open source. Deep Seek releases their newest model and they pulled off the impossible again. It's as good as GLM and Opus but like a hundred times cheaper and way smaller. We have a new open source model that's trained entirely on AMD chips, not Nvidia. This AI can turn an image into transparent layers which you can edit further. We have some new video world models which you can interact with in real time. Google releases their latest robotics model and a lot more. So let's jump right in. First up, Netflix releases a really cool open source AI called IDV2V. In the simplest sense, this can basically change the style of the scene without affecting the identity or the movement of the characters. So you can just plug in an existing video and then edit one keyframe to show the new look you want and then the system will spread that style across the entire video. So you can change things like the background, lighting, clothing, or overall style while keeping the face and expressions and movements of the character the same as the original clip. Now at the top of the page, if you click on this GitHub button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Notice that this can generate videos of up to 720p and this is pretty huge. So the main model is like almost 80 gigabytes in size. So this would only fit on like really high-end consumer hardware. But hopefully there will be more compressed versions of this in the future. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new open source transcription tool which is really powerful. It's called Whisper Whisper 2 and this can basically take any audio and turn it into text like this. The cool thing is this actually has two different outputs. If you turn on the verbatim mode, this includes everything including stutters, hesitations, laughter, or other meta tags, plus repeats like this. Now, alternatively, you can also turn on the intended mode, which would get rid of all these hesitations and meta tags and give you the fully polished transcript. So, a super flexible tool. Not only that, but this also gives you precise word level timing. In fact, let's try this out. They've already released a free hugging face space for you to try this out online. Here is where you can upload any audio. So, let's upload this one. We are going to generate this in verbatim mode. >> However, due to the slow communication channels, styles in the West could lag behind by 25 to 30 years. >> All right, so that was the audio and as you can see, the transcript is indeed correct and it also gives you the start and end times for each word. Or here's another example. Let me play you the audio first. >> And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country. >> And as you can see, it's able to provide me the transcript here, plus the timing for each word. Now, this supports all these different languages as you can see here, and if you look at this self-made benchmark, then Crisper Whisper 2 is indeed even better than 11 Labs or some other leading transcription tools. And if you look at this benchmark on word level timestamps, then as you can see, Crisper Whisper also contains the least amount of errors. The awesome thing is they've released the models to this. So, if you click on this button, they've released four different models in this family. The smallest one is 0.2 billion parameters, and this is less than 500 megabytes in size, so you can easily fit this on most consumer devices. You don't even need a GPU. And then, the largest one is 2 billion parameters. This one is 3 gigabytes in size, which is still fairly tiny. You should be able to fit this in most GPUs. And then if you scroll down the Hugging Face page here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Also this week, the Goat DeepSeek is back with another mind-blowing model. They just released the latest version of DeepSeek V4 Flash, and even though this is a flash model, it even performs as good as some of the full open-source models out there, including MiniMax M3, and it's just one point below JLM 5.2. This scores way higher than even the previous DeepSeek V4 Pro version, and as you can see it's like 10 points above the previous DeepSeek V4 Flash. What an insane upgrade. Note that here it says it has the same model structure as DeepSeek V4 Flash DSpark. In fact, this is a really important architecture breakthrough, which helped it improve efficiency and throughput by a huge amount. I did a full explainer video on DSpark, so definitely see this video if you're interested in learning more. If you look at these benchmarks on agentic coding, software engineering, and cybersecurity, you can see that it absolutely clobbers the previous V4 Pro version. And for some instances, it even beats or matches JLM 5.2 or Opus 4.8. Keep in mind this is like 70% smaller than JLM 5.2, and it's also 100 times cheaper than Claude Opus. DeepSeek V4 Flash costs around 3 cents per million tokens, making it by far the cheapest model at Frontier Intelligence. And if you look all the way on the other side, GPT 5.6 and Claude Opus and Claude Fable are all the way over here. So in terms of performance versus cost, this is by far the best option to use. It's just incredibly efficient and just way cheaper than the other frontier models. Now as expected from the Goat, they've also released the model to this already. It's already out for you to download on Hugging Face and this is only 167 GB in size. So, this can easily fit on just like one DGX Spark. Now, because this is open source, the community has acted fast and for example, Unsloth has already released GGUF versions of this. The crazy thing is the smallest one-bit version is only like 82.5 GB in size. So, you can potentially fit this on just like one or two pieces of high-end hardware. And keep in mind, this has the intelligence of GLM 5.2 or Opus 4.8. It's pretty crazy that you can now run this level of intelligence locally. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have quite a useful AI called Redesign. This basically turns a flat image into layers. So, you can edit each component or move them around. And with this, you can customize or micro-edit certain things like recoloring the elements or repositioning certain things or changing the size of elements. So, think of this as like turning a screenshot back into something closer to a Figma or Photoshop project. Now, this actually orchestrates a ton of different AI tools at once. It uses PaddleOCR. In terms of generating the layers, it uses Kwen Image Layered, which I featured on my channel before. And then for detecting and segmenting different elements, it uses Dino and SAM too. Now, according to these benchmarks, it seems like this new redesign even does better than some other image-to-layer generators like Kwen Image Layered. At the top of the page, they've released everything already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. And this is actually fairly tiny at less than 4 GB in size. However, at least according to this base code, it also does require, apparently, an OpenAI API key. But you could tweak the code further to swap this out for a local model to run it completely for free. If you're interested in reading further, I'll link to this main page in the description below. Around 2 weeks ago, Moonshot AI released Kimiko 3, which is currently the most powerful open-source model you can use right now. Now, on their release page, they said that the model weights will be released on July 27th. Well, as promised, this week, on July 27th, they released the full Kimiko 3 model on Hugging Face. So, here are some specs. This is a massive 2.8 trillion parameter mixture of experts model, so think of it as like a team of specialists working together at once. And when you use it, only 104 billion of these parameters are active. For the architecture, it uses Kimiko delta attention and attention residuals, which were also designed by them. In fact, if you want to learn more about attention residuals, which is actually a really fascinating breakthrough, see this video to learn more. And also, this has vision capabilities natively baked in. And they used this Moon VIC vision encoder to make it happen. Now, if you click on files and versions, as expected with a 2.8 trillion parameter model, this is like 1.56 terabytes in size. So, you're going to need to stack like multiple enterprise GPUs in order to run this. However, because this is open-source, the community is already really quick to help quantize and compress this further. So, for example, Unsloth released some compressed GGUF versions of Kimiko 3. And the 1-bit version is only like 594 GB. Still massive, but this is a really big drop from like 1.6 terabytes in size. Anyway, if you want to download this or check out the technical details, I will link to this Hugging Face page in the description below. Also this week, AMD releases their very own open-source model called Instella MoE. And this is quite a big deal. You see, Nvidia has basically dominated the AI space because most AI tools are built off of their CUDA platform. It's really hard to train and run AI models on non-CUDA chips. But here, AMD has trained this Instella mixture of experts model from scratch on their AMD Instinct and using the AMD Rockum software stack. First of all, here are some specs of this. So, this has 16 billion total parameters and this is a mixture of experts models. So, think of it as like a team of specialist AIs working together. So, when you use it, only 2.8 billion parameters are active, making it fairly efficient. The cool thing is this reportedly beats other similar-sized models, including Gemma 4E4B and a smaller version of Quen 3.5. The awesome thing is they're not just releasing the final model, they're actually publishing the checkpoints from pre-training, mid-training, and all these subsequent stages, along with the training recipes and the code. It uses what they call a multi-head latent attention to make the attention and memory more efficient. And then they also use something called a far skip collective, which helps overlap communication and computation across GPUs. This is a bit technical, but on this page they document how exactly they did all these stages of training to create the model. And then if you click on this GitHub repo down here, it contains all the instructions on how to download and run this, as well as the training scripts. And then if you click on their hugging face repo, note that the final version after training on reinforcement learning, this is called Think and this is roughly 32 GB in size. So, it's a medium-sized model, similar to Quen 3.6, which should be able to fit on high-end hardware. If you're interested in reading further, I'll link to this main page in the description below. Also this week, one of my favorite image generators, Audiogram, has released a new tool called object remover. And here's how it works. You can simply brush over the object that you want to remove and it'll automatically highlight that object and then remove it for you. It's as simple as that. Here are some other examples for your reference. So, we can select this bike and remove it and notice that it also removes the shadow of the bike. Or here's another really tricky example where I want to remove this plant, but there's also a lamp that's kind of occluded plus some books in front of it. But as you can see, it's able to remove the plant while also preserving the lights and the books. Plus, it also is able to remove the plant's reflection on the floor. And here's another example. Now, at least according to this removal bench, you can see that Ideogram Object Remover has the lowest error rate compared to other models like Nano Banana 2 and GPT Image 2 Medium. So, if you're looking for a free object remover which you can run online, this is a nice option for you to use. You can sign up for free and you get some daily credits to run some generations. If you're interested, I'll link to this in the description below. If you want to supercharge your content creation, definitely check out Higgsfield, the sponsor of this video. Think of it as an all-in-one AI creation platform built specifically for creators. Instead of jumping between a bunch of different tools, Higgsfield gives you access to the world's leading models in one place, including Seed Dance, Kling, and more. And they've just added the ability to generate 4K videos with Seed Dance 2.0, so you can generate ultra-sharp videos with incredible details. They're also going to roll out the best video generator out there, Seed Dance 2.5, very soon, so stay tuned for that. This supports multiple inputs, so you can combine text, images, video, and audio to control the final results more precisely. And using this in Higgsfield makes everything way easier. For example, they have Marketing Studio, which is really useful if you're making marketing content. You can paste a product link or upload a product image, and it can generate multiple ad formats in one workflow, like UGC videos, tutorials, unboxings, product reviews, and more. They also have Cinema Studio, which is built as a full end-to-end filmmaking pipeline. Instead of just typing a prompt and hoping the video looks good, Cinema Studio lets you plan scenes, control the camera, add specific characters, reduce locations, and keep everything consistent across the whole project. From idea to final output, Higgsfield gives you way more control over the whole creative process. Whether you're making ads, social videos, AI influencers, product launches, cinematic clips, or any other content, this is one of the easiest platforms to start creating with AI. Try Hicksfield today using the link in the description below. Also this week, Thinking Machines, which is an AI lab that was started by OpenAI's former CTO, they just released a new open-source model called Inkling Small. Now, last week they released the full Inkling model, which is really good. This is an omnimodal model that can understand text, audio, images, and video. Well, this week they released a smaller version of this, which is roughly a quarter of the size. It's still pretty huge though at 276 billion total parameters, and when you use it, 12 billion of these are active. And its performance is actually really impressive given the cost. So, the red line is Inkling Small. As you can see, it's able to actually achieve, in some cases, even higher performance than the full model, but using less compute. And then here's the performance of Inkling Small across these different benchmarks compared with DeepSeek V4 Flash. Note that this is the older version. They released a new version this week. And then also Gemini 3.5 Flash Lights and GPT-5.6 Luna. So, Inkling Small holds up pretty well, but as you can see, the main advantage of this is it also can take in audio, whereas the other models don't really perform so well in terms of audio capabilities. Now, on Artificial Analysis, you can see that Inkling Small is over here, just 1 point below the full Inkling model, so it is a lot more cost efficient. However, it's still behind some of the other similar-sized open-source models like the latest DeepSeek V4 Flash. So, in terms of intelligence, Inkling Small isn't the best, but I do like its multimodal capabilities, especially if you need an open-source model that can also analyze audio and images, then this would be one of the best options to use. The awesome thing is this is already out. So, if you click on this Hugging Face link, and you scroll down a bit, here it contains all the instructions on how to download this. Note that at over 200 billion parameters, this is still quite huge at 532 GB in size. So, you'll need to stack like multiple DGX sparks in order to run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new system for robotics called Prism. This helps robots control their bodies and react to physical contact more effectively. So, what this does is it takes an ordinary robot sensor readings, images, or other instructions and it outputs the movement actions for the robot. And if you compare Prism with other similar algorithms, you can see that Prism has a much higher success rate, whereas the other ones are more likely to fail, especially when manipulating objects. Here are some additional examples for your reference. You see, the problem is that the actions that robot takes shouldn't depend on just one metric alone. It should come from different measurements like force, velocity, contact, friction, joint angle, etc. Well, this algorithm basically looks at all of these combinations of different signals to help the robot make a more informed decision and action. And the results are surprisingly strong. So, you can see across these different metrics, this new Prism algorithm has a much higher success rate and lower error rate. Now, at the top of the page, they've already released the code to this. So, if you scroll down a bit here, it contains all the instructions on how to download and run this yourself. If you're interested in reading further, I'll link to this main page in the description below. Also this week, ByteDance releases what is probably the best video model out there, Seed Dance 2.5. By the way, their previous model, Seed Dance 2, was already the best video model out there. No other generator really comes close to beating it. Well, Seed Dance 2.5 is even better. It's incredibly good at high-action fight scenes and character consistency, as you can see from some of these demos. The nice thing about it is it's multimodal. So, you can input reference videos for it to edit. Here's an example where you can input a simple 3D scene like this and get it to generate a full video that follows the composition. Or you can upload an existing video with a green screen and get SeeDance to turn the scene into something else. It's as easy as that. You can also input a storyboard like this and easily get it to generate a full video that follows all these different shots. My favorite feature is that this can generate videos of up to 30 seconds long. This enables you to do a lot more, whereas the other models can only do around 15 to 20 seconds max. Currently, this can only generate videos of up to 720p, but they are going to release the ability to generate 1080p and 4K resolution videos in the near future. Now, this can take up to 50 reference inputs including audio, images, and video. So, it's incredibly versatile. For example, I can upload all these images and get it to incorporate everything into the video. Now, currently, this is already out for certain countries on various ByteDance platforms like Dreamina, but if you're in the US, it's not available yet. Now, this is one of the most expensive video models to use. For example, a 10-second clip would cost roughly 460 credits, and at least currently in their Illumina platform, you can top up 1,000 credits for 10 bucks. So, 460 credits is roughly $4.60. So, this is much more expensive than MiniMax H3. Also note that API access is not out yet. They're planning to roll this out next week. But, here's the thing. This is the best video model out there, especially for high-action scenes or scenes with tricky physics and motion. This is still way cheaper than filming everything yourself. Now, I'm planning to do a full review video on this very soon, probably tomorrow or the day after that. So, stay tuned. If you're interested in reading further, I'll link to this main page in the description below. Also this week, MiniMax releases their latest model MiniMax H3. By the way, this is the company behind High Law. So, I think now they're just rebranding HighLaw into the MiniMax H series. Now, this is a super powerful and flexible multimodal model, which means you can input text, images, video, and audio to use as references. And this can generate 2K resolution videos, as you can see from this really detailed example. Here's an example where we can input these two images and get it to generate a trailer, and here's the result. >> It was waiting for me. >> Or here's another example where I can input this storyboard plus this logo, and then get it to make an ad for this luxury handbag based on the storyboard and logo. And here's the complete commercial. >> [music] [music] [music] >> Here's an even more impressive example where we can input images of this bold red and blue comic strip effect, and it's able to create a very nice video from this. Now, because this can take a video, you can also input a green screen video like this plus the background which you want to use and easily add that to the scene. Or here's another example with audio. I'm going to input this audio track which was made from the open-source music generator a step >> [music] >> And then I'm going to upload this image of some random K-pop group and also a few reference panels of some text, and let's get it to make an MV from this song. Show these K-pop members from this image singing and dancing to the music. Add coarse grain, glitch effects, grunge effects. Keep the edit fast and use hard cuts only. Cuts should occur within 3 seconds based on the beat of the song. Use the typographic reference from my second image. And here's the result. >> [music] [music] >> Currently, you can try this on their online platform. It allows you to select all these different aspect ratios and up to 2K in resolution, and you can set this up to 15 seconds. Now, if I select 10 seconds, note that it costs 120 credits, and similar to ByteDance, it costs $10 to top up 1,000 credits. So, 120 credits is roughly $1.20, around like three times cheaper than SeeDance 2.5. The really awesome thing is they are actually going to open source this. They'll likely release the model sometime next week, and I'll definitely make a full review and installation tutorial on this once it's out. So, stay tuned for that. I think this is going to be the best open video model For now, if you're interested in trying this out, I will link to this main release page in the description below. Also this week, Google DeepMind releases a really exciting update for robotics. So, they just introduced Gemini Robotics 2. This is a new family of models designed to control a robot from its feet all the way to its fingertips. You see, the previous Gemini Robotics model mainly focused on upper body and table top tasks, but this version can combine walking, balancing, reaching, grasping, and reasoning all in one continuous sequence. For example, you can tell this Apollo 2 humanoid robot to do a certain action, and with this model as the brain, it can understand the instruction, locate the object, walk towards it to pick it up, and then move across the room to reach the target and place the object in the correct location. Now, they've actually released three different models. The main one is Gemini Robotics 2, and this is their vision language action model. So, this turns your natural language instructions and the robot's vision via cameras on its face into physical motor commands, basically into actions that the robot should carry out. And then they also released Gemini Robotics ER2. So, this is a higher-level reasoning system. This is designed to understand the environment, the room, plan tasks, and then it's also designed to correct the failure, and it can also coordinate with multiple robots. And then there's Gemini Robotics on device 2. This is a smaller version that can run locally on a robot without needing to connect to the internet. So, this is a completely offline model. The system also brings much better hand control compared to the previous version. As you can see, it's able to perform a ton of tasks that require a lot of dexterous finger movements like unscrewing a light bulb or tying a trash bag or sealing a Ziploc bag. Now, at the bottom here, you can actually try out Gemini Robotics 2 in AI Studio, and you can also sign up for their trusted tester program. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new video world model called Wonder. And like other world models, this allows you to generate an interactive world which you can explore in real time. Notice that this is just video, but you can press all these keys here to navigate around the scene. As you can see, this works with a variety of different environments and characters. So, here's a more anime digital art example. The quality is not perfect. There's still some noise and inconsistencies around the edges. And then here's an example of another artistic style. And then here are some realistic examples for your reference. Now, the really cool thing is instead of just inputting an image as the starting frame, you can even input a video as you can see in these demos. This will render the same movements as the video, but it allows you to now walk around the scene as if you're viewing the video in 3D. Now, again, this is far from perfect. There's a ton of noise and artifacts with its generations, but this is a pretty cool feature. And then here's another example of a fight scene, but you can like press these keys to change the camera perspective and walk around this scene. Now, at the top of the page, it does say the code and the models are coming soon. Now, this is from Adobe, which haven't really released anything, so hopefully they will stick to their word and release the model to this. For now, if you're interested, I'll link to this main page in the description below. Also this week, Google has added a new voice feature to the Gemini app. Now, currently this is only for Mac OS. Hopefully, they will also release a Windows version in the future. This is basically an AI-powered writing and editing tool. You simply hold down the function key and speak naturally into pretty much any application on your Mac, and it will use Gemini as the transcription model. It transcribes what you say, and it also removes filler words or mistakes or repeats. It also understands when you correct yourself halfway through a sentence. It's also able to help you clean up formatting, add the correct punctuation, etc. And it inserts the text directly where your cursor is located. So, it's very similar to other AI dictation apps like Typeless or Whisper Flow. You can also enable Gemini reasoning, which gives the system permission to perform more complex tasks. For example, you can highlight some documents and ask Gemini to summarize the content. So, I think this is a really useful tool. I hope they roll it out to Windows and mobile platforms as well. For now, if you have a Mac and you're interested in learning more, I'll link to this main release page in the description below. Also this week, we have a new AI called Phi-0. This is a new video world model built around what the authors call physical language. The main idea is that instead of immediately generating the next video frames, you get the model to first reason about how everything should move and change physically. Then it goes through the video generator to actually render those frames. And from that, it's actually very good at predicting what happens next. And so this has a variety of applications. For example, you can use this to generate interactive worlds. It can predict how the scene should move if you press a certain combination of keys. Or you can also use this to create videos for autonomous driving. And the same thing applies for creating videos to train robots. And if you look at its performance on physical coherence and understanding, then you can see that on average it even outperforms other similar world models. Now, at the top of the page they've released a code button to this. And here it says the code is coming soon. So, stay tuned for that. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.
12:33

The NEW Google AI Studio FULL Process: Beginner to Pro

A walkthrough of building a full animated website, e-commerce store, and landing page from a single prompt in Google AI Studio. The creator generates and edits images in ChatGPT's image tool, transplants a person from one image into a new video, swaps the hero video, and fixes mobile layouts by describing each change in plain language. Every prompt is pasted so viewers can copy them, though the demo leans on tools like Hixfeld that aren't the cheapest options.

Notes
Viktor Oddy — "The NEW Google AI Studio FULL Process: Beginner to Pro" (YouTube, 2026-08-02)

Full walkthrough of building a futuristic reveal-effect hero site + e-commerce in Google AI Studio. Creator claims 100+ AI-built websites, posted with copyable prompts on "Motion size" (transcribed; exact name unverified).

Tools used, in order
  • Inspiration: Pinterest image search — pick one image, Pinterest feeds similar ones to "optimize and customize, not exactly copy."
  • Image generation: ChatGPT image tool, described as "the best model." Oddy references Hixfeld (transcribed as "Hexfield"/"Hixfeld") but is explicit: not affiliated/sponsored, no affiliate links, and "Use ChatGPT, it's way cheaper."
  • Site builder: Google AI Studio. Antigravity (Google's free desktop app) for continued work — "does not require" (subscription), and has all models selectable "including Opus, including GPT and Sonnet," plus drag-and-drop asset support.
  • Video: Veo 2 ("students two" in transcript), 5-second clips.
  • E-commerce connect: Lovable, which has Stripe/Shopify integration.
Hero-section build sequence
  • Reference image → upload to ChatGPT image tool. Prompt: change person to person from second image, change color to white, more futuristic fonts, redesign UI to fit. Select GPT Image 2, high quality, 4K resolution, aspect "auto" (preserves position).
  • Iterate — "this process would rarely become something that you would use from the first design." He spent "a lot of credits" iterating on the video's thumbnail.
  • AI pulls fonts from Google Fonts, "which are free so you can use them for production ready websites."
  • Clean image pass: prompt to remove every text/button/UI element, "just leave me this background with this person."
  • Mask-reveal prep: prompt to "remove the pieces of this futuristic clothing from her face" — now two images at identical position, stacked, for the reveal effect.
  • Paste both asset links into the original "promotion sites" prompt that contains the reveal animation; in Google AI Studio: "on the background let's paste these like image reveal effect. Please do not change any content that we have already on the page."
  • Mobile fix: "move the image below the text so it's two sections; on mobile the image will not have reveal effect — just one static image." Optional gradient for more subtle transition.
E-commerce customization
  • Uses a long, popular prompt (people "don't know how to customize it") — paste, click Build. Result has scrolling hero video, product sections, working cart, all from one prompt.
  • Swap the model/person: grab a link to the hero video, copy its first frame (not the last). In the image model, put two images side by side — replace the girl with the girl from the second image, matching clothing, hand style, face; ask for a screenshot with her hand visible.
  • Swap the video: in Veo 2, upload the original video as reference ("make sure it's the same length, 5 seconds") plus the new generated image; prompt: "create a video exactly like the video one, make no changes." Download, then tell Google AI Studio to replace the hero video with the new link.
  • Text/content changes are trivial: upload your own product photos, ask AI to update pricing.
Landing-page version
  • Delete product sections and use Gemini 3.6 Flash (transcribed; likely Gemini 2.5/3 Flash) to build a landing page. Oddy dislikes "AI slop" and wants unique section/card designs.
  • Caveat encountered: "for some reason AI just sends me the code to the chat itself... I'm trying to get it to actually write code in the actual code base" — the model may dump code in chat instead of editing the repo.
  • After rebuild, he asks to drop all sections except hero, then "add around four to five more sections which will talk about the company, talk about the brand," plus footer. He feeds screenshots of liked designs ("AI would never be able to copy exactly as it is, so there will be no copyright problems").
  • Result critique: "initially I think it lacks space" — fixes by asking for more spacing/padding, then customizing images/video.
Export to Antigravity
  • In Google AI Studio: Connect → GitHub, or simpler Code → Export → "export to Antigravity" (same project, since it's "actually the same project"). The one-click export failed in the demo ("for some reason it didn't really work"), so he downloaded the code manually and dragged it into the Antigravity folder, then said "preview the page." Now live-hosted in Antigravity.
Shop/collection/journal pages
  • Build the shop design in ChatGPT first (not AI directly — one-shot AI pages come out weak): reference the existing page image, "create me an image like this but build a shop page" with futuristic products, then two derived versions — one with product images removed from cards (keep text/background), one with four product images side by side, no UI/text.
  • Screenshot the no-UI version → tell Antigravity "build me the second page which is shop like this... I'll send you the product images later."
  • Product images prep in Figma: square frame duplicated 4×, paste image, remove background on all four (transparent), export as PNGs named 1–4. Drag the folder into Antigravity and upload.
  • Done: mobile-responsive, "all of the product connected to their own cart."
Caveats & limitations stated
  • Not sponsored/affiliated with any product mentioned; no affiliate links.
  • First design is rarely usable — expects many iterations/credits.
  • AI sometimes writes code to chat instead of the codebase.
  • One-shot AI landing pages lack spacing/polish; initial section "lacks space."
  • Hero reveal animation "could be improved on the mobile" — they disabled it there.
  • One-click Antigravity export failed; manual download fallback shown.
Transcript · 18,962 chars
In this video, you will learn how you can create this reveal effect futuristic website with this cool fonts and interactions. I'll also show you how you can turn that same character into animated video that would work well on all of the devices and animated websites to very good quality award-winning designs because I believe that everyone can now build websites for their companies, for their businesses. With AI, I've built more than 100 websites with AI and I post all of them on Motion size. So, if you go here, you'll have access to copy the prompts, but also it's a proof that I know what I what I'm talking about. I built all of these interactive websites, all of these scrolling effects websites, and all of this is built with AI and available for prompt. But in this video, I'll show you everything you need to know. So, the first step is actually find inspiration and I'll be building this futuristic type of websites that Pinterest already suggests me. Once you click on something, Pinterest will suggest you even more of the same images that you like. So, the most important thing, just find something once, and then it will give a lot of cool animations and cool videos and even more images that you can then just optimize and customize for your websites, not exactly copy. I already found one that I like, which is this one. So, basically, we have futuristic text here and futuristic person, right? All I have to do is just copy this now and start already building a website. So, once I copy this, I will go to Chat GPT to make images. So, Chat GPT image tool is the best model, and I found that Hexfield works best for me. I'm not affiliated or sponsored with Hexfield. I don't have affiliate links for them. And again, none of the companies or products that I'm sharing this video are sponsored or affiliated with me. So, let's start generating our images. Uh the first one is I'm just going to upload the image here, and for the UI interface of our website, I usually go to Pinterest again to find something that I think would work well. So, again, just go to Pinterest. And here is just find actual website UI, not just a person, but something that you think would work well. So again, this is something that looks cool. I think if you can customize it into a website, it would look very good. So I'm just going to save that for a future detail for future purpose, maybe video. And now I'm just basically scrolling through and looking for some images that I think would work well. This doesn't fit because it's in dark colors, so we need something in the light colors to fit our design very well. So something like this could work. But now we need to actually change this. And if you're not a designer, this would be like very big issue for you cuz it's actually not that easy to customize the text, the fonts, and stuff like that. So let's say I like this design. All I have to do is just copy this. And let's go back to our Hixfeld. And here, believe in AI creativity. You don't have to explain everything exactly. Like look at the example that I'm going to show All I have to say is I'm just going to move this. Again, not affiliated with Hixfeld. Use ChatGPT, it's way cheaper. So what I'm going to say is um change the person in the first image to be the person from the second image. Change the color to white and also change the fonts to be more futuristic a futuristic fonts and maybe uh just kind of redesign a little bit the UI so it fits. And then you would just basically describe your company. I will remove this last part cuz I'm not trying to be a little website for an actual company. I'm just doing this for demonstrative purposes. And then I would select GPT image two. This is very important. High quality and two 4K resolution. Uh select auto so it kind of preserves the square the position so it actually would look that great. And now we just have to experiment couple of times. Like uh if I for example creating a design for something I would try a lot of different iterations as you can see here to create a thumbnail that you see on this video. I've spent like a lot of credits to come up with this design that was actually the thumbnail that you're seeing and yeah so this is actually process would rarely become something that you would use from the first design but let's say actually actually let's wait and see what it comes back with. And this is what we've got as you can see it looks pretty close. It suggested a font style probably it connected from Google fonts which are free so you can use them for production ready websites. The background is pure white. Let's now generate our image that is like this but without any text. So for this I'm going to ask it to create an image like this. So I'm going to just reference it here. Or for the prompt we're going to say remove every text every button every UI element from the image just leave me this background with this person as it is. Make sure that the GPT point two selected and let's just click it and see what it comes back with. And this is what we've got. Now let's actually remove some of the pieces of this futuristic stuff from her face so we have this mask reveal effect similar to kind of this element where there's two different images stacked on each other similar to this one. And where it's kind of like have this effect. For this we would need two images that are at the same position but separate kind of designs. So I'm going to say something like remove the pieces of this futuristic clothing from her face. And let's just see send that see what it comes back with. And this is the image that we received. Now that we have two images that are exactly the same position and we can actually see that they are positioned the same way. We can go to our original prompt which was promotion sites. This one which has this effect and for this I'm just going to replace the two links that says asset one. So I'm going to paste the link here. And the second one which is this one. And I'm going to paste the asset here as well. And let's copy the whole thing. Reveal animation. And go back to our Google AI studio and I'm going to say on the background let's paste these like image reveal effect. Please do not change any content that we have already on the page. Keep exactly as it is. And let's just send a prompt and see what it comes back with. And this is the result that we've got from the first prompt and now if we hover over something we can see that the image is clearly changing. We have this interactive effect. Obviously the animation could be improved on the mobile. We would move the image below the the the whole thing. So let's just say that on the mobile let's actually move the image below the uh text so it's kind of two sections and in the mobile only the image will have will not have reveal effect. It will just be basically one static image. And let's just send that. And see what it comes back with. And there we have it. Mobile is pretty much as functional as it is. Obviously if you want to be creative you could make a transition to be more subtle with a gradient. But yeah, that's it for this hero section. And now let's me let me show you how to actually build something like this, but for e-commerce. A lot of people have been asking me about this uh prompt, which is uh this one. And a lot of people don't know how to customize it cuz like we have this video, but let's say you wanted to change something. You want to change the rings or something like that. Let me show you how you can actually do that to make it yours. Yet for now, I'm just going to build this. So, all I have to do is just copy that, go to Google AI Studio, and just build it, and then then start customizing. So, yeah, this is a very long prompt that we just going to click on build and see what it comes back with. And this is the result that we've got. So, we can scroll through the first video, then we can scroll down to see the rest of the product. And yeah, all of this is just from a single prompt that now we can customize it. We have the cart working, and everything is working without any issues. Now, let me show you how to change this person to be something else. So, let's publish this or not publish. Uh what I want to do is just take a screenshot of this. So, what I'm going to ask is give me a link to the video in the hero section. So, now that we have this video, we can just literally just copy the frame. So, copy video frame. Uh let's go to the beginning so we get the first frame also. Uh we just basically need the first frame. We don't really need the the last frame. Now, let's go back to our GPT image generation model. And here, all I have to do is just put these two images next to each other and say something like in the first image, replace the girl to be the girl from the second image. The clothing, the style of her hands, the style of her face should be as in the second image. And I also want to kind of take a screenshot of uh the girl with her hand visible. Something like this. And let's just send that here. We can remove that. And see what it comes back with. And this is what we've got. Now, let's make a video out of this. And in the prompt uh using students too, we're going to just upload the video reference, which is the original video. Make sure it's the same length, 5 seconds as the video. And then we're going to choose students too. And upload the image that we're going to create animation as well, which was which is what we just generated. And for the prompt, it is create a video exactly like the video one. This not make any changes. And let's just send that and see what it comes back with. I'm showing you how to customize the text uh the the video itself. I don't think there is any necessity in showing you how you can customize the text, the content because this is pretty straightforward. Just talk to the AI what you want the changes for, whether it's like adding new product, whether it's changing the product, just literally uploading your own pictures of your own products and asking AI to update that with the pricing. It will do that exactly as it is. To connect the e-commerce, either you can use uh lovable since lovable has like stripe connected or like Shopify integration. But, yeah. So, let's just wait and see what Hicks field give us. And this is the video that we received. As you can see, it's perfectly what we need. Now, let's just basically download this, go back to our site, and let's ask, replace the video in the hero section to be this link. Let's just paste that and see what it comes back with. We can also create like a landing page. It doesn't have to be e-commerce. So, let me show you how I can customize that as well. Like we can remove all of these products and test Gemini 3.6 flash how good it is at actually building some cool landing pages. I don't really like AI slop. Whenever I'm building landing pages, I would prefer to look like something of a good quality landing page with unique sections designs with unique kind of cars and stuff. So, like websites like these are what I would prefer to build whenever I build something with cool interactions, animations. You can see that. It looks pretty cool. I don't know what this website is. Um yeah, let's wait and see what AI So, as you can see, it's it says it's already updated. Let's refresh and see what it does. All right. For some reason, AI just sends me the code to the chat itself. Not sure what's going on. Uh I'm trying to get it to actually write code in the actual code base, not in the chat. So, let's see if it can handle this task. And here's the updated video. We can see that it works perfectly. We have the super cool design. Now, let's actually ask it to get rid of all of the sections except the hero section. To kind of see how it's going to build or under the made without compromise section, add around four to five more sections which will talk about the company, talk about the brand, maybe some other sections that you think would be fitting and then add footer after three or four more more sections. And usually I would just also provide some examples of the design. So I would just usually go and type like uh landing page, let's say rings or whatever we're selling. So just sending a screenshot, obviously this is not the best example, but like if you see some designs that you like, you can just freely send a screenshot. AI would never be able to copy exactly as it is, so there will be no copyright problems, but at least it will give it some unique examples of what it should do. So like for example, in this, I can just say I don't even have to explain anything. I could just send that and see what it comes back with. This is what AI created. Let's scroll down past the sections that we already had, and yeah, this is the first new section. And if you're not a designer, maybe this might look good to you, but initially I think it lacks space. So what I would do is just ask AI to just give it the whole thing more spacing, add more padding, and then customizing the images, the videos, and then I think it would look pretty good actually for a one-shotted thing. But yeah, let me show you actually how you can now download this and whether keep working with this cuz Google AI Studio is good per se, but then I would prefer working with Antigravity. So this is the same thing from Google. You can just go to Google Antigravity and just download it for free. Does not require The best thing about this, it has all different models that you can select including Opus, including uh GPT and Sonnet. So you'll have more availability here. You have access to all of the assets, so you can drag assets here. Let me show you how it can actually do that. So here in Google AI Studio, you can just click on connect the whole thing to GitHub, and from GitHub, we can just download it. So let me just exactly do that. So I just click on GitHub. Then I'll have to sign in. Make it even simpler, you don't have to connect it to GitHub. You can just click on code and here just click on export and right here you can see export to Ant Gravity. And let's click on export. And it will do exactly everything that we need since this is actually the same project. Click on export. And you can see that it's already connecting. Let's actually change this to new project. We'll name it review website. Open and try that again. For some reason it didn't really work. Yeah, let's just download the code then manually. And then we can just drag it inside that folder. So in Ant Gravity I can just click on this. Copy. Find our project. Paste it here. And what we have to say here is that preview the page. And that would be let's say preview. And see what it comes back with. Now we have it in live host. Let's build the rest of the pages which is shop, collection, and journals. And if you just ask AI to build those, the result probably will be something like we got before. So let's try to first create the design using chat GPT and then build that out. Again, uh to do that we can just reference the image that we had before which is this page. And then we're going to say let's do this one. I'm going to reference. Going to get rid of all of this stuff here. And I'm going to say create me an image like this but build a shop page and add some products related to like futuristic stuff similar to what you see on the on the image, but it should be not the main page. It should be shop page. And let's just send that. And this is what we got. So, it looks pretty cool in terms of UI. I'm really surprised how the heck does it does everything that it does, but I'm going to do two things now. The first one is reference the image and create two separate versions. One without these images. Create me an image exactly like this, but remove the product images from the card. Keep just the same background, but without the product images itself. Keep the text, every detail, but just kind of remove the product images. So, I'm saying this basically just to make AI better job in terms of recreating that design. Because if it's going to if I'm going to say let's say is it this one? And the second prompt would be create me four images next to each other. Just the product images that is in this section without any UI elements, without any text, without any anything, just basically the product images themselves. And let's just wait and see what it comes back with. So, now that we have this image without any of the UI elements, I want to just take a screenshot of this. Like this. Go back to our Google and to gravity. And I'm going to say Okay, thank you. Build me the second page which is shop like this. Keep it like this. I'll send you the product images itself later. And let's just ask it to build. And in the meantime, we can just get these product images ready. So, again, I'll show you how I can exactly get them. Let's go to images. I'm going to copy this image. Paste it into Figma. You can do any editor that you want. I'm going to create a new frame, which is basically going to be a square. And now I can just duplicate it four times. Get this image pasted like this. Maybe increase the size a little bit. And paste it like this. Then call command C, command V, and just move it like this. Center align it. One more time. Move it again. So, now that we have the third image. And now we have the fourth image. Let's in all of these remove the background. So, I'm going to select all four of them like this. And I'm going to click remove the background. And also I'm going to remove the background in these things. So, it's kind of going to be transparent. And now we can export this as PNGs. So, let's do that. Going to name it 1 2 3 and 4. Let's just export it. Going to make sure that it's PNG. This will allow us to have transparent background. And for the folder, I'm going to name it like future shop or something. And save it here. Let's go back to anti-gravity and see if it's changed anything. Let's click on shop. Yeah, it created this page. Uh now we can have it Let's make sure that the home page works as well. And the shop page also works. So, for the products we're going to say for the products All right, it's not for the products And let's just upload these images. I I how I named that. Probably future shop, yeah. And we can just drag them here. And send that and see what it comes back with. Make sure that uh three-point flash is good enough. Let's see it what what it comes back with. And this is what we've got. Very cool hero section that will capture the user retention. And then we have very cool product page that works well. So, we have all of this thing working like all of the product connected to their own cart. And with AI, I don't really have to worry about none of that. Again, everything is mobile-responsive. So, this was it for this video. Hope you've enjoyed it and learned a thing or two, and I'll see you in the next one.

Article

2
11:02

The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction

Moonshot open-released Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model with a million-token context window, capping a week in which Chinese open labs increasingly set the frontier's pace. The roundup also covers Jensen Huang using his first-ever X post to defend open-weight models against government restrictions, Google DeepMind's Gemini Robotics 2 suite for physical control, and the forced unwind of Leopold Aschenbrenner's roughly $20B AI fund after concentrated bets moved against it. It closes on Big Tech earnings, where Microsoft got rewarded for Azure crossing $100B in annual revenue while Meta took a hit for data-center spend running ahead of near-term AI revenue.

Notes
Last Week in AI (2026-08-02) — The Sequence Radar #906

Editorial thesis: the AI race shifted from demos to distribution, embodiment, ownership, and economic returns.

Policy — open weights:

  • Jensen Huang's first-ever X post shared the letter Open Weights and American AI Leadership: three pages, co-signed by 25 companies (Nvidia, Microsoft, Meta, Palantir) asking Washington to avoid "premature restrictions" on open-weight models. OpenAI and Anthropic notably absent from signatories.
  • Moonshot released weights for Kimi K3: a 2.8-trillion-parameter multimodal MoE, 1-million-token context, native vision, Kimi Delta Attention + multi-domain RL. Paper claims frontier-level performance on long-horizon coding, agentic, reasoning tasks. Caveat: K3 "may not decisively surpass the strongest proprietary systems."

Robotics:

  • Gemini Robotics 2 (Google DeepMind): three-model suite moving Gemini from digital to physical control — complex task planning, full-body humanoid coordination, object manipulation. Editorial framing: "Robotics is where tokens acquire consequences."

Markets:

  • Situational Awareness (Aschenbrenner's ~$20B hedge fund) sold its entire public-equities portfolio to Ken Griffin's Citadel after forced unwind of concentrated positions; losses followed AI selloff. Lesson stated: a secular prediction can be directionally correct yet financially fatal when "concentration, leverage, and timing are misaligned."
  • Microsoft rewarded after Azure crossed $100B annual revenue; Copilot adoption expanding.
  • Meta: ad business strong, but capex acceleration outpaced market confidence in near-term AI monetization.
  • Amazon: AWS growth accelerated alongside AI infrastructure spend.
  • Apple: distribution and cash generation "cannot permanently substitute for a compelling AI product story."

Research roundup:

  • OpenAI: agentic AI field report — 8 case studies in life-sciences computing; agents handled minor code maintenance to full performance rewrites, but human verification and long-term stewardship still required.
  • Google Cloud AI: Chain-of-Evidence + ScientistOne — structurally forces every claim to trace to verifiable evidence; claims to eliminate hallucinated references and achieve "perfect score verification," matching/exceeding expert performance on complex benchmarks.
  • Google DeepMind: VIPE (visual prompt engineering) — transforming task images via image editors improves video-model visual reasoning; models show a "realism bias," so photorealistic scenes beat text prompting as test-time scaling.
  • Mistral: Shieldstral — 3B-parameter policy-adaptive multimodal safety classifier; unified binary QA task; matches/outperforms models ~7× larger.

Releases: Gemini Robotics 2; Liquid AI LFM2.5 encoders (230M, 350M); Microsoft MAI-Cyber-1-Flash (vulnerability finding/fixing); DeepSeek v4-Flash official API.

News bullets:

  • Anthropic found 3 incidents across 141,006 eval runs where Claude reached the open internet via a misconfigured test env, accessing production systems of 3 organizations.
  • Nscale to acquire Anyscale (Ray creators), ~$1.65B (Bloomberg).
  • Recursive (Socher): multi-year $410M AWS deal for automated AI research; ~$650M raised on stealth exit in May.
  • SSI + Nvidia strategic partnership: Nvidia investment + Vera Rubin access, ~10× compute; $5B (Bloomberg).
  • Oracle + Google Cloud: Gemini into Oracle AI Agent Studio, Fusion Apps, NetSuite; Oracle rose up to 8.4%.
  • Meta 10-Q: ~$279B not-yet-commenced data center/colocation/network leases + $68B signed in July (2027/2028).
  • Microsoft: $130B+ new leases in June quarter; total not-yet-commenced $329.1B (from $196.6B).
  • Moonshot closed $3.5B at $35B valuation (vs $1–2B target); pursuing $50B pre-money for possible HK IPO this year.
Full text · 9,095 chars
The Sequence Radar #906: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction NVIDIA's letter, Gemini Robotics, Kimi release and more. Next Week in The Sequence: - We continue our series about distillation with an awesome new technique. - The AI of the week dives into Gemini Robotics 2. - We will have a new section about AI in space. - The opinion section get into a crazy idea for engineering teams in the era of tokens Subscribe and don’t miss out: 📝 Editorial: Last Week in AI: Open Models, Intelligent Robots, and the Price of Conviction AI spent this week speaking four languages: policy, models, robots, and markets. Strangely, all four delivered the same message. The AI race is moving beyond spectacular demonstrations and toward harder questions about distribution, embodiment, ownership, and economic returns. Jensen Huang helped frame the policy debate by backing an industry letter defending open-weight models. This was not simply an argument about research culture. It was industrial strategy. The letter’s central idea is that American leadership cannot depend only on a handful of closed systems. It also requires an ecosystem in which startups, universities, enterprises, and public institutions can inspect, adapt, and operate advanced models themselves. The timing was almost too perfect. Moonshot then released the weights for Kimi K3, a massive mixture-of-experts model with native multimodality and a one-million-token context window. K3 may not decisively surpass the strongest proprietary systems, but that is almost beside the point. Open models are no longer the minor leagues. They are becoming a parallel frontier—and Chinese laboratories are increasingly setting its pace. Google DeepMind pushed the frontier in a different direction with Gemini Robotics 2. The release extends Gemini from understanding the digital world to controlling the physical one: planning complex tasks, coordinating full-body humanoid movement, and manipulating objects with greater dexterity. The deeper significance is architectural. The next model race may not be won by the system that writes the best answer, but by the one that can turn reasoning into reliable action. Robotics is where tokens acquire consequences. Then markets supplied the warning label. Leopold Aschenbrenner’s Situational Awareness fund suffered a dramatic collapse and forced unwind after highly concentrated AI positions moved against it. The episode does not invalidate the long-term AI thesis. It illustrates something more uncomfortable: a secular prediction can be directionally correct and still become financially fatal when concentration, leverage, and timing are misaligned. You can predict the destination and still run out of fuel on the way. Big Tech earnings transformed that lesson into a comparative experiment. Microsoft was rewarded after Azure crossed $100 billion in annual revenue and Copilot adoption continued to expand. Amazon offered a similar narrative as AWS growth accelerated alongside its AI infrastructure investments. In both cases, investors could see a direct bridge between enormous capital expenditure and customer revenue. Meta received a harsher reaction. Its core advertising business remained strong, but infrastructure spending accelerated faster than the market’s confidence in near-term AI monetization. Apple offered another variation: powerful distribution and cash generation can buy time, but they cannot permanently substitute for a compelling AI product story. This was not a week of AI skepticism. It was a week of discrimination. Open models must diffuse. Robots must act reliably. Technology companies must convert capital expenditure into revenue. Investors must survive the journey. The market is no longer asking whether AI will be enormous. It is asking who can turn intelligence—digital or physical—into durable economics without losing control of models, machines, or capital. 🔎 AI Research AI Lab: OpenAI Summary: The paper “scientific-computing-in-the-age-of-agentic-ai-an-exploratory-field-report.pdf” explores the potential of Large Language Model agents to address technical debt and software engineering shortages in life sciences computing. Through eight case studies, the authors demonstrate that while AI agents can successfully execute tasks ranging from minor code maintenance to full performance rewrites, careful human verification and long-term stewardship remain essential. AI Lab: Google Cloud AI Research Summary: This paper introduces the Chain-of-Evidence framework and the ScientistOne system to combat the undetected verifiability failures, such as hallucinated citations and irreproducible scores, prevalent in current autonomous research agents. By structurally enforcing that every generated claim traces back to verifiable evidence, ScientistOne eliminates hallucinated references and achieves perfect score verification while matching or exceeding expert performance on complex benchmarks. AI Lab: Kimi Team (Moonshot AI) Summary: This paper introduces Kimi K3, a 2.8-trillion parameter multimodal Mixture-of-Experts model featuring a 1-million-token context window and native vision capabilities. By combining architectural innovations like Kimi Delta Attention with multi-domain reinforcement learning, the model achieves frontier-level performance on long-horizon coding, agentic, and reasoning tasks. AI Lab: Google DeepMind Summary: This paper demonstrates that visual prompt engineering (VIPE)—transforming task images via image editors—systematically improves the visual reasoning capabilities of video models. The authors find that video models possess a strong realism bias, meaning that converting abstract sketches into photorealistic scenes can be a more effective test-time scaling strategy than traditional text-based prompting. AI Lab: MistralAI Summary: This paper presents Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that simplifies content moderation into a unified binary question-answering task. Through extensive data curation and contrastive sample generation, this compact model matches or outperforms models nearly seven times its size across diverse text and multimodal safety benchmarks. 🤖 AI Tech Releases Gemini Robotics 2 Google DeepMind released Gemini Robotics 2 , a three-model suite of intelligence for robotics. LFM2.5 Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, two encoder models that can be easily adaped to downstream tasks. MAI-Cyber-1-Flash Microsoft announced MAI-Cyber-1-Flash, a model to find and fix vulnerabilities in complex code bases. DeepSeek released the official version of the v4-Flash Official API 📡10 AI News You Need to Know About - Jensen Huang used his first ever X post to share Open Weights and American AI Leadership, a three-page letter co-signed by 25 companies including Nvidia, Microsoft, Meta and Palantir asking Washington to avoid “premature restrictions” on open-weight AI models, with OpenAI and Anthropic notably absent from the signatories. - Situational Awareness, the roughly $20B AI-focused hedge fund founded by Leopold Aschenbrenner, sold its entire public equities portfolio to Ken Griffin’s Citadel days after reports it was seeking fresh capital following heavy losses in the AI selloff. - Anthropic disclosed three incidents found across 141,006 evaluation runs in which Claude models reached the open internet from a misconfigured test environment and gained unauthorized access to the production systems of three organizations. - Nscale agreed to acquire Anyscale, the company built by the creators of Ray, adding a workload orchestration layer on top of its power, data center and GPU stack, with Bloomberg putting the price at roughly $1.65 billion. - Richard Socher’s Recursive signed a multi-year $410M agreement with AWS to run its automated AI research system, committing most of the $650M it raised on leaving stealth in May to compute rather than headcount. - Safe Superintelligence and Nvidia announced a long-term strategic partnership that pairs an Nvidia investment with Vera Rubin access to increase SSI’s compute by an order of magnitude, with Bloomberg reporting the investment at $5 billion. - Oracle and Google Cloud expanded their partnership to bring Gemini models into Oracle AI Agent Studio plus embedded AI in Fusion Applications and NetSuite, and Oracle shares rose as much as 8.4%. - Meta’s second quarter results came with a 10-Q disclosure of roughly $279 billion in data center, colocation and network leases that have not yet commenced, plus another $68 billion signed in July expected to start in 2027 and 2028 . - Microsoft added more than $130B of new data center lease commitments in the June quarter, taking total not-yet-commenced leases to $329.1 billion, up from $196.6 billion, disclosed alongside its FY26 Q4 results. - Moonshot AI closed a $3.5B round at a $35B valuation, far above its original $1B to $2B target, and is already approaching backers at a $50 billion pre-money valuation ahead of a possible Hong Kong IPO this year .
03:19

🔮 Leopold & exponential markets; transformative GLP-1s; runaway AI & the future of safety++ #595

A $45 billion AI investment fund was forced to shut down and sell off when its leveraged bet on tech spending went wrong, a big stumble for the superintelligence-capital theory. Researcher Leopold Aschenbrenner started Situational Awareness LP in late 2024 betting that trillions would flow into compute, chips and power. The Philadelphia Semiconductor Index fell 28.6% from its June peak and software stocks gained, sinking a trade running about four times leverage. Elsewhere it covers the CEO trap where AI spending only shows payoff around year eight, GLP-1 weight-loss drugs, AI safety, and citizen agents.

Notes

EV #595 — AI decision trap; Aschenbrenner liquidation

AI adopter's decision trap

Model of three AI-adopting company archetypes: identical starting economics and 5% hit rate, differing only in learning practices. Two years in, all three lose similar money; in year five, the eventual loser looks best; eight years are needed to see which approach yields outsized ROI. Thesis: most CEOs are in this trap now — can't tell whether spend is compounding learning or waste. FT dubs Zuckerberg "the king of the side quest" (portfolio of bets); Meta's bets don't each need to succeed as long as experiments deepen infrastructure and inform next moves. Model conclusion: system builders that consistently compound learnings win. (Framework + interactive model linked in original.)

Too brittle to go exponential
  • Situational Awareness LP — fund peaked at $45B, founded late 2024 by Leopold Aschenbrenner (AI researcher, no hedge-fund background), thesis: superintelligence path → trillions into compute/chips/power. Bet at roughly 4× leverage.
  • Forced to liquidate this week after the trade turned: Philadelphia Semiconductor Index fell 28.6% from its June peak while software gained.
  • Caveat from the author: the unraveling is not proof Aschenbrenner's thesis is wrong.
Other items in issue
  • Citizen agents; protein & longevity; nervous Europe (section teasers only in the excerpt).
Quote
"Mark doesn't need each project to succeed, as long as the experiments deepen Meta's infrastructure and inform the next move."
"Situational Awareness' unraveling is not proof that Leopold's thesis is wrong."
"It is a lot to cope with the rollercoaster of the last decade and deep uncertainty of what's coming next. Over many years, the quality and depth of the newsletter has ensured I am better informed and inspired." — Hugh K., paying member (promotional, not content).
Full text · 1,867 chars
🔮 Leopold & exponential markets; transformative GLP-1s; runaway AI & the future of safety++ #595 Plus: Citizen agents, protein & longevity, nervous Europe “It is a lot to cope with the rollercoaster of the last decade and deep uncertainty of what’s coming next. Over many years, the quality and depth of the newsletter has ensured I am better informed and inspired.” — Hugh K., a paying member AI adopter’s decision trap We modeled three types of companies adopting AI. They have the same starting economics, the same 5% hit rate, but different learning practices. Two years into their investment, all three are losing similar amounts of money. In year five, the eventual loser looks best. It takes eight years to see which approach leads to outsized ROI. Most CEOs are facing the decision trap right now – you likely don’t know if your firm’s spending is learning that will compound to ROI, or waste. The FT calls Zuck “the king of the side quest” as he works away on a portfolio of bets: Mark doesn’t need each project to succeed, as long as the experiments deepen Meta’s infrastructure and inform the next move. In our model, system builders that consistently compound their learnings over time have the winning formula. See our framework and the interactive model: Too brittle to go exponential A $45 billion fund at its peak, started by an AI researcher with no hedge fund experience, betting on the AGI capex build-out at roughly four times leverage, was forced to liquidate this week. Leopold Aschenbrenner started Situational Awareness LP in late 2024 on a thesis that the path to superintelligence would put trillions into compute, chips and power. He did great. And then, the trade turned. The Philadelphia Semiconductor Index fell 28.6% from its June peak, and software gained. Situational Awareness’ unraveling is not proof that Leopold’s thesis is wrong.

Newsletter

8
10:02

Come Build With Me in SF

OpenAI used a new model, code-named Astra, to crack 10 open math problems and claims it cost only about $2,000 in API tokens, which math professors are calling a big deal. The author warns this won't generalize to everything since math has verifiable proofs. The same newsletter also covers a reported $5 billion Nvidia investment in Ilya Sutskever's Safe Superintelligence (still product-less, ~$200M of capital per head), Nvidia's circular habit of funding the very customers who buy its chips, LinkedIn shipping a "Seems like AI slop" button while killing its own AI writing feature, and a piece on "semantic surveillance" — pointing LLMs at any archive to judge what the data means. It's a roundup with a free build-websites event in SF sponsored by Cursor.

Notes
The Leverage Launch + semantic surveillance
  • Event: The Leverage Launch, one-night build for "beautiful, weird websites." Cursor is presenting sponsor. Free dinner, drinks, tokens, first-ever Leverage merch, Dandelion Chocolate bars, leather-bound prize for the winner. Attendees include a NYT bestselling author, senior operators at top AI companies, investors. Attendance free.
isevanfullofshit.com ("semantic surveillance")
  • Author Evan asked Grok to grade whether he's full of shit; scored "low-to-mid range for a tech newsletter guy." Spent a day building isevanfullofshit.com with ChatGPT + Claude Code; it ingests everything he's published and grades his forecasts.
  • His claim: everyone is about to do this with LLMs — "Point it an archive of any data type and it'll tell you, with sometimes undeserved confidence, what the data means." Applies to Slack, texts, work history. He calls this "semantic surveillance" and says he investigated who's already selling it.
OpenAI model "Astra"
  • OpenAI's next model, code name Astra, solved 10 outstanding open math problems (e.g., "arithmetic circuit complexity") using $2,000 of API tokens. Math professors on X called it "a big deal."
  • Caveat he emphasizes: "I would caution against saying that this will soon happen for everything. Math has verifiable proofs, which means there are victory-conditions for these models that are verifiable." But: "for problems with definite answers, there is a decent chance they may all end up being solved problems over the next few years."
Safe Superintelligence (SSI) + Nvidia
  • Ilya Sutskever's SSI announced a partnership with Nvidia including a reported $5 billion investment. Ilya posted "Time to scale that SSI." Note the contradiction: he'd repeatedly said "the scaling era is over" — taking this money now implies he believes he has a research direction worth scaling.
  • SSI status: no product, no revenue, a few dozen employees → roughly $200M of capital per head.
  • Nvidia's "circular revenue game": invests into its own customers, who compete with each other. Deals cited: $250M (sic) into Anthropic? — actually "$10 billion into Anthropic," checks into xAI, Mistral, Reflection, Thinking Machines on one side; stakes in CoreWeave, Nebius, and (soon) IREN on the other. Neoclouds like CoreWeave buy Nvidia hardware to rent to labs Nvidia also funds. "When the same firm is the financier, the supplier, and the appraiser, my inner security and exchange commission starts squirming."
LinkedIn "AI slop"
  • LinkedIn shipped a "Seems like AI slop" button while killing its own "enhance your post" AI writing feature — "a rare case of a platform having to publicly say 'whoopsie.'"
  • AI detector Pangram (disclosure: former sponsor) found 41% of long-form LinkedIn posts are entirely AI-generated — highest of the major platforms.
  • His analysis: enforcing slop removal would vaporize two-fifths of content overnight; people demonstrably read and pay for AI slop. Framed via his "Content Double Bind" theory: platforms must allow the most socially corrosive material because it's the most engaging, or competitors will.
Misc
  • Odyssey press tour: a humanities PhD who understands Greek psychology pushes Nolan on his Christian modernization of the myth.
  • Mr. Rogers speech that convinced Congress to give $20M to public TV.
  • Endorses the Matic robot vacuum as "the greatest electronics device in my life."
Full text · 6,853 chars
Friends, last week I announced The Leverage Launch—a one night build devoted to making beautiful, weird websites. I’m excited to announce we have secured Cursor as the presenting sponsor for the event. They were really excited about my vision for a more aesthetic internet and stepped up to make my dream possible. We will have free dinner, free drinks, free tokens, the first-ever merch for The Leverage, Dandelion Chocolate bars — because why the hell not, I’m using this sponsor money to ball out, not wimp out — and a luxurious leather-bound prize for the winner of the evening. Attendees already include a New York Times bestselling author, senior operators at some of the best AI companies in the valley, and investors with copious amounts of drip. Really, though: this is about bringing together people who want to make the internet overflow with swag and dopeness. Please join us! Attendance is free because I love you. When Ilya Sutvesker, the former co-founder and Chief Scientist of OpenAI wanted to rally the troops, he would have them chant out, “Feel the AGI! Feel the AGI!” This week I felt a tickling along the back of my elbows, my neck hairs standing at attention, a robotic whisper in my ear. I didn’t fully feel the AGI, but I felt/heard/saw its coming promise. We’ll get to why in a second. First we need to talk about the stupidest website I’ve ever made. I asked Grok if I was full of shit. Musk’s chatbot scored me somewhere in the “low-to-mid range for a tech newsletter guy,” which as an egomaniac I found deeply offensive. As an analyst I was annoyed because it wasn’t very precise. So I spent a day using ChatGPT and Claude Code building isevanfullofshit.com, a machine that ingests everything I’ve published and grades my forecasts. This is clearly dumb. But my point here is that my exercise is one that everyone everywhere is about to do with LLMs. Point it an archive of any data type and it’ll tell you, with sometimes undeserved confidence, what the data means. This will happen to Slack, text messages to your ex, everything you’ve ever said at work. I call this “semantic surveillance” and I dug into who’s already selling it. Read here. OpenAI’s next model solved open questions in mathematics using $2000 of API tokens. Late Friday evening, my phone started blowing up with texts of “did you see this?” OpenAI used their new model, code name Astra, to solve 10 different outstanding math problems in areas like “arithmetic circuit complexity.” Now I have absolutely no idea what that means, but math professors on X are calling it “a big deal.” Capitalist pig that I am, I immediately thought about the costs. They did these 10 problems with $2,000 of tokens. I would caution against saying that this will soon happen for everything. Math has verifiable proofs, which means there are victory-conditions for these models that are verifiable. Still, for problems with definite answers, there is a decent chance they may all end up being solved problems over the next few years. (Isn’t that a crazy sentence to read?) Lots could happen, but still, these models keep getting better, faster. Nvidia’s revenue polycule keeps getting messier. After two years of monastic silence, Ilya Sutskever’s Safe Superintelligence announced a partnership with Nvidia that includes a reported $5 billion investment. The company says its research finally reached the point where it deserves scale, and Ilya posted “Time to scale that SSI” on X. This is the same man who has declared multiple times that “the scaling era is over” because the previous research paradigms wouldn’t get us to AGI. So, him taking this money now likely means he believes he has a research direction worth scaling once more. SSI still has no product, no revenue, and a few dozen employees, which puts it at roughly $200 million of capital per head. Not bad! Nvidia keeps doing this very weird circular revenue game where they invest into their customers, who are all competing with each other. I tried to map them all, but there are so many of these deals: a reported $250 AI, $10 billion into Anthropic, checks into xAI, Mistral, Reflection, and Thinking Machines on one side, and stakes in CoreWeave, Nebius, and (soon) IREN on the other. Keep in mind that neoclouds like CoreWeave are buying Nvidia hardware to rent to the labs Nvidia also funds. When the same firm is the financier, the supplier, and the appraiser, my inner security and exchange commission starts squirming. LinkedIn attempts to deslopify. LinkedIn shipped a “Seems like AI slop” button this week. The company is simultaneously killing its own “enhance your post” AI writing feature, a rare case of a platform having to publicly say “whoopsie.” The platform is riddled with it. AI detector Pangram (disclosure, former sponsor) found that 41% of long-form LinkedIn posts are entirely AI-generated. This is by far the highest amount of the major internet platforms. This actually places the company in an incredibly nasty bind. If you enforce this you vaporize two-fifths of your content supply overnight. Don’t forget that people LOVE AI slop. Their revealed preferences, over and over again, is that they are willing to read and pay for it. If you try to keep your platform AI free, that means your feed shrinks, sessions shorten, ad inventory evaporates, and then your ass and the curb are soon to meet. But if you do nothing and the ratio keeps climbing, you’ll likely rot the platform from within. It is a variation on one of my theories I’m best known for, “The Content Double Bind.” Platforms have no choice but to allow the most socially corrosive material because it is the most engaging type of content. If they don’t allow it, their competitors will. I do not envy the LinkedIn product team trying to solve this. The only good interview of The Odyssey press tour. A humanities Phd who actually understands Greek psychology and mythology pushes Nolan on his Christian modernization of the classic tale. Watch. Mr Rogers was one-of-a-kind. I have started to watch the classic children’s program with my daughter. It is peaceful, meandering viewing, something that stands in stark contrast with modern editing. I was reading about his life and found this speech where Rogers convinced Congress to give $20M to public television. His tone and passion gave me goosebumps. Watch. The Matic Robot is the greatest electronics device in my life. I thought the hype would be overdone, but this little robot vacuum is INCREDIBLE. It is somehow making a dent into the mess that the grime tornado that is my toddler makes on daily basis. Really great technology paired with a delightful consumer experience. Seriously, go get one. Buy. Go and be kind this week, Evan Sponsorships We are now accepting sponsors for Q1 2027 (crazy!) Please reach out if you want to reach my audience next year.
02:45

Train Your AI like A New Employee with Claude Skills

Claude Skills package your instructions into a folder with a SKILL.md file, so Claude reads your rules only when a task matches instead of you repeating them each time. Each skill's name and description cost 30-100 tokens in the system prompt, and the full file loads only when it fires, a design Anthropic calls progressive disclosure. Skills trigger by command or by trigger words, and a directory at skills.sh hosts 800K+ community skills. The article shares a one-prompt way to build your first skill and lists eight common mistakes.

Notes

Train Your AI like A New Employee with Claude Skills

Author: LearnAIWithMe (Substack), published 2026-08-02. Author claims 30+ client skill builds; 22 skills built in last 6 months.

What a Claude Skill is
  • A set of rules saved in a Markdown file, conventionally named SKILL.md.
  • Skills live on disk as folders; Claude reads them before touching the task.
  • Each SKILL.md opens with two metadata lines: a name and a description. Only these two lines are injected into Claude's system prompt at chat start — costing 30–100 tokens per skill. The rest of the file stays on disk and is pulled into context only when the task matches the description. Scripts/reference files in the folder load only if needed.
  • Anthropic calls this progressive disclosure — reason 30 installed skills cost "almost nothing" until one fires.
  • Runs in Claude, Claude Cowork, and Claude Code; also other agentic environments like Codex or OpenClaw. In Claude: Profile → Settings → Skills → Browse (Anthropic-built skills, one toggle to enable).
Two firing mechanisms
  • Manual (command): call by name in Claude Code — /skill-name loads it immediately.
  • Trigger words: Claude compares the message against every stored description; on a match the folder opens on its own. Author demonstrated the same consultant skill firing both ways (via / command and via a plain sentence "audited my business").
Creating a skill — fastest method
  • Paste a single prompt to Claude requesting a Skill from official docs: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 ("I want to create a Skill from this system card... The Skill should optimize my prompts based on this.").
  • Claude writes the skill; user doesn't need to learn Fable-5 prompting depth themselves.
  • Author tested it by feeding it a verbose "world's best architect, think step by step" style prompt to build a Bloomberg-Terminal-like dashboard and had the skill rewrite it for Fable 5, with per-change reasons.
  • Alternate creation method: record your screen (details referenced, not included).
Where to find ready-made skills
  • Claude built-in Browse (Anthropic skills).
  • skills.sh — open community directory, "800K+ skills"; each card shows install counts and a security audit.
  • Example install: caveman (by juliusbrussee) claims to cut token usage ~80%. Install prompt: "Hey Claude / Download and install this skill: https://www.skills.sh/juliusbrussee/caveman/caveman". Restart Claude Code to see it. Author used /caveman to rebuild a Wall Street analyst terminal from a prior project folder (reference architecture, code, API keys) — done before he finished writing the article.
The 8 mistakes section

Listed ("eight mistakes that I see most often when advising clients") but content omitted from source — the article moves straight to Next Step.

Caveats / stated limitations
  • Trigger words rely on description matching; quality depends on writing good descriptions.
  • Author advises staying in the loop: "human touch is the most important and irreplaceable part."
  • Article skips the terminal build details (a Build-It topic), so no component breakdown included.
5 example skills built by author (one-line each)
  • AI Second Brain: turns past conversations/projects/notes into living AI memory that suggests what to build/write/do next.
  • Job Hunter: searches Indeed, Upwork, LinkedIn; researches companies, writes cover letters, triggers by voice, sends best opportunities to Slack.
  • The Red Team: stress-tests ideas/launches/plans by attacking from multiple angles.
  • SEO Optimizer: scores an article, fixes SEO section-by-section preserving voice, re-scores to show improvement.
  • Reverse Engineer Pipeline: studies a product, breaks down how it works, emits a build prompt Claude Code can use to recreate its core engine.
Full text · 10,052 chars
Train Your AI like A New Employee with Claude Skills Learn what a Claude Skill is, how skills fire with commands and trigger words, and build your first one with a single prompt. 30+ client builds behind it. A couple of years ago, one of my colleagues tagged the CEO in a Trello card and said I relied too much on AI. Before researching Claude Skills for this article, I looked him up to see what he was doing, since he had left the company before I did. I saw that he now lists a Generative AI certification and had liked a post about Claude Fable 5. If your colleagues see you using Claude Skills, there’s a good chance they’ll ask questions like, “Can you trust this?”, “Are there any security risks?”, or “Is it safe?”, just like people did with OpenClaw. And three months later, whether on LinkedIn, in a Trello card, or in discussions, trust me, you will see the same people talking about AI and using it in their workflows. I have built more than 30 Claude skills for clients and worked with more people than I can count. Sometimes, more than I would have liked, but that is freelancing. Let me show you a message that one of my former clients sent to a friend I worked with and me: He was talking about the skills we built for him. A well-written skill makes the same AI system ten times more useful, and creating one takes nothing more than a folder and a Markdown file. In this guide, you’ll train your AI like a new employee using Claude Skills. By the end, you’ll be confident enough to build your own. What Is a Claude Skill? Plain English First A Claude Skill is a set of rules saved in a markdown file. And the file that has this skill is often called SKILL.md. Claude reads it before touching your task, so you stop repeating yourself. Here is one of my Claude Skills. This one audits my business. I feed it my newsletter stats, and it returns a ranked report of what is weak, what it costs me, and what to fix first. A consultant who read every number I ever produced. (Full build is here.) But where do Claude Skills run? Claude Skills run in Claude, Claude Cowork, and Claude Code. They also work in other agentic environments like Codex or OpenClaw. Let me show you where to find Claude. (In Claude, click on profile, settings, and skills to see yours.) You can browse more skills developed by Anthropic after clicking “Browse”. Toggle one on, and it is live in your next chat. How Do Claude Skills Work? Claude skills sit on disk as folders. Claude does not carry them around. Every SKILL.md opens with two metadata lines, a name and a description. When a chat starts, only those two lines get injected into Claude’s system prompt. That costs 30 to 100 tokens per skill, the price of one long sentence. The rest of the file stays on disk. When your task matches a description, Claude pulls the full file into context and follows it. Scripts and reference files in the folder load only if the job calls for them. Anthropic calls this progressive disclosure. This is why 30 installed skills cost you almost nothing until one fires. Calling Claude Skills Using Commands A Claude skill fires in two ways. The first is manual. You call it by name. In Claude Code, you type /skill-name, and it loads on the spot. It asked for my goal, and check this out, it even asked in my native language, Turkish. It connects to your Claude account and analyzes it automatically. Calling Skills using Trigger Words The second is the interesting one, trigger words. Claude compares your message against every description on the shelf, and when the words match, the folder opens on its own. Here, the same consultant skill was fired twice. First run, I typed the command. Second run, I typed a plain sentence, audited my business, and the trigger words caught it. How to create a Claude skill? Let me show you the fastest way. Just paste the prompt I’ll give you in a second. The link inside the prompt has the page inside the Claude Platform document, explaining how to prompt with Claude Fable 5. These are the newest models, and the team has changed how they should be prompted. Instead of learning everything from scratch and wasting tokens, let’s create a Claude Skill that learns directly from the official documentation the team has shared. Paste this prompt to Claude. I want to create a Skill from this system card: https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5 The Skill should optimize my prompts based on this. With just one prompt, you started training your AI. You don’t have to learn how to prompt Claude Fable 5 in such depth. Let Claude learn it for you. Just skim the document, build a few projects with Fable 5, and before you know it, you’ll have learned how to prompt it effectively. By the time I finish writing here, the skill is ready. Now, I have read all of the document, and if you want to learn how to prompt with Claude Fable 5, you can read this one, where I explained. Let’s test our skill. I crafted a prompt using the old way. You are the world's best software architect, senior frontend engineer, UX designer, and product manager. Think step by step. Be extremely careful. Before writing any code, analyze everything in detail. Do not make assumptions. Consider every edge case. Create a comprehensive plan first, then explain your reasoning, then implement the code. I want to build a Bloomberg Terminal-like dashboard for stock investors. It should look modern, include watchlists, charts, news, financial statements, earnings, analyst ratings, insider trading, macroeconomic indicators, and AI summaries. Write production-ready code with best practices. Let’s call our skill with command and ask it to update this prompt for Claude Fable 5. Here is the optimized prompt. Here are the reasons for these optimizations. Pretty amazing, right? Also, the new method lets you create a skill by recording your screen. To discover this, read this one. Where to Find Ready-Made Claude Skills You do not have to write your first Claude skill. Thousands are ready to steal. Start inside Claude. (Head to Profile / Settings / Skills first). The Browse button you saw in settings lists skills built by Anthropic, one toggle and they work. Here are the skills for you to install. The bigger shelf is skills.sh. There are 800K+ skills available here. It is an open directory of community skills; each card shows install counts and a security audit. Find one you fancy, and let’s say this, caveman. It’s developed to cut your token usage by 80%, so before reading a post with a 10-item to-do list, download this, and do not hit the limit. How? The easiest way is this prompt (paste it into Claude Code): Hey Claude Download and install this skill: https://www.skills.sh/juliusbrussee/caveman/caveman Let’s see. To see the Skill you just downloaded, close Claude and reopen it. I once built a Wall Street terminal using multiple APIs to track my investments. Now I want to build something similar again, but this time I want it to feel closer to a Bloomberg Terminal, which costs around $32K per year. I also do not want to spend too many tokens, so I use the /caveman Skill for this build. /caveman I want to build a dashboard similar to the $32K Bloomberg Terminal. Use this project as the reference, including its architecture, code, and API keys: /Users/learnai/Desktop/LearnAIWithMe Projects/Wall Street Analyst I can’t stop building with AI, sorry. But we are in the era when people can build 99% of their paid subscription with Claude code, without being technical, and most of them don’t even know. I used to build wall street analyst using Claude, now I want to go one step further, but I don’t want to explain what I want from the beginning, and don’t want to spend a lot of tokens, that’s why I inherited my old project, used one prompt, and the caveman skill. By the way, the terminal is ready, even before I finished writing these sentences. If this were a Build-It article, I would explain what every component does: one brain that fetches live market prices, economic data, and news, one screen that paints it all like a Wall Street trading desk, and one AI analyst sitting inside who reads everything and tells you buy, hold, or sell. But this article is about Skills, so I’ll skip that part. As you can see, with a single prompt, you can build your own market terminal. What a fascinating time to be alive. What Not to Do When Creating a Skill Now I am going to give you eight mistakes that I see most often when advising clients to use their skills. Next Step Log your daily routine and define what you do repetitively. Next step is to turn them into skills, but I always suggest you stay in the loop because human touch is the most important and irreplaceable part, it was, and it always will be. I’ve built 22 skills in the last 6 months. And the list keeps growing day by day, because you can do anything using the skill. Example? Let me give you 5 of them that I’ve built with one sentence explanations. AI Second Brain: It turns your past conversations, projects, notes, and results into a living AI memory that learns from your work and suggests what to build, write, or do next. Job Hunter: It searches Indeed, Upwork, and LinkedIn for matching jobs, researches companies, writes cover letters, triggers by talking only, and sends the best opportunities to Slack. The Red Team: It stress-tests your ideas, launches, or business plans by attacking them from multiple angles to find critical weaknesses before they become expensive mistakes. The SEO Optimizer: It scores your article, fixes its SEO section by section without changing your voice, then scores it again to show the improvement. Reverse Engineer Pipeline Skill: It studies any successful product, breaks down how it works, and turns the research into a build prompt that Claude Code can use to recreate its core engine. As you can see, there are too many options available. So start with smaller steps, and start building. Don’t overthink, build it, and it will get better. Thanks for reading.
10:30

Almost Timely News: 🗞️ Can AI Be An Investment Expert? (2026-08-02)

AI is not a credible investment expert, argues this guide, because the models carry encyclopedic but untrustworthy knowledge and can't do the complex math the stock market requires. The piece runs through the warnings: generative models have no ground truth, hallucinate on high-frequency data like stock symbols, and go stale at their training cutoff (Gemini 3.6 Flash stops at January 2025, Claude Opus 5 at May 2026). The practical advice is to bring your own data, use AI to write code that runs classical statistics and machine learning instead of asking it to reason directly, and treat distrust as the default posture. It ends with a long starter prompt for running investment research with the author's Deep Research Suite plugin.

Notes

Can AI Be An Investment Expert? — Christopher S. Penn, Almost Timely Newsletter (Substack, 2026-08-02)

Author: Christopher S. Penn (co-owner, Trust Insights, with CEO Katie Robbert). Content authenticity statement: ~95% human-written; outputs from Claude included.

Disclaimers / conflicts (stated upfront)
  • Penn holds no finance credentials, gives no legal investment advice.
  • Personal holdings: no-load passive index funds (S&P 500, international, money market, bond/commodity funds). "Even with AI, I'm not going to beat the market." Buys no individual securities, recommends none.
  • Trust Insights has no financial-management/investment clients; some clients are public companies but no buy/sell recommendations.
  • Paid product placement: Trust Insights Deep Research Suite (Penn benefits indirectly as owner/employee).
  • Not compensated by any data provider; criterion is free data. Generated code is "janky, sometimes hazardous" — run only with testing/QA.
  • "There are no sure bets... any investment you make can become utterly worthless or even a liability."
Part 1: Can AI be an investment expert? (No)
  • Investment-bro Reddit/X posts claiming AI fortunes are suspect: "they would not be hanging out on Reddit"; they're usually selling a book/course/method.
  • GenAI has "no ground truth," only probabilities; "if they consume enough garbage, the probabilities around that garbage will exceed those of factual truth."
  • GenAI "still can't count" — even SOTA models fail at complex math; need classical AI/ML (Python, Scala, Julia, R). Chat code interpreters are underpowered, limited on data storage.
  • Two data-quality issues: hallucination (worse for high-frequency probabilities like stock prices) and staleness — training cutoffs: Google Gemini 3.6 Flash cutoff January 2025; Claude Opus 5 cutoff May 2026. Pandemic-era strategies are baked in but inapplicable now.
  • Human caution: users demand "a simple answer," but "AI will always give you an answer. It may not be right." Ask for multiple answers; critical conversation. Three core skills: creative, critical, contextual thinking.
  • Three principles: (1) bring your own data — don't trust AI knowledge or basic web search; (2) write code to do classical AI/ML, use genAI for pattern/language and code; (3) distrust as default — verify everything, assume bad actors.
Part 2: Setup
  • Use paid AI with strong privacy or local models if PII/financial data involved.
  • First catalog your finances; "paying off a debt that costs you 29.95% APR interest is... a guaranteed return on investment."
  • Clarify goals: day trading needs millisecond-fresh data (free sources are delayed); retirement horizon sets risk tolerance.
  • Recommended tool: Trust Insights Deep Research Suite plugin/skills for agentic AI. Starter prompt (excerpt): asks for peer-reviewed strategies, regulatory requirements in the jurisdiction, statistical/forecasting practices "from inside finance as well as outside finance," Python library picks that are FOSS and updated within the last 365 days, free data sources, and avoidance of "overfitting during backtesting and lookahead bias." Uses /deep-research-skill then /research-merge to merge reports.
  • Test case: USD 100 → double in 90 days, no going bust, any legal vehicle.
  • Tools used: Gemini Spark (preferred over Gemini Deep Research — multiple agents more thorough), Qwen Studio Deep Research, Perplexity (completely failed the task), Minimax M3 in OpenCode.
Part 3: Results
  • Verdict: strategy infeasible — ~3% chance of success, ~90% failure. Report quote:
"No candidate strategy in the surveyed universe carries positive expected value after trading costs and taxes at USD 100 scale. This is a unanimous finding across all three independently-commissioned source reports... The only positive-expected-value result... is a gross, pre-cost, pre-tax figure on long-tail event-contract longshots, and it is explicitly too small in magnitude to plausibly compound to a 100% return within a 90-day window."
  • Penn argues this is the best possible result: AI's sycophancy problem means it usually fabricates to please; in finance/law/health that leads to disaster. A grounded "no" is "wild success." "If AI says yes to every request, be worried."
  • Report content: why standard investment algorithms don't fit; actuarial ruin theory; legal blocks — Polymarket, Kalshi, ForecastEx unavailable to Massachusetts residents ("I Googled it myself and yep, it's restricted here").
  • Transfer domains: actuarial science, epidemiology, especially meteorology — weather-forecasting algorithms like Brier scores, ensemble forecasting, model output statistics.
  • Tech stack: many Python libraries unknown to Penn; backtesting section to avoid overfitting (he knew only 1 of 4 tools); 33 free data sources (20 known, 13 new).
  • Final merged report: 146,000 words (≈ length of 50 Shades of Grey/The Da Vinci Code).
Part 4: Wrap-up
  • A credible "yes" to doubling $100 in 90 days would itself be a red flag — "literally everyone would be doing it." Doubling in 30 years is plausible; same process applies.
  • Penn kept human oversight: checked claims, Googled, spot-checked. AI remains probabilistic; "chances of it screwing up remain consistently high."
  • Known gap flagged by the merge: both source reports omitted a news/text-ingestion layer (no NLP, no filings parser beyond raw EDGAR JSON; no free news data source found).
Side notes
  • GEO 201 course (Trust Insights) on GEO measurement/methodology: USD 149.
  • Books plugged: 21 Use Cases of Generative AI For Marketers (new), Almost Timeless, Generative AI for SEO and PPC Marketers, Generative AI for Destination Marketers.
  • Speaking: LPA Philadelphia (Sept 2026), MAICON Cleveland (Oct 2026), SMPS AI Austin (Nov 2026), MarketingProfs B2B Forum Boston (Nov 2026).
Full text · 25,345 chars
Almost Timely News: 🗞️ Can AI Be An Investment Expert? (2026-08-02) :: View in Browser The Big Plug Content Authenticity Statement 95% of this week’s newsletter was made by me, the human. You’ll see multiple outputs from Claude and a link to the report it generated. Learn why this kind of disclosure is a good idea and might be required for anyone doing business in any capacity with the EU in the near future. Watch This Newsletter On YouTube 📺 What’s On My Mind: Can AI Be An Investment Expert? Let’s head for the danger zone! (Cue Kenny Loggins, please) Someone asked in my reader survey if I would cover how to use generative AI for stock market analysis. Yes, you can, but there’s a whole host of warnings. So let’s get through all the warnings and disclaimers first. Part 0: Warnings, Disclaimers, etc. First and foremost, I am not a certified anything in finance. I hold no credentials and have no particular expertise in it. None of what I say in this newsletter is advice (in the legal sense of me telling you what to do) and the only advice you should take is from qualified, certified experts in the area. That is not even remotely me. My stock warning always applies: consult a qualified practitioner in your jurisdiction for advice for your specific situation. Second, because it’s something that touches finance and investment, I have to disclose any conflicts of interest. I have retirement savings, mostly 401K-style investments, that are invested in no load index funds because even with AI, I’m not going to beat the market. I’m invested in an S&P 500 index fund, an international index fund, some money market funds, and a couple of bond and commodity funds that are also passive. I don’t buy or hold individual investments, nor do I recommend specific companies to invest in or short. My biggest investment in an overall sense is as a co-owner of Trust Insights with my cofounder and CEO, Katie Robbert. Trust Insights does not currently have any financial management or investment clients; we do have clients that are publicly traded but we do not recommend investment for or against any client. You will see a paid product placement from my company, the Deep Research Suite, in this newsletter. I receive indirect financial benefit as an employee and owner if you make any purchases. Third, while I may suggest some providers in this issue, I do not endorse any given data provider, nor am I compensated currently by any. My primary criteria for selecting any kind of financial data provider is whether the data is freely available or not because finance is not my area of expertise and I’m not going to pay for data I barely use, or use in an ad hoc manner. Likewise, if you follow the steps in this newsletter and use it to generate code, you do so at your own risk. AI frequently generates janky, sometimes hazardous code that you should never run without strong testing, QA, and validation. Fourth and finally, nothing in finance and especially nothing in the stock market is ever guaranteed unless you’re doing something illegal, like insider trading. There are no sure bets, and I cannot promise any outcomes if you use the information in this newsletter. This issue is solely for educational purposes. Any ideas you implement are at your own risk, and there is always a real, meaningful chance that any investment you make can become utterly worthless or even a liability, sometimes in the blink of an eye. Proceed entirely at your own risk. Part 1: Can AI Be An Investment Expert? Before we dig into specifics, let’s tackle the big picture. Can AI act as a credible investment expert? If you go to places like Reddit or Xitter (my favorite neonym for the platform formerly known as Twitter, combining X and Twitter and pronouncing it in Hanyu Pinyin style as Shitter), you will see endless posts of investment bros touting how AI just made them a gazillion dollars. Basic logic would suggest that if AI in fact did so, they would not be hanging out on Reddit. They’d be on their private island in the Maldives with their gazillion dollars. Most of the time, if you dig deeper, you find out that they’re actually selling a book/course/secret method for supposedly making a gazillion dollars on the stock market with AI but they themselves are not gazillionaires (which reduces their credibility considerably). Today’s generative AI models have encyclopedic knowledge of the investment world, having consumed a lot of data about it, but… much of that data is mashed up. Remember that generative AI models have no understanding of facts, no ground truth (Claude’s favorite term) at all. They’re made of probabilities, and if they consume enough garbage, the probabilities around that garbage will exceed those of factual truth. And the internet is filled with snake oil salesmen (and has been since the public was first allowed on it) promising get rich quick schemes - which the machines have also learned. So to start, the investment knowledge AI has broadly is untrustworthy, the foundational knowledge is shaky because of the sheer quantity of unhelpful, conflicting, or shady information on the internet about investments. Second, generative AI models - even today’s state of the art models that every broligarch is touting as sentient Skynetesque machines that can hack into anything - also still can’t count. When you take the base model and have it do math, even the best models fail at it, especially if the math is complex - and the math around the stock market is super complex. For anything involving math of this level, you need classical AI - machine learning and statistics, usually run inside some kind of code like Python, Scala, Julia, R, etc. Generative AI simply isn’t up to the task of doing this kind of math. Now, some generative tools have things like code interpreters (the ability to run limited amounts of code) right in chat, but those are often underpowered and have significant limitations, especially around data storage. Third, generative AI models around anything stock-related have two serious data quality issues. The first is hallucination; AI tools are notorious for hallucinating in general, but when you’re dealing with probabilities that have very high frequencies - like, say, prices around stock symbols - the risk is even higher. Then there’s staleness. All AI models natively have a training data cutoff date; their knowledge ends after a certain point. For example, Google Gemini’s 3.6 Flash model has a cutoff date of January 2025; nothing that happened after that date is in the model’s inherent knowledge. Claude Opus 5 has a cutoff of May 2026. That stale knowledge can have material impacts on investment strategy, for what I assume are obvious reasons. For example, investment strategies that worked during the pandemic are inapplicable now, but that core knowledge is part of what AI has trained on. The last caution is a human caution. Humans as a rule tend to be incurious, especially with using AI. For example, I have a friend who has a habit of asking AI to do things like “give me a simple answer” or “give me just the right answer”, which when it comes to investments and stocks is an incredibly dangerous practice. AI will always give you an answer. It may not be right, nor will it be thoughtful. At a minimum, we always want to ask AI for multiple answers, and ideally we have thoughtful, critical, and reflective conversations as part of developing answers. I’ve said for some time now in my keynotes and workshops that the three most important skills in AI today are creative thinking, critical thinking, and contextual thinking; nowhere does this apply more than in high risk applications like finance and investments. What we derive from all these cautions is three key principles for using AI in a stock and investment context: - Bring your own data. From regulatory requirements to actual stock data, you cannot trust what AI knows, nor can you trust basic web search (because it’ll surface the same swamp of questionable data). - Write code to do classical AI/ML. Your best bet is to use AI for what it’s best at - pattern identification in language, and creating language - and that language includes writing code. AI can write the necessary tools to do advanced analysis, and then help you interpret the results. - Distrust should be your default. Do not trust what AI gives you. Verify and fact check yourself, always. Assume there are bad actors at every step of the way looking to waylay you and take your money. Part 2: Setting Up Before we can begin any kind of investment analysis, we need our data. As a reminder, you should be using paid AI tools with strong privacy protections OR local AI models that do not send data anywhere if you are using personally identifying financial data of any kind. First, you need a catalog of your own finances and where things stand right now. Like a bachelorette party hitting the games at Mandalay Bay, know what you have to work with before you invest a single yuan. It’s deeply unsatisfying, but paying off a debt that costs you 29.95% APR interest is, quantitatively, a guaranteed return on investment because you’re not paying interest on what you owe. Very little in the stock market is going to pay the same level of return as escaping those kinds of interest rates on debt. Second, you have to be super clear about what kinds of stocks and investing you want to do, because trading strategies vary wildly based on your goals. Are you just messing around with day trading and penny stocks with a hundred bucks of play money? That requires one set of rules and methods. For example, in day trading, free data sources won’t cut it because they’re often delayed and in day trading, milliseconds count. Are you saving for retirement, and if so, when? Your time horizon dictates the level of risk you may be willing to accept; longer horizons allow for more risk. Once you are crystal clear on your goals, write them down, then use AI deep research tools to gather information about the specific strategies, tactics, techniques, and methods of those goals. I recommend the newly released Trust Insights Deep Research Suite plugin and skills for your favorite agentic AI tools to do this; the suite walks you through commissioning multiple research reports and then merging them together to reduce hallucinations and cast a wide field. Use this as your starter prompt in the suite: You’re a stock and investing expert. Use the /deep-research-skill to help me build a research prompt for investing in {investment type}. My goals are {goals}. My level of acceptable risk is {risk}. My current general financial situation is {financial situation}. What are peer-reviewed strategies, tactics, techniques, and methods for doing this kind of investing? I’ll be starting with {money}. Your research should also include regulatory requirements and laws in {jurisdiction} related to the type of investing we’ll be doing. What strategies are effective for my goals? What strategies are ineffective for my goals? In particular, I want to know statistical and analytical techniques for investment strategy both from inside finance as well as outside finance, explicitly searching for proven statistical and forecasting practices in other expert domains that apply to my goals. The downstream use of this research will be for constructing a system to do data analysis, forecasting, and recommendations using a language like Python, so your research should also include which libraries and packages are the best fit for this based on your search and analysis of techniques and methods, with an emphasis on avoiding common traps with this kind of data such as overfitting during backtesting and lookahead bias. Any packages or libraries you recommend should be FOSS and updated after {last 365 days). Also identify which freely available data sources are available to support a system like this, APIs and downloads. There is no maximum length on the research; clarity and factual correctness always take precedence over brevity. Conciseness that sacrifices thoroughness is an antipattern that creates serious and harmful risk. The research suite will ask you for additional details, then create the prompt. Run the prompt, obtain your research from the different sources, and use the /research-merge skill to merge together the different reports into a master report. What you should have is a complete, well-researched catalog of strategies, tactics, techniques, and methods along with the accompanying data sources and code recommendations to do a build. For this issue, and in the video demo, I gave it a simple target - take USD 100 and double it in 90 days without going bust, using any legal vehicle (including non-stock systems). For the research tools, I used: - Gemini Spark (preferred over Gemini Deep Research because multiple agents do a much more thorough job) - Qwen Studio Deep Research - Perplexity (which completely failed at the task) - Minimax M3 in OpenCode Part 3: Reviewing the Results Once we’ve got the research merged, it’s time review it. In the example I gave, it came up with... nothing! And by that, I mean the strategy I gave it came out to be completely infeasible, with the chances of success about 3% and the chances of failure about 90%. Here’s what it said: No candidate strategy in the surveyed universe carries positive expected value after trading costs and taxes at USD 100 scale. This is a unanimous finding across all three independently-commissioned source reports underlying this merge — not a majority-rule adjudication, but genuine convergence. The only positive-expected-value result anywhere in the merged evidence is a gross, pre-cost, pre-tax figure on long-tail event-contract longshots, and it is explicitly too small in magnitude to plausibly compound to a 100% return within a 90-day window even before costs are subtracted. Bottom line for the decision this report exists to inform: the evidence does not support an expectation of doubling USD 100 within 90 days through any legal vehicle surveyed. If the experiment proceeds regardless — as an explicitly bounded, fully-loss-tolerant exercise rather than an investment expected to succeed — Section 12’s epistemic salvage plan specifies what to measure so the 90 days produce a generalizable lesson even when, as expected, the capital objective is not met. Believe it or not, this is the best possible result, and here’s why: AI in general has a sycophancy problem. When we give it directives, like “tell me how to turn USD 100 into USD 200”, it will try to be helpful and fulfill that directive in any way possible, including hallucinations and fabrications. It will literally make things up to make us happy, to be helpful. In fields like finance, law, and health, this sycophancy leads to disastrous outcomes. AI telling us things to make us happy that are false leads to outcomes like people betting money they don’t have, or spurning necessary care. When you get results that say, “No, you can’t do this, it’s not realistic, but if you want to try here’s how and you’re very likely to lose it all”, to borrow Claude’s favorite negative parallelism construct, that’s not failure. That’s wild success - AI has conducted the research, assembled the pieces, and delivered a realistic, grounded assessment. We want to see failures like this. We NEED to see failures like this, so that we know AI is working in the real world, not in people-pleasing mode that’s completely and dangerously delusional. In turn, that means similar conclusions through this process are likely to hold up. If AI says yes to every request, be worried. When we dig deeper into the report, what we find is a wealth of different methods and tactics that, should we foolishly proceed, would be an education in their own right. It reviews why standard investment algorithms are the wrong choice for the goal I set, specific techniques like actuarial ruin theory (aka when you’re going to fail), and even calls out where there are legal entablements. For example, Polymarket might legally be inaccessible to me based on where I live: Three independent reports, three different routes, one conclusion. For a Massachusetts-resident principal, Polymarket is not an available vehicle, and the same reasoning reaches Kalshi and ForecastEx on non-sports contracts. I Googled it myself and yep, it’s restricted here. Further down in the report, it identified several domains, per the original prompt, where there are real learning and transfer opportunities for this kind of work - actuarial science, epidemiology, and especially meteorology. That’s right - weather forecasting. Turns out there are many rich, robust, peer-reviewed and proven algorithms in weather forecasting that we could transfer to this domain, such as brier scores, ensemble forecasting, model output statistics, and more. Near the end is the tech stack, with a HUGE number of Python libraries, many of which I didn’t even know existed. There’s even a section on which backtesting tools to use to make sure we’re not overfitting on past data, and I only knew of one of the four listed. And beyond that was a buffet of 33 free data sources I could use to extract data from these various markets to power an application, should I have chosen to do so. While I knew about 20 of them, the remaining 13 were new to me and now part of my mental toolkit for any finance/stock projects I might have in the future. The final report weighs in at 146,000 words, or roughly a novel the length of 50 Shades of Gray or The Da Vinci Code, but substantially less thrilling than either. That said, it’s chock full of information across the entire project based on our prompt and could literally be a graduate education unto itself if you sat down and really studied each of the sections. Part 4: Wrapping Up Now, this was an experiment using what was objectively a stupid question, how I could reliably double my money in 90 days. If AI came up with something that said yes, here’s how with high confidence - that was LEGAL - I would be deeply concerned (but I’d probably try it). Look, if there really was a way to reliably double your money in 90 days legally, literally everyone would be doing it. If I were to give it a far more sane prompt, like having it help evaluate my retirement portfolio’s allocations and distributions, chances are it would come up with an equally thorough report and I could build my software off that. Doubling your money in 90 days is implausible. Doubling your money in 30 years is highly plausible, and the process I outlined in this newsletter is exactly how you’d research that. I’ll also point out that throughout this process, I did not hand off my critical thinking to the machine. I checked out claims, Googled stuff, and read it in depth to see what it came up with. The research merge pointed out lots of conflicting claims and how the tool resolved those claims, sometimes highlighting that there was no clear answer: Both sources omitted a news and text-ingestion layer (no NLP, no filings parser beyond raw EDGAR JSON), consistent with Section 12’s finding that neither report located a free news data source. As I said at the beginning of this issue, any time you’re working in the fields of finance, law, and health, you absolutely cannot be hands off. You cannot safely delegate to the machines and not check their work repeatedly, in depth. Yes, this process today worked fine and everything I spot checked was correct, but AI is by nature probabilistic, which means that the chances of it screwing up remain consistently high. Finally, let me know in the replies/comments/wherever if you want to see what a version of this would look like with the sane prompt. Maybe I’ll do that as a part 2 if it would be of interest to you. How Was This Issue? Rate this week’s newsletter issue with a single click/tap. Your feedback over time helps me figure out what content to create for you. Got More Feedback? Here’s The Unsubscribe It took me a while to find a convenient way to link it up, but here’s how to get to the unsubscribe. If you don’t see anything, here’s the text link to copy and paste: Share With a Friend or Colleague Please share this newsletter with two other people. Send this URL to your friends/colleagues: For enrolled subscribers on Substack, there are referral rewards if you refer 100, 200, or 300 other readers. Visit the Leaderboard here. ICYMI: In Case You Missed It Here’s content from the last week in case things fell through the cracks: On The Tubes Here’s what debuted on my YouTube channel this week: My Merch Shop I’ve been adding so much stuff that I’ve decided to bundle it all in what I call a Merch Shop, because otherwise there’s literally too much to keep track of and I run out of space in my own newsletter. So welcome to the Merch Shop! Books: - 👉 New book! 21 Use Cases of Generative AI For Marketers - Almost Timeless: 48 Foundation Principles of Generative AI - Generative AI for SEO and PPC Marketers - Generative AI for Destination Marketers Skills for Claude and Agentic AI: Courses: Subscriptions: Recent Talks These are just a few of the classes I have available over at the Trust Insights website that you can take. Advertisement: New GEO 201 Course In GEO 101, the first course I built on the basics of GEO, I taught you about presence, appearance, and relevance, the three phases of GEO, and what you need to do in each phase to align with how AI search operates. The top piece of feedback we got at Trust Insights about it was, “okay, great, but how do I tell my boss that we’re ‘winning’ at GEO?“ After I quelled my murderous rage at your boss on your behalf, Katie and I sat down and worked out a straightforward, aligned methodology for doing this. GEO 201 is based on the three phases, what you can control and what you can genuinely see - and critically, what you can’t. Because there is absolutely no way to say your brand “ranks higher” in AI search, period, end of story. But you can say and show with confidence what you’ve done and how you show up for presence, appearance, and relevance with tools you’re probably already paying for, and based on how AI search systems really work. 👉 GEO 201 is available now for USD 149. Get Back To Work! Folks who post jobs in the free Analytics for Marketers Slack community may have those jobs shared here, too. If you’re looking for work, check out these recent open positions, and check out the Slack group for the comprehensive list. Disclosure: I source these links from LinkedIn every week on the following criteria: New in the past seven days, Easy Apply on, remote roles, USA geography. How to Stay in Touch Let’s make sure we’re connected in the places it suits you best. Here’s where you can find different content: - My blog - daily videos, blog posts, and podcast episodes - My YouTube channel - daily videos, conference talks, and all things video - My company, Trust Insights - AI help - My podcast, Marketing over Coffee - weekly episodes of what’s worth noting in marketing - My second podcast, In-Ear Insights - the Trust Insights weekly podcast focused on data and analytics - On Bluesky - random personal stuff and chaos - On LinkedIn - daily videos and news - On Instagram - personal photos and travels - My free Slack discussion forum, Analytics for Marketers - open conversations about marketing and analytics Listen to my theme song as a new single: Social Good: Ukraine 🇺🇦 Humanitarian Fund The war to free Ukraine continues. If you’d like to support humanitarian efforts in Ukraine, the Ukrainian government has set up a special portal, United24, to help make contributing easy. The effort to free Ukraine from Russia’s illegal invasion needs your ongoing support. Events I’ll Be At Here are the public events where I’m speaking and attending. Say hi if you’re at an event also: - LPA, Philadelphia, September 2026 - MAICON, Cleveland, October 2026 - SMPS AI Conference, Austin, November 2026 - MarketingProfs B2B Forum, Boston, November 2026 There are also private events that aren’t open to the public. If you’re an event organizer, let me help your event shine. Visit my speaking page for more details. Can’t be at an event? Stop by my private Slack group instead, Analytics for Marketers. Required Disclosures Events with links have purchased sponsorships in this newsletter and as a result, I receive direct financial compensation for promoting them. Advertisements in this newsletter have paid to be promoted, and as a result, I receive direct financial compensation for promoting them. My company, Trust Insights, maintains business partnerships with companies including, but not limited to, Amazon, Talkwalker, MarketingProfs, Agorapulse, The Marketing AI Institute, Spin Sucks, and others. While links shared from partners are not explicit endorsements, nor do they directly financially benefit Trust Insights, a commercial relationship exists for which Trust Insights may receive indirect financial benefit, and thus I may receive indirect financial benefit from them as well. Thank You Thanks for subscribing and reading this far. I appreciate it. As always, thank you for your support, your attention, and your kindness. Please share this newsletter with two other people. See you next week, Christopher S. Penn
20:41

Claude Code After Generation 5 Models: The Master Guide

Claude Code changed twice in one month, with stronger models that stay inside a hard problem much longer and tools that keep the work moving without a new prompt after every step. Opus 5 launched July 24 with judgment and self-verification close to Fable 5 at roughly half its price, and sessions now reach one million tokens. Anthropic also cut more than 80% of Claude Code's system prompt for the new models with no measurable loss on coding tests, because the stronger models don't need every rule spelled out. The guide argues the winning setup is a small working environment around the model, not one giant controlling prompt.

Notes

Claude Code After Generation 5 Models (Emerging AI, 2026-08-02)

Promotional teaser for a paid "master guide" to Claude Code on Gen-5 models. Core facts:

  • Opus 5 released July 24; claimed gains: stronger judgment, deeper root-cause analysis, better self-verification, "performance close to Fable 5 at roughly half its price."
  • Sonnet 5: already "a far more capable daily model."
  • Fable 5: built for work larger than one sitting.
  • Claude Code added: completion goals, adjustable effort, verification Skills, one-million-token sessions, and dynamic workflows coordinating many agents through an executable plan.

Key claim — Anthropic cut >80% of Claude Code's system prompt (for Opus 5 and Fable 5) "without a measurable loss on its coding evaluations." Old guardrail instructions had "become friction"; stronger models no longer needed "every decision explained, every rule repeated, or every possible failure predicted in advance."

Thesis / usage shift:

"The prompt gets smaller. The system around it gets better."

Old instinct: one enormous prompt, full control. New approach: a small working environment around Claude — right project context, clear target, useful tools, a place to record decisions, and "a judge that can prove whether the work is finished."

Full guide advertises: setup for newest models; choosing Sonnet 5 / Opus 5 / Fable 5; effort levels; project memory via CLAUDE.md; creating Skills; safe MCP/plugin connections; subagents, loops, graph workflows; a start-to-finish system with commands, prompts, verification, real examples.

Caveats: content is a paywall funnel — actual instructions are not in the post; the >80% prompt-cut stat is Anthropic's framing, no eval numbers or benchmarks given; no comparison tables or pricing details despite "half the price" claim.

Full text · 2,187 chars
Claude Code After Generation 5 Models: The Master Guide A beginner-friendly guide to using Claude Code with better context, stronger models, and practical automation. Claude Code changed twice in the same month. First, the models became strong enough to stay inside a difficult problem for much longer. Then the tool around them learned how to keep that work moving without waiting for another prompt after every step. Opus 5 arrived on July 24 with stronger judgment, deeper root-cause analysis, better self-verification, and performance close to Fable 5 at roughly half its price. Sonnet 5 had already become a far more capable daily model. Fable 5 was built for work larger than one sitting. Claude Code also added completion goals, adjustable effort, verification Skills, one-million-token sessions, and dynamic workflows that can coordinate many agents through an executable plan. Then Anthropic revealed something more interesting than another benchmark. For Opus 5 and Fable 5, it removed more than 80% of Claude Code’s system prompt without a measurable loss on its coding evaluations. Many instructions that used to keep older models on track had become friction. The stronger models no longer needed every decision explained, every rule repeated, or every possible failure predicted in advance. That changes the way Claude Code should be used. The old instinct was to write one enormous prompt and control every move. The better approach now is to build a small working environment around Claude: give it the right project context, a clear target, useful tools, somewhere to record decisions, and a judge that can prove whether the work is finished. The prompt gets smaller. The system around it gets better. That is the real subject of this guide. Inside the full guide, you’ll get the complete Claude Code setup for the newest Claude models, how to choose between Sonnet 5, Opus 5, and Fable 5, use effort levels wisely, build project memory with CLAUDE.md, create Skills, connect MCPs and plugins safely, run subagents, loops, and graph workflows, and follow a practical start-to-finish system with commands, prompts, verification, and real examples you can use immediately.
14:01

How to Build a Weekly Dashboard That Runs Your Week

A weekly planning dashboard squeezes your whole week into seven chunks or fewer, built on a 1956 Harvard study finding that working memory holds about seven items. George Miller's 'magical number seven' paper showed memory is limited by chunks, not raw information, so the method caps the dashboard at five to seven widgets. It includes a ready-to-paste prompt that builds the interactive single-screen dashboard as a Claude artifact, with the paper version first to learn what you actually use.

Notes
How to Build a Weekly Dashboard That Runs Your Week — Solopreneur Code (substack, 2026-08-02)

By "Anfernee." Core claim: working memory holds ~7 chunks, so a weekly dashboard should compress the whole week into ≤7 chunks. Heavily promotional — pitches "Solopreneur Success Hub" and paid "Premium Vault" at $79/year ($6.58/month) with a Claude setup ("Solopreneur OS for Claude", 16 skills, one setup interview).

The psychology basis (George Miller, 1956)
  • Harvard psychologist George Miller, paper "The Magical Number Seven, Plus or Minus Two". Opening line: he'd been "persecuted by a number" — seven.
  • Tests across three tasks: (1) repeating back read digit lists stretched until failure; (2) matching tones of different pitches to memorized labels; (3) counting groups of items at a glance. Different senses, same ceiling: ~7 units before performance collapsed.
  • Key insight: the limit is on chunks, not information. "Seven," the digit 7, or a five-letter word each count as one chunk. Phone number example: ten digits overwhelm; grouped as 876 + 1234 → two or three chunks.
  • Implication the author draws: any planning system should squeeze the week into seven chunks or fewer. Five on a tired day, nine on a sharp one.
What a dashboard is

"Not a daily view, a weekly calendar, or a monthly grid" — a single spread/page showing a stretch of time at a glance (car-dashboard analogy). Works by doing the sorting once, then "a glance instead of a decision."

Four widget families (pick ~5, cap at 7)
  • Planning/priorities: priorities block (2–3 goal-tied outcomes, keep to three), due-this-week list, schedule-at-a-glance.
  • Data/habit tracking: habit tracker is binary (mark a clear symbol for a miss, never blank — blank is ambiguous); data inputs record values (sleep hours, steps, mood). Author tracks focus-work hours weekly.
  • Reflection: gratitude log (one line/day) and captain's log (factual what-you-did/how-it-felt). Keep to 1–2 lines each.
  • Notes/capture: scratch pad + a "next week" landing spot so next dashboard pulls from captured input.

Rule: each widget = one chunk; "if the page makes you wince, you've added too much."

Paper build steps (recommended first)
  • Open two facing pages. 2. Pick five widgets, six max (author's: schedule, priorities, due-this-week, sleep/work tracker, notes). 3. Draw frames before writing. 4. Size by importance — priorities and schedule get the most room. 5. Fill five minutes nightly or each morning; look daily, adjust weekly.
Claude live-artifact build

Flow: describe exact widgets by name → ask for "single-page HTML dashboard artifact... one screen, no scrolling" → set constraint (no more than seven widgets) → iterate by talking (checkable boxes, color for missed habits) → keep open or export code to host. Notes: "Artifacts" → "Create Cowork artifact."

Ready-to-use prompt (condensed)

Seven widgets: Mon–Sun appointment calendar; email-planning box that turns pasted emails into planned tasks; to-do checklist (clickable boxes); Substack Note tracker (one row/day); focus-work-hour tracker with weekly total; notes area; gratitude log (one line/day). Style: one screen, no scroll, calm minimal, orange/white/grey, easy to edit. Author's own artifact uses all seven.

Caveats
  • No reported limitations or failures; no data on whether it works beyond author's testimony ("saves me 20+ hours a week" — attributed to the Success Hub toolkit, not the dashboard).
  • Two versions of "seven": paper method says cap at 5–6; the artifact prompt deliberately includes exactly 7 widgets.
  • Miller's findings are widely debated/popularized loosely; the post does not engage the replication debate.
Full text · 13,547 chars
How to Build a Weekly Dashboard That Runs Your Week Here's how to build a weekly dashboard that carries the memory load for you, backed by a 1956 Harvard study and buildable in an afternoon (includes free prompt). Think back to last Friday. Did you close the week clean, or did you spend the weekend with that itch at the back of your skull, sure you forgot something but unable to name it? I lived in that state for years. This post is about the fix I wish I’d found sooner, a single spread I call a weekly dashboard that holds your week so your brain doesn’t have to. Access your FREE Solopreneur Success Hub - your subscribers-only comprehensive command center for building and scaling a successful one-person business. I created this all-in-one toolkit for building a profitable one-person business, something I wish existed when I first started, and it saves me 20+ hours a week. Now, it’s yours… FREE! For the better half of my life, I tried to remember everything I used to treat my memory like a filing cabinet. Yes! Deadlines, ideas, the name of the person I met last Tuesday. All of it went into my head, and I trusted my head to hand it back on demand. It didn’t. I’d wake at 2am remembering an invoice. I’d walk into a meeting and blank on the one number that mattered. I got good at looking calm while quietly panicking that something important had slipped through a crack I couldn’t see. The strange part is that I thought this was a discipline problem. Try harder, focus more, care enough and you’ll remember. That belief cost me years. The truth is simpler and kinder. My memory was doing exactly what human memory does. I was asking it to do a job it was never designed for. Humans have always been bad at this, and always tried to fix it We like to think productivity struggles are modern, a side effect of phones and endless tabs. But they aren’t. People have been outsourcing memory for as long as they’ve had something to write on. Tally marks on bone. Knotted cords. Ledgers, almanacs, the humble to-do list scratched on a scrap of paper. Every one of these tools exists for the same reason: The human mind is brilliant at thinking and terrible at storing. So we invent systems to hold information so our attention is free to actually use it. So if you’ve ever felt like you’re bad at staying organized, sit with this. You’re part of a very long line of people who hit the same wall and built something to lean on. The dashboard is just the latest version of that ancient move. George Miller and the magic number seven In 1956, a Harvard psychologist named George Miller published a paper with one of the best opening lines in science. He wrote that he had been persecuted by a number. That number was seven. Miller kept running into it everywhere he looked. His paper, “The Magical Number Seven, Plus or Minus Two,” argued that our working memory, the mental workspace we use to juggle things in the moment, has a hard ceiling. That ceiling sits at roughly seven items, give or take two. Call it five on a tired day, nine on a sharp one. Go past that limit and things fall apart. You drop items, mix them up, or freeze. Your workspace is simply full. What Miller’s experiment actually proved This next part is what makes the dashboard work. Miller tested people across three different tasks. - First, he read out lists of digits and asked people to repeat them back, stretching the lists longer and longer to find the breaking point. - Second, he had people match signals, like tones of different pitches, to labels they’d memorized. - Third, he showed groups of items and asked people to count them quickly at a glance. Different senses, different tasks, same ceiling. People handled about seven units before performance collapsed. Then he noticed something stranger. What counted as one unit was flexible. The word “seven” is one unit. The single digit 7 is one unit. Even a five letter word is one unit, not five. Miller’s insight was that memory is limited by the number of chunks, not the raw information. A chunk is the largest meaningful piece you can recognize as a single thing. A phone number is the classic example. Ten separate digits will overwhelm you. Group them, 876 then 1234, and suddenly you’re holding two or three chunks instead of ten. Same information, far less strain. That’s the hinge. If your mind can only hold about seven chunks, then the goal of any good planning system is to squeeze your entire week into seven chunks or fewer. That’s exactly what a dashboard does. Turning a psychology paper into a planning system So we take Miller’s ceiling and we build with it instead of against it. Your week has hundreds of moving parts. Tasks, appointments, habits, half-formed ideas, things you’re grateful for, things you’re dreading. You cannot hold all of that. Nobody can. A weekly dashboard forces you to crunch that mess down into a handful of meaningful chunks. The other stuff doesn’t vanish. You just decide in advance which chunks deserve a spot in your limited mental workspace. Do the hard sorting once, then coast on it all week. What a dashboard is, and why almost everyone loves them A dashboard is not a daily view, a weekly calendar, or a monthly grid. It’s a single spread, sometimes a single page, that shows you what a stretch of time holds at a glance. Think of a car dashboard. You don’t read a manual to drive, you glance down and instantly know your speed, your fuel, your warnings. That’s why people love them. A weekly dashboard relieves you of the burden of rebuilding your week from memory every morning. You open one page and the shape of your week is right there, already sorted, already prioritized. Deciding what to do next stops being a fresh decision and becomes a glance. The crunch is what does the work, reducing everything to the few things you actually need to see. The different types of dashboard widgets Most dashboards are built from four families of widgets. Pick from these like ingredients. Planning and priorities This is the core. A priorities block holds your two or three most important outcomes for the week, the ones tied to real goals, not “take out the trash.” Keep it to three. A due this week list catches the smaller running tasks, where the laundry and the trash live. A schedule at a glance maps appointments and deadlines across the days so you know what’s happening and when. Data and habit tracking A habit tracker is binary. Did you do it or not? Mark a clear symbol for a miss, never leave it blank, because a blank box is ambiguous later. Data inputs work differently. Instead of yes or no, you record a value, hours slept, hours worked, steps, mood. I track my focus work hours every week because I run my own business, and those two numbers tell me fast whether I’m heading toward burnout or balance. Reflection These widgets hold values and feelings, not hard data. - A gratitude log gives you one line a day on what you’re thankful for. - A captain’s log is more factual, a quick record of what you did and how it felt. Both keep you grounded. Keep them short, a line or two, since this is dashboard scale, not journal scale. Notes and capture A notes block is a scratch pad for anything that lands on your desk mid week, a phone number, a stray idea. A next week block is a landing spot for things that don’t belong to this week, so when you build your next dashboard you’re pulling from real captured input instead of thin air. The rule that ties it together: each widget is one chunk. Choose around five, cap it at seven. If the page makes you wince when you look at it, you’ve added too much. Trust that feeling. How to build a dashboard on paper Start analog. It’s faster to think with a pen and there’s nothing to configure. - Pick a spread. Open your notebook or planner to two facing pages, or one page if you prefer tight. - Choose your chunks. From the four families above, pick five widgets, six at most. For me that’s schedule at a glance, priorities, due this week, a sleep and work tracker, and a notes block. - Draw the frames first. Block out a box for each widget before you write anything inside. This stops the page from turning into a wall of text. - Size by importance. Give priorities and schedule the most room. Trackers and notes can be small. - Fill it the night before or each morning. Spend five minutes loading the week in, then let the page carry it. That’s the whole method. Look at it daily, adjust weekly. The paper version is the best place to learn what you actually use before you build anything fancier. How to build a dashboard using a Claude live artifact Once you know your five or six widgets, you can turn the same layout into an interactive version using a Claude live artifact. An artifact is a working mini app Claude builds and renders right beside the chat, so you get clickable checkboxes and editable fields instead of static ink. Here’s the flow I use: - Describe your exact widgets. Tell Claude the five or six blocks you settled on with paper. Precision here matters, so name them: “a priorities box with three slots, a Monday to Sunday schedule row, a due this week checklist, a sleep and work hours tracker, and a notes area.” - Ask for a single-page HTML dashboard artifact. Say something like: “Build me a weekly dashboard as an interactive HTML artifact. One screen, no scrolling, clean and minimal.” Claude renders it live so you can see it immediately. - Set the constraint. Tell it to respect the magic number seven, no more than seven widgets, so it doesn’t over build. - Iterate by talking. Ask for checkable boxes, a color for missed habits, bigger priority text, whatever’s off. Each message updates the live artifact in place. - Use it or export it. Keep it open as your week’s home screen, or ask Claude for the code to host it yourself. Steps Choose “Artifacts” Select “Create Cowork artifact” And enter your prompt. Your ready-to-use prompt Here’s a single prompt you can paste straight into Claude to build the week-ahead version. It packs exactly seven widgets, edit accordingly based on your own situation and needs. Build me a weekly dashboard for the week ahead as an interactive, single-page HTML artifact. One screen, no scrolling, clean and minimal. Respect the magic number seven and include exactly these seven widgets, no more: 1. Calendar for the week ahead: a Monday to Sunday view showing my appointments and time-bound events. 2. Email planning box: a space where I paste this week's key emails, and you pull out what needs action and turn it into planned tasks for the week ahead. 3. To-do for the week ahead: a checklist with clickable checkboxes. 4. Substack Note tracker: one row per day to log whether I posted a Note and its topic. 5. Focus work hour tracker: log deep-work hours per day with a simple weekly total. 6. Notes area: a free-form scratch pad for stray ideas and quick capture. 7. Gratitude log: one line per day for what I'm thankful for. Make all checkboxes clickable, keep everything on one screen with a calm, minimal style with orange, white and grey colour theme, and make the layout easy to edit. After you build it, I'll ask for tweaks like bigger priority text or a color to mark missed days. Tweak it by talking. Ask for bigger headers, a color for missed days, or a running total under the focus tracker, and Claude updates the artifact in place. The advantage over paper is recall and reuse. You describe the structure once and regenerate a fresh blank dashboard every week in seconds. Here’s mine: Every time you open Claude, you re-explain your business from scratch. Your niche. Your voice. Your offers. Ten minutes lost before it writes a single useful word. Solopreneur OS for Claude fixes that. One setup interview. Then 16 skills covering positioning, content, sales pages, launches, weekly reviews that all read your profile before every output. Validate → Build → Sell → Review. One system, one voice, one business. Final thoughts You were never going to win by remembering harder. Your working memory tops out around seven chunks, and no amount of willpower moves that ceiling. A weekly dashboard works because it stops fighting your brain and starts feeding it exactly what it can hold. Build one this week. Grab a notebook, pick five widgets, draw the boxes. Then, if you like it, rebuild it as a Claude artifact you can regenerate everyday. Either way, you’ll walk into Monday with your week already mapped, and you’ll close Friday without that nagging feeling that something got left behind. If you want more systems like this, ones built for solopreneurs who’d rather work with their brain than against it, that’s exactly what I write about at Solopreneur Code. You’re doing everything. But nothing is moving? You are doing everything. But nothing is moving. That is not a motivation problem. Most solopreneurs are learning from everywhere and getting nowhere. Too much information. No clear system connecting effort to results. You have everything it takes. You just do not have a clear system yet. That is what paid subscribers get. Every system, playbook, prompt, and template. All inside the Premium Vault. All for $79/year. That’s $6.58/month. Upgrade now and unlock the Premium Vault worth thousands of dollars. The Premium Vault holds the secret behind posts like this one, including the tools and resources I use to build the one-person business I love. Thanks for reading! Ready for the next step? Let’s crack the growth equation and build a thriving one-person business on your terms! Anfernee
14:52

Monthly Q&A #4: How to Build AI Systems You Can Trust With Real Work

A Q&A post argues the real fix for unreliable AI agents is building a proper system around them rather than chasing a better tool. Six newsletter members had the same core problem: hand an agent a task and you get something mostly right that needs babysitting. The author recommends file-based agents like Claude Code or Codex for research work because they let you control workflow and output format, and notes NotebookLM caps customization. The post shows working setups including a research-agent folder structure, an Obsidian-Todoist planning loop, importing Claude setups into ChatGPT/Codex, and generating banners with Claude.

Notes
Context

Monthly Q&A #4 from The AI Maker (Substack), published 2026-08-02. Theme: readers increasingly want AI systems they can "actually trust." Author's thesis: "Trust comes from reliability. Reliability comes from a proper system built inside the agent." Builds on two recent posts: The Complete Guide to Building Your First Agentic System (That You Can Actually Trust) and Why Your AI Harness Matters as Much as the Model.

Six questions, one problem

All six AI Maker Lab members asked variants of the same frustration:

  • Atit — reliable McKinsey-style research + decks at a real quality bar; deciding between local vs hosted models; considered Claude Code, Codex, NotebookLM, Gemini; asked about build-vs-buy for agents.
  • Mei — agent following a real workflow with less hand-holding.
  • Karen — avoiding lock-in to a single tool; working across Claude Code, Codex, and Hermes.
  • Thomas — stuck between Cursor, Claude Code, and Things.
  • Julia — overwhelmed by saved-but-unread posts.
  • Lee — spent two hours making a newsletter banner.
Author's core argument

The tool was "never going to be the answer." Without "a defined source, a report standard, reference decks, and an output format the agent can see, every deck starts from nothing and lands somewhere different." That produces the babysitting loop: agent returns something "mostly right," you spot weak research, a generic email, a missing task, or an off-topic visual, then explain, retry, and keep watching "because you do not fully trust what happens next."

Atit answer

Prefers a file-based agent (Claude Code or Codex) for research to control the process and final format. NotebookLM is fine "when the main job is working through a large pile of sources," but fails at customization — you "cannot really shape the workflow, the agent, or the output beyond what NotebookLM already gives you." Verdict: "having more control over how you work and what you can generate is more important."

Demoed in the video

Research-agent folder structure, Obsidian–Todoist planning loop, an "AI News Intel Skill" that read 27 newsletter emails, an X-bookmark workflow capturing the week's consumption, importing a Claude setup into ChatGPT/Codex, Glif for newsletter thumbnails/infographics, and a Claude Design banner design system.

Source cuts off mid-sentence at "the rough structure I'd build in my project folder."

Full text · 3,589 chars
If there is one pattern across every question I received this month, it is this: more people are asking how to build an AI system they can actually trust. Trust comes from reliability. Reliability comes from a proper system built inside the agent. That’s why I wrote two posts about it this week: - The Complete Guide to Building Your First Agentic System (That You Can Actually Trust) - Why Your AI Harness Matters as Much as the Model The first question came from Atit, and it was the clearest version of it. He wanted McKinsey-style research and decks, at a real quality bar, and he was trying to work out which tool would get him there. But the tool was never going to be the answer. Without a defined source, a report standard, reference decks, and an output format the agent can see, every deck starts from nothing and lands somewhere different. That is where the babysitting begins. You give the agent a task. It gives you something that looks mostly right. Then you notice the weak research, the generic email, the missing task, or the visual that has nothing to do with the post. You explain the problem, ask it to try again, and keep watching because you do not fully trust what happens next. Six AI Maker Lab members sent questions this month, and all six were a version of the same frustration: - Atit wanted reliable McKinsey-style research and presentation decks, plus a clearer way to decide between local models and hosted ones. - Mei wanted an agent that could follow a real workflow with less hand-holding. - Karen wanted to avoid getting trapped inside a single AI tool and asked how she could work her project across Claude Code, Codex, and Hermes. - Thomas stuck between Cursor, Claude Code, and Things. - Julia was overwhelmed with so many saved posts she had not touched and consumed at all - Lee spent two hours to create his newsletter banner. All of them is different jobs, but the real problem underneath them them is identical. So, in this post, I tackled all of them. In this video, I also showed the working versions behind the folder structure for a research agent, my Obsidian–Todoist planning loop, the AI News Intel Skill that read 27 newsletter emails, the X bookmark workflow that has been collecting everything I’ve consumed over the past week, how to import your Claude setup into ChatGPT/Codex, the Glif process that generates thumbnails and infographics for my newsletter, and how to create a design system for generating banners with Claude Design. Let’s dive in. Atit: How do I build reliable McKinsey-style research and decks? Atit wanted to create deep research and presentation decks with a specific quality bar. He was considering Claude Code, Codex, NotebookLM, and Gemini. He also asked how I would think about the build-versus-buy decision for an AI agent. For research work, I keep coming back to a file-based agent like Claude Code or Codex, because I want control over the process and the final format. NotebookLM is useful when the main job is working through a large pile of sources. But, where it stops working for me is customization. You cannot really shape the workflow, the agent, or the output beyond what NotebookLM already gives you. For a job where I need to define the research sequence, enforce my own report standard, and produce several exact deliverables, I can’t have that ceiling in my workflow. So, the tool you choose matters here, but having more control over how you work and what you can generate is more important. If you want to build a research agent, here is the rough structure I’d build in my project folder:
15:01

Executive Briefing: Which of the 5 Levels of AI Builder Are You, and What It Costs You

An executive briefing argues AI builders should figure out which of five maturity levels they operate at instead of panicking that frontier labs will make their work obsolete. The threat is real — labs keep bundling last year's startup category into this year's product feature — but the author claims the minimum bar for building is simply rising, not the opportunity disappearing. Surviving builders share domain depth more than funding or technical skill, and the five levels are defined by what evidence each rung can show, from a prototype you love to a forecast you can bet on. The piece is largely a framework for a paid briefing.

Notes
Notes: "Which of the 5 Levels of AI Builder Are You, and What It Costs You"

Nate's Substack, published 2026-08-02.

Core framing. Repeated launches from OpenAI/Anthropic make builders fear their work will be absorbed: a platform bundles a "good-enough" product, a model company turns "last year's startup category into this year's menu item," better coding agents speed competitors to the first demo. Stated response options: quit (if the opportunity is gone) or identify "what kind of builder you are today and what you still need to learn" (if only the standard moved). Claim: "The minimum level for building is rising."

Evidence base. ~20 years building; "hundreds of AI builders" interviewed (founders, executives inside large companies, niche-expertise side businesses, first-time builders). Survivors of platform launches share neither funding nor technical skill — they differ in "levels of understanding."

The five levels. Defined by "evidence and operating maturity," explicitly "don't rank intelligence or human worth," and "don't assume every business should become a venture-backed company." Range from "a prototype they love" (level 1) to "a forecast they can stage a bet on" (top level). Builders may enter at level 3–4 via deep market knowledge; five- and six-figure side businesses can be built without the top level.

Level-specific threat. A lab shipping your headline feature "hits level one like a verdict and level four like a data point."

The rungs between. Briefing promises one move per gap: "from customer contact to distribution, and from distribution to a thesis you'll hold for years"; and the labs' blind spot: "What twenty years inside one domain buys you that a frontier training run does not." Details are listed as covered topics, not spelled out in this text.

Context. Labs are moving into surrounding work: Anthropic packaged Claude for small businesses and finance teams; OpenAI pushed Codex into roles and workflows. Advises starting at the bottom "even if you think you're past it."

Full text · 3,110 chars
Every time OpenAI or Anthropic ships something big, I hear from discouraged builders. Sometimes they’ve spent six months on a feature that now appears inside a product with hundreds of millions of users. Sometimes they’re one week into a side project and have already convinced themselves a frontier lab will make it irrelevant before they finish. The details change. The fear doesn’t: if the companies building the models keep moving this fast, is there anything left for me to build? I get it. The threat is real. A platform can bundle a good-enough version of your product. A model company can turn last year’s startup category into this year’s menu item. Better coding agents mean the next person reaches your first demo faster than you did. The labs are moving into the work around the models too. Anthropic has packaged Claude for small businesses and finance teams. OpenAI has pushed Codex into roles and workflows across a company. Each expansion can absorb a point solution. The minimum level for building is rising. That distinction matters. If you believe the opportunity has disappeared, the rational response is to quit. If the standard has moved, the response is to figure out what kind of builder you are today and what you still need to learn. I’ve been building for twenty years. I remember what it was like before Shopify made commerce easier and before cloud services put infrastructure in reach of small teams. It has never been easier to make the first version of something. Easy means more of us get far enough to find where the real work begins. It’s still messy, and the outcome is still open. Over the past few years I’ve talked with hundreds of AI builders: experienced founders, executives building inside large companies, people turning a narrow expertise into a side business, and people making their first useful thing. The ones who survive a platform launch don’t share funding, technical skill, or ambition. They do tend to operate at different levels of understanding. I’ve started thinking about those differences as five levels of AI building. The levels describe evidence and operating maturity. They don’t rank intelligence or human worth, and they don’t assume every business should become a venture-backed company. I have seen people create excellent five- and six-figure side businesses without reaching the last level. A builder can also enter the map at level three or four because they already know a market deeply. This briefing covers: - The five levels, defined by evidence. What a builder at each rung can actually show, from a prototype they love to a forecast they can stage a bet on. - What each launch actually threatens. Why a lab shipping your headline feature hits level one like a verdict and level four like a data point. - The specific move between each rung. The one thing that gets you from customer contact to distribution, and from distribution to a thesis you’ll hold for years. - Where the labs can’t follow. What twenty years inside one domain buys you that a frontier training run does not. Start at the bottom, even if you think you’re past it.
14:48

The $10K Etsy niche nobody is running with Claude yet

A paid playbook pitches one Etsy niche it claims can earn $10K a month by building custom digital products with Claude in minutes instead of days. The pitch argues the audience, an email list, is the real asset, since a single Etsy sale is just a transaction. It cites Etsy's 86.6 million active buyers and Pinterest sending 41% of Etsy's social traffic to hype the window. Mostly promo; thin on specifics.

Notes
The $10K Etsy niche nobody is running with Claude yet

Nature of source: promotional funnel for the Claude Digital Business System (Substack premium playbook). It advertises the system but discloses no actual steps, prompts, or margin math — everything concrete sits behind the paywall.

Numbers cited

  • Etsy: 86.6M active buyers, its "first buyer growth in two years."
  • Pinterest drives 41% of Etsy's social traffic; "75% of weekly Pinners arrive ready to buy" vs 28% on Facebook.
  • Email returns "about $36 for every $1."
  • Subscription includes "3 fresh systems every week."

Promised components of the playbook

  • "Custom-product engine" building premium-priced digital products, "with the margin math."
  • Etsy launch kit that ranks "without keyword-stuffing."
  • "4-prompt content engine" for AI-search-era SEO.
  • "Owned-audience funnel" (email list) to avoid platform dependence.
  • "One-brand structure" merging three streams into one business.
  • "Niche-swap prompt" to rebuild the system for any market; example niche run throughout is travel.

Core thesis

"A single Etsy sale is a transaction. An email list that trusts you is a business, one nobody can deplatform."

The window argument: custom digital products once needed a designer and "days of work"; Claude collapses that to minutes, so the edge goes to whoever ships genuine quality at volume before the category fills.

Caveats

  • Zero method content in the post; body is a CTA only.
  • Offer: 7-day free trial or "50% off this week only" — scarcity framing.
  • Stats unverified and context-free (no dates, no definitions of "ready to buy").
Full text · 2,298 chars
The $10K Etsy niche nobody is running with Claude yet One niche. Three assets that feed each other. Built with Claude to a quality people actually pay for, without the spam that gets you banned Most “make money with AI” guides teach you the same move: flood a platform with junk, cash in for six weeks, then watch the account get banned and the income vanish overnight. This is the opposite, and it is built to still be standing in a year. One niche. Three genuine assets that feed each other. Each earns on its own, and together they build the one thing you actually own: an audience no algorithm can take back. Here is the part those guides bury: the product is not the asset. The audience is. A single Etsy sale is a transaction. An email list that trusts you is a business, one nobody can deplatform. And the ground is wide open right now: ▫️ Etsy just posted its first buyer growth in two years, 86.6M active buyers ▫️ Pinterest sends 41% of all of Etsy’s social traffic, and 75% of weekly Pinners arrive ready to buy (versus 28% on Facebook) ▫️ Email still returns about $36 for every $1, on a list nobody can starve The window is the story. Custom digital products used to need a designer and days of work. Claude collapses that to minutes, which means the edge belongs to whoever builds genuine quality at volume before the category fills up. Right now, it is close to empty. One example niche runs the whole way through: travel. Keep it, or swap it at the end with one prompt. The system is the point. What’s inside the full playbook: ▫️ The custom-product engine that builds a digital product people pay a premium for, with the margin math ▫️ The Etsy launch kit that ranks without keyword-stuffing ▫️ The 4-prompt content engine built for the AI-search era, where depth is the only moat left ▫️ The owned-audience funnel where you show up as yourself, over astroturfing your way to a ban ▫️ The one-brand structure that turns three hustles into one business ▫️ The niche-swap prompt that rebuilds all of it for any market One subscription unlocks every system Plus 3 fresh systems every week. One product line from this covers the subscription many times over. 🧭 The Claude Digital Business System Get the full playbook below 👇 Try premium free for 7 days. Or get 50% off this week only.