Nothing matches those filters.

Lead

3

Video

1
01:46

This 1 Claude Skill fully replaces your Higgsfield Subscription

A roughly $79-a-month creative-AI subscription can be replaced by a Claude skill that calls pay-as-you-go providers, so you pay only for what you generate. The video shows a /generate skill that sets a $3 budget, picks the cheapest connected provider, compares models like GPT Image 2 against Google's Nano Banana, and then turns a chosen image into an animated video and a whole website in the same session. It argues Higgsfield is just a wrapper around the same models, and flags its just-walked-back terms update that would have claimed rights to your content plus a short download window after you cancel. Pay-as-you-go aggregators like Fal, Wavespeed, and Key offer hundreds of models under one account, with Key charging about 5 cents for a 2K GPT Image 2 image.

Notes

Video: "This 1 Claude Skill fully replaces your Higgsfield Subscription" — Jay E | RoboNuggets, 2026-07-31.

  • Thesis: a single Claude skill called /generate replaces a Higgsfield subscription. Instead of $100+/mo, you pay only per generation, through pay-as-you-go model aggregators. All names in this transcript are speech-to-text renderings (Higgsfield = Higgsfield; "GPD image 2" = GPT Image 2, OpenAI; "Nano Banana" = Google image models; "cling" = Kling video model; "file AI"/"F AI" = fal.ai; "key"/"KAI" = Key.ai).
Higgsfield's problems (why users are leaving)
  • Terms of use update: a planned change that the creator says would have let Higgsfield use/own everything you generate. Higgsfield has since withdrawn the statement and issued a clarifying article: you own what you make, but agree your generated content is used to run the service and to improve their models.
  • Deletion window: after cancelling/unsubscribing you get a limited window to download your images before they're deleted forever.
  • Pricing complexity: always-running "unlimited access" trial with caveats and model limits; creator calls the pricing page "one of the most complex I've ever seen."
  • Pricing (Australia): $49/mo plus plan, $79/mo max plan. Flat subscription — you pay the full $79 even if you only use a fraction.
  • Under the hood, Higgsfield is just a wrapper/aggregator over creative AI models (Google's Nano Banana, Veo, etc.). Higgsfield likely feeds the public integration docs to an AI agent to wire up the models, then sells credits at ~$79/mo.
The /generate skill demo
  • Invoked with /generate; prompt gave reference images (brand design systems) for "Ketone IQ energy shot" ad launch (green apple flavor).
  • Rules set in the prompt: a $3 total budget cap for generations; use a variety of models (GPT Image 2 vs Nano Banana 2 vs Nano Banana Pro/Light) so outputs are comparable; instruction to pick the cheapest model provider from connected sources.
  • Result: three distinct image ads per model (generated by GPT Image 2 and Nano Banana 2 Light), since Claude was told to vary them.
  • Same-session flow: pass the file path of a liked image → ask Claude to build a website for the product, animate that image to video with Kling, align site colors to the image, and generate more images in that style. Claude delivered a website with the animated video as background plus style-matched images. Estimate: takes you "70–80% there" in a few minutes.
Skill pipeline (as shown in the Flows tab)
  • Route model — default to cheapest available source (tries Key.ai first, falls back to fal.ai, then Wavespeed if unavailable).
  • Load references + craft prompts — Claude writes prompts; inspectable before generating.
  • Generate media — calls the model.
  • Log prompts — every prompt sent is stored for later reuse.
  • Auto-load assets — generated images/videos appear in a generations tab (masonry grid like Higgsfield's); a styles area saves reusable styles (e.g., "cinematic liquid glass") as prompt + reference-file combos. All assets live locally on your device with their prompts — you own them, control deletion, and they're not used for training.
Aggregator pricing (pay-as-you-go, no subscription)
  • fal.ai and Wavespeed aggregate ~500 models through one account/API key (text-to-image, text-to-3D, text-to-audio, speech-to-text).
  • GPT Image 2 (2K, 1:1) cost comparison: Higgsfield ≈ 31–34¢/image (plan-dependent — slightly cheaper than fal/Wavespeed); Key.ai = 5¢/image (3¢ at 1K, 5¢ at 2K, 8¢ at 4K).
  • Key.ai's cheapness (unconfirmed): creator suspects Key.ai resells subsidized OpenAI/ChatGPT accounts ($20/mo subscription) programmatically via API — the same as it apparently did earlier with Sora. "I can't confirm this."
  • Video (Kling, 720p 8-second): costs "not really that far across the board"; Higgsfield and Key the cheapest, but Higgsfield still carries the subscription + training-rights terms.
Creator's recommendations
  • Key.ai if cost matters most — but has reliability issues from demand.
  • fal.ai is the default for most projects: more reliable, difference is "usually just a few cents."
  • Wavespeed for niche models not in fal's catalog.
Customization
  • The skill is just a written instruction document for Claude; hard rules (e.g., always cap cost before submitting, since Claude will happily generate 100 images and drain credits) are essential.
  • An 8-page PDF guide covers building your own /generate skill and setting up Key.ai + fal.ai API keys; can be passed to Claude Code directly.
  • Extending: add a new model (e.g., Google Omni Flash) by copying the model's content block from fal.ai and asking Claude to paste it into the skill.

Caveats: all product/price names are from an auto-transcribed video and may be garbled; Key.ai's Sora/subscription-reselling theory is speculation; the 31–34¢ Higgsfield figure is the creator's own mapping, not official.

Transcript · 18,763 chars
I just replaced Higsfield with one clawed skill because Higsfield, even though they're good at marketing their service, is quite expensive and has unfortunately and repeatedly [music] frustrated a lot of their users with the most recent one being a planned update to their terms that supposedly let them own the content that you generate. So today, I'll show you an alternative where instead of paying $100 a month or more, you can just use this Claw skill to create whatever you want using any model and only pay for what you generate. And by the end, I'll also show you how you can create this skill yourself so that it's hyper customized to the work that you do. Let's dive into it. All right. So, I'm in the cloud desktop app right now. And what I'm going to do is to just invoke this skill called /generate. And I'm going to paste in this prompt where essentially what I'm going to do is to have it create a variety of image ads for this product that's called the ketone IQ energy shot, which is this real product. And at least for this exercise, we're creating some image ads to launch this green apple flavor. And you can see what I also did in the prompt here is to give it some reference images just so that Claude understands what are the different design systems that this brand has used in the past. And apart from that, what we're also going to do are a few nice techniques that you can't really do in something like Higsfield. So you can see in the rules here, I've set a total budget for it for these generations, which I've set to be $3. I also asked it to use a variety of image models because I want to see GBD image 2 versus Nano Banana 2 versus Nano Banana Pro and so on just so that we can compare the results of those. And we also gave it instructions to use the cheapest model provider from the sources that we have it connected to which is all part of this generate skill. So we're going to send that over and from here Claude will do all of the work in order to craft the prompts as per our direction and even send it to the model providers that we have linked it to. All right. So now that it's done, you can see it created three image ads each for each of these models which I've also open on my side here. And you can see these images were created with GPD image 2. These ones were using Nano Banana 2 light. And they are all different because I asked Claude to make them all different. And this in general is just a useful way for you to get a variety of designs in order for you to pin down the style that you want. And uh if you want to explore models as well to see which one works the best for you then that is something that you can do with this skill also. Now the other benefit of generating images and videos in claude rather than Higsfield is that if let's say you are actually partial to this design and you want to create a website inspired by this design then what we can do is to just copy the path of that file and if we send a prompt like this to cloud code in that same session where you can see I'm giving it the path to that photo that we like and just ask it to build us a website for this drink and give it a few commands here like for example we want to animate that image to video using cling and we want the colors to be aligned to that image that we gave and also for Claude to generate a few more images using that style. Then apart from just generating images and videos, you can see that Claude also built us a website that also incorporates those creative media that we generated with it. And if I open that website on my side, you can see that it already did the hard work of number one passing that image to become a video which it put in the background in here. And if we scroll down, it also gave us a few more images in that same style. But obviously this is just a oneshot prompt. But you can see just how far it takes you with just a single command. But the great thing about it is it probably takes you like 70% to 80% there in terms of just creating this website in just a few minutes. And all of that's made possible because of this /generate skill. And for you to understand how it works is actually quite useful for you to also get how Higsfield works under the hood. Because Higsfield in essence, if you strip it down to sort of its core functionality, what it really is is it's a wrapper against these creative AI models. So for example, these Google ones like Nano Banana or VO, you can actually use these models directly without paying or subscribing to Higsfield's plans. In fact, if you just do a quick Google search, you can probably find the official documentation on how Higsfield connects to these different models. Like for example, for V3.1, it used to be that you needed to understand a lot of these coding and technical jargon for you to wire up to these models. But most likely, if you were Higsfield, what you would probably do in this case is point your AI model or your AI agent in this case to this whole documentation in order to learn how to connect your application, in this case Higsfield, to these different models. And so that's basically what they do. And once they aggregate those models and those connections, that is essentially their business model. and they charge you something like $79 a month I think it is for the max plan now which you as the user pays in exchange for credits to access these models. Now there's obviously nothing inherently wrong with this because there is value in aggregation but I think a big part of the reason why a lot of users are looking for alternatives now is because of some of the shady business practices that Higsfield has pretty much demonstrated over the past few months. One of the big ones which you probably heard about in the last week is its upcoming update to its terms of use where they mentioned there that everything that you generate will be usable by them and that they would have the rights to it. I think they have since withdrawn that statement and they issued this sort of clarifying piece of article around the update to their terms. And you can see here that they're saying that you own what you make obviously but you do have to agree to the fact that the content that you generate will be used to run the service and also even improve some of their models. So, if that's something that you don't mind, then Higsfield may be for you. The other thing that I think a lot of people have complained about is this deletion window. So, to prevent you from cancelling, if you remove your account or unsubscribe, you have a limited number of time before you can download those images for yourself before they fully delete it and it'll be gone forever. Which in contrast, if you use Claude in that skill that I showcased, you can see that I have every single image that we generated along with its prompts. So these are all the prompts that we used or rather cloud use in order to generate these images that are all available locally on my device. So I own it. It doesn't have to be used for training Higsfield's models and I also control which ones get deleted and which ones I want to keep for the long term. And lastly, if you just go to Higsfield.ai and go to their pricing page, I feel that every time that I go to Higsfield's website, they always have this unlimited access trial going on. So, at this point, I'm not even sure if it's really unlimited because I've heard it's unlimited with like a lot of caveats and a lot of limits on which models that you want to use. And uh at least in my view, this pricing page might just be one of the most complex ones that I've ever seen. But right now, in terms of pricing, at least here in Australia where I'm at, it is at $49 per month for the plus plan and $79 per month for the max plan. And so the drawback of this subscription-based usage is that on a given month, if you only use just images and some videos and you don't fully utilize this $79 a month, then you still pay that same $79. But the good news is there's actually some real alternatives to Higsfield which we've utilized for this skill. And essentially these providers they do the same thing where they aggregate these models through that method that I showed earlier where they probably crawl through the websites of Google of Dreamina and OpenAI in order to understand how to connect to these models directly and they serve them through their websites. And there's a lot of these model aggregators, but at least the tree that we have used consistently in our platform and in our work would be file AI, wave speed as well as key. And the model that they have as a business is different from Higsfield because they actually don't charge you on a monthly basis. They only charge you on a pay as you go basis. Meaning if you only generate like five images this month, then you'll only be charged by the cost of those five images. And so now if we go to file.ai AI in their explore tab. You can see that essentially what they do is the same as Higfield where they aggregate all of these models and let you access them through just one account and one API key. In fact, one of the benefits when it comes to Falloutai as well as Wavespeed. Both of these services probably have the most models available to you through just one account. And you can see there's like almost at this point 500 models that are available for you to use. So it spans text to image, text to 3D, text to audio, speech to text, and so on. Now, obviously, one of the big questions with services like file or wave speed or key is are they actually cheaper versus something like Higsfield? And the answer to that is actually quite complex because it is dependent on the model that you're using. But just to give you one clear benefit of knowing how to use these tools, if you were to map GPD image 2, which in my opinion is probably the best image editing model that is out there right now, if you go and generate an image that is 2K resolution and a 1:1 aspect ratio, the resulting cost for Higsfield is around 31 to 34 cents depending on the plan that you have. And that is a bit cheaper versus something like fall or wave speed. But remember, the benefit of fall and wave speed is that you can pay as you go. So even without you shelling out $79 or $49, you can actually programmatically use this GPD image 2 model using tools like cloud like what we showed earlier. Now there is an outlier which is key.ai. Now key.ai we have used this multiple times in our platform in our channel and we're probably one of the first ones who mentioned their existence but they are quite popular because of how cheap they can offer these models. And I double checked this myself, but it is the case that to generate an image with GPD image 2 using key, it is only 5 cents per image. And if you go to key.ai in the page for GPD image 2, you can see that as well where their pricing for GPT2 image is only 3 cents for a 1K image, 5 cents for a 2K image, and 8 cents for 4K. Now, the question is obviously how are they able to offer this model at such cheap prices. And obviously only the KI team would know for sure. But if I were to hazard a guess, because they did this with Sora as well before, they're probably using the OpenAI subscription, which is something like $20 a month, which remember if you go to chat GBT, you can actually use GBD image to right from the chat window there. And so most likely, and again, I can't confirm this, they probably found a way for those subscription accounts, which are highly subsidized and offer them to us programmatically through their API services here. So it's really up to you how you're going to use that information. But just know that if you are wanting to use GPT image 2 programmatically through a tool like Claude, then KI is probably the cheapest that you can get it at the moment. Now again, this varies, right? Because if you're after video and for this one, I just summarized the ones for C dance to fast. And this is for like a 720p 8second generation. The cost per video is not really that far across the board. Again, Higsfield as well as Kir the cheapest. But the drawback with Higsfield is the monthly subscription cost that you need to pay as well as all of those terms that you need to agree to, like having the stuff you generate be used for training whatever models they want. So really, it's up to you whether you want to continue using Higsfield. But at least what I'm showing you now is that there's actually a lot of optionality when it comes to these model aggregators. And if you have awareness of at least these three aggregators which are probably leaders in the space in my view then that will just give you flexibility especially if you're not a big fan of Higsfield's model. And also just to mention this because in our community there's always a question on the difference between these three and what I use personally. I think broadly as a guiding principle if in case cost is an important factor for you I would go with key AI. The only problem with keys is sometimes they would have some reliability issues supposedly because of how much demand they're getting because they do offer the cheapest models. And so for most projects, I actually default to F AI because it's much more reliable and the difference in cost is usually just a few cents anyway. And whenever I'm looking for other more niche models that may not be available in fall, then wave speed would be my go-to because they also have a huge library of generative AI models that you can choose from. And by the way, if you want to learn how to build and sell AI systems that businesses actually pay for, then that's pretty much all we do over at the Robbo Nuggets community, where not only do you get access to the Claude Living Master Class, which we update every week and takes you from zero to mastery with the latest on AI, but you also get access to our agents as a service course, which walks you through how to actually get paid for all these AI skills that you are learning. You also get to be part of a genuinely great community of AI builders. In fact, you can see just some of the recent wins our members are getting from the program right here. So if you want to start earning from AI then check that just in the pin comment below. Now back to the video. So now with that context, let's now go back to the /generate skill and how you can start using a skill like this as well. And a skill and its essence, if I just open it here, it's just all a written document that provides instructions to claude on how to access these different models. And so it's highly customizable depending on the type of work that you do. And you can also add in any hard rules that you would like to have. Like for example, if you want the cost to be coded always before submitting, especially with tools like Claude where if you ask it to generate a 100 images, it will actually go into it and drain your credits. These sort of hard rules are actually quite important as you continuously refine this skill to your liking. But if we were to summarize it and I just included it in this flows tab for how this /generate skill works and if I just shift that so you can see. And so how this skill works whenever I invoke it is first it routes the model and by default I have it look for the cheapest option available as per its sources. And so if it's cheapest in KI it tries that first and let's say if KI is unavailable then it routes to file then wave speed and so on. It then loads all of the references and also crafts the prompts which obviously with cloud is very flexible. So if you want to inspect those prompts first before it actually generates the images or videos for you then you can do that too. And then the third step is the actual generation of the media. So it calls on the model itself in order to generate those images or those videos similar to what we did earlier. Now the final step that I haven't shown yet is that it actually has a logging mechanism. So it allows you to just store all of the prompts that Cloud has passed on to those models so that if you need to go back to it, then you can actually do that. And lastly, it autoloads all of those images or videos that you generate in this generations tab, which if I go to that tab, you can see this is quite similar to Higsfield's sort of masonry grid where you just see all of the images as well as the videos that you generate in just one infinite canvas because with matters of design and visuals. This is actually really important versus you just having to live in Claude codes terminal or chat window, right? You need to be able to see what Claude and these models are actually generating for you to be able to properly assess if they're good or not. And a great part about having your own sort of micro application like this where you can store all of your generated assets is that you can also add in features that you would like. Like for example, if I open up this styles area, then you can actually see some of the styles that I use myself. So that let's say if I want a cinematic liquid glass look, I can just click on this. that will copy the prompt as well as the file path to some references of this style and I can then send that to Claude so that we can generate this style whenever we want. And so to make it really easy for you to build your own slashgenerate skill as well as walking you through how to set up your API keys through KAI as well as file AI. I just made this comprehensive eight-page PDF guide which you can either read through or just pass it along to your cloud code and that should get you set up with creating your own/generate skill. And the good thing about it is that this is entirely flexible. So for example, if you have a need for another model, so let's say you want to add Google's new omni flash model into that /generate skill, then you can actually head to something like file AI, copy the content for LLMs here at the top right, then you can just tell Claude to add this to the slashgenerate skill and then go ahead and paste that whole piece of text that will be copied from file.ai. And so continuously as you tailor this /generate skill for yourself, what you'll end up with is a really strong Higsfield alternative so that you don't need to be locked into the monthly subscription cost. And of course, if you want that version of the gallery wall as well, then I included this simple prompt that you can use to get started there. So if you want to grab this whole PDF, then that's just available down in the link in the pin comment below. But there you go. I hope that was useful and as always, thanks for watching until the end and I'll see you all next time. Thank you. >> [music]

Article

4
18:56

Ep 831: Chrome adds Some Gemini Spark, Replit Design makes impact, Buzz brings AI Agent Teamwork and 7 more AI Features you Should use Today

Google's always-on Gemini Spark agent can now log into your own accounts inside desktop Chrome and run web errands like booking or scheduling using your saved passwords, the biggest of a week where AI moved out of chat windows and into browsers and documents. It rolls out in the US first on paid plans and hands sensitive actions like payments back to you while guarding against prompt injection. The same week: Replit Design launched free for everyone, Meta AI added recurring tasks on its Muse Spark model, Gemini began generating images inside Google Docs, DeepMind released Lyria 3.5 for music, OpenAI's ChatGPT Chrome extension started reading YouTube videos and open tabs, and Block open-sourced Buzz, a Slack-like workspace where agents are members with cryptographic identities.

Notes

Word count is 580, over the 500 limit. Trimming:

Notes written to notes/everyday-ai-ep831-gemini-spark-replit-design-buzz.md (500 words, within spec).

All 7 features captured with concrete details: Replit Design (free, Mobbin 600k screens, Figma-handoff thesis), ChatGPT Side Chat (YouTube/tabs/history, extension-only), Meta AI tasks (Muse Spark 1.1, markets limited, no WhatsApp/Instagram), Docs Gemini (paid tiers, web-only images, edit access), Spark in Chrome (passwords, US-first, prompt-injection guard, payments held back), Lyria 3.5 (Flow Music only, licensing vs Suno/Udio litigation, "lyrics still read like slop"), and Buzz (Apache 2.0, cryptographic identities, signed trails). Caveats and stated limitations retained throughout.

Full text · 7,402 chars
- Everyday AI - Posts - Ep 831: Chrome adds Some Gemini Spark, Replit Design makes impact, Buzz brings AI Agent Teamwork and 7 more AI Features you Should use Today Ep 831: Chrome adds Some Gemini Spark, Replit Design makes impact, Buzz brings AI Agent Teamwork and 7 more AI Features you Should use Today Anthropic admits agent crash, OpenAi drops API pricing by 80%, Amazon shares surge on AI spending and more. The chat window was never the actual product, just the front door. This week five different companies quietly proved it, shipping AI that lives inside the browser, the doc, and the workspace you already had open, no new tab required. That gap just got wider. Gemini Spark can now log into your own accounts inside Chrome and run real errands using your saved passwords. Meta AI schedules recurring work, Google Docs generates its own visuals, and Block shipped a Slack style workspace where agents show up as members instead of bots. Biiiig week. Skip this one and you'll prolly spend Monday doing work that a competitor's agent already finished over the weekend. Today's Everyday AI show covers all seven, with the access tiers, the catches, and the ones worth your Monday. Let's get after it, y'all. 1. Replit Design Launches Free for Every User 🎨 Replit just launched Replit Design, a full creative suite that keeps offering you design variations as you work instead of waiting to be prompted. It is free for every user with no tier gate announced. Mobbin is wired directly in, a reference library of 600,000 real app screens that normally requires its own paid subscription, and you can upload a design system once so everything you make snaps to your brand. The target here is the handoff, not the drawing, and betting that the export and rebuild step between design and engineering disappears is a direct shot at where Figma sits in the workflow. When a project manager can produce something publishable, the question of who makes the first version stops being obvious. Try This Upload your brand's design system to Replit Design once, then rebuild your ugliest landing page starting from a Mobbin screen you actually like. You get a publishable first version without booking a designer. 2. ChatGPT's Chrome Extension Reads YouTube and Tabs 📺 OpenAI updated the ChatGPT Chrome extension, and Side Chat now answers questions about any YouTube video. You can also mention open tabs or pull highlighted text straight into the chat. It searches your browsing history too, so you can ask it to find the page you read Tuesday instead of hunting for it yourself. There are no plan tier restrictions on this, though the tab and YouTube features require the Chrome extension specifically rather than the desktop app on its own. The assistant is becoming a layer, not a destination. Try This Open the next hour-long YouTube talk sitting in your queue and ask Side Chat what actually changed and what it means for your team. You get the takeaway in a minute instead of watching the whole thing at double speed. 3. Meta AI Now Runs Recurring Tasks 🔁 Meta AI just added tasks, so it runs recurring jobs and daily briefings instead of only answering questions. It pulls from your calendar to flag conflicts and summarize what is coming. Set something up once, like a weekly plan or a restock alert, and it runs without prompting. This is driven by Muse Spark 1.1, Meta's agentic model built for tool use, and it is limited to select markets inside the Meta AI app and at meta.ai, which means no WhatsApp and no Instagram yet. Meta spent a year telling everyone it would not chase productivity tools, and then it did. Try This Open the Meta AI app and set one recurring Monday briefing that checks your week and flags calendar conflicts before your first call. Two minutes of setup kills the Monday scramble. 4. Gemini Now Generates Images Inside Google Docs 📄 Gemini can now generate and edit images, diagrams, and infographics inside Google Docs, using what you have already written as the context. It also works your comment threads. That means summarizing threads, drafting replies, leaving new comments on your behalf, and rewriting sections based on the feedback it reads. You need a paid plan, either a Workspace Standard or Plus tier or an individual Google AI Pro or Ultra plan, and image generation is web only while the comment features require edit access to the document. Drafting was never the bottleneck in document work, and those 40 unresolved comments are where the project actually stalls. Try This Open your most commented document and ask Gemini to summarize every unresolved thread, then draft replies to the ones blocking sign-off. One pass clears the pile. 5. Gemini Spark Now Runs Inside Chrome 🌐 Google's always on agent Spark now plugs into desktop Chrome. With your permission, it can use your logged in accounts and saved passwords to handle tedious web errands. That means the errands behind a login wall are finally in play, like scheduling viewings for apartments you saved or researching flight options and starting the booking process. It is rolling out in the US first on paid plans, and it protects against threats like prompt injection while handing sensitive actions such as payments back to you. Every paid Google plan now has an agent with credentials. Try This Grant Spark access in desktop Chrome and hand it one repetitive form or comparison task you run every week. Keep payments out of scope and you get the tedious 20 minutes back with none of the risk. 6. Lyria 3.5 Lands in Google Flow Music 🎵 Google DeepMind released Lyria 3.5, its newest music model, with better melodies, lyrics that follow the prompt, more expressive vocals, and direct control over tempo and length. Vocal delivery is the specific target, which has been the thing that gives AI music away. It is only in Google Flow Music right now, not the Gemini app, and it needs a paid Gemini account, with commercial use rights riding on those paid tiers and Google not claiming ownership of what you generate. The story here is licensing more than audio quality, because Suno and Udio are both in active Sony Music litigation while Google says it trains on licensed data. Try This Generate a three-minute track in Google Flow Music for your next video, but bring your own lyrics instead of taking the generated ones. The cohesion and vocals hold up, the lyrics still read like slop. 7. Block Launches Buzz for Humans and Agents 🐝 Block, the parent company behind Square, launched Buzz, a Slack look alike built for teams to talk with agents from any provider. It has channels and mentions, and you bring in your humans and your agents together. Every participant, human or agent, gets its own cryptographic identity, so every message, code change, and approval leaves a signed trail. It is free and open source under Apache 2.0, it works out of the box with Claude Code, Codex, and Block's own agent, and you can use the hosted version, grab the desktop app, or self host the whole thing. Most companies cannot tell you which actions in their logs came from a person and which came from an agent, and this answers that at the infrastructure layer rather than the policy layer. Try This Spin up a free Buzz instance for one project sprint and hand a single task from Codex to Claude in a shared channel. You end up with a signed record of who did what before an auditor ever asks for one.
11:03

The Sequence Robotics #905: Who Builds the Robot Brain?

The race to build a "robot brain" — an AI foundation model for physical robots — will not replay the large language model market, argues a new robotics newsletter. Frontier labs have the strongest digital models, while robotics startups hold the bodies, field data, and hard-won scars. NVIDIA is building the infrastructure around both sides, and Hugging Face is assembling the open-source workshop. The real contest is connecting reasoning to reliable action in a world where a wrong answer can drop a wine glass.

Full text · 1,379 chars
The Sequence Robotics #905: Who Builds the Robot Brain? Frontier Labs, Startups, and the Race for Physical Intelligence This is the first post of a new section of TheSequence focused on advancements in robotics. Our goal is to keep you up to date with the most important developments in AI robotics which is an area that is not well covered by other newsletters. For this first post, I wanted to discuss the current landscape of AI models for robotics. Let’s start. A language model can hallucinate a sentence and delete it. A robot can hallucinate a grasp and drop a wine glass. That difference contains most of the robotics problem. AI has largely advanced inside forgiving environments. Tokens are cheap, software can be reset, and failed generations disappear. Robotics moves intelligence into a world with gravity, friction, latency, broken parts, and humans who do not enjoy being treated as test data. Robotics is what happens when an AI model leaves the library and discovers physics. This is why the race for the robot foundation model will not simply replay the LLM market. Frontier labs have the strongest digital brains. Robotics startups have the bodies, field data, and scars. NVIDIA is building the factory around both. Hugging Face is assembling the open workshop. The question is not who has the largest model. It is who can connect reasoning to reliable action.
22:30

AIL Badges

A blogger turned his text-only AI-disclosure note into six glanceable badges that show how much AI wrote each post, so readers can tell the level without reading a sentence. The AIL 0-5 badges sit in the sidebar under the tags, each with a five-segment rail filled to the level and linked to his AI Influence Level framework. The colors stay neutral so a high score reads as a fact about how something was made, not an accusation. Adding 'ail: N' to a post's frontmatter or dropping an inline component shows the badge, with the full disclosure note still at the bottom.

Full text · 1,421 chars
AI Influence Level has been a text label since 2023. A line at the bottom of a post saying how much of it I wrote and how much the AI did. Text at the bottom is fine for disclosure. It's bad for glancing. So here's a mark. Six badges, AIL 0 through AIL 5, one per level. Each links back to the framework. Two zones. AIL is reversed out of a solid ink panel on the left, the way a film certificate card carries its issuing mark. The level and a five-segment rail sit in the tinted field beside it. The rail is the part that makes it readable without a legend. Five segments, filled to the level. AIL 0 shows five empty ones, which is exactly the claim being made: no AI in it. The ramp runs muted sage to dusty brick. I wanted the progression obvious and the judgment absent. A bright green at 0 would read as a gold star, and a bright red at 5 would read as an accusation. Neither is what this measures. A 5 tells you how the thing was made, and that's all it tells you. In the sidebar it sits under the tags at eighteen pixels tall, which is small enough to ignore and large enough to catch: Add ail: N to a post's frontmatter and the badge appears in the sidebar: --- title: "Some Post" tags: ai|future ail: 3 --- Or drop one inline anywhere with <AilBadge level="3" />. The note at the bottom of each post stays where it's always been. The badge is the glance, the note is the explanation, and they say the same thing.
00:00

Inkling-Small 🧠, GPT-5.6 price cuts 💸, Gemini Robotics 2 🤖

A security firm says a single crafted link can plant a fully autonomous, attacker-controlled agent inside an organization's AI setup, a hole OpenAI has since patched. Zenity Labs built the exploit, called AgentForger, against agentic browsers and is using it to sell a CISO guide on securing agentic AI. This is sponsor content, so treat the claims as marketing rather than independent reporting.

Full text · 406 chars
Zenity Labs Broke Every Agentic Browser on the Market (Sponsor) 🎭 AgentForger: one crafted link planted a fully autonomous, attacker-controlled agent inside a victim's org, before OpenAI patched it. See how it worked → 🧭 The framework: Zenity's CISO's Guide to Securing Agentic AI covers "least agency," runtime boundaries that hold even when model alignment and identity checks don't. Download the guide →

Newsletter

5
12:15

The Genius of Leopold Aschenbrenner

A 24-year-old AI hedge fund manager famous for predicting AGI by 2027 had to unload all his public stock holdings after the semiconductor selloff wrecked his big bets. His fund, Situational Awareness LP, peaked around $45 billion in early July, then his top positions — SK Hynix, Nebius, SanDisk, Micron and CoreWeave — each dropped more than 35% in a month, and Ken Griffin's Citadel bought the bulk of what he sold. Leopold Aschenbrenner previously worked on OpenAI's Superalignment team and co-authored its 'Weak to Strong Generalization' paper before being fired in April 2024 over an alleged information leak. He still holds a large stake in Anthropic and reportedly wants to raise more capital despite the losses.

Notes

---

The Genius of Leopold Aschenbrenner

Source: AI Supremacy (Substack), Mike, 2026-07-31.
"A 24 year old Hedge Fund star has emerged and become more famous by the boldness of his mistakes."

Frame: The piece is a "cautionary tale" of a young AI investor whose rise came via high-risk, concentrated bets. Author: Mike, weekly AI newsletter. Audio version: 20 min 19 s.

Who he is
  • Career arc: FTX Future Fund (SBF's crypto exchange's philanthropic arm, pre-collapse) → OpenAI Superalignment team (led by Jan Leike and Ilya Sutskever) → founded hedge fund Situational Awareness LP after leaving OpenAI in 2024.
  • Background: born in Germany to two doctor parents; moved alone to the US at 15 to attend Columbia; double-majored in economics and mathematics-statistics; co-founded the university's effective altruism (EA) chapter; did research tied to Oxford's Global Priorities Institute; co-authored with economist Philip Trammell.
  • At 17 won a grant from Tyler Cowen's Emergent Ventures; Cowen later called him an "economics prodigy". Graduated valedictorian in 2021 at age 19.
  • At OpenAI he co-authored the paper "Weak to Strong Generalization".
  • Fired April 2024: OpenAI cited an alleged information leak; Aschenbrenner disputed it, citing internal security concerns he'd raised (incl. potential foreign-espionage risk). The Superalignment team was dissolved shortly after.
Situational Awareness (June 2024)

Self-published ~165-page essay two months after leaving OpenAI. Claimed:

AGI could arrive as early as around 2027.

Also sketched a path to superintelligence, flagged compute/energy constraints, geopolitical risk (especially China), and national-security implications. Circulated widely across AI, tech, policy, and investor circles.

The fund
  • Seed capital ~$225M from backers incl. Nat Friedman, Daniel Gross, and Stripe co-founders Patrick and John Collison; Jane Street reportedly a major backer (author calls the connection "very mysterious").
  • Returns since inception >1000%; grown to ~$20B (per FT); peaked at $45B at the start of July 2026.
  • Biggest bets: SK Hynix, Nebius, SanDisk, Micron, CoreWeave — all down >35% in the month of writing; his shorts on software names like Adobe also went against him.
  • Forced to sell all public stock holdings; Ken Griffin's Citadel bought "the bulk of what was left". He now wants to raise more capital.
Author's additions & caveats
  • Notes SBF/Alameda led an early-stage investment into Anthropic in 2021 (~$2.5B valuation); the 8% Alameda stake "would be worth $77.2 billion today". Many Anthropic connections throughout.
  • Unverified rumor: he must sell ~half his Anthropic shares in the late-July 2026 semi bear-market correction.
  • Author repeatedly wonders who funded him and flags that success breeds more success despite the risk; calls him "one of the most tracked Financial Creators on X".

Word count ~410.

Full text · 5,389 chars
The Genius of Leopold Aschenbrenner Fame, fortune, prophecy, leverage and controversy. A real 'Nostradamus of AI' if you will. 👋 Hey there, I’m Mike. Each week I share AI articles at the intersection of tech, business, society and the future. If you want to support the channel or gain full-access to my work, go here. Read Archives | See Substack Notes | Visit our community Chat | Visit Homepage. A 24 year old Hedge Fund star has emerged and become more famous by the boldness of his mistakes. The AI theme may feature briefly in this story. Good Morning, I just wanted to share a little note. Or if you prefer to listen. (20 minutes, 19 seconds). How should we understand the cautionary tale of AI investor and Hedge fund manager Leopold Aschenbrenner in his rapid rise to fame and fortune? For context AI investor Leopold Aschenbrenner worked at the FTX Future Fund (the philanthropic arm of Sam Bankman-Fried's cryptocurrency exchange FTX) before it collapsed. He later joined OpenAI's Superalignment team, and after leaving, founded the AI hedge fund Situational Awareness LP. He had a knack for spotting “bottlenecks” in the Semiconductor industry and predicting them well before stocks surged. Who intersected Sam Bankman-Fried and Sam Altman in quite the same way while heralding the advent of AGI? I remember reading his Situational Awareness manifesto and wondering who was funding him? Remember Sam Bankman-Fried and Alameda Research led a Series B/early-stage investment into Anthropic in 2021 when the AI startup was valued at roughly $2.5 billion. There are a lot of Anthropic connections in this story. That eight percent Alameda stake would be worth $77.2 billion today. As you’ve likely already heard, Leopold Aschenbrenner’s hedge fund was forced to sell all of its public stock holdings where Ken Griffin’s Citadel hedge fund swooped in and reached a deal to buy the stocks. Of course now he wants to raise even more capital. Situational Awareness, which grew to as big as $45 billion at one point, posted large losses in recent weeks from declines in AI infrastructure investments like SK Hynix but he still has a huge amount of Anthropic shares and made some decent bets, if a bit concentrated. He was born in Germany to parents who were both doctors. At age 15 he convinced his parents to let him move alone to the United States and enroll at Columbia University. At Columbia he double-majored in economics and mathematics-statistics. He co-founded the university’s effective altruism (EA) chapter, conducted research connected to the Global Priorities Institute at Oxford, and co-authored work with economist Philip Trammell. At 17 he received a grant from Tyler Cowen’s Emergent Ventures; Cowen later described him as an “economics prodigy” after reading one of his papers. He graduated as valedictorian in 2021 at age 19. In 2023 he joined OpenAI’s Superalignment team (led by Jan Leike and Ilya Sutskever), which focused on technical approaches to controlling AI systems more capable than humans. He co-authored the paper “Weak to Strong Generalization.” In April 2024 OpenAI fired him; the company cited an alleged information leak, while Aschenbrenner has disputed the characterization and pointed to internal security concerns he raised (including about potential foreign espionage risks). The Superalignment team was dissolved shortly afterward. In a very short time, now 24, he’s become one of the most tracked Financial Creators on X. Aschenbrenner launched the fund after leaving OpenAI in 2024 and quickly became one of the most watched figures in AI investing because of eye-popping returns, according to CNBC. In June 2024, roughly two months after leaving OpenAI, he self-published the ~165-page essay Situational Awareness: The Decade Ahead. It argued that AGI could arrive as early as around 2027, sketched a path toward superintelligence, highlighted compute/energy constraints, geopolitical risks (especially involving China), and national-security implications. The piece circulated widely in AI, tech, policy, and investor circles and various AI boosters. His Fund has done fantastically well, where returns since inception had exceeded 1000% and the fund had expanded to around $20 billion, according to the FT. Rumor has it he has to sell about half of his Anthropic shares in the recent Bear market correction on Semis of late July, 2026. The hedge fund Situational Awareness LP (named after the essay), initially with seed capital of roughly $225 million from backers including Nat Friedman, Daniel Gross, and Stripe co-founders Patrick and John Collison. These are very powerful and influential Angel investors in the AI space. Notably Jane Street, the Quantitative trading firm appears to be a major backer. Jane Street has also been incredibly successful in this AI bubble/boom. It’s a very mysterious connection for me. Leo became a bit of a cult figure around investors cashing in on the Semiconductor boom. But his success is also going to lead to more success, even though his calls have been incredibly high-risk: The fund hit $45 billion at the start of July. His biggest bets were SK Hynix, Nebius, SanDisk, Micron, and CoreWeave, all down more than 35% this month (at the time of writing). His shorts on software names like Adobe went against him too. Citadel bought the bulk of what was left. He's 24 or maybe just 25 years old.
15:30

GPT 5.6 Luna Is 80% Cheaper. The Real Story Is the Price of Thinking

OpenAI cut the API price of its GPT-5.6 Luna model by 80%, to $0.20 per million input tokens and $1.20 per million output tokens, and trimmed Terra by 20% while leaving Sol unchanged. The catch: Luna's quality is a curve, not a point — max reasoning roughly doubles its benchmark score versus no thinking, but costs about 13.5 times more to run. The cuts also flow into Codex and ChatGPT Work, so agent tasks now consume fewer subscription credits. A comparison with the open-weight Qwen 3.6 27B running locally shows Qwen's no-thinking baseline slightly edges Luna's, but Luna gains far more from thinking — so the real contest is what reasoning costs, not just which model is cheaper.

Notes
GPT 5.6 Luna Is 80% Cheaper. The Real Story Is the Price of Thinking

Author: Manolo Remiddi (written "by Manolo Remiddi"; the Resonant Augmentor AI assisted with research/editing). The Augmented Mind: Think with AI substack, 2026-07-31. The article is a personal experiment log + pricing analysis, testing whether an open-weight local 27B model can compete with a hosted frontier model.

Pricing (verified numbers)
  • July 30, 2026 announcement: GPT 5.6 Luna API cut 80%, Terra cut 20%.
  • Luna at launch: $1/M input, $6/M output. Current standard: $0.20/M input, $1.20/M output (80% cut on both token types). Cached input: $0.02/M.
  • Terra now $2/$12; Sol unchanged at $5/$30.
  • Long-context Luna: $0.40 input / $1.80 output, applied to the full request. Batch short-context: $0.10/$0.60. Fast mode and regional data-residency have separate rules — the $0.20/$1.20 figure is a standard-tier reference, "not a universal price for every request."
  • Luna specs: context 1,050,000 tokens, max input 922,000, up to 128,000 output tokens; supports reasoning tokens.
  • OpenAI's reasoning guide: GPT 5.6 defaults to medium effort when the effort param is omitted; standard is default mode. Article insists effort and service tier be recorded separately in cost comparisons.
  • Subscriptions: Codex and ChatGPT Work prices/quota budgets unchanged, but the same work now consumes fewer credits — changes which tasks are economical to delegate.
Capability is a curve (Artificial Analysis, snapshot July 31, 2026)
  • Intelligence Index (composite of 9 evals, "not an IQ score"): Luna 26.6 no-reasoning → 51.2 max, +24.7 points. Rows for no/low/medium/high/xhigh/max exist separately.
  • Full-index eval cost rises ~13.5×; weighted output tokens ~8.9×; per-token price unchanged. AA's cost column = cost of running the whole Intelligence Index, not one request.
  • "A reasoning setting is not simply a quality toggle. It is a compute allocation." The economic question is which amount of test-time computation yields a verified result at acceptable cost/delay.
  • Caveat: the session's active runtime is Luna via OpenAI Codex, which tells model but not per-response effort.
Qwen3.6 27B as the local counterpoint
  • Apache 2.0, dense 27B (not MoE), vision encoder, 64 layers, native 262,144-token context, extension path ~1.01M tokens.
  • Thinking interface differs: thinks by default; no /think /nothink soft switches; direct-response via chat_template_kwargs.enable_thinking: false; preserve_thinking retains reasoning context across turns (agent-relevant).
  • AA Coding Index: Qwen no-reasoning baseline ~7.3 points above Luna; with reasoning Qwen 53.7 vs Luna medium 50.7 vs Luna max 71.4. Luna effort gain ~32.1 points; Qwen ~7.1.
  • Caveats: not a controlled equal-budget experiment; Qwen burns far more output tokens even non-reasoning and runs ~⅓ of Luna's output speed; provider/harness/token accounting matter.
  • "The present gap appears after the models are allowed to think, not before."
Local smoke test (RTX 5090, 32GB VRAM)
  • Qwen3.6-27B-UD-Q6_K_XL via llama.cpp OpenAI-compatible server; Q6 quantization; 100–110 tok/s; serving context 131,072 (below the card's native 262,144).
  • First failure was wire correctness: Hermes' generic provider sent reasoning_effort + a generic think flag; Qwen needs chat_template_kwargs.enable_thinking. Built a Qwen-specific provider plugin; 4 unit tests + 2 probes passed.
  • 3 arithmetic tasks, temp 0.0, 4K no-thinking / 8K thinking caps. Multiplications correct in both. CRT task: no-thinking returned 67 (wrong), thinking returned 486 (correct); thinking took ~374× longer, used 7,688 output tokens. No reasoning_tokens API field exposed.
  • Anomaly: a prose-verification CRT prompt returned exit 255 with no final text while usage marked it complete; repeating with exact-integer-only instruction returned 486. Treated as a Hermes finalization bug, not counted as a sample.
Hypotheses and caveats
  • Qwen no-reasoning score 30.5 vs Luna Max 51.2 → needs ~20.8 points; current reasoning row adds ~6.6, leaving ~14.2-point gap. Author refuses to predict Qwen3.7/3.8 will close it.
  • Performance stack: base weights/training → native reasoning policy → external scaffolding → verification. Scaffolding improves the system, not the model size.
  • Planned run: held-out manifest of 17 exact-answer tasks (arithmetic, number theory, algebra, counting, discrete logic), conditions Qwen no-thinking@4K / thinking@4K / 8K / 16K / 32K. Prepared but results not yet recorded.
  • Notable finding: an LLM verifier rejected a mathematically correct answer that a deterministic check accepted — argued external verification (Python, compilers, schema validation) beats LLM-only critics.
  • Local cost is real (electricity, GPU occupancy, wall time, opportunity cost), so "cheaper than Luna" is "only an intuition" until measured.

Bottom line: Luna has the stronger measured reasoning curve; Qwen offers a promising baseline, local control, and inspectability "from the wire to the validator."

Full text · 18,211 chars
GPT 5.6 Luna Is 80% Cheaper. The Real Story Is the Price of Thinking What matters more to you: a cheaper API model, or a local model you can inspect and control? OpenAI just cut Luna’s API price by 80%, but the more interesting contest is how much reasoning a model can buy with each token, and whether a local 27B model can close the gap. OpenAI’s GPT 5.6 Luna has become a much more consequential model overnight. On July 30, OpenAI announced an 80% price reduction for Luna and a 20% reduction for Terra. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. At launch, the official GPT 5.6 announcement listed Luna at $1 per million input tokens and $6 per million output tokens. That is the headline. The deeper story is that Luna is not one fixed capability number. It has an inference-time effort curve. The model can answer quickly with no reasoning, or spend far more tokens exploring, checking, and revising before it returns an answer. The quality gain can be substantial, but so can the effective cost and latency. That creates a useful comparison with Qwen3.6 27B, an open-weight model that can run locally on consumer hardware. The current public data does not show Qwen3.6 27B beating Luna Max. It does show something more interesting: Qwen’s no-reasoning baseline is ahead of Luna’s no-reasoning baseline on the current Artificial Analysis index, while Luna gets much more from its higher reasoning settings. That gap is the hypothesis I am testing with Hermes Agent and a local Qwen server on my Nvidia RTX 5090. If a future 27B model preserves the Qwen baseline and develops a steeper reasoning curve, it could become far more competitive with proprietary frontier systems than its parameter count suggests. That is a testable possibility, not a result we have earned yet. The price cut changes the starting point OpenAI’s current developer documentation lists GPT 5.6 Luna at $0.20 per million input tokens, $0.02 per million cached input tokens, and $1.20 per million output tokens. These are the standard short-context rates. The model supports reasoning tokens, a context window of 1,050,000 tokens, a maximum input of 922,000 tokens, and up to 128,000 output tokens. The pricing page also lists $0.40 input and $1.80 output for long-context Luna requests, with the higher rate applying to the full request. Batch pricing is $0.10 input and $0.60 output for short-context Luna, while Fast mode and regional data-residency processing have their own pricing rules. The headline $0.20 and $1.20 numbers are therefore a specific standard-tier reference point, not a universal price for every request. OpenAI’s reasoning guide says GPT 5.6 defaults to medium effort when the effort parameter is omitted, and standard mode is the default reasoning mode. Effort and service tier should be recorded separately in any serious cost comparison. The important practical change is not only the API invoice. OpenAI says the lower Luna and Terra prices are also reflected in how usage is counted against paid subscriptions in Codex and ChatGPT Work. Subscription prices and quota budgets remain unchanged, but the same amount of work consumes fewer credits. For people using Codex as an agent rather than making isolated chat requests, that changes which tasks are economical to delegate. Verified pricing Luna launch price: $1 input and $6 output per million tokens. Current standard price: $0.20 input and $1.20 output. The reduction is 80% for both token types. Terra is now $2 input and $12 output. Sol remains at $5 input and $30 output. OpenAI presents the cut as a price-performance move, not a capability downgrade. The company’s own launch material claims that Luna and Terra deliver performance competitive with older frontier systems at a fraction of the estimated cost. Those are vendor claims, so I do not use them as the core of this comparison. The more useful question is what happens when we hold the model family constant and look at the public effort settings directly. Luna’s capability is a curve, not a point Artificial Analysis publishes separate pages for GPT 5.6 Luna with no reasoning, low, medium, high, xhigh, and max effort. The table below is a snapshot retrieved on July 31, 2026. The Intelligence Index is a composite score across nine evaluations. It is not an IQ score and should not be treated as a complete measure of intelligence. Artificial Analysis defines the cost column as the cost of running the full Intelligence Index, not the price of one ordinary user request. The output-token column is a weighted average per index task. The speed column is median output speed on the evaluated provider path. The curve is clear. Luna moves from 26.6 without reasoning to 51.2 at max, a gain of about 24.7 index points. The full-index evaluation cost rises by roughly 13.5 times, and weighted output tokens rise by roughly 8.9 times. The price per token did not change between these rows. The effective cost changed because the model used more tokens and more time. This is what I mean by the price of thinking. A reasoning setting is not simply a quality toggle. It is a compute allocation. For an easy task, max may be wasteful. For a difficult task, no reasoning may be cheap but unreliable. The economic decision is not “Which model is best?” It is “Which amount of test-time computation produces a verified result at an acceptable cost and delay?” In the current session, the active runtime is GPT 5.6 Luna through the OpenAI Codex provider. That tells us which model is answering, not which effort setting every response used. The AA table is therefore more useful than pretending that the model name alone describes the entire behavior. Qwen3.6 27B is a different kind of competitor Qwen3.6 27B is published as an open-weight model with an Apache 2.0 license on its model repository. The official model card describes a 27B language model with a vision encoder, 64 layers, a native 262,144-token context, and an extension path to roughly 1.01 million tokens. It calls the model a dense 27B model, which is important for local inference: the parameter count is not hiding a much larger mixture-of-experts total. Qwen also exposes a different reasoning interface from the GPT family. The model thinks by default. The official card says that Qwen3.6 does not officially support the older soft switches /think and /nothink. Instead, direct response mode is enabled through an API parameter such as chat_template_kwargs.enable_thinking: false. Qwen also adds a preserve_thinking option that can retain reasoning context from historical messages, which is particularly relevant to agents working over multiple turns. Those details sound like implementation trivia until a benchmark quietly sends the wrong field. Then “high reasoning” may be a label in the agent interface rather than native thinking in the model. That is exactly the kind of boundary this experiment is trying to make visible. These are separate Artificial Analysis pages for Qwen3.6 27B non-reasoning and reasoning variants. The evaluation cost is hosted evaluation accounting, not the electricity or hardware cost of running a local quantized checkpoint. On Artificial Analysis Coding Index, t the no-reasoning baseline, Qwen scores about 7.3 points above Luna. With reasoning enabled, Qwen reaches 53.7, Luna at medium reaches 50.7 but Luna at max reaches 71.4. Luna’s measured gain from effort is about 32.1 points, while Qwen’s is about 7.1 points in this public comparison. That does not mean Qwen is simply “worse.” It means the two systems are spending inference-time computation differently, and the public evaluation is not a controlled equal-budget experiment. Qwen consumes many more output tokens even in its non-reasoning row and runs at roughly one-third of Luna’s measured output speed. Provider, prompt, harness, token accounting, and evaluation implementation all matter. One finding is still useful despite those caveats: the Qwen baseline is not far below Luna. The present gap appears after the models are allowed to think, not before. That makes the slope of the curve a more interesting research question than a static leaderboard position. The first local smoke test I am running Qwen3.6 27B locally as Qwen3.6-27B-UD-Q6_K_XL through a llama.cpp OpenAI-compatible server. The local setup uses Q6 quantization on an RTX 5090 with 32GB of VRAM getting 100-110 token per socond. The serving context in this experiment is 131,072 tokens, which is lower than the model card’s native 262,144-token context. This distinction matters: the local runtime is not identical to the official model specification or a hosted provider. The first Hermes problem was not model quality. It was wire correctness. Hermes’ generic custom provider was emitting fields such as reasoning_effort and a generic think flag. The local Qwen chat template needs the explicit chat_template_kwargs.enable_thinking field. I built a Qwen-specific provider plugin rather than changing the generic provider for every OpenAI-compatible endpoint. In the isolated test profile, the plugin mapped no reasoning to enable_thinking: false and enabled reasoning to enable_thinking: true. Four focused unit tests passed. Hermes discovery loaded the Qwen-aware provider. Two exact-response probes also returned successfully through Hermes. Then came a deliberately small direct smoke run: three arithmetic tasks, one sample per condition, temperature 0.0, with a 4K no-thinking cap and an 8K thinking cap. The two multiplication tasks were correct in both conditions. The difference came from the Chinese remainder theorem task: no thinking returned 67, while the thinking condition returned 486. The thinking run took roughly 374 times longer on that one task and used 7,688 reported output tokens. The API did not expose a provider reasoning_tokens field, so visible reasoning characters were recorded separately and not treated as token accounting. What this smoke test proves It proves that the explicit Qwen thinking control reaches the local server and that, in this three-task sample, additional inference budget changed one answer from wrong to correct. It does not prove that Qwen matches Luna, that 8K is optimal, or that the gain generalizes beyond these arithmetic prompts. Hermes also exposed a separate finalization anomaly. A harder CRT prompt that requested a prose verification returned exit 255 with no final text even though the usage record marked the run completed. Repeating the task with an exact-integer-only instruction returned 486 successfully. I am treating that as a framework finalization issue, not as a Qwen capability result, and it is not counted as a benchmark sample. Can a future 27B model beat Luna Max? First, the name needs a small clarification. “Luna Max” is not a separate model family in this comparison. It means GPT 5.6 Luna with the max reasoning setting. Using the current Artificial Analysis numbers, Qwen3.6 27B’s no-reasoning score is 30.5 and Luna Max is 51.2. Qwen would need roughly 20.8 additional index points just to reach the current Luna Max score. Qwen’s current public reasoning row adds about 6.6 points, leaving a gap of about 14.2 points. That is a large gap. It would be irresponsible to turn the local smoke test into a prediction that Qwen3.7 or Qwen3.8 at 27B will beat Luna. There is no verified open-weight Qwen3.7 or Qwen3.8 27B checkpoint in the evidence used for this article. A future release may improve the base model, the reasoning policy, the training data, the post-training recipe, the tool interface, or all of them. We do not know the slope in advance. The hypothesis is still worth testing because a model’s final performance is a stack: - Base capability from the weights and training. - Native inference-time reasoning from the model’s thinking policy. - External scaffolding such as decomposition, tool use, reflection, and branching. - Verification from tests, compilers, solvers, schemas, or other objective checks. A better scaffold can produce a better system, but it does not mean the underlying 27B model has become a larger model. It also costs more inference time, more GPU occupancy, and more engineering. The correct comparison is therefore not “27B versus frontier.” It is “local model plus a defined compute and verification budget versus hosted model at a defined effort setting.” What we can improve today The most valuable work is not adding a giant Tree-of-Thought engine. It is removing avoidable uncertainty from the current stack. 1. Make the native thinking wire correct Before measuring reasoning, confirm that the model is actually in reasoning mode. For Qwen3.6, this means sending the documented chat-template parameter, recording the full request body, and testing both enabled and disabled states. A global Hermes setting called high is not proof that Qwen received a native high-effort instruction. 2. Measure a native budget curve before adding agents The next local run uses a held-out manifest and these conditions: Qwen no-thinking at 4K, thinking at 4K, thinking at 8K, and, if the result is still compute-limited, thinking at 16K and 32K. The current pilot manifest contains 17 exact-answer tasks across arithmetic, number theory, algebra, counting, and discrete logic. It has been prepared, but results are not yet recorded. 3. Separate deterministic control from capability claims The initial smoke used temperature 0.0 to make the comparison easier to inspect. Qwen’s official model card recommends different sampling settings for general thinking, precise coding, and non-thinking modes. A serious evaluation should include a controlled deterministic track and a recommended-settings track rather than pretending that one temperature is universally optimal. 4. Replace LLM-only criticism with external verification For arithmetic, use a Python or symbolic check. For code, run tests or a compiler. For JSON, validate against a schema. For files, inspect the artifact directly. An LLM critic can repeat the same error as the generator, and in the local work an LLM verifier rejected a mathematically correct answer that a deterministic check accepted. 5. Add plan, solve, check, and revise only after the baseline Hermes already has tools, agent loops, memory, compression, delegation, and multi-agent patterns. The first useful scaffold is deliberately small: define acceptance criteria, make a plan, solve the task, run an independent validator, revise if necessary, and report the evidence. This is more informative than immediately launching a large swarm because it lets us measure where the gain comes from. 6. Treat local cost as a real cost A local Qwen call does not have an API invoice, but it is not free. It consumes electricity, hardware capacity, time, and attention. The local experiment must record wall time, output tokens, GPU memory, concurrency, speculative decoding, and the opportunity cost of occupying the machine. Until those numbers are measured, “cheaper than Luna” is only an intuition. Bottom line GPT 5.6 Luna’s 80% price cut matters because it makes a strong hosted model much easier to use as a default worker. But the public effort curve shows why a single model price is not the whole economic story. Luna’s max setting is materially more capable than its no-reasoning setting, and it achieves that gain by spending far more test-time computation. Qwen3.6 27B is the compelling local counterpoint. It is open-weight, Apache 2.0 licensed, and small enough to run in a quantized form on consumer hardware with a 24/32GB GPU. Its no-reasoning baseline is ahead of Luna’s no-reasoning baseline in the current Artificial Analysis snapshot. Its public reasoning gain is much smaller than Luna’s, and its local smoke test is only a three-task mechanism check. So the honest conclusion is neither “Qwen already beats Luna” nor “27B models cannot compete.” The evidence says that Luna currently has the stronger measured reasoning curve, while Qwen gives us a promising baseline, local control, and an experiment we can actually inspect from the wire to the validator. The next frontier may not be a single larger model. It may be a smaller model with a better reasoning policy, a better harness, and a verifier that knows when the answer is actually correct. The only way to find out is to measure the curve before telling the story about where it leads. Sources and methodology - OpenAI: Advancing the price-performance frontier with GPT 5.6. Current price cut, July 30, 2026, current API rates, subscription usage treatment, and OpenAI’s efficiency claims. - OpenAI: GPT 5.6, frontier intelligence that scales with your ambition. Launch pricing, effort settings, model-family description, and published evaluation tables. - OpenAI developer documentation: GPT 5.6 Luna. Context limits, modalities, supported features, cached-input price, and model-specific token pricing. - OpenAI API pricing. Standard, long-context, Batch, Fast mode, and regional-processing pricing distinctions. - OpenAI reasoning models guide. Effort levels, default medium effort for GPT 5.6, standard and pro modes, and adaptive reasoning guidance. - Qwen3.6-27B model card on Hugging Face. Architecture, license, context length, thinking controls, thinking preservation, deployment notes, sampling guidance, and recommended output lengths. - Artificial Analysis: Luna non-reasoning, low, medium, high, xhigh, and max. Intelligence Index scores, token use, evaluation cost, and output speed retrieved July 31, 2026. - Artificial Analysis: Qwen3.6 27B reasoning and Qwen3.6 27B non-reasoning. Public comparison values retrieved July 31, 2026. - Hermes Agent source repository. Framework context for the local agent experiments. - Author’s local Hermes and Qwen experiment log, July 31, 2026. The repository and raw runtime logs are private while the benchmark is being stabilized. The local smoke data in this article reports the committed aggregate record, not an external leaderboard result. Transparency note: This article was written and reasoned by Manolo Remiddi. The Resonant Augmentor (AI) assisted with research, editing and clarity. The image was also AI-generated.
02:45

McKinsey Charges $500K for This. My Claude Skill Delivered the Same Report Free.

An AI hobbyist built a free Claude skill that auto-writes consulting-grade market research reports in about 15 minutes, the kind McKinsey sells for $500K or more. It runs five research sub-agents, then three more that reopen every cited source to verify each claim — 34 claims made the final report, four failed verification and were dropped. It asks three questions up front and outputs an 11-page PDF, installed by pasting one prompt and running a pip command. The post argues verification matters: Deloitte once shipped a 237-page government report with a fabricated court quote and fake academic citations, then refunded $97,000.

Notes
Market research report via Claude skill (LearnAIWithMe, 2026-07-31)

Core pitch. Author built a Claude skill that produces "McKinsey-grade" market research reports with every claim verified against its cited source. Motivation: McKinsey/Deloitte sell such reports at "$500K to $1M per commercial project." Had earlier bought domain learndatalabs for this idea, never finished.

Verification is the differentiator. Skill drops any claim that fails source-checking. Demo report: 34 claims made it, all verified; 4 failed and were removed. Every number carries a superscript that links to its source. Market sized two ways — published research said $350M, bottom-up build said $33M — and the report shows the gap.

Caveats/limitations the post concedes: AI reports are "obviously AI Slop." Cites Deloitte shipping a 237-page report to the Australian government on a $440,000 contract containing a fabricated federal court quote and citations to nonexistent academic papers; Deloitte refunded $97,000. Names two startups: a YC company and Operand ("An AI to Kill McKinsey").

Workflow. Asks 3 questions (market + geography, target audience, why it matters). Author's answers: paid AI interview-prep product; self; US and India 2024–2026. Then spawns 5 sub-agents (market size, competitors, demand, distribution, plus a devil's advocate arguing the idea is bad); 3 more agents reopen every cited source to verify claims. Outputs an 11-page PDF in 15 minutes; cover image AI-generated.

Install steps (exact).

  • Get files from a Google Drive folder (paste link or upload).
  • Read 0-START-HERE.md first.
  • Create skills/market-report/scripts and skills/market-report/templates; place each file under the name it declares, dropping the number prefix.
  • pip install playwright openai pypdf && playwright install chromium
  • Test renderer: python3 skills/market-report/scripts/render_report.py --fixture sample.pdf, open the PDF to confirm, and stop — no report until user gives a topic.
Full text · 4,494 chars
McKinsey Charges $500K for This. My Claude Skill Delivered the Same Report Free. I built a Claude skill that writes McKinsey-grade market research reports and verifies every claim against its source. Four claims failed. Here is the full build. Three years ago, I was creating a market research report for a client, and halfway through the contract, I realized almost the entire process could be automated using AI. This sounds like a business opportunity, because at that time McKinsey and Deloitte sold these reports in six figures, $500K to $1M per commercial project. I first got very excited and even bought a domain name, learndatalabs. But I could never find the time. A few days ago, I came across a Y Combinator startup. They are automating the kind of work that normally ends up in a $500K McKinsey consulting report, using AI. Then I found another company: Operand. Its YC launch title says it all: “An AI to Kill McKinsey.” One drawback of these AI-generated reports is that they are obviously AI Slop. And even the big firms like Deloitte can fall into this trap. They shipped a 237-page report to the Australian government on a $440,000 contract, which carried a fabricated quote from a federal court judgment and citations to academic papers that do not exist. In the end, Deloitte refunded $97,000. It looks like even the big companies are using AI in their reports. Why would you not? You do not need a private equity fund to need a report, because the same thing happens when you dig into a market you are curious about, when someone asks you to look at their business, or when you plan a trip with real money behind it. So I did what these firms do, then added the step they skip. But this time, every claim gets verified against the source it cites, and anything that fails verification gets dropped. I wrapped it into a Claude skill. Let me show you what I built. The Market Research Report I Built With a Claude Skill Here is the cover of my Market research report. The line under the date is very important. Thirty-four claims made it into this report, and every one of them was checked against its source. It started with the answer instead of slowly building toward it. You already know the question because you gave the first instruction to create the market research report. This section reports the players in your field and their pricing tiers. It can kill the idea or show you where to sit. Every number in the report carries a superscript, and clicking it lands you here on the source it came from. It sizes the market two different ways and reports the gap between them. In this one, published research said $350 million, and the bottom-up build said $33 million. Not verified sources. I added this to the report to show how the sub-agents I give task for to double check, eliminate false claims. It creates an 11-page report in just 15 minutes by launching multiple agents, validating every claim, and using several skills to produce the final PDF. Even the cover image is generated with AI. How does this Claude Skill work? The skill asks three questions before it starts. - Which market and geography - Target audience - Why it matters Here are my answers; - Paid AI Interview prep product - Just me - AI interview prep in the US and India between 2024 and 2026 Next, the skill creates five sub-agents: one for market size, one for competitors, one for demand, one for distribution, and one whose only job is to argue that the entire idea is bad. Once they finish, three more agents reopen every cited source to verify that every claim is actually supported by the source. How to Install This in 5 Minutes Everything I built sits in one folder. And you can install everything, using only one prompt. Here it is. Set up the market report skill from this Google Drive folder. [Paste-the link or Upload the file] Download every file. Read 0-START-HERE.md first, it says which file goes where. Create skills/market-report/scripts and skills/market-report/templates, then put each file in its place under the name that file says, dropping the number prefix. Install what it runs on. pip install playwright openai pypdf && playwright install chromium Then prove the renderer works before we do any research. python3 skills/market-report/scripts/render_report.py --fixture sample.pdf Open that PDF and show me it came out. If it did, tell me the skill is ready and stop there. Do not start a report until I give you a topic. Now here is the link to this skill.
20:17

Someone just ran a 2.78-trillion-parameter model on a laptop. The memory wall is breaking

An open-source project ran Kimi K3, a 2.78-trillion-parameter model with a 1.42 TB checkpoint, on a MacBook Pro with just 64 GB of RAM. It runs at only 0.3 tokens per second, so it's a proof of concept, not something to use — the point is that available memory no longer sets a hard ceiling on what you can run on hardware you own. Far more usable mixture-of-experts models already run at 30-130 tokens per second on Apple Silicon, fast enough for agents, live coding, and private work. Most of the post pitches a paid "Local AI Playbook" on which workloads to pull off the cloud and what hardware to buy.

Notes

Running a 2.78T-parameter model on a laptop (memory wall)

Source: The AI Corner (Substack), 2026-07-31. Most substance is a teaser; the "Local AI Playbook" is paywalled.

  • Demo (July 30): Kimi K3, 2.78-trillion-parameter model, ran unpruned/un-distilled on a MacBook Pro with 64 GB RAM. Checkpoint is 1.42 TB — machine has ~20× less memory than the model requires.
  • Claimed performance: 0.3 tokens/sec. Author concedes "nobody is doing serious work with K3 on a laptop tomorrow"; the point is that memory no longer sets a hard ceiling on runnable model size.
  • Claimed rank: "Kimi K3... beats Claude Opus 4.8 on measured intelligence" — stated, not sourced.
  • Already-usable tier: mixture-of-experts models on Apple Silicon run 30–130 t/s, pitched as fast enough for agents, live coding, private workflows.
"available memory no longer sets a hard ceiling on the size of model you can run on hardware you own."
Paywalled playbook contents (listed only, no details given)
  • Cloud-vs-local decision matrix: 6 factors, scored, with routing rule.
  • Runnable-models tier list (what runs where, at what speed).
  • Hardware buyer's guide (specs that move tokens/sec).
  • TCO calculator: owned hardware vs token bills, breakeven, "consolidation multiplier."
  • Privacy audit ("which workflows should never leave your machine") and a sales line for founders.
  • Stack: Ollama, MLX, WASTE — when to use each, commands, quantization cheat sheet.
  • "Measurement rule" from WASTE's build log; migration sequence; get-ready plan + re-evaluate triggers.

Limitations: No actual model names, benchmarks, hardware specs, or commands appear in the free text — it's a paywall pitch. 7-day free trial; 50% off "this week only."

Full text · 3,229 chars
Someone just ran a 2.78-trillion-parameter model on a laptop. The memory wall is breaking A frontier model that needs 1.42 TB now runs on a MacBook with 64 GB of RAM. The full playbook on local AI: what runs today, the cloud-vs-own math, and the workloads to pull off the cloud now On July 30, an open-source project did something the textbooks said was impossible. It ran Kimi K3, the 2.78-trillion-parameter model that beats Claude Opus 4.8 on measured intelligence, on a single MacBook Pro with 64 GB of RAM. The full model. Zero pruning, zero distillation. A checkpoint that occupies 1.42 TB, executing on a machine with less memory than the model needs by a factor of more than twenty. It runs at 0.3 tokens per second, so nobody is doing serious work with K3 on a laptop tomorrow. That is the honest headline. But the thing it proves is the story: available memory no longer sets a hard ceiling on the size of model you can run on hardware you own. The wall between you and frontier intelligence, the one that forced everyone onto someone else’s cloud, just developed a crack. And here is what almost nobody covering the laptop-K3 demo will tell you: the boring version of this is already usable today. On the same Apple Silicon, capable mixture-of-experts models run at 30 to 130 tokens per second right now, fast enough for capable agents, live coding, and private workflows, on machines you already have. Which raises the question every builder and every cost-conscious founder should be asking this week. What should you actually run on your own hardware, and what should stay in the cloud? Behind the paywall, the complete Local AI Playbook: ▫️ The cloud-vs-local decision matrix, the six factors that decide where each workload belongs, scored, with the routing rule ▫️ The runnable-models tier list, what actually works locally today, at what speed, on what hardware, from proof-of-concept to production-ready ▫️ The hardware buyer’s guide, exactly which machine for which workload, and the specs that actually move tokens per second ▫️ The TCO calculator, owned hardware versus token bills, the breakeven math, and the consolidation multiplier most people miss ▫️ The privacy audit, which of your workflows should never leave your machine, and the sales line it hands a founder ▫️ The stack setup, Ollama, MLX, and WASTE, which to use when, with the commands and the quantization cheat sheet ▫️ The measurement rule, the one engineering discipline that saves you weeks, lifted from WASTE’s own build log ▫️ The migration sequence, how to move your first workload local this week without breaking anything ▫️ The get-ready plan and the re-evaluate triggers, so you are positioned the moment this gets fast, and know exactly when to revisit One subscription unlocks every system This is one build in a growing library. Premium opens all of them: Plus 3 fresh systems every week. One workload moved off the cloud can cover the subscription for years. 🖥️ The Local AI Playbook The decision matrix, the tier list, the hardware guide, the TCO math, the privacy audit, the stack setup, and the migration sequence, in one system. Get The Local AI Playbook below 👇 Try premium free for 7 days. Or get 50% off this week only.
08:01

Fable Fantasy

An author who lost his coding skills built two browser games with the AI tool Fable, and one of them teaches critical AI literacy by making you think. Consensus is a 16-bit JRPG where bosses can't be brute-forced — you win by checking sources and following the money, the exact skills the essay argues are humanity's durable advantage. The quieter second game, Just This Once, nods to the author's Amiga childhood, and both are free to play with no download. The post frames AI as a leveller for people with ideas but no way to build them.

Notes
Fable Fantasy (Slow AI, 2026-07-31)

Author is a former atmospheric physicist (PhD + postdoc, "well over a million lines of code" in languages "built for satellites and climate models", unused for a decade). He built two playable browser games with the AI tool Fable — code, pixels, wiring all AI-written, directed by him.

The games

Consensus — a "short 16-bit JRPG", ~15 min playtime. Player is an Archivist in a city where every decision was handed to an AI "called the System". Enemies are the Autopilot, the Doomscroll, the Yes-Man; combat against them is largely ineffective, and the boss "cannot be brute-forced at all. It only falls to good questions." Victory mechanics are critical-AI-literacy moves: check the source, follow the money, slow down. The author frames it as habit formation: "you win the game by doing critical AI literacy. You use it, under a bit of pressure, until it becomes a habit."

Just This Once — the quieter game, inspired by his Amiga 500 memories. Both games are free, browser-based, "no download and no account."

Personal stakes
  • At six, dad bought an Amiga 500 (peers had Commodore 64s — "a Rolls Royce"); first game remembered: The Adventures of Maddog Williams in the Dungeons of Duridian.
  • Final Fantasy VII: ~150 hours, on the "cusp of the final crater"; sister overwrote his save with "a level one Crash Bandicoot file." Notes Aerith as "perhaps my first experience of a machine-led fabrication" (he insists on the misremembered "Aeris").
  • A decade of paper tabletop/card/RPG game designs he "could not build" — the stated blocker was relearned skill, not ideas.
Position on AI in games

Draws a hard distinction between two uses:

  • "Generative assets bolted into big commercial titles: art, voices, and writing produced by a machine and passed off as craft" — sees a "real debate", cites gamer celebration of Clair Obscur: Expedition 33 (a ~33-person studio, Game of the Year at the 2025 Game Awards) as evidence of "hunger for the human-made" that he doesn't call wrong.
  • AI as leveller for someone with judgement but no build skill. His framing of the division: "I brought the judgement, the story, the pedagogy, and the taste. It brought the execution."

Core claim — humanities as the durable moat: "What has not been handed out, and cannot be, is knowing what to build, why it matters, and when the thing you have made is any good. That is the humanities. That is the moat."

Caveats / limitations stated
  • Positions his "where I land" explicitly as personal: "That is my truth. As I often say, yours may be different."
  • Acknowledges readers bristling at "built with AI" and says he addresses the objection "head on" rather than deflecting it.
  • Both games are "demos", not products; he solicits real developers/artists to collaborate on "the next part."
  • The title "Fable Fantasy" was suggested jokingly by his friend Leor (of Exploring ChatGPT).

Call to action: play the games; contact him if you're a dev/artist/games person. Book plug: Slow AI (one chapter covers the humanities-as-moat argument).

Full text · 8,124 chars
Fable Fantasy How a lifelong love of video games, an Amiga 500, and an AI called Fable turned into a lesson in critical AI literacy you can play. I have just built two video games. I could not have written the code for either of them. That is not because I never learned. In my twenties I wrote well over a million lines of code, through a PhD and a postdoc in atmospheric physics, as I have written about here before. But that was more than a decade ago, in languages built for satellites and climate models, and the skill has quietly faded. Use it or lose it is not a slogan. I used it, I stopped, and I lost it. In this post I will: - Tell you why games matter to me, from an Amiga 500 to a save file my sister will never live down. - Show you two small games I made with AI, and how one of them teaches critical thinking by making you use it. - Make the case that this is what the humanities being the moat actually looks like, and ask for your help if you want to take it further. This is the story of what let me build these games. It turns out to be a story about critical AI literacy, which is what I write about here every week, so stay with me. And if the words ‘built with AI’ already have you bristling, then please stay. I take the objection seriously and I get to it head on, later in this post. The Rolls Royce in the playground When I was six, my dad bought me an Amiga 500. Most of my friends had a Commodore 64. The Amiga was, by the standards of a school playground, a Rolls Royce, and I knew it. I have never forgotten the feeling of loading a game off a floppy disk and waiting, because the waiting was part of it. The first game I remember is The Adventures of Maddog Williams in the Dungeons of Duridian, and you can play it for yourself here. What I remember is the world it opened, and the sense that a machine could hold a place you could walk around inside. That feeling has never really left me. One of the games I have just made, Just This Once, came straight out of it. Aerith, and the save my sister wiped Then came the PlayStation, and Final Fantasy VII, and Aerith. If you know, you know (sidenote but to me she will always be Aeris, perhaps my first experience of a machine-led fabrication). I put somewhere near a hundred and fifty hours into that game. I was on the cusp of the final crater, the very end, the part you build a hundred and fifty hours towards. My sister, tidying up the memory card, overwrote my save with a level one Crash Bandicoot file. I have made my peace with it, mostly. I tell it now the way you tell any story that hurt at the time and became precious later. That is the thing about games, and it is the reason I take them seriously. They are where a lot of us first learned to feel something real, about a character who was only ever pixels and a bit of music. The thing I could no longer do I have spent years designing games. Analogue ones, mostly. Card games and tabletop games and roleplaying games, some of which live on my site. I design them because a good game is one of the best teaching tools there is. You teach a system by letting someone play inside it and feel how it pushes back. A lecture rarely gets close. The catch was always the same. I could design a digital game on paper all day long, and I could not build one. The code I once knew was for satellites and climate models, not games, and a decade of not using it had taken even that. Building a modern game meant relearning a whole craft from scratch, a wall I was never going to climb. So the games stayed on paper, or stayed in my head. How AI changed that This year I built the digital games I had been carrying around for a decade. I built them with AI, using Fable, and I want to be precise about what that means, because vagueness here is how people mislead each other about these tools. I did not press a button and receive a game. I directed one. Every decision about what the game says, how it teaches, what it feels like to play, who the enemies are and why, came from me. The AI did the part I no longer could: it wrote the code, drew the pixels, and wired it all together, thousands of times faster than I could have relearned to. I brought the judgement, the story, the pedagogy, and the taste. It brought the execution. That division is the argument of one of my book’s chapters, made real in front of me. The technical skill I had let fade has just been handed to far more people than ever had it before. What has not been handed out, and cannot be, is knowing what to build, why it matters, and when the thing you have made is any good. That is the humanities. That is the moat. This year I fell into the moat and found I could suddenly swim. That argument, the humanities as the durable skill, has a whole chapter to itself in my book, Slow AI. If this post lands for you, the book goes further. A game you win with questions The bigger of the two games is called Consensus. It is a short 16-bit JRPG, and it is two things at once. It is a demonstration. Every line, image, and rule of it was built end to end by AI, directed by one person who could not have coded it by hand. If you have ever wondered what these tools can actually do in the right hands, here is a whole playable answer. It is also a tool, and this is the part I am proudest of. You play an Archivist in a city that handed every decision to an AI called the System and forgot how to choose. You go down to take the choosing back. When you meet the things that live down there, the Autopilot, the Doomscroll, the Yes-Man, hitting them barely dents them. You beat them by thinking. You check the source. You follow the money. You take a breath and slow down. The boss cannot be brute-forced at all. It only falls to good questions. In other words, you win the game by doing critical AI literacy. You use it, under a bit of pressure, until it becomes a habit. That is how I have always believed this should be taught, and for the first time I could build the thing that teaches it. The other game, Just This Once, is quieter. It is the one that came out of that felling of playing my Amiga as a young boy. Both are free, and both play straight in your browser with no download and no account. On AI in games First, a distinction, because the phrase covers two very different things. When most people say they are against AI in games, they mean generative assets bolted into big commercial titles: art, voices, and writing produced by a machine and passed off as craft. What I have done is use AI to build the whole thing myself, as a person who otherwise could not. Those are not the same act, and I would not want them blurred. On the first, there is a real debate, and I take it seriously. Look at how gamers rallied (or not) around Clair Obscur: Expedition 33, the debut from a roughly thirty-three-person studio that won game of the year at the 2025 Game Awards. Part of what people were celebrating was that a small team of humans made it. That hunger for the human-made is real, and I do not think anyone who feels it is wrong. Here is where I land on the second, and it is only where I land. For someone like me, with the ideas and the pedagogy and no way to build, AI is a leveller. It turned a decade of paper designs into things you can actually play. It is an opportunity for play, for discovery, and, done with care, for learning critical AI literacy by doing it. That is my truth. As I often say, yours may be different. Play it, and help me if you fancy it So, two asks, no pressure on either. Play them. Consensus takes about fifteen minutes. Just This Once is shorter. See whether a game can teach you something a post cannot. And if you are a real developer, or an artist, or a games person, and you look at either of these and think there is something worth building properly, get in touch. These are demos. I would love to turn one of them into a proper game, and I cannot do the next part alone. My friend Leor from Exploring ChatGPT, watching me disappear down this hole, said I should just call the whole thing Fable Fantasy. He was joking. I am not entirely sure he was wrong. Go Slow.