Nothing matches those filters.

Lead

1

Video

1
13:41

Create HYPERREALISTIC AI Characters That INTERACT | FREE & LOCAL

You can build a cast of photorealistic AI characters that act and interact in the same scene, all free and running on your own computer. The technique is a LoRA, a small model trained on a set of your character's photos, and the video lays out five rules for doing it right: pick a trigger word, vary shots so the character generalizes, and write captions that leave details changeable later. It covers three open models, Create 2, Ideogram 4, and LTX 12.2 for video, where Ideogram 4 lets you draw bounding boxes so each character lands exactly where you place it. Everything runs inside ComfyUI, and the creator provides automated workflows that build your training dataset for you.

Transcript · 37,122 chars
You can now create a whole cast of AI characters and make them act and interact in the same scene. Thanks to a brand new generation of open models that are far more photorealistic than ever, rivaling and in my opinion even surpassing close source tools like ChatGPT and even Nano Banana Pro. They work in any style and are super controllable. I mean, look at this. You can draw a bounding box for each character and object and the model places them exactly inside it. This gives you full control over your composition without needing to rerun the prompt 20 times. So, in this video I'm going to show you how you can build your own perfect data sets for each character, train your own AI models on them, and keep them all consistent across every single shot. So, whether you want to create AI influencers, movie protagonists, or AI animation characters, all of it can run locally on your own hardware for free. So, let's start building your cast. Everything comes down to one technique, a LoRA. A LoRA is a small file that sits on top of a big model. The base model already knows what humans in general look like, but it doesn't know what your specific person looks like. And we can train your LoRA by just showing it a bunch of images of your specific character. Here are five rules and concepts you need to know to train a good one. Number one, the trigger word. The trigger word is a made up word you put in every caption for your character. During training, everything that's not explicitly mentioned in the caption gets attached to this word. Your trigger word becomes your character. Number two, vary things you want to generalize. For example, if you train a LoRA only on close-up images of your character, the model starts to learn that your character only really exists in a close-up position. So, definitely put some wide shots in there as well. The same goes for lighting and poses and clothes and everything else, which brings us to rule number three. Decide what should stay consistent and what you want to be able to change later. If your character is only supposed to wear one set of clothes, it's fine to train on the same set of clothes in each image. But if you want to maybe dress up your character later, you should vary this in the data set. But also use rule number four, the caption decides whether you can change something later. So, for example, you have a set of 20 images of your character with the same clothes. As long as you add a detailed caption, you are probably able to change it later because of the Laura learns that these clothes are not attached to your trigger word. So, the more detailed your caption, the more flexible you are later. But, also the longer your prompt has to be because you have to recall all these elements that you want to your character to have. And this is because of rule number five. A caption is basically a reverse prompt. And that's why you always need to create these captions in the same style you would later create an image prompt for that image model. For example, Create 2 wants like long natural language prompts and Ideogram 4 wants structured JSONs with bounding boxes. You're probably thinking, "Well, this sounds like a lot of work creating these data sets." But, we created automated workflows that help you set it all up and make character creation a lot of fun. Once you have these data sets, you are flexible to train LoRAs for any AI image model. But, because they're quite new and pretty cool, I will show you how to create models for Create 2, Ideogram 4, which is that model that allows you to draw these bounding boxes and really precisely place your character in the scene, and also 12.2, which is not an image model, it's a video model, so you can also create some cool scenes with your characters. All of this runs inside of ComfyUI, a free node-based interface for AI models. If you've never installed it, we have a free step-by-step guide on our website. Once ComfyUI is running, you always set up the workflows in the exact same way. Drag and drop the workflow file into the ComfyUI interface, open the ComfyUI manager, and click install missing custom nodes. Select them all and install them, and then restart ComfyUI. Then, it's time to load the models used in this workflow. And you can always find the models we use in these nodes next to each model loader node. So, you just click on it, download the model, and place it in the mentioned folder. And once that's done, either restart ComfyUI or click R to refresh and select the correct model. So, let's open the new consistent character the the simple version. We are going to work from left to right. So, let's zoom in on the left here. As always, here you can find a little guide and here you can find the model selection, the model loader group. Next to these model loaders, you can find the corresponding node telling you where to download the model and where to put it in your ComfyUI folder structure. These two groups down here you can ignore. These are utility groups. So, let's move up here to the settings. First, we need to name our character and I usually name the character the same word that the trigger word will be. Next, you can set the resolution. Don't change anything here. And finally here, we can set if you want to save out the data set and I recommend that set that to two here to preview mode first. That way you can refine your data set and once you're happy, you just switch that to one, click run the workflow again and it will save out your final images. So, now let's import our character here and I want to create a realistic character. So, I import this image right here. Usually, I would recommend to import like a high-quality, high-resolution, full-body image. I don't have that of this yet, so I only import this one right here. The workflow will still work. But if you load like a close-up like this, you need to be a bit more specific in the prompt. Clothes would usually be extracted from the full-body image, but since we only have like the upper body here, we should name the rest. Otherwise, the model will just fill it in. So, let's add something like wearing black loafers. Oh, and I forgot, I want to go to preview mode right here. So, this is basically all you need to set up. Right now, you can already zoom out and click run. So, now a few things are happening. First of all, a close-up of the face will be extracted. Now, in this case, it's not that important because the input image is already a sort of close-up, but if you have like a wide-body shot, it's really helpful for the model to understand your character if it also sees this close-up shot. And next, when we zoom out here, you can see the workflow is already generating a data set. A few minutes later and our data set is done and you can see we have a really nice collection of varied images of our character in different situations, close-up, wide shots, different lighting conditions, emotions, failing to look at a frog. Now, the cool thing about this workflow is that it's completely modular. So, if you want to change something, for example, you can come to this group right here and maybe change the prompt. Maybe she's sitting on a bench in a desert at night with a crashed UFO in the background holding a sandwich. Now, if you only want to rerun individual images, you can always click on that save node above here and click on this play button and it will start calculating again. But, you can also create your own groups because these are completely modular. You could just select one of these groups and copy them anywhere you like. And now you can think of which reference images you want to connect. And you can see here you have up to four inputs. And right now I'm referencing the T-pose and the original input face. But, of course, we could reference any other image that we created here. But, usually the T-pose and the full face is giving you a lot of flexibility. So, let's create another pose right here. So, a wide shot of the character sitting in a cafe, maybe. And this looks pretty good. Now, of course, you can edit all these prompts and make them more specific to your characters. I just tried to keep them like really general so they work with any kind of character. But, for this character, I already tested training a Laura and it worked really well. But, I realized that her necklace changed a lot. So, what we could do now is just copy this group right here and put it somewhere. Doesn't really doesn't matter where. And let's try something like this. And now we have a close-up shot of her necklace. Though, I think it looks a bit different. So, what we could do now is create a load image node and just use the snipping tool and do this. Go back here, copy that in there, and also use this one as a third reference image. And now it looks much better and much closer to the original. If you want even more images of this, we could copy this group. And you can see that there's a set node right here. So, we could call that necklace and then right here, instead of the T-pose, we could we could reference necklace and then create a close-up side view. And you can see here during image generation, it's trying to generate shoes. And that's because they're in the general prompt. So, you can not only choose which um reference images you want to choose, but you can also choose the prompts that you want to reference. This gets the style prompt and the clothes and feeds that into the prompt right here. But, you can also only connect this manual prompt right here, like so. And this is looking a lot better. You don't have to create these additional images, but you can see how much fun it is. And sometimes it's necessary if your character has like really specific detail that you want to show the model. Okay, so now that we have these images, let's just go here to the beginning and save out this data set by selecting one and clicking run again. Once you're done, you can find all your images in ComfyUI output CCC and you will see a folder that is named like your character. Here are all the images that you generated during the testing phase. And here you can see your final data set with the selection of your final images. But, before we start training the LoRA, let me show you how easy it is to train other styles as well. So, for example, maybe you remember this guy. This is Hans Schmankerl, the first ever character created with one of my consistent character workflows, I think like 2 years ago or something. Drag and drop this full body image right here. I'm calling him Hans Schmankerl. So, then I will put in a different prompt. I will put in a 3D animated Bavarian man, CG rendering. And for the clothes, I will put in nothing because we already have like the full clothes here. And I will stop talking this way. It automatically extracted the face and started working on the images. Yeah, it just works really well for any kind of style. He's also managing to actually look at that frog in that prompt. I just noticed a mistake. Sorry, let me fix that. Ah, yes, much better. So, I forgot these um necklace images right here. So, now he's wearing that beautiful necklace. Uh let's remove these new uh prompts. So, yeah, this came out really beautiful. Speaking of beautiful, this video is sponsored by our amazing Patreon supporters. If you want to get your hands on advanced workflows, the example data sets, and Laura's we trained for this video, and our amazing Discord community, the link is in the description. As an additional thank you to all our power Laura trainers out here, our advanced supporters on Patreon can also get the advanced version of the character creator, and it adds two stages. First of all, the upscaler. This one upscales all your images using Seed VR 2 to 2K resolution, and create really amazing detail without changing the look of your character. So much fun to just uh use the compare node right here. It's just amazing. And the second one is the face detailer. I noticed that in some shots when the the face is really small, for example, like in full-body T-poses, the face changes a little bit because the model doesn't have enough resolution to generate. So, I created this face detailer group right here. But, this one will not just upscale every image because in some images the face is already filling up almost the whole image. So, we created this custom node here that will pre-select the images where your face is too small, and only upscales them using the face detailer. And look at these amazing results. Like here, the style doesn't even matter. It works for any type of character, and really helps to pull your face closer to your original input face. You can also get your hands on beta versions of versions of the character workflow. I'm always testing it with new models. For example, right now I'm testing a version based on Crea 2. There was a community patch by Austris to kind of turn this into an edit model, and I have some really promising results, though they are not as good as the one with the current version, but we'll see. Maybe in a few weeks it'll be the main version. But, now let's continue with the free versions of the workflow. And you can use this workflow, the data set tagger, to caption your whole data set. For this, copy the link to your data set folder, and also add a Laura trigger word up here. And for pretty much every model except for Ideogram, you want to put in txt here. Next, you can click on this preview image node and click this play button to check if all your images loaded correctly. Next, scroll down here. This is the prompt studio, and here you can select a bunch of presets for different models. And we want to create a Laura for Korea 2. And a few minutes later, it's done. Right here you have this preview node where you can check out all the different prompts it created. So, remember it's using this detailed caption style. You can also just adjust this prompt right here, right in a line. It's basically an LLM, so a large language model. You can write in here, "Don't mention the clothes", for example, and that typically works really well. But for now, this data set looks good. And if we now check out the folder, you can see that there are corresponding text files for each image. This is the perfect data set we can start training with. You can use any Laura trainer with this data set, but I'm going to use AI Toolkit by Austris, because I think it's the simplest one. And also Austris, the creator, has like this really cool YouTube channel where he shows you how you can train all sorts of different Lauras. So, if you want to dive in deeper, definitely check out his channel. You can run AI Toolkit fully locally, and Austris has done a lot of work getting these large models to work on like older GPUs. But what I'm going to do is I'm going to use the RunPod template, because especially with the newer models, you really have to have the newest GPUs to train Lauras fast, or it will just take forever. So, I'm going to sign in to RunPod. By the way, we also have an affiliate link that if you use that to sign up, you get a credit bonus between 5 and $500 for the first $10 you spend on the platform. So, once you've added a few dollars, you can go to hub and search for AI Toolkit. And you want to use this one here, the official one. Click configure pod, select a GPU. Here you can see all the GPUs that it's working with. You can see a 5090 is $1 an hour, but because I need a Laura fast, I'm going to use the RTX 6000. I click deploy pod. Now you need to log in and if you didn't set a new password, the password will just be password. First, we need to upload our data set. So go to data sets, click new data set and I will just call it again like the trigger word. Now you can click add images and just drag and drop that here. And here are all our images and the corresponding captions. So now we can go to new job. Here we need to give it a name and I will again call it like this and maybe like a version number behind that and next we need to select the model architecture. And if you want to train a Laura for Korea, we need to select Korea 2 raw. Because I have like a 6000 Pro, I don't need to use the low VRAM option. I can just deactivate that. We could go lower for the linear rank here. I think we could do 16, but I'm just going to leave that. 3000 steps might be too much, but let's just keep it because you can download all the versions in between and select our data set. Now during training you can create samples and here for example, it will create samples every 250 steps. So you can see how your Laura is training. I think I will just keep that as it is and down here you have the example prompt. So what we could do here is we could adjust them. For example, we could put in our trigger word. For example, Emily demo, a woman with red hair. No, she does not have red hair. You can also delete some here. We maybe we don't need that many. Here's my selection of prompts. I try to keep them as simple as possible to really see what kind of qualities get attached to my trigger word. So now I can just click create job. This will then take us to this job and here we can just click this play button right here and it'll start. After roughly 80 minutes, the training has been completed. During training, you can always go to samples and see how your model trains. In the beginning here, you see the model has no idea what our character looks like and then iteration after iteration, more features of our character start showing up. And then, roughly around here, it's starting to look really close to her. This is at step 2,500. Speaking of steps, go back to overview, and here to the right, you can download the last versions of your model. This will be the last one at step 3,000, but the last one doesn't mean it will be the best one. So, download this one, but also download some of the other ones here and test them out with an image generation workflow to find the best version. You can test out this workflow with any CreA 2 workflow, but I also wanted to give you this very basic free one right here that I used to create all the images. Again, you can find all the models you need right here in the description. Down here in this node, you load in your character Laura, and you can see I'm loading in the previous version of Emily here because the other one is still training. And I think you usually find the best version of your Laura between step 1,250 and 2,000, actually, at least in my experience. So, you can see here I'm testing it with 1,500 and this simple prompt. And I would say this is working quite well. Now, CreA really loves these super detailed prompts. So, I recommend writing a prompt like this one time or having an LLM write it for you and then just adjusting it for the different scenes. You can create some really amazing stuff with it. We could also add like um an amateur style Laura, so it looks a bit less uh fashion professional photo uh shoot style, but I think that's pretty cool. Okay, but what about multiple characters? Let's give her some friends. So, I ran the character creator three times, one time for every character. We have this cool model right here. For her, I also created some additional close-ups of her makeup and her ear piercings and uh makeup again. You already know her. For her, I created some more close-up shots of that necklace and the mesh top without the blazer, a single shot of that necklace, and with different lighting conditions, side view extreme close-up. And then we also have this guy right here. And for him I created some additional necklace shots as well. I threw them all together in one folder. But then the trick is to also include images of these characters together in the same shot. So the model can learn that these are separate entities. And to create these, it's super simple. You can use any edit-based image model. But one way to do it is just to duplicate one of these groups in character creator. And then I created two load image nodes for each character, and always loaded the wide shot and a close-up, and concatenated them together like this with this node, resized them to make them a bit smaller, and then I fed them all as references into the sampler group, creating a prompt like this. The three characters are standing together in an image, blah blah blah blah blah. Like this, I created a few variations, sometimes with two characters, sometimes with all three. And not just standing next to each other, but sometimes also interacting. Make sure to throw these in into your full data set folder like this. And now let's caption these images. For this, again, I load the data set tagger workflow, and this time I leave the set Laura trigger word empty. And for this, I'm going to select Korea 2. And because this is just a large language model, you can add your own tagging rules, your own captioning rules. So for example, I just added this. Blonde woman with a bus cut is Berlin model, the black man with a gray turtleneck is male model number one, and the woman with curly hair is Emily number five. If there are multiple characters in the same image, add their trigger words and describe where they are in the image and how they are interacting. And now, since we're still training a Korea Laura, let's add.txt and just click run. And it worked really well. Like Berlin model and Emily stand side by side on on beach. Like it just worked. So, now we can go back to RunPod. We can add a new data set and this time I called multiple realistic. Add images, select all, and drag and drop them in here. Again, it loaded all these images correctly. And now we can create a new job. Select Rhea Raw. I'm still on that good GPU. And now for three characters, I think what I will do is I will bump this one up to 64 linear rank 64. Other than that, make sure to select the correct data set. And now for the prompts, what I usually do is I use a large language model of my choice. Of course, this can also be like a local one. And now we have these final prompts and I'm just going to copy them over. Let's click create job and add that to the queue. So, I downloaded all the different versions and I think my favorite one is the at step 2,200. So, I loaded that in and you can see again I had an LLM create this really long prompt for me that I usually just describe the next scene and it will use this prompt format and this will create some really beautiful images. This one is like a model fashion photo shoot in like a an a hotel room and that worked so well. Like, his necklace is correct, her necklace is correct. Her piercings are correct. Now and now we could just like copy this prompt format, put these characters in an old castle or wherever you want them to be, describe the scene and it will adjust everything around that and the hotel room changed to this weird castle. Again, really good image. Uh except how he's touching this curtain is a bit awkward, but other than that, really, really cool. And just so you know how well this works, this is just my output folder. This was the first image. This one was the next one. And because I had the LLM make him buy this water bottle, that that appeared a few more times in the following shots, which is a a bit weird. Um >> [snorts] >> in the club, but still hydrated. Um, so I asked to remove that water bottle and you can see how you precise you can prompt then. I brought them together, waited some uh solar shots. That worked really well. So, this is no cherry-picking. These are all the images that came out of the model. Here I tried something more complex. That is a prompt issue that now Emily is there two times in the shot. I fixed that, then it worked well again. So, here you can see it worked and here is the hotel, a museum, and uh an empty swimming pool. I'm really impressed by these images and I think you could push this even further if you add like a realism Laura, for example. And the cool thing about Creia is that you can prompt really precisely. So, you can really tell the model where each character is placed and what kind of composition you want. And this of course also works with any style. So, let's uh revisit uh Hans Schmunkel and his wife Anya Schmunkel and you can see um it was just so much so much fun to prompt different situations and it just worked it just worked so well. Like the only thing that didn't really work at first was I wanted to make them ride these cows, but then I got more precise with the prompt and from then on it just uh worked. And even though there are not many other clothing styles in the data set, it was still able to give Hans this really cool bathing uh outfit here. But now let's talk about Ideogram 4. Ideogram 4 is a really cool new open-weight model that allows you to prompt with bounding boxes. So, you have full control over the placement of your character. And the process for the creating a data set and training the Laura is really similar, but there are some differences. So, let's go through it. First, let me copy this data set and call it uh Ideogram. And now we still have all the the text captions in there. I'm just going to delete them, so we only have the images and let's go back to the tagger. I load that in and give it the path to this folder. This time the JSON file extension is correct and we also have the Ideogram 4 model selected right here. This is a much longer prompt, but you care about this part right here. This is pretty much the same thing as with the Korea thing. You just add which character has which caption, which trigger word. Once you have added this, click run. And once it's done, you can not only see the JSON files that it created, you can also see a preview of the bounding boxes. And don't worry if they're like multiple, Ideogram will create extra bounding boxes like for example here for the tree or there is a little frog down here and it saw that. So, that's pretty cool. Sometimes the bounding boxes are not correct though. So, to fix this, you have multiple options. One is our tag and review workflow. And then you copy in the path and click run right here. And here you can go through all these images and if there's a wrong bounding box like this, you can just change that. Just make sure the whole character is covered like this. And here you can see the multiple people and here it worked really well. And once you have corrected all your images, you can just run the workflow and this will then actually overwrite these JSON files. But you can also correct and create these prompts right in AI Toolkit, which is especially handy if you have like an older GPU because large language models can also take quite a long time. So, for this, I'm going to create a new data set and again I'm calling this multiple realistic Ideogram. I select everything and just drag and drop that in here. So, here is our data set and you can see that the captions are missing. And that's because up here you need to change that to JSON. The cool thing about AI Toolkit is that you can also click on an image and this will then also show these bounding boxes and you can also edit them and save them. And this is honestly a really cool interface to edit your JSON prompt with. You can also auto caption your images. So, this is pretty much the the exact same process as our ComfyUI workflow. For this, you just upload an image only data set and then you go here, select um Ideogram for Captioner. This is also using the same model that we are using Gwen 3, and then you can click Add to Queue, and it will do the exact same process, but right in AI Toolkit. If it doesn't track your element, you can also add a new element and draw a new bounding box for it, and I really recommend taking the time and uh working through these images. It really helps when generating images later. Once you're done, you can click New Job. Now, before we can start training, we need to get access to the Ideogram model. This is a gated model with a special license, so you have to log into Hugging Face, create an account, and go to this Ideogram page, and click that you accept the license. Then, what you need to do is go to Access Tokens, and click here Create New Token. A read token is fine. Give it a name, create token, and then you have this key here, and you want to copy this. Go back to AI Toolkit settings, and enter your Hugging Face token here. And now you can give it a name, select Ideogram for. Again, I don't need the low VRAM option. For three characters again, I'm bumping this up to 64. Also, select the correct dataset, and now for the prompts. What I like to do for the test prompts, I like to go to datasets and select an image where all of these characters are in frame, and I'm just going to copy this description. Then, I'm using a large language model of my choice, copy in the prompt, and say Create 10 example prompts using this format. And one problem that happens is that the bounding boxes are are rotated, but this actually correct. That's just some Ideogram weirdness. So, we have to tell Claude, "No, this is actually correct." So, if your large language model tells you this is the wrong format, no, this is actually the correct format. And then again, I just copy over all these prompts. And make sure if you select the um correct data set to change the caption extension to JSON. And that's it. You can now click create job and add it to the queue. Once training is finished, you can use my example Ideogram workflow. It looks like this. To the left here, you can select the preset. Pretty much always want to use maximum quality. You can set the resolution and megapixels. I would leave that as it is. And now you need to add your Ideogram Laura two times. And you can see I used iteration 2000. So, even though it continued on to 3000, these came out looking a little overbaked, a little bit weird. So, I'm usually using something around 2000 2250 steps. Now you need to download all these models. And I just realized I forgot to add the download links. I will do that. So, when you download the workflow, you will find the loader nodes right here. And then we are using this Ideogram prompt builder by Keijai. First of all, you start with the high-level description of your scene. And down here, you add a description of your background. You can select your style, describe your photo, the aesthetics, the lighting, the medium. You can describe all of that. You don't have to describe it. And then the cool feature about it is that you can manually create these prompt boxes right here. In this on this canvas, you can just drag for a new bounding box. And then you have a prompt for each of these bounding boxes. But let's now delete this and show you what I've done here. So, I created the first bounding box. And here, I put in a prompt with the trigger word for each character. So, this is male model, this is Emily, and this is Berlin model. And you can see I've added the trigger word as well as a short description. Then for example, to show you how much you can control this, I also added this additional bounding box here to show you that you can make him hold up uh his hand holding a walnut. And then I also placed animals on top of each uh head. So, for example, he has a brown squirrel sitting on his head, a bluebird sitting in her hair, and a mantis insect climbing on her head. When I click run, this JSON prompt is created that will then be fed into the sampler. And um yeah, this is the final image. Just for fun, I added this draw bounding box node, so you can actually see how precise this that is. That's so cool. Like, for example, let's let's add another mantis. So, we could just like maybe copy this, put this here, and say a mantis insect climbing on climbing on her leather coat. Let's try that. And you can already see another mantis appearing here. That's really cool, but also creepy. And uh here are the bounding boxes again. So, this is a bit heavier. Don't expect it to work like the first time you try this. It definitely needs some time to get used to and figure out the best format. But, once you've got a hang of it, it's like so much fun. So, yeah, this image is kind of cool. I added these two characters like to see if they I can generate a back view of them, if that works. That worked fine. I generated her, and then I added this uh fish, and then I added this red head, which was uh a lot of fun. And I feel like this model even responds well to these weird poses right here. Like, this is an awkward camera angle and an awkward pose, but still it worked really well. Okay, but what if you want to train a video model? Of course, you could take these images and use like LTX, for example, or a pay tool like Kling or Seed Dance, but you can also train your own video model with it and prompt for your characters. And this has some really great results, some really good consistency, and it's really easy to do. Like, once you figured out all this stuff, it's really easy. Again, you can use the exact same data set that you created before, like with the different characters all in one folder. And if you want, you can import the data set tagger, and here from the presets select one video because one is a video model, but you can train it purely on images, which is really cool. But we already created like a data set for Korea, and this is very similar. So, I'm just going to try if I can train the one model on the Korea data set. So, back to run pod, I click new job, and now for characters, you can train either a one 2.2 low noise model, or you can train a one 2.1 one model. And the cool thing about this one is that it works with one 2.1 and one 2.2. And since I often use tools like our AI rendering workflow, that one works with a version of one that was trained on one 2.1. I usually like to train one 2.1 models just because I'm more flexible that way. I'm going to select this one, bump up the rank a little bit. I'm going to select the multiple realistic data set. And down here, since this is a video model, this will try to generate video as demo, but this just takes a really long time, and I don't want that. So, here on the number of frames, just set that to one, so it will generate sample images instead. You can go to your large language model of choice, and I'm just going to write I already created these JSON prompts, so now I'm just telling it convert the JSON prompts to video prompts for one 2.1. So, let's copy over these prompts. They look pretty good. We need the text extension for this one, and yeah, that's all we need to do. We can click create job and add it to our queue. Once it's done, you can add the last free workflow for today to ComfyUI. This one is just a one 2.2 video generation workflow, and as always, you can find the links to the models in these descriptions. You can see I'm loading in like a Laura stack right here. So, I'm loading in the Light X2V Laura that allows you to run this workflow a lot faster. You can find all the links to these Laura's right here. For this one, I tested out the different Laura's and I found that at step 2,500, it was the best one. So, I load that in at full strength, not just for the high noise Laura stack, but also for the low noise Laura stack. So, it's right here. And then for the prompt again, it's really good to create like a sort of preset and then tell a large language model, "This is the format. Please adapt this format to the next scene." Since we already created these example prompts, I will just copy this one all three subway platform. Let's load that in. Make sure that you add these keywords here for the Laura's to activate. And now all that's left to do is click run. And yeah, there it is. First try, this is what came out of it. But I think it looks really cool. So, yeah, that's it. I'm really impressed with how far we've come with multiple characters in one scene. I remember when I first tried to train a flux Laura on multiple characters, it was like horrible merging and stuff like that. Or then you had to use like regional prompting techniques that were super heavy on your computer and not easy to use. Um so, this is just a lot more fun. Personally, if I would recommend this workflow to newcomers, I would suggest you start with the Korea 2 workflow because Korea 2 is kind of light on your computer. It's pretty fast compared to the other models. It has some good prompt adherence. It has some good support by the community, some amazing style Laura's already. So, if you want to push it more towards realism, there are some cool Laura's for that already. Um yeah, but other than that, that's it for this one. And um again, huge thanks to our amazing Patreon supporters for making these videos possible. Remember, if you want access to our amazing Discord community and the Laura's that I trained, additional example files, and the advanced workflows, consider supporting us. It really makes spending weeks trying to figure all this stuff out um possible. So, thanks a lot for that. Thanks for watching and see you next time.

Newsletter

2
17:55

Kimi K3: The Open-Source Model That Just Cracked the Frontier Moat

Moonshot AI shipped Kimi K3, the first trillion-parameter model with fully open weights, and it ranks near the top frontier models at a fraction of the cost. It's a 2.8-trillion-parameter mixture-of-experts model, so only about 50 billion parameters run per question, and it scored 57.11 on the Intelligence Index, fourth overall behind Fable 5 and GPT-5.6 Sol. It took first place on several coding and writing benchmarks, but those scores come from harness-specific tests some evaluators won't rank. It's slow and verbose, hallucinates about half the time, and ships with no content guardrails. Full weights drop July 27, and its per-task cost is low enough that community-distilled versions should soon run on consumer hardware.

Notes

Kimi K3 (Moonshot AI) — launch notes

Moonshot AI launched Kimi K3 on July 16, 2026: a 2.8T-parameter Mixture-of-Experts model, 896 experts with 16 active per token (~50B active params per forward pass). Architecture built on two innovations: Kimi Delta Attention (KDA) and Attention Residuals. Moonshot claims ~2.5x scaling efficiency vs K2 (unverified independently). Text+image input, text output; context window 1,048,576 tokens with flat pricing across the whole window and free automatic context caching. Weights promised July 27, 2026; API live now.

Benchmarks
  • Artificial Analysis Intelligence Index: 57.11 (#4) — behind Fable 5 (59.86), GPT-5.6 Sol max (58.89), GPT-5.6 Sol xhigh (57.65); above Claude Opus 4.8 (55.69), Grok 4.5 (53.83), GLM-5.2 (51.09). Coding Index 76.24; Agentic Index 50.07.
  • Frontend Code Arena: #1 at 1679 Elo (up from K2.6's #18), first in 6 of 7 domains; Program Bench #1 (77.8); SWE Marathon #1 (42.0); Louie's Writing #1 at 2840 Elo (first open-weight model to top it); AutomationBench-AA 52.7; BrowseComp 91.2% (300K context compaction) / 90.4% (1M, no management); DeepSWE 67.5 (3rd); FrontierSWE 81.2 (2nd).
  • Vals Index 74.7 — below Fable 5, above GPT-5.6, Sonnet 5, Opus 4.8, Muse Spark.
  • BenchLM declined to rank K3 — its coding rows rely on harness-specific signals (DeepSWE, FrontierSWE, Program Bench), no weighted SWE-bench Pro or LiveCodeBench.
"No rank is better than a rank built from one flattering corner of the table." — BenchLM
Cost per task (Artificial Analysis, actual eval token usage)
  • Design prompt: K3 $0.03 vs Fable 5 $0.38, GPT-5.6 $0.11 (~12x and ~3x cheaper). CS:GO clone: $3.24 vs $10 / $6. Cheapest full evaluation run, near top score.
  • Runtime: 62 tok/s, 1.99s TTFT — slow but low latency; generated 130M tokens in eval vs 63M average (extremely verbose; cost implication flagged).
Hardware / distillation path

Full precision: 16–18 DGX Spikes (~$4–5K each, >$80K total); 1-bit quant: 4–6 units. Q4_K/Q5_K quant → 20–25 GB VRAM, fits RTX 5090 (32 GB). Expected timeline post-weights: first GGUF within 48h; fine-tunes/distill datasets in week 1; production-ready 14B–30B distilled students in 3–6 weeks (typically retain 80–90% of teacher capability, 5–10x faster).

Caveats
  • No content guardrails and no query redirection — the model you call is the model you get.
  • Hallucination rate 51% (vs Fable 5 at 55%, GPT-5.6 Sol at 89%) — still >half, a material risk in high-stakes use.
  • Blind spots in medical imaging and visual detection; slow; verbose.
  • No CyberGym score released; possible US export-control trigger if it clears thresholds.
  • K3's lead may be short-lived: GLM 5.5 (1T), MiniMax Pro (1T), Qwen 4 on the horizon.

Note: written by Manolo Remiddi with AI research/editing assistance; image AI-generated.

Full text · 11,227 chars
Kimi K3: The Open-Source Model That Just Cracked the Frontier Moat The world’s first open-weight model in the trillion-parameter class scored 57.11 on the Intelligence Index, behind Fable 5 and GPT-5.6 Sol, but at a quarter of the cost and with no content guardrails. Something New Is Here On July 16, 2026, Moonshot AI launched Kimi K3, a 2.8 trillion parameter model with native multimodal support, a 1 million token context window, and pricing that makes the numbers look like a disruption play rather than a status-quo move. The weights are scheduled for full release on July 27, 2026. Until then, the API is live and available now. Kimi K3 is not an incremental update. Moonshot calls it their most capable flagship model to date, and the independent benchmarks suggest the claim has some teeth. But the real story is not just the benchmarks, it’s what they mean for the open-source AI ecosystem and the corporate moat that Western labs have been building for years. What It Is Kimi K3 uses a Mixture-of-Experts architecture with 896 experts, activating just 16 per token. It’s built on two architectural innovations, Kimi Delta Attention (KDA) and Attention Residuals, both designed to help information flow through longer sequences and deeper models. Together with the extreme sparsity, K3 achieves roughly 2.5x the scaling efficiency of K2, converting compute into capability more effectively. The model supports text and image input, outputs text, and has a context window of up to 1,048,576 tokens. Context caching is automatic and free. Pricing is flat across the entire window, no tiered increases for long contexts. The Benchmarks And What They Actually Mean Kimi K3 scored 57.11 on the Artificial Analysis Intelligence Index, placing it at #4 overall, behind Claude Fable 5 (59.86), GPT-5.6 Sol max (58.89), and GPT-5.6 Sol xhigh (57.65). That’s above Claude Opus 4.8 (55.69), Grok 4.5 (53.83), and GLM-5.2 (51.09). It also scored 76.24 on the Coding Index and 50.07 on the Agentic Index. But the Intelligence Index is only one composite. The more interesting story lives in the breakdown: - Frontend Code Arena: #1 at 1679 Elo — a 17-place jump from K2.6’s #18, past Claude Fable 5. First place in six of seven frontend domains. - Program Bench: #1 (77.8) - SWE Marathon: #1 (42.0) - Louie’s Writing benchmark: #1 at 2840 Elo — the first open-weight model to top it, surpassing Fable 5. - AutomationBench-AA: leads the board at 52.7 - BrowseComp: 91.2% with context compaction at 300K tokens; 90.4% without context management at 1M tokens. - DeepSWE: 67.5 (third overall) - FrontierSWE: 81.2 (second overall) The Vals Index puts K3 at 74.7 — below Fable 5, but above GPT-5.6, Sonnet 5, Opus 4.8, and Muse Spark. BenchLM, one of the more cautious benchmark analysts, notes that K3’s coding rows consist of harness-specific signals like DeepSWE, FrontierSWE, and Program Bench rather than weighted SWE-bench Pro or LiveCodeBench. Their position: “No rank is better than a rank built from one flattering corner of the table.” BenchLM remains skeptical and has not ranked K3 in their weighted leaderboard. That’s an important data point, not every evaluator is treating these numbers as straightforward. Artificial Analysis also measured runtime characteristics: 62 tokens per second output speed and 1.99 seconds to first token. The model is slow but has good latency. More importantly, it is extremely verbose, it generated 130M tokens during the full Intelligence Index evaluation, compared to an average of 63M. This verbosity has real cost implications for production use. The Cost Comparison That Actually Matters Raw token pricing tells only part of the story, because models differ enormously in how many tokens they consume to complete a task. A more honest comparison looks at the cost per task: These figures come from Artificial Analysis, calculated from actual token usage during the Intelligence Index evaluation. K3 is the cheapest way to run the full evaluation, and it scores near the top. That’s the data point worth repeating. The same pattern shows up in real-world demos. In a design prompt comparison, K3 cost 3 cents versus Fable 5 at 38 cents and GPT-5.6 at 11 cents, a 12x and 3x difference respectively. For a CS:GO clone, K3 cost $3.24 versus $10 for Fable 5 and $6 for GPT-5.6. The gap is not theoretical. The Distillation Wave: What Comes After July 27 Kimi K3’s full weights release on July 27, 2026, is the API equivalent of a starting gun. Because the model will be open, the community will begin distilling, quantizing, and optimizing it almost immediately. This is what happened with Llama, Qwen, and DeepSeek — and it happens faster each cycle. Understanding what that means requires understanding the MoE architecture. K3 has 2.8 trillion total parameters, but only 16 of its 896 experts are active per token. That means roughly 50 billion parameters are being used in each forward pass. That distinction between total and active parameters is the entire reason this model matters for consumer hardware. What Fits Where At full precision, K3 requires 16-18 DGX Spikes at roughly \$4-5K each — over \$80K in hardware. At aggressive 1-bit quantization, that drops to 4-6 units. These are enterprise numbers. But distillation changes the equation entirely. Community quantization at Q4_K or Q5_K should bring the active parameters — the 50B that actually process each token — down to 20-25 GB of VRAM. That fits on an RTX 5090 (32 GB). With the right KV cache management and speculative decoding, inference becomes viable. Total throughput will be lower than the API, but you’re running a model that scored 57.11 on the Intelligence Index on hardware you can actually buy. Distilled student models will go further. A 14B-30B parameter model distilled from K3’s reasoning traces, coding outputs, and agentic workflows will run comfortably on consumer GPUs. These distilled variants typically retain 80-90% of the teacher model’s capability on the tasks they were distilled for, while running 5-10x faster and using a fraction of the VRAM. What It Means Three things shift when a model at the Fable 5 / GPT-5.6 level becomes available as open weights: - Distillation targets improve. The quality of a distilled model is bounded by its teacher. Until now, the best open teachers for distillation were models in the Llama 405B or Qwen 72B range. K3 raises that ceiling by an order of magnitude. Every model distilled from K3 will inherit more of its reasoning depth, coding ability, and knowledge work quality. - Local reasoning becomes frontier-adjacent. Running a distilled K3 “student” on an RTX 5090 or DGX Spark means you get chain-of-thought reasoning, long-context understanding, and agentic tool use without sending data to any API. That matters for privacy, sovereignty, and cost. - On-premises sovereignty stops being a startup fantasy. For governments, hospitals, legal firms, and mid-size companies, K3 gives a path to running intelligence comparable to Opus 4.8 and close to GPT-5.6 entirely in-house. No per-token fees. No guardrails from a third party. No data leaving your network. That’s not incremental, it’s structural. The timeline, based on the pattern from every previous open-weight release: expect the first GGUF quantizations within 48 hours of the weight release. Community fine-tunes and distillation datasets within the first week. Production-ready distilled student models (14B-30B range) within 3-6 weeks. By the time August ends, running a local K3-derived model on consumer hardware will be a solved problem. The question isn’t whether this ecosystem will materialize. It already exists for smaller models. The only variable is how quickly the community can optimize a 2.8T parameter MoE architecture for inference on hardware that wasn’t designed for it. The Strategic Shift And Why It Matters Two things about Kimi K3 are genuinely unprecedented: First, no guardrails. Unlike proprietary models that route queries to less capable versions or refuse certain topics, K3 has no content filtering or query redirection. The model you call is the model you get. For researchers, this means consistent performance across domains, including medical, legal, and security-adjacent work. For users, it means the model doesn’t suddenly downgrade when it detects sensitive topics. Second, the weight release. Moonshot has committed to releasing the full weights by July 27, 2026. If that happens, organizations can run a 2.8 trillion parameter model entirely on-premises. This isn’t a 7B parameter toy, this is a model that scores in the same league as Fable 5 and GPT-5.6 Sol, available for anyone with the compute to run it. For a government agency, a research institution, or a mid-size company, this is a genuine path to AI sovereignty, running frontier-level intelligence without touching a Western API. Moonshot has not released a CyberGym score, and if K3 surpasses certain thresholds, it could trigger US export controls equivalent. The geopolitical implications of this model are real, not rhetorical. The Risks and the Caveats Kimi K3 is impressive. But it’s not perfect, and the hype cycle around it has already run ahead of the evidence. BenchLM is right to be cautious. The coding scores, while strong, come from harness-specific benchmarks. We don’t have weighted SWE-bench Pro or LiveCodeBench numbers for K3 yet. The model remains unranked in BenchLM’s weighted leaderboard, and that absence is meaningful. The 2.5x scaling efficiency claim is Moonshot’s own metric, we haven’t seen an independent verification. The full technical report, which should accompany the weight release, will be necessary before we can assess the training methodology and data recipes. Then there’s the question of what comes next. The open-source race is accelerating: GLM 5.5 (1T parameters expected), MiniMax Pro (1T), and Qwen 4 are all on the horizon. K3’s lead, if it exists, may be narrow and short-lived. The hallucination rate also deserves attention. K3 sits at 51%, better than Fable 5 (55%) and much better than GPT-5.6 Sol (89%), but 51% is still more than half. In high-stakes applications, that’s a material risk. The Bottom Line Kimi K3 is the most capable open-source model to date. It scores near the top on independent evaluations, costs less per task than any competing frontier model, and will be fully available on-premises in eleven days. It is slow, verbose, and has real blind spots in medical imaging and visual detection. Its coding lead is strong but comes from benchmarks that not all evaluators treat as equivalent. And the open-source race it’s part of is accelerating faster than anyone expected. What K3 has done, however, is prove something that was theoretical until now: a trillion-parameter model, open weights, running on “commodity hardware”, can compete with the best proprietary models at a fraction of the cost. That changes the game. The question now isn’t whether organizations will adopt this technology, it’s how quickly they can afford not to. Transparency note: This article was written and reasoned by Manolo Remiddi. The Resonant Augmentor (AI) assisted with research, editing and clarity. The image was also AI-generated.
00:31

How to Change Your Life: The 5 Forces That Shape Who You Become

A solopreneur newsletter argues real life change comes from thousands of small decisions, not one dramatic moment, and maps it to five forces: fate, luck, environment, character, and knowledge. It's the written companion to a paid masterclass, and most of the piece pitches the author's paid subscription and free toolkit. Thin on concrete detail and heavy on self-promotion.

Notes

How to Change Your Life: The 5 Forces That Shape Who You Become — Solopreneur Code (2026-07-17)

Companion piece to a live masterclass in the Solopreneur Mastery Club; mostly promotional (free "Solopreneur Success Hub" toolkit, claimed to save "20+ hours a week," plus a $79/year = $6.58/month Premium Vault upsell).

Core claim: real change is "the accumulation of 1,000 small decisions and 10,000 tiny actions," not one dramatic moment. On the "big win / lucky break / rock-bottom" theory:

"That moment rarely comes. And when it does, it's usually not the cause of change. It's the result of change that was already happening quietly underneath."

The 5-forces framework (labels borrow Chinese concepts):

  • Fate (命) — "unchangeable blueprint of your birth": genetics, family wealth, era of birth, innate talents. The one force you cannot touch.
  • Luck (运) — cyclical opportunity: market cycles, global events, meeting the right people at the right time.
  • Environment (风水) — physical/social energy: living space, chosen city, immediate peer circle.
  • Character (积阴德) — "good deeds quietly without expecting immediate reward": integrity, helping others, ethics, positive reputation.
  • Knowledge (读书) — "active pursuit of knowledge, skills, cognitive growth": schooling, reading, trades, mindset updating.

Remaining four are "yours to shape."

Diagnosis: solopreneurs "learning from everywhere and getting nowhere" — too much information, no system linking effort to results; framed as a systems problem, not motivation.

Caveats: despite promising "the science," no studies, sources, or data are cited; the model is presented as assertion. Chinese labels (命/运/风水/积阴德/读书) are mapped loosely to Western psychological concepts without provenance. Actual "how-to" steps are absent — frameworks only, with action deferred to the paid vault and video recording.

Full text · 3,644 chars
How to Change Your Life: The 5 Forces That Shape Who You Become Real change isn't one big moment. It's 1,000 small decisions, and most of them are yours to make. Quick note before we start. This one is a little different from what I usually share. I recently ran a live online masterclass inside the Solopreneur Mastery Club, and instead of my usual topics like solopreneurs, systems, and AI, I wanted to talk about something more personal, the science of changing our lives. This post is more of the companion piece to that session. Access your FREE Solopreneur Success Hub - your subscribers-only comprehensive command center for building and scaling a successful one-person business. I created this all-in-one toolkit for building a profitable one-person business, something I wish existed when I first started, and it saves me 20+ hours a week. Now, it’s yours… FREE! The recording holds the full conversation, and what follows here is the written breakdown so you can revisit the ideas, skim the frameworks, and take action without scrubbing through the video. If you were in the room live, treat this as your notes. If you’re reading it fresh, treat it as the map, then watch the recording for the full walkthrough. When we talk about wanting to “change our lives,” the goal can feel overwhelming, vague, or almost magical. Most of us secretly wait for one dramatic moment to flip a switch and turn everything around. A big win. A lucky break. A rock-bottom day that finally forces us to move. That moment rarely comes. And when it does, it’s usually not the cause of change. It’s the result of change that was already happening quietly underneath. Real transformation isn’t a single event. It’s the accumulation of 1,000 small decisions and 10,000 tiny actions. The good news is that this makes change learnable and achievable. If you want to know how to change your life, you don’t need a miracle. You just need a blueprint. Everything we experience and become can be traced back to five core forces: Fate, Luck, Environment, Character, and Knowledge. 🧬 1. Fate (命) - Definition: The unchangeable blueprint of your birth. - Key Factors: Genetic makeup, family wealth, era of birth, and innate talents. 🌊 2. Luck (运) - Definition: The cyclical waves of opportunity and external shifts over time. - Key Factors: Market cycles, global events, and meeting the right people at the right time. 🏡 3. Environment (风水) - Definition: The physical and social energy of your surroundings. - Key Factors: Your living space, the city you choose, and your immediate peer circle. ☯️ 4. Character (积阴德) - Definition: Doing good deeds quietly without expecting an immediate reward. - Key Factors: Integrity, helping others, ethical behavior, and building a positive reputation. 📚 5. Knowledge (读书) - Definition: The active pursuit of knowledge, skills, and cognitive growth. - Key Factors: Formal schooling, reading books, learning trades, and updating your mindset. One of them you can’t touch. The other four are yours to shape. Let’s break down each one, look at the science, and pull out what you can actually do with it. You’re doing everything. But nothing is moving? You are doing everything. But nothing is moving. That is not a motivation problem. Most solopreneurs are learning from everywhere and getting nowhere. Too much information. No clear system connecting effort to results. You have everything it takes. You just do not have a clear system yet. That is what paid subscribers get. Every system, playbook, prompt, and template. All inside the Premium Vault. All for $79/year. That’s $6.58/month. Upgrade now and unlock the Premium Vault.