Transcript · 29,078 chars
Last week, the best open- source AI video generator, Miniax H3, was released, and this is an absolute beast. It can do so many things. It has incredible world knowledge, and it's super flexible. Now, because this is open- source, the community has built on top of it, and now you can do even more things with it. So, in this video, we are going to go over some faster, better, and more interesting ways to run Miniax H3. This is a more advanced tutorial that continues from this video. So, if you haven't installed Miniax H3 yet, or if you don't know what it is, definitely see this video first. Anyways, let's get started. Now, this video assumes you already have Comi installed. The first thing you need to do is to update Comfy UI because they've added some additional features which are required to run some of the tools and nodes I'm going to show you in this video. So, simply click into the update folder and then click on update comi.bat and it'll proceed to update to the latest version of comi. All right, afterwards let's press any key to continue and then we can start up Comfy UI. All right, after you've opened this up, let's pull up one of the default miniax workflows. Let's just go with image to video. And a quick recap from the previous video. If you click on this workflow here, there are a few ways you can speed this up by inserting a few nodes after this load diffusion model step. So for example, one way is to use this patch sage attention node and connect this between the diffusion model and then basic guider and basiculer. And then over here we need to set sage attention to auto. Again, if you're not sure what I'm talking about, definitely see this video first on how to set this up. Another way to speed this up is to use this spectrum node. Well, they've actually updated this to a newer version which has even lower quality loss. So what you should do is also update this to the latest version. So in your Comfy folder, if you click on Comfy and then custom nodes, you should see your Spectrum folder over here. So let's double click into this and then at the top, type in cmd, which would open your folder up in command prompt. And then simply type get pull to update this to the latest version. I've updated this yesterday, so it says already up to date. And then afterwards, after restarting com UI, simply double click anywhere on your canvas and then type in spectrum to add this latest spectrum node. And now it should look like this. So, simply connect this after the model and whatever comes after it. Make sure you update to the latest spectrum node because it's even better. Now, here's another really cool tip. Let's say you want to preview the video as it's generating so that if it doesn't look good, you can just cut the generation halfway instead of waiting for it to complete. Well, the goat called Key released a really useful way to make this happen. This requires the model preview override node in KJ Nodes, which again we've installed in the previous video. Assuming you do have KJ nodes installed, you probably need to update it first in order for this new method to work. So simply go into your comfy UI and then custom nodes folder and then look for KJ nodes and then double click into the KJ nodes folder and then at the top type in cmd and for this folder again we are going to type in get pull to update to the latest version. All right, so afterwards let's exit out of this and then you also need to download this TAE model. I'll link to this in the description below, but if you click on files and versions and you click on VAE approx, here's where you can download this TAE saved tensor file, which is pretty tiny. It's less than 10 megabytes in size. So, let's click download and this goes in comi in models and then in VAE approx. Make sure you download this in VA approx. All right. So, let's click save. And then afterwards, back in our CompuI workflow, we just simply need to add this preview node after the load diffusion model. So, let me double click here and then type model preview override and you should see this node. So, let's select this and then place it somewhere here. Note that you do need to update to the latest KJ nodes in order to see this tiny VAE field. All right. So afterwards, you just connect this with the model and then whatever goes after it. And then here the fields are pretty self-explanatory. So the max resolution to preview this is 1024. JPEG quality is 80%, preview frames is how many frames you want to preview here. Since for this example, we are generating a video of 5 seconds at 24 frames per second. Let's set it to something like 124 preview frames to span the whole 5 seconds. And then this is the preview frames per second. Let's just set it to 12. And then for the tiny VAE, here is where we need to select our newly downloaded TAE H3 safe tensors file. If you don't see this, make sure you press R to refresh your model list after downloading this file. All right, so that's about it. Next, let me proceed to click run. And you'll see that it's going to give me a preview even before it finishes generating the video. So you can preview the generation, and it gets clearer and clearer as it proceeds from step to step. And if you don't like the look of it, you can just press X to cancel this like halfway so you don't continue wasting compute. Now, this is just a really lowresolution video at 12 frames per second, but after it's finished generating the video, it'll still save and preview the final video in this save video node. So, that is how you can use this preview method, which I think is super useful. All right, next let's talk about how to add Loras to your workflow. The awesome thing about open-source models is that people can build on top of it. They can fine-tune luras for whatever effect or style or camera motion or action or character they want. For example, someone has created this whispering Laura. Another person has creating this looping sketch anime sticker Laura. Another one has created this fictional woman, etc. And then Foul AI has also created this really useful realism people Laura, which if you look at the comparison without and with Laura, you can see that the videos look a lot more realistic. So, especially if you want to create some imperfect casual amateur shots, then this realism people Laura is super useful. So, that's what I'm going to demonstrate in this video. How do we add this realism people Laura into our workflow? First of all, you'll need to download this Laura. So, I'll link to this page in the description below. If you click on files and versions, simply download this realism people safe tensor file. It's only 131 megabytes in size. So, let's click download. And this goes in Comfy UI in models and then in Lauras. Let's click save. All right. Afterwards, let me just show you a simple text video example. So, I'm going to pull up the text video workflow. And let's do something like an amateur realistic lowquality video of a woman sitting at a cafe. She places her cup down and looks out the window. Now, to add the Laura, simply click on this corner to expand the workflow. And you just need to place the Laura after this load diffusion model. In fact, let me also pull up these acceleration steps first to make it even faster. All right, so let's add the Laura node somewhere here. Simply double click on the interface and then type in Laura. And there's actually a ton of load Laura nodes you can use. I'm just going to use this one, load Laura, and then place it here. And if you've just downloaded the realism people Laura, simply press R to refresh your model list. And then in the dropown, you should see the realism people Laura over here. So let's select this. The strength here is basically how much influence you want this slower to have. One would be 100%, you can make this a bit weaker. So let's go with like 08. And then let's connect this with the model and then whatever comes after it. So that's pretty much it. Now there's one more thing you need to be aware of, which is that most Loras have trigger words. You need to add the trigger word into your prompt for it to work. So for this example, the trigger word is realism. So let's also add this into our prompt. And that's pretty much it. Let's press run. And here's our generation. So, as you can see, it does look a bit more realistic and non-professional. Now, speaking of luras, here's another really useful way to speed up your generation by like four to five times. Now, you can use this for whichever workflow you want. I'm just going to show you an example of an imagetovideo workflow with this start frame. And we just need to download and add a Turbolora over here. So, let me show you that really quickly. And what a Turbolora does is it basically lets you decrease the number of steps to generate the video from 20 to only like four steps. So, it's up to like a five time speed up. Now, there are a ton of these turbo luras available for Miniax H3. One of the earliest ones is by Larry VR. This actually is pretty robust. It works pretty well. We also have some turboloras from the awesome light x2v. They've created turboluras for previous open video models as well like ltx and one. And we also have another one from Joyox. Here it says that this is for the BF16 base model. So you can use the Joyox one if you are using the BF-16 base model which is 40 GB in size. I think most of you are just using the pruned FB8 or IN8 version which is a lot smaller. But if you are using the BF16 version, then this Joy Fox one is actually pretty good. It seems to be even higher quality, especially for high action scenes compared to the Light X2V Laura. Now, I think for me their most robust option is to go with the light X2V version. And it looks like today they've also released some versions that are compatible with Comfy UI. So there's a four-step version here and an eightstep version here. In addition, KI also released his own version of the Lite X2V Laura which is compatible with Comfy Y, especially if you have low VRAM or RAM. This resized version is much smaller at only 315 megabytes in size, whereas the normal Turbo models are almost 2 GB in size. There's a lot to choose from. I'll link to all of these in the description below, but for this video, I'm going to demonstrate using this smaller resized version from KI. So, let's click on download. And this goes in Comfy UI in models and then in Lauras. Let's click save. All right. So afterwards, back in our Comfy UI workflow, if you haven't added a load Laura node after the diffusion model, load that first. And then afterwards, let's click down here and select the Turbo Lura, which I just downloaded. Now, each Laura has different optimal configurations. If you're using the keya here, he seems to suggest using four steps at 0.75 strength with erde or SA solver. So let's set the strength here to 75. And then for the algorithm in this dropown, let's select er. The annoying thing about this is it's not in alphabetical order. So ess. And then for theuler, you can set this to either beta or I heard beta 57 also works very well. Now for steps, four is the minimum number of steps you can use with the current turbo models. But if you find that the quality is too low, you could bump this up to like six or eight steps. So at least for me, let's set this to like six steps. Another thing to note is that if you use erd as the sampler, it turns out that this spectrum node doesn't actually work with erde. So, let's also click on this and press Ctrl +B to bypass this node. And then let's press run. All right. And here's our result. Not too bad, especially since it allows you to speed up your generation by up to five times. Now, at least for me, I don't actually like using a Turbo Lura. So, here are some additional ways to speed up your generation or to fit Miniax with even lower VRAM. So, first of all, let me get rid of this Laura node. Let me reenable this spectrum node, which I think is very useful. Now, for spectrum, I tend to use res multistep as the sampler. And then for theuler, I tend to set it to simple. But feel free to play around with different combinations to see which one works best for you. And since we got rid of the turbola, let's set the step count back to 20. All right. Now, instead of using a turbo model, what we can do instead to make this faster or fit in lower VRAMm is to use smaller models. So, there's also this experimental folder from KI and he released several even more compressed models for Miniax. So, for these two up here, he basically released a W4 A8 version. This one is for text to video and image to video. This one is for reference to video. And as you can see, each of them are just like 12 GB in size. Whereas even the smallest pruned F8 or INT8 version is 21 GB in size. So this is almost half the size which is great. You can probably run this on just like 8 GB of VRAM. So for me I'll just show you a texttovide example. So I'm going to download this one. And this goes in comfy UI in models and then in diffusion models. Let's click save. All right. Afterwards note that KI also released a more compressed video VAE which is only 3 GB in size. If you look at the official Miniax H3 video, this is like 5.2 GB in size. So again, downloading this will save you like 2 GB of memory. Let's click download. And this goes in Comfy UI in models and then in VAE. Let's click save. If you want to supercharge your content creation, definitely check out Higsfield, the sponsor of this video. They've just added the most capable video generator out there, Seed Dance 2.5. The biggest improvement from this model is that you can now generate up to 30 seconds of video in a single pass with multiple shots and an actual narrative with audio built-in. You can also extend an existing generation with new shots while keeping the same characters, locations, pacing, and overall look consistent. What really stands out is the reference system. You can feed it up to 50 references at once, including 30 images, 10 videos, and 10 audio files. So you can provide your characters, environment, visual style, motion, and soundtrack all in one generation. You also get much more control over editing. For example, you can specify exactly what happens during different timestamps, change just one section without affecting the rest of the video, or move the same performance into a completely different environment, or even change the camera angle while preserving the characters and action. Seed dance 2.5 supports text to video, image to video, video to video, and other references, giving you ultimate flexibility on your video creation. And right now is the best time to try Seed Dance 2.5 on Higsfield because they are offering unlimited Seed Dance 2.5 for up to 33 days. Terms and conditions apply. And if you're up for the challenge, check out the Higsfield Global Film Festival, which has a massive $1 million prize pool. Simply create any original film inside Higsfield. It can be any story or genre and submit it for a chance to win cash prizes and other rewards. Try Seed Dance 2.5 in Higsfield today using the link in the description below. All right. Afterwards, simply press R to refresh your model list. And then down here for the unit, you can select this new V4 A8 version, which is a lot smaller. And then for the video VAE, you can now select this int8 version, which is again a bit smaller. And with this setup, plus with spectrum and patch sage attention, you could potentially run this pretty quickly with just 8 GB of VRAM. To make this even faster, you could probably reduce the step count to something like 16. I find that often times you don't even need 20 steps. So, let's try an example. I'm going to upload this car. And then let's write an orbit shot of the car zooming in and out quickly. And let's press run. And here's our result. Not too bad. All right. So that sums up some additional ways you can speed up your generation or fit this workflow with even lower VRAM. Now if this official workflow looks pretty complicated to you with all these nodes and noodles, there's actually a much easier and well-designed workflow for you to use and this is called Comfy Miniax H3 Easy by NKXX188. This is actually incredibly simple. You can basically run text, image or reference to video all in just one workflow. So first we need to clone this repository. So simply click on this green button and then copy this URL and then in your comy folder in custom nodes at the top here type in cmd to open this up in command prompt and then let's type get clone and then paste in the URL over here. Now for me I've already downloaded this. That's why it says it already exists. But if it's your first time it should proceed to download this Miniax H3 easy folder in your custom nodes. So, simply open this up and then double click on this workflow folder. And there are two workflows here. They're quite similar. I'm just going to show you this first one. So, let's drag this onto my Comfy UI interface. And you should see this. This is pretty much it. Look how simple and minimalist this is. So, first of all, let's load up our models. I'm going to use the smaller W4 A8 model, which I just showed you. Same with the reference one. And then for video VAE, I'm going to use this smaller int8 one. And then down here, actually, let me move this node up here. This is where I can choose either image to video or first frame last frame or reference to video. So if I choose image to video, then it's going to automatically select this model. If I choose reference to video, it'll automatically select this model. So everything is just built into one workflow. So I don't need to switch between different workflows. Let me set this back to ITV. And then here's where I can select the resolution. Again, this looks way cleaner. I don't need to refer to this table. It's just the regular like 360, 480, 1080p, which I'm familiar with. And then here's the aspect ratio, the duration. That's pretty much it. You can also expand some advanced settings like frame rate over here. And then if you do choose image to video, then here is where you would upload an image. Now, the nice thing about this is, let's say you're doing reference to video, which allows you to add different types of media onto here, right? You can add images, video, and audio. Well, all you need to do is just input another node. For example, let's do load video and put it over here. And you can also just connect this to the media input. Or you can also do load audio. So, let's insert a load audio node. And we can also connect this directly to the media output. Again, this is way cleaner than the official reference to video workflow, which forces you to drag your input references into one of these connections. This is just way simpler. Another thing I really like about this is you can actually refer to your references within your prompt. So for example, I can write characters in and then if I press the add sign, I can select from my input references. So I want to say the characters in my video are wearing the headphone in this image and dancing to the music from this audio file. And then if you want, you can also add some additional acceleration methods here. And then here's where you would set the sampling algorithms as well as the step count. And that's pretty much it. It's a lot simpler to use. So if you're interested in this workflow, I'll also link to this in the description below. Honestly, I think the Comfy team should hire this guy to design their official workflows. All right. Next, here are some interesting things that you can get Miniax to do. It turns out you can just prompt it to generate audio. This can be like a music or sound effect generator. So let me exit out of everything and start from scratch. What you first need to do is use the imagetovideo workflow. And for the load image node, we can just press Ctrl +B to bypass this so that there is no image. For the resolution here, we can decrease this all the way to the lowest value, which is 0.1. And here is where we can enter our sound effect or music that we want to generate. For example, let's write cinematic orchestral music for an epic battle scene. Now if you expand this corner, we can keep all the settings the same. But basically over here, we can simply click on this create video node and press Ctrl +B to bypass it. We can also bypass the VAE decode node. All we want to do is take the audio, right? So simply double click anywhere on the interface and then search for save audio and you should see this from Comfy Essentials. So let's click on this and place it in here. And we simply need to link the audio output into this. And here you can select different formats. Let's select MP3. And that's pretty much it. So, what we're doing is we're essentially using like the lowest megapixels available for the video because we don't actually need it. And we're only going to extract and download the audio from the generation. Let's press run. Now, when you run this, it's going to say save video is missing. That's because we disabled the previous steps for creating the video. But you can ignore that because it'll still generate our output. >> [music] >> So that is indeed some epic orchestral music for a battle scene. Here I'm only showing you a demo of 5 seconds, but remember over here for duration we can increase this to whatever you want. Here's another example. Let's try something like a man screaming AI never sleeps. And this week has been absolutely insane. Epic rock background music. And then for the duration, let's set this a bit longer to like 8 seconds and then press run. All right, here's our result. >> AI never sleeps and this week has been absolutely insane. And there you go. It follows my prompt pretty well. All right, so that's how you can use Miniax to just generate audio. What about using it as an image generator? It turns out you can do that, too. So, I'm going to link to this page in the description below from this user, and they've created a pseudo image generation workflow for Miniax. So simply download this workflow file anywhere on your computer and then afterwards drag and drop the workflow onto your Comfy interface. So you might see some missing models here. Let's first go ahead and select the appropriate models for each of these nodes. And then afterwards here is where we would enter our prompt for the image. And let's try something like this. Anime style. A young woman with long silver hair and fox-like ears wearing a shrine maiden outfit, red ribbon sash. She's kneeling beside an ancient mosscovered stone lantern, gently placing small paper offering while soft fireflies drift around her. And then here is where we can select the aspect ratios. So sure, let's go with 2 to three at 1 megapixel resolution. And here is where we select how many frames for miniax to generate. Now since I just want an image, I'm just going to set this value to one. And then here are all the regular settings. And over here, because I set the length to one, the batch index should also be just one. And that's pretty much it. Let's press run. And here we go. Here is indeed a shrine maiden with white hair and fox ears next to this temple thing with fireflies around her. So that's how you can also use Miniax to generate images. Now, it's not really a frontier image model. I think Crea 2 and Adio are much better at this. So at least for me, I wouldn't really use this as just a dedicated image model. Next, I also want to talk more about how to prompt Miniax correctly because a good prompt makes a huge difference in your video generations. In fact, the Miniax team has released a full prompt writing guide for Miniax H3. So, it actually follows a very fixed format and this varies for text to video, image to video, and reference to video. And if you follow this format, your videos are going to look a lot nicer. I'll link to these guidelines in the description below, but for example, for text to video, they recommend that you include these three fields. So, there's an integrated multimodal description. So, it describes the visuals, actions, shots, speaker, dialogue, etc. And then the overall soundsscape. So, what other ambient sounds or physical action sounds or other non-verbal sounds that should occur in the video? And also any background music that the characters cannot hear, but only the audience can hear. For example, we can take this simple prompt. A woman walks through a rainy Tokyo street at night and opens her umbrella and we can enhance it to fit the following format. So, for example, the first section should just describe the scene and all the actions that occur including the camera motion, what exactly she does. So, here she should open a red umbrella and then the overall scene. So, there should be like neon reflections simmering on the wet pavement. A train passes in the background. And then next we move on to the overall soundsscape, which should be this plus any background music. And here's our result. So you can see it's much better if you prompt it this way and you describe all the actions including the background music, the camera movement, the overall scene, plus the overall soundsscape and the background music. So that's text to video. And then for image to video, it depends if you're using your input image as the first frame or if you're inputting two images, one for the first frame, one for the last frame, or if you're just using the input as the last frame. Let's do a quick first frame example. So here it says the description should first establish the style, subjects, composition, scene anchors in the image, and describe the next action. Character identity, clothing, colors, key objects, etc. should remain consistent. So the recommended structure is the first frame anchor action onset continuous development result or reaction. So what I tend to do is just copy this into chat GPT and then I would upload my input image and then say give me a prompt for the video generator based on the below instructions and this input image and then I just pasted the instructions from the Miniax prompt guide. So, it gave me this. And then in my image to video workflow, instead of just writing a car driving down a highway, I'm going to copy and paste the prompt from chat GPT in here and then press run. All right. And here's my result. [music] As you can see, this looks much better than my earlier video where my prompt is just a car driving down a highway. This looks a lot more cinematic. The camera movement, the music, everything is just a lot better. So, those were some examples of how to prompt better for text to video and image to video. They've also released another prompting guide for reference to video. I'll link to all of this in the description below. Next, if you're wondering about GGFs for Miniax, we already have a ton available. For example, this user Abi Ray has released some GGUFS and the Q3 version is only 15.6 GB in size. Unsloth also released his own GGFS for Miniax H3 and the Q2 version is only 6.7 GB in size. So you could potentially even fit this on like just 4 GB of VRAM. However, at least if you're using Comfy UI, it's actually no longer recommended to use GGFS. It's quite technical, but basically Comfy UI is actually more optimized if you just run a quantized Conrot version. Even though these are larger models with the right optimizations, you can fit them in much lower VRAM. And in general, the quality of these Convert versions is much better than a GUF, which tends to have more compression and sacrifice in quality. But in case you're still interested in running GGFs for whatever reason, then I mean the Unsloth version is super tiny. The Q4K version is only 6.7 GB. So I'll also link to these GGFs in the description below. So, that sums up my more advanced tutorial on how to run miniacs better and faster and some additional things you can tweak. The open- source community is incredibly fast. There's so many updates that come out every date, it's really hard to keep track of everything. If you're interested in more minax updates, be sure to check out my ex or Twitter where I will continue to post some interesting findings. Hopefully, this video was helpful for you. If you run into any errors, welcome to copy and paste the exact error message that you see in the comments below, and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So, to really stay uptod date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.