Transcript · 25,126 chars
We have a new open source video generator that you can run locally on your computer. It's called Min Max H3 and this is by far the best one available. The best thing is you can even run this with as low as 5 to 6 GB of V RAM. Plus it has audio built in and it's insanely good at character consistency, world knowledge, instruction following, and even generating and syncing to music. It's just way better than any other open model we've seen so far. In this video we're going to go over all the incredible things that it can do. Plus of course I'm going to show you how to install it on your computer so you can run it for free and unlimited times offline. Let's jump right in. First of all, here are some ridiculous generations from the open source model. As you can see it has a ton of existing knowledge built in. These were just generated with text prompts and no image references. But as you can see it can do cartoons, it knows all these different characters, TV shows, and movies. It has no problem doing regular scenes of people talking or moving or even high action scenes. It's really good at understanding even unusual actions or prompts. Here's another really tricky prompt of a realistic photo but with hand drawn animations. It's very tricky but as you can see it's able to follow this very well. So it's absolutely exceptional in terms of prompt understanding. It's also great at rendering different animation styles. So here's a clay mation example or here's an anime example. In fact the awesome thing about Min Max is that it's multi modal. So this can take in images, video, and even audio to use as references in its output. Here's an example where we can put these three anime characters together and as you can see it's able to render a video of them very consistently. Plus the jiggle physics are also top notch. Or here we can easily get it to create a commercial from these reference images. Here's another example where we can generate a commercial from just this reference of a Nike shoe. Or here's a short vampire romance drama with these two reference characters and this background. >> Hello. >> You should not be here. >> Who are you? >> The man who owns this house. >> You're not human. >> And you smell like war. >> I'm not afraid of you. >> Then you should be. >> It also seems to be very good at generating text and interfaces. So, here's an example where I upload a screenshot from this website and it's able to make an animation from this very well. This is also surprisingly good at creating gameplay scenes. We can just upload a screenshot of this first-person shooter and get it to generate some footage of the game. So, this can potentially be very useful for game design and ideation. Here's a more complicated example where I can upload a video of this man walking down the street plus some voxel image references and it's able to transform the scene with these voxel elements. Or here's another example where we can copy all the actions from a reference video on to a new video with different characters based on a reference image. You can also add or remove characters and objects in an existing video, as you can see from this example. It's able to understand and detect one of these characters and make a clone of them. Or here's another example where we can prompt it to replace the scenery with this image after the person opens the blinds. And it's able to execute this flawlessly. It's basically like an open-source Gemini Omni. So, those are some demos of all the incredible things that it can do. Hopefully, this gives you a sense of how flexible and powerful MiniMax is. I really can't believe they're just open-sourcing this for us for free. Anyway, next let's go over how to install this. So, in this video, we are going to use a platform called ComfyUI to run MiniMax H3. This is one of the most popular platforms for running open-source image, video, and audio generators locally on your computer. So, if you're not familiar with ComfyUI, definitely see this video first where I go over how to install and use it. Now, assuming you do have ComfyUI, the first thing you should do is within your comfyUI folder, double click on this update folder, and then run this update comfyUI.bat file. This will update your comfy to the latest version. All right, after updating, it says press any key to continue, and it'll just automatically exit the terminal. Next, we can start up comfyUI. All right, after opening up comfyUI, simply click on templates on the left sidebar, and then search for Minimax at the top, and you should be able to see a few different workflows. Now, in case you don't see these workflows for whatever reason, I'll also link to this page in the description below, where you can manually download the workflow and then just drag and drop it onto your comfyUI interface. Note that some of these have the API tag, which means you're actually running Minimax from the cloud instead of locally, so ignore those ones. What we are going to focus on are these bottom ones: text to video, image to video, and reference to video. Let's go over text to video first. So, let me click on this to open up the workflow, and here it says it has found two errors. So, let's view the details, and here it says it's missing some models. So, first of all, we need to download all the models for this to work. So, I will link to this page in the description below, and you just got to click into each of these folders and download the appropriate model for your computer. Let's click into diffusion models first, and here this first Minimax model called FL2VA, this is basically for text video and image to video. And then, this bottom one, Ref2VA, this is for reference to video. So, I'm going to show you some text video and image to video examples first. So, we're going to download one of these ones. Now, the full model is 66 GB, which would probably not fit for most of you. There's also an INT8 version, which is 34 GB, and then a pruned FP8 version or a pruned INT8 version, which is 21 GB. The nice thing is some users have reported that these models work on as low as just 12 GB of VRAM. And if you have even lower VRAM, don't worry, I'll show you some other options you can use later in the video. Now, note that the largest full model is the highest quality, and as it gets more compressed, there is some sacrifice in quality, but it's still very good. Anyways, download the version that works on your hardware. Since I have a more recent Nvidia GPU, I'm going to download this FP8 one. And this goes in ComfyUI, in models, and then in diffusion models. Let's click save. All right, afterwards, let's go back to the root folder, and then next we also need to download a text encoder. So, this uses Qwen-3-VL-32B as the text encoder. Again, you are given three different versions with different compression. Choose the one that would fit on your device. Now, for the text encoder, this doesn't have to fit on your VRAM. This can also just fit on your RAM. Anyways, this is what I'm going to download, so let's click on download. And this goes in ComfyUI, in models, and then text encoders. Let's click save. All right, finally, we also need to download the VAE. So, let's click on this, and then we need to download both the audio VAE and the video VAE. So, let's proceed to download both. Both of these go in ComfyUI, in models, and then VAE. Let me also download the second one in the same folder. All right, after downloading everything, let's go back to our workflow, and let's press R to refresh our model list, and then over here is where we can open up each of these drop-downs to select the model that we just downloaded. So, for example, for the UNet, I'm going to select this MiniMax FP8. For the clip name, I'm going to select Qwen-3-VL-32B. And then for VAE, I'm going to select this one. For audio, let's go with this one. And that's pretty much it. It should remove all the errors and the red outline from this node. And then afterwards, here are some additional settings you can set. So, you can choose from all these different aspect ratios. For megapixels, you can refer to this table over here. So, for a 480p video, it'll be 0.4 megapixels. And you can set the duration over here. Now, if you click on this button to expand this workflow, notice that it actually looks like this. And here are some additional settings you can set like the sampler and the scheduler, which are basically the algorithms used to generate the video. I would just leave this at the default, and then this is the number of steps it takes to generate the video. Again, I would just leave this at the default of 20 steps for now. All right, so to collapse the workflow, simply click on the parent component up here, which takes you back to this compressed workflow. And that's pretty much it. Let's press run. All right, so that wasn't too bad. It took around a minute per second, and here's the generation. Notice that this has audio built in. >> [music] >> And note that this is automatically saved in your video folder within the output folder. All right, so that's text to video. Next, let's go over an image to video example. So, again, let me pull up templates and then search for Min Max, and then let's select this one, image to video. Again, make sure you select the one that does not have an API tag. So, the workflow will look like this. Now, since we've already downloaded the models from the previous step, we just need to load the models down here. Let me select the model that I downloaded. For the text encoder, it'll be Qwen VL 32B, and after selecting the models, this red outline should be gone. And then here's where we upload an image. So, let me upload this image, and then for the aspect ratio, let's set it to 16:9, and then again at 480p. You can also choose the method on how you want to rescale the image. For me, I just tend to leave it at the default of nearest exact, which works very well. And then again, over here, I can set the duration. And let me just add a simple prompt here, and that's pretty much it. Let's press run. All right, so that again took around 5 minutes, so around a minute per second. Here's the result. All right, so that's image to video. Next, let's go over the third and final workflow, which is reference to video. This is the most powerful workflow you can use. So, let me first click into the workflow, and once you click into it, again, you might see some errors and missing models. So, let's go ahead and fix all these red outlines. Now, for this reference to video workflow, for the load diffusion model, make sure you download one of these four models. The lowest prune versions are like 21 GB, and some users have reported to even successfully run this on as low as 12 GB of VRAM. Anyways, for me, I'm going to download this FP8 version, so let's click download, and this goes in ComfyUI, in models, and then in diffusion models. Let's click save. All right, so afterwards, back here, let's press R to refresh our model list, and then in this drop down, I can select this ref to video model. And then afterwards, for load clip, we can just select the Gemma 32B model, which I downloaded for the previous workflow. And then for the video and audio VAE, looks like it has already selected these by default. All right, so the model selection is done. Next, let's go over the workflow. So, what reference to video does is this can take in images or even video or audio to use as input. Let me first show you an image example. So, over here is where you can plug in multiple images over here for it to use as a reference. Now, by default, this workflow has two load image nodes. If you just need one, you can just simply click on this and press control B to bypass it to just load one image. But for me, I'm going to show you an example of two images, so let me un-bypass this, and then let me upload some images. All right, so I'm going to upload this character plus this car, and then down here is where I would enter the prompt. Note how you would refer to your attachments in the prompt. So, you would type something like picture one or picture two. All right, so here's my prompt. The woman in picture one walking to the side of the car in picture two, she opens the door and gets in the car in a dark, misty forest at night. Here's where you would select the aspect ratio. So, again, you can choose from all these different aspect ratios, and here's where you can choose the resolution. So you can refer to this table over here. 0.4 megapixels is roughly 480p. And then here is where you would select the duration of the video. So right now it's set at 5 seconds. And that's pretty much it. Let's press run. All right, and afterwards here is our result. >> [music] >> All right, now instead of just using two images, let me show you an example where we can upload a reference video instead. So I'm going to get rid of this load image node, and double click anywhere on the canvas, and then search for load video. Now this one load video by comfy essentials does not work. You'll need to use this load video upload node instead. So let's put this over here, and then let me upload this green screen video. And then let's connect this to the ref video zero input up here. And for the format, I'm going to set it to none. And then for the image, let me change this to a dark background like this. For my prompt, I'm going to write replace the background in video one with the dark bamboo forest in photo one. And that's pretty much it. Let's press run. All right, here's our result. Now instead of video, you can also upload audio. In fact, its audio and music understanding is amazing. So let me get rid of this load video node, and next let me click anywhere on the canvas, and search for load audio. And we need to use this one load audio upload. So let me click on this and add it anywhere here. And for the audio, let me upload this track. >> [music and singing] [music] [singing] >> All right, so afterwards let's drag this to the reference audio node over here. For the duration, this should be 14 seconds. And then for the image, let's upload this image of a woman singing. Now, because this is 14 seconds, let me also set the duration to around 14 seconds. And then for the prompt, I'm going to write a music video of photo one singing this song, audio one. She is singing passionately on a windy cliff. Her hair is blowing in the wind. Let's press run. All right, and here's the result. >> [music and singing] [music] [music and singing] >> So, as you can see, you can easily create music videos from this. In fact, this is a pretty bad example. You can probably prompt it further and add more elements and different transitions and cuts to make this even more epic. All right, so that sums up how you can use all these different inputs including images, video, and audio using this reference to video workflow. All right, next let me show you how to speed up your generations even more. Now, like I said, at least for me, the base int8 or fp8 model can generate at around 1 minute per second. However, there are many hacks you can use to speed this up by like 30 to 40%. So, I'll link to this page in the description below. The first thing you need to do is to download Sage Attention. Now, it's kind of a pain, but I'll make this as beginner-friendly as possible. So, in order to download Sage Attention, you need to download the version that matches your PyTorch and CUDA version. So, to check your PyTorch version, in your ComfyUI folder, note that I'm using the Windows portable version, you should see this Python embedded folder. So, double-click on this, and then at the top here, type in cmd to open your Python embedded folder up in your command prompt. And you just need to paste in this line here. I'll put this in the description or a comment so you can just easily copy and paste it. But basically, this uses Python to run this snippet of code. It imports torch, and then it prints out the version of torch. So, let's press enter and you can see here I am using torch version 2.6 and it's using CUDA 12.4. And one thing I forgot to mention is you also need to figure out the Python version that ComfyUI uses. So, again, in your Python embedded folder at the top here, type in CMD to open this up in command prompt and then type in So, you can see for me I'm using Python version 3.11. All right. So, now that you have your PyTorch and CUDA and Python versions, simply click on this page, which I will also link to in the description below, and over here simply look for this SageAttention that fits your CUDA version plus your torch version plus your Python version. Once you found the right one, simply click on it to download it to wherever you want. All right. After downloading the pre-built SageAttention wheel, you can just install the wheel using pip install in your terminal. If you do have the Windows portable version, then simply click into your Python embedded folder and then at the top here type in CMD to open this up in command prompt. And then afterwards, we need to type python.exe {dash} m pip install. Then you would basically copy the path to your downloaded file and paste it in here. So, after pressing enter, it should proceed to install SageAttention for your system. Now, for me I see this message because I already have it installed. The next step is you also need to install KJ nodes. So, I'll link to this page in the description below. Here are the instructions on how to install this. The first step is to clone this repo into the custom nodes folder. So, going back into our ComfyUI folder, simply click on ComfyUI and then custom nodes and then at the top type in CMD to open your custom nodes folder in command prompt. Afterwards, on this KJ nodes page I'm going to click on this green button and then copy this URL and then back in my terminal I'm going to type in git clone and then paste the URL in here. So, this is going to proceed to clone this repository into a folder within your custom nodes folder. And then afterwards, you also need to install the dependencies for this. So, in your ComfyUI Windows portable root folder, at the top here, simply type in CMD to open this folder up in your terminal, and then copy and paste this line into here. So, this is going to proceed to install all the requirements for KJ nodes. All right, that's all you need to do. Simply restart ComfyUI, and after restarting, let me show you how to apply this to either the text-to-video workflow or the image-to-video workflow. It's exactly the same thing. Again, the workflow looks like this. Simply click on this corner to expand the workflow, and what you need to do to speed this up is to place some additional nodes after this load diffusion model node. So, let me double-click anywhere on the canvas and search for patch, and you should see this patch sage attention node after you've downloaded the KJ nodes. So, let me select this and place it onto the canvas. And for this sage attention setting, let's set it to auto. And we basically need to connect the model to here, and then connect the outputs to where the model normally goes. In our case, it would be the basic guider and the basic scheduler. So, that's one way on how you can speed up the generation by like 20 to 30%. Another quick hack is you can double-click anywhere on the canvas and then type in easy cache, and you just need to select this one, easy cache, which should already be available. This is part of the default Comfy nodes. So, let's put this anywhere on the canvas, and this basically can go in between these two nodes. So, let me reconnect the model to here, and then reconnect the output to here. So, by adding this additional easy cache node, it can also speed up the generation by an additional 5 to 10%. And you can stack both of these together. Here's another hack on how you can speed this up even further. So, here's another node called ComfyUI Spectrum, and this is specifically designed for MiniMax H3. If you scroll down a bit, here is how you can install this. So, simply go into your custom nodes folder, and then at the top here, type in CMD to open this up in command prompt. And then afterwards, I will copy this line and paste it in here. And this is going to clone the repo into our custom nodes folder. All right. Now, after restarting ComfyUI, simply double-click anywhere on the canvas and then search for spectrum apply and you should see spectrum apply Minimax H3. So, simply add this node to the canvas and you can put this anywhere before the basic guider and basic scheduler. So, let's just link it over here. So, here are three different methods you can use to speed up your generation. So, that's how you can speed up text to video and image to video. Now, the reference to video workflow looks a bit different, right? It looks like this. But, it's the same logic. So, we basically need to add the nodes after this load diffusion model. So, let me show you that really quickly. I'm going to search for batch stage attention and then add it here. I'm going to connect the model to here and then connect the outputs to basic guider and basic scheduler. And then afterwards, let's also add easy cash and then put it between these two nodes. So, that's how you can also speed up the reference to video workflow. All right. Now, like I said, the pruned models are reported to even work on 12 GB of VRAM. But, if you have even lower VRAM, well, you can use another platform called 1 to GP. Here, the author says that you can run Minimax H3 with as low as 5 GB of VRAM at 480p resolution. So, that's some insane optimization. Now, I've already gone over how to install 1 to GP in a previous tutorial. It's basically the same instructions as before. So, I'll link to this page in the description below if you're interested in installing 1 to GP. Now, it's still quite early. They've only released Minimax H3 for a few days, but we already have lower support for this. If you're not familiar with the term lower, this is basically a fine-tuned model which allows you to generate a certain character or style or action or effect. The awesome thing is one of the best toolkits for training Laura's called the Austras AI toolkit has already added support for Minimax H3. So, you can use this to train your own Minimax H3 Laura's. Now, this is quite technical and beyond the scope of this tutorial, but if you are interested, I will link to this page in the description below. Finally, it's important to also talk about the licensing of this. So, note that this is extremely permissive. Here it says you can basically use this for anything you want including commercial usage as long as your commercial products and services do not generate more than 20 million USD or equivalent in yearly revenue. If it does, then you'll need to contact them to work out an agreement. However, note that all these territories including the EU, UK, Korea, and the US are excluded from this Minimax community license. If you are in one of these places, you cannot use or run or modify this open model or its outputs unless you apply for a separate license. Now, this might sound extremely restrictive, but it's not. They just want to be legally safe because they are facing lawsuits from like I think Disney and some other companies. I mean, these places are basically legal landmines, so they got to be cautious. The nice thing is it's actually fairly easy for you to just fill out this form and get approved, so you can actually run the model if you're located in any of those places. In fact, I'll link to this page in the description below where it explains more about why they are doing this. And if you scroll down a bit here, it contains the application form which you can fill out so you can actually use this if you're in the US or the other excluded places. All right, so that sums up my installation tutorial of Minimax H3. This is by far the best and most flexible open-source image generator you can use right now. It's incredibly knowledgeable. It can generate all these different characters and art styles and different actions. It's amazing at following your prompt plus its multimodal capabilities are a godsend. Let me know in the comments what you think of this, what other cool or impressive things were you able to get it to generate. If you run into any errors with the installation, welcome to copy and paste the exact error message that you see in the comments below and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.