Transcript · 20,175 chars
Minemax H3 is the best free video AI king ever, and it's not even close. >> [screaming] >> Hello humans, my name is Kay Ovrlod, and boy oh boy, do I have some mind-blowing stuff for you today. Because yes, you heard it right, we have a brand new video AI model that was released called Minemax H3. That is simply the best free video AI model ever made, capable of generating videos from text, images, as well as from multiple other references at the same time. It is just incredible. So, today I'll show you how to install it, how to run it locally, and on RunPod, and show you how to get the best results possible. So, that being said, sit back, relax, and let's go. And to install Minemax H3, you have two ways. The first is of course by using my one-click installer that is available for my Patreon supporters. Just double-click on the installer, and then it will automatically install ComfyUI and all the models and nodes that you need to run Minemax on your computer. And the second way is to rent a GPU on a website like RunPod, and use my special RunPod installer, and run Minemax as if this was running on your local computer. And then once you have ComfyUI up and running, for this video I prepared a special Minemax Ultra workflow that you can find on my Patreon. Then you're going to drag and drop inside ComfyUI. And now we can finally have some fun. Now, before we begin, try to explain the workflow, what makes this Minemax H3 model so special, and why is everybody and their grandma talking about it. Well, Minemax H3 is a 33 billion parameter model that can generate videos from text or images in high resolution and with native audio, and that can run on your local computer. But that is not all. This model is extremely powerful and versatile thanks to its ability to do video editing and using multiple references at the same time to generate your final video. Now, obviously, Minemax H3 is a very thick boy, and you really need a lot of VRAM to generate videos at high resolution at a decent speed. But don't worry, even if you have 8 GB of VRAM, the model will work as well. Okay, so enough talking, let's actually show what this model can actually do and see how good the video generations are really are. Okay, so first, let's start with the text-to-video workflow, which is by far the easiest to understand. Here, once again, all of the most important part will be located in the first column right there. And before we start generating, we need to look at the available speed-up options that are available for the model right now, which are stage attention and spectrum. Now, basically, all of these two allows you to generate your video much faster at the expense of a small hit in video quality, especially if you're using spectrum. Now, there's also plenty of other nodes that can do that, but stage attention and spectrum are by far the best. Oh, and by the way, speaking of speed-ups, this is entrepreneur in the future because as I'm editing this video, we actually already have the release of the MiniMax Turbo Laura. That's right. Now, obviously, right now, it is still a Laura in training. It is still not done. As of right now, the only version available is the one at 500 steps. And to make it work, you need to use the special comfy UI version. And I've already tried it and it is very good for a Laura that is not even done training. And although in this video, you will only see videos generated without the Laura, just know that if you're watching this video right now, the workflow on Patreon will already be updated to support this Laura. And you can either choose to run it without the Laura. And basically, the differences between the workflow is very simple. Without the Laura, we are basically using the res multi-step sampler with simple scheduler and 20 steps. Whereas, if we are using the MiniMax Turbo Laura, we will be using the Euler sampler with beta scheduler at eight steps. That's it. That is pretty much the only difference. So, instead of generating a video with 20 steps, you are now able to generate your video in around eight steps or even less. And of course, as I said, this Turbo Laura is not even done training, so in the upcoming days, we should have an even better Turbo Laura that we can use to generate absolutely amazing videos. So, yeah, there you go. Okay, so then right here, this is where you're going to input your prompt for the video generation. However, keep in mind that MiniMax H3 has a very, very specific prompting style. And I'm not going to go into details on how it works and everything. If you want more info, you can simply either read the note I have input right there, or if you want to simplify your life, you can simply go there and then copy and paste this entire note inside software like ChatGPT, and there you go. And now you can simply just ask ChatGPT to write you a prompt for whatever video you want. Oh, and also, right before you generate, let's also have a talk about the resolution, because here you have actually two options to choose the resolution of your video. Either you use the resolution selector that will automatically select the resolution following the aspect ratio guide right there, or you can simply just leave this option enabled and then choose yourself the resolution that you want. Now, if you don't want to type anything yourself, it is much easier to just use the resolution selector. And also here, I have highlighted all the resolutions that I recommend you to use, and it is simply either 480p, 720p, and 1080p. So, depending on your GPU, you might want to choose either one of these three, because these resolutions are more standard, so they are much easier to upscale. But obviously, if you have like a 5090, you can definitely go to the 1080p and then generate your huge video from scratch. And then here, finally, this is where you input the video duration in seconds. And of course, the longer the video, the longer it will take for the video to be generated. So, yeah, I mean, very simple stuff. If you've done it once, you've done it a thousand times. So, let me actually just generate a video. Let me write a prompt. Then I'm going to leave the 0.9 megapixel resolution, use 10 seconds instead, and then I'm going to click run, which in the end will give us something like this. >> I'm dead, Jerry. The city says I'm dead. Well, that should clear up the parking tickets. Good news. I got you a box. >> So, yeah, I mean, it's pretty good. Now, as you can see, like this is very, very decent. Now, this was generated in one single shot. This is not cherry-picked or anything. I mean, it is really, really good. Actually, I even generated another version with the same exact prompt, but this is not at 1080p. Take a look. >> I'm dead, Jerry. The city says I'm dead. Well, that should clear up the parking tickets. >> [laughter] >> Good news. I got you a box. >> So, as you can see, in terms of quality, except resolution, there aren't that many differences between the two generations. Hence why it's probably better to generate at a lower resolution, like 720p, and then upscale it to a higher resolution later. And of course, the model can do a lot of very cool stuff, like a Breaking Bad parody, for example, or memes, kind of like this. >> Jesse, we have to cook. >> Yeah, there you go. Like, super, super funny uh Breaking Bad, Baking Bread parody. We all know that. It's not new, but it's still really, really fun. And of course, it can do more than just show parody or real-life stuff. It can generate anime parody as well. Like, take a look. >> I'm going to become the king of the hook again. >> Yeah, I mean, this is really, really incredible. Like this is insane. Like you can make your own anime parody or anime stuff already immediately right now. Like we've just a simple text prompt. Now this video was generated at 1080p, so you get like the best quality possible. But if you don't want to generate at such a high resolution, you can also generate at 720p. And if you do, you get something like this. >> I'm going to become the king of the hook. Okay. >> So yeah, I mean as you can see, very decent as well. Maybe not as good as the 1080p resolution, but it's still really, really fantastic. I mean it's really, really cool. So yeah, I mean I could spend hours and hours generating new videos to show you how good it is. But I'm sure that by this time that you're watching this video, you've already seen multiple examples anyway. So let's move on and see what else this model can do. Now the second thing that this model can also do is that instead of generating videos from text, you can generate a video from a image, which is definitely something that I prefer. I do like to generate my images on a separate model like Create 2, and then using an image to video model to bring it to life. I find this much, much better. So here basically the principle is exactly the same, except that now you have the choice of inputting two different images. Either you input one single image and you just make that image into a video, like I'm going to show you right now. Let's say I input this image, this selfie of a warrior on a battlefield, then I'm going to write my prompt. Then here I either have the possibility of using the original image size, or if I don't want that because it's going to take too much time, I can just use the normal resolution selector and then generate at the resolution written right there, which is going to be 720p instead. Input the video duration and now if I click run, which in the end will give us something like this. >> Well, I guess I'm going to be late for bingo tonight. >> And here, I mean as you can see, like this is fantastic. This is really, really good. This is great. I mean, huge, amazing quality. The generation is fantastic. The people behind him walking, you know, sitting down and everything. I mean, this is really super, super impressive. Super realistic, super good. I mean, yeah, I mean, this is really, really good. And of course, this is not the only thing that you can do with this particular workflow. And that is because you also have the ability to input a second image. Because this workflow can also work as an in-between video generation. So, if you put a first frame and a last frame, you will be able to generate a video between those two beginning and end frames. So, like for example, if I enable this last frame and then I input frame one with the warrior sitting on the ground, the last frame where he is shouting, if now I write my prompt and I click run, which in the end will give us something like this. >> [screaming] >> So, yeah, I mean, as you can see, this is really fantastic. I mean, we used only two separate frames, one beginning and one end, and with one single prompt, Minimax made a video in between those two frames. And I mean, I mean, what do you want me to say? It looks really, really good. It's it's fantastic. I mean, LTX could do it, too, but Minimax is really just on another level. It is really super, super good. But of course, once again, this is not the only thing that this model can do. Because not only can do all of that, but more, because Mini Max also has a very specific video model that allows you to use several different references to make your video. And the way it works is very, very similar to the image to video workflow, except that this time you have three different types of references. You can actually use image reference, video reference, and also audio reference, which really makes this workflow and model extremely powerful and extremely versatile. Now, the way you do this is the same exact principle as what we did with the image to video, except that this time it's kind of up to you to decide how many references you want to use. And I think the limit is around nine references that you can use at one single time, which is really just insane. Now, by default, I have input here like two different references for the images, one reference for the video, and one reference for the audio, but if you want to add more or disable different references, you can do so very easily as well. So, like for example, let's say that I want to add an additional reference image, all I have to do is just click on one node, then press control C on your keyboard, then control V to copy and paste that node, and then what you're going to do is that you're going to click and then connect it to that node right there. And you're going to connect it to the reference image node right there. And as you can see, now we have three different reference images that we can use. And you can do the same thing with video, audio, it's the same exact principle. And if you don't want to use a certain reference, you can just click on the node, and then click on this little button to bypass the reference, as well as right there. So, like for example, let's say that I upload an image of this character sheet right there, then another character for reference image two, and then for the third image, the interior of this castle right there. And now, if I write my prompt, which once again, I made it with ChatGPT, I'm not going to write all of that myself. Like, no way. And now, if I click run, which in the end gives us something like this. >> Wow. So, it can generate multiple characters inside the video? >> Yep, it looks like it. >> Amazing. >> Yeah, it sure is, buddy. I mean, this is really just incredible. I mean, it's insane. I mean, >> [laughter] >> I mean, I'm I'm kind of shocked. I'm going to say it's It works so good. It worked better than I thought it's going to be. It really like took the characters from the sheet, the character sheet, all the images, and then made video like this with full dialogue interacting with two different characters, two different voices, and I mean, it looks so good. It's It's amazing. I mean, yes, it was generated at a lower resolution, so you can have even higher quality if you generate at a higher resolution, or if you upscale it. But, I mean, just like that already, it is It is really, really good. I mean, it's It's incredible. I got to say, it's insane. It's fantastic. You don't need any LORAs, you don't need any training. Everything is already done for you. Just Just incredible. And also, of course, there's plenty of other ways of using this, using different types of references. Like, for example, let's say that I want to use a video reference. Let's say I upload this video of this woman kind of like turning around, and I want to like edit this video. I want to replace that woman with a different character. Let's say I want to replace it with that girl. I'm actually going to upload that girl into image zero. Then, I'm going to disable all the other references. I'm going to write my prompt, check the video direction. And now, if I click run, which gives us something like this. I mean, listen. I mean, is it amazing or what? What do you want me to say? It's It's incredible. I mean, we literally just like took a video reference and then replace it with a character from an image. No Laura training, no anything. We just like put everything together and it just worked. It just works. All right, it's Huh? It It's I'm I'm blown away. I'm blown away. If you can tell, I'm blown away. It is incredible. I mean And obviously, this is once again not the only thing that we can do. We can do much, much more than that. We can do so much more. I can spend hours and hours and hours showing you everything that you can do. Like for example, you can even like simply use like one single video as a reference. If I use like this video from, you know, Soprano with Paulie talking, which I'm not going to play because I kind of want to avoid any, you know, copyright issues if there is any. But what I can do instead is just use that video as a reference and then make him say something completely different with my prompt. And because the video not only can reference the video itself, but also the audio, we should also be able to copy his voice fairly well. So now if I click run, so that in the end we get something like this. >> Wow, this model is incredible. If Tony hears about this, he's going to be pissed. >> So yeah, I mean, what what do you want me to say? It's It's great. It's fantastic. It's It's everything that we've ever wanted from a video model. It's It is that simple, guys. It is that simple, okay? Like Why Why Why are you still here? Just Just stop listening to me and, you know, like and and do your thing. Do your thing. Try it out. It It's It's amazing. It's amazing. Oh, and also one little trick. If you want to upscale your video from a lower to a higher resolution, I highly recommend using my LCX 2.3 ultra workflow version 3 because in that workflow I have added a video enhancer upscaler, which is actually really really good at enhancing and upscaling at the same time your video. So, you can literally just upload a video made with Minimax. This is a video that was made and generated at 720p, and then you can like input a resolution of full HD, for example, and then if you click run, it will use LCX 2.3 to take this video and then upscale it at a higher resolution and enhance it at the same time. So, that in the end we go from this, a 720p video that looks like this, there we go, to a 1080p video. Yeah, I mean, there is really a huge difference in quality between those two videos. Hence why this workflow is really insanely good because not only we are upscaling, but we are also enhancing the original video. And obviously, it takes way less time to generate at 720p and then upscale it at 1080p rather than generating the full video at 1080p instead. So, yeah, I definitely recommend you to use this upscaler whenever you want to upscale one of your videos made with Minimax. Now, the one a little bit of a weird stuff with this model is the fact that the license is very very strange. Basically, it says that you cannot really use the model if you are in the European Union, United Kingdom, Korea, or the United States, which is, you know, very very strange. But, that is because Minimax is currently in a lawsuit against Disney, and they kind of want to avoid any issues with the model. Now, if you want to make sure that everything is fine, you can simply click on this little link and just like sign a waiver giving you the right to use this model however you want. It takes like 30 seconds to do and then you're good to go. Now, obviously, if you are like a normal person, you don't necessarily need to do that. I you know, you're fine. But in my case and a lot of people's cases, since we do have companies and we do represent companies ourselves, just to make sure that everything is fine, we need to kind of go through this, sign that waiver, and everything is going to be fine. So, yes, YouTube, if you're watching this video, just know that I did receive confirmation and I do have the right to use this model. Thank you very much. I love you. No sarcasm intended. So, yeah, there you go. This has been Min Max H3, simply the best open weight video AI model ever made. An absolute beast of a model that you can run on your computer right now. And once again, in this video, I barely scratched the surface of what this model can do. And even with my version one of the Min Max workflow, you can pretty much do absolutely insane stuff. But don't worry, this is definitely not the last time that I'll be making an update video about this model. There is really too much to say and too much to do. So, yeah, like once again, stop wasting time, download this workflow, download this model right now either locally or on RunPod, and just use it. It is incredible. Insane. So, that you and I can finally make our dreams come true. >> [music] >> And there we have it, folks. Thank you guys so much for watching. Don't forget to subscribe and smash the like button for the YouTube algorithm. Thanking also so much to my Patreon supporters for supporting my videos. You guys are absolutely awesome. You people are the reason why I'm able to make these videos, so thank you so much, and I'll see you guys next time. Bye-bye.