Transcript · 18,942 chars
This is currently the best open- source music generator you can use. You can run this with 4 GB of VRAM or less. Plus, you can even do cover songs and edit existing songs. And according to some benchmarks, this even beats Sunno V6, which is pretty crazy. So, this is called UA 2. And first of all, here are some demos. So, on the left is the prompt dictating the style and genre of the song. And on the right are the lyrics. First, here's a demo of a smoky jazz song. >> [music] [music] [music] >> The glasses clink in a quiet tune [singing] [music] under the hum of a fading >> [music] >> Lights flicker low. Whispers in bloom. [music] Neon smiles and a quiet laugh. [music] Friendships build [singing] on a fragile draft. Before I sleep, I'll keep [music] this far. a glowing memory [singing] in the dark. [music] Or here's another demo of a boogie woogie style of the 1930s. >> Hey, can you hear it? That shuffle through [music] the door. Old shoes on a new floor. My feet itching [singing] for more. Hey John, can you hear it? Boogie woogie. This rhythm [singing] turns me on. Let's go dancing [music] soon. I'm ready. I can't stand still. Come on. Hey John, can you hear it? Boogie [music and singing] woogie. Heartbeating like a tune. Spin me around this crowded room. Let's go dancing soon. [music] [music] Now, this also supports different languages. So, here's a flamingco example in Spanish. comfort [singing] [music] and truly. [singing] Hey [music] >> [music and singing] [singing] [music] [singing] [music] >> Or here's a folkronica example in Russian. Yeah. [music and singing] [music] No way. [music] >> [music] [music] >> or here's an emo example in Korean. [music] [singing] [music] [singing] >> [singing and music] >> up. Now, the really cool part about this is you can also input another song as a reference. So, for example, it can use the melody of an existing song, but you can change up the style. For example, here is old lang sign, but we are going to turn this into a groovy jazz funk style instead. [music] [music] >> [music] >> Should all acquaintance [music] be forgot and never brought [music] to [singing] mind? Should all acquaintance [music] be forgot and days of old [music] lang for >> or here's an even crazier example where we can input this Beethoven track. Once I play it, you'll probably recognize the melody here. We are going to turn this into a theatrical hard rock and then the lyrics are just going to be Where's my wallet? Where's my wallet? [singing] Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? [music] Where's my wallet? Where's my wallet? Where's my wallet? Where's my [music] wallet? Where's my wallet? Or here's another example where we can input the Jingle Bell song as the reference audio, but instead let's turn this into a minor version. [music] [music] Dashing through the snow in a one-horse open sleigh. [music] Over the fields we go, laughing all the way. Bells on bobtails [music] ring, making spirits bright. What fun it [singing] is to ride and sing a slayighing song tonight. >> Jingle [music] bells, jingle bells, jingle all the way. Oh, what fun it is to ride in a one-horse [music] open sleigh. >> So, this is a very flexible tool and the only open- source one so far that can generate some good sounding covers. Or here's another cool thing you can do. You can even just link this to an AI agent like GBT Astra or GLM and it can do the music generating or editing for you. So for example, let's start with this pop song. Let me play this for you first. [music] >> [music] >> And then afterwards you can iterate this further. For example, we can tell it to keep the melody and lyrics but make it more jazzy through reharmonization, >> but it still doesn't really sound jazzy enough. So let's tell it to remove the guitar. [music and singing] All right. So this is just a really quick and simple example of how you can get an agent to generate and then iteratively edit the song just with text prompts. Now, before we go over the installation, it's worth noting how this actually works. So, instead of just turning your text prompt into a complete finished song, what it does is it actually writes out an editable score first. So, here's an example where it writes out the notes of both the vocal track and the instrumental track, plus of course the key and tempo. This is essentially a basic sheet music for the song, and then it uses this as the backbone to actually generate the full song. And because of this, it gives you some very flexible editing capabilities. That's why it's very easy to generate cover songs from this or to change up the lyrics or even change the key from major to minor. And get this, if you look at these benchmarks in terms of this song quality index, not only does it beat the other open source models out there like Miniax Music and AEP 1.5, but it even beats some closed models like Sunno V6 and Sunno 5.5, which is pretty crazy. However, it does perform slightly under Sunno V5, which interestingly, at least according to this benchmark, sounds better than version 5.5 and version 6. If you do any kind of content creation, definitely check out Luma, the sponsor of this video. Think of it as an aentic AI workspace that can work alongside you through your entire creative process instead of just giving you the results of a single prompt. For example, let's say I want to create an entire marketing campaign for a new product. Instead of jumping between a bunch of different AI tools, I can just get Luma agents to autonomously do everything. It can develop the concept, generate the visuals, and shape the project all within the same workspace. What makes it especially interesting is that its Luma agents understands things like motion, physics, and 3D space. Rather than completely taking over the creative process, you can continuously guide the agent, change direction, and refine the results as you work. And one of the most powerful features is Luma's skills. You can basically create reusable skills for workflows you do all the time. Basically, you give Luma a set of instructions once and then you can run that same workflow on different assets whenever you want. For example, I can create a skill where I input any product photo and it'll output some UGC videos of an influencer talking about the product. Or here's another example of a skill where I can upload the product photo and it'll drop it into water like this. Luma basically gives you an intelligent creative co-pilot that can consolidate all your creative workflows into one place. Whether you're creating marketing campaigns, branded content, product visuals, or social media content, Luma is one of the best platforms you can use. Try Luma today using the link in the description below or by scanning the QR code here. If you're interested in trying this out, next, let's go over how to install this. Now, if you click on this GitHub link at the top and you scroll down a bit here, it does contain all the instructions on how you can run this. But the default code is like this where you need to work with some Python code which might not be intuitive for everyone. So instead, we are going to run this in a visual interface called Comfy UI. This is the most popular platform for running open-source image, video, and audio generators locally on your computer. In fact, if you're not familiar with Comfy UI, I highly recommend you watch this video first where I go over how to install and use it. Now, assuming you do have Comfy UI, the first thing you should do is update it to the latest version. So, in your root comi folder, simply click into this update folder and then double click on update Comfy.bat. So, this will proceed to update your Comfy to the latest version. Afterwards, let's press any key to continue to exit the window. And then next, let's start up a fresh session of Comfy UI. All right, after loading up Comfy UI, what you need to do is drag the U2 workflow onto your interface. So, I will link to this page in the description below where you can download the full workflow. Simply click on this link so that it will download this JSON file. Now, you can save this wherever you want. I'm just going to save it in my comfy root folder. Afterwards, simply drag this UA2 workflow onto your Comfy interface, and it should magically open this pre-built workflow for you. So you don't have to build out anything yourself from scratch. Now when you first load this, you might see this error message involving missing models. So let's proceed to download the missing models. So I'm going to link to this page in the description below. You need to download the audio encoder and the checkpoint for this. Let's first click into the audio encoders folder and we need to download this file which is 1.4 GB in size. Let's click download. And this goes in Comfy UI in models and then in audio encoders. Let's click save. And then afterwards, we also need to click into the checkpoints folder. And here you can choose to download either the full BF-16 version, which is 7.8 GB in size. Or if you have like less than 4 GB of VRAM, then you can download this smaller quantized version, which is only 3.96 GB. For me, since I do have enough VRAM, I'm going to download this full version, which should be able to fit in like 8 GB of VRAM. So, this goes in Comfy UI, in models, and then in checkpoints. Let's click save. Now, back to our Comfy UI interface, simply press R to refresh your model list. And then for this load checkpoint node, simply open the dropdown and select the model you just downloaded. In my case, I'm going to click on this BF-16 model. And that should get rid of all the errors that you see. Now, this workflow has two components. The first component at the top here is just turning a text prompt into a full song. And then at the bottom here, you have the option of inputting an audio for reference. So you can make cover songs with this feature. Let's go over the text to music workflow first. At the top here is where you would enter the style prompt. So basically this would describe the genre, the style, the pacing, the instruments or other things that you want to specify about the song. For example, let's do something like future bass, modern, energetic, inspiring. And then at the bottom here is where we would enter the lyrics. Note that this takes in metatags. So you can specify like verse one, verse two, intro, outro, bridge, chorus, pre chorus, etc. For me, let me just show you a simple example with one verse and one chorus. Next, this would be fed through this generate ABC node, which basically generates the notation of the song. After I press generate, you can actually see a preview of this notation over here. It's basically sheet music like this, which dictates the vocal melody as well as the instrumental melody throughout the song. And then afterwards, this notation would be plugged through this node along with your style prompt and lyrics to actually generate the music. Here is where you can specify the maximum duration of your song in seconds. So right now it's set at 360 seconds, which is 6 minutes. Now, this is just the maximum duration. If your lyrics are shorter than that, then it's just going to generate a shorter song. And then next, it gets plugged through this K sampler to actually generate the music. Note that the seed is basically the unique ID of every song. Right now, it's set at seven and fixed. That means if you use the exact same prompts and the exact same settings as before, you're going to get the exact same song as before. So, if you want to generate a completely different song while keeping the same lyrics and the same style prompt, then you need to change the seed to another value. Or you can also set this to randomize afterwards. And then the number of steps is like how many steps it takes to generate the song. In general, the more steps you have, the higher quality the song will be, but it's going to take slower. And then if you use fewer steps, it's going to generate faster. I just tend to leave it at the default of 32 steps. The CFG is like how literally you want the AI to follow your prompts. So if you get a song that doesn't really follow your style prompt, or if it has some errors with the lyrics, then you could set this CFG to a slightly higher value to follow your prompt more literally. And then the sampler anduler are basically the algorithms used to generate the song. I just tend to leave it at the default values. And then afterwards, this goes through the decoding step. Now, if you look at this note here, it says if your GPU has enough VRAM, so I would assume like over 12 GB of VRAM, then you can just use this regular decode method, which is a lot faster. If you have less than 12 GB, then it's best to use this tiled version. So, since I do have over 12 GB, I'm going to connect this one instead. So, first I need to take this input and connect it over here. So, what I'm going to do is hold down shift and then click on this connection and drag it over here. And then afterwards, I just need to connect this audio output over here. And then for this one, I can just press Ctrl +B to bypass or disable it. And then finally, it will generate my song over here. So, let's press run. All right. Afterwards, let me pull up the stats for you. So, this was pretty quick. I'm using just my laptop with an RTX 5000, which has 16 GB of VRAM, and this took just over 2 minutes to generate. Next, let me actually drag the style prompt and the lyrics over here so you can see them. And let me play the generation for you. I used to wait for [music and singing] the world to change something [music and singing] on it way. But every road that I never chose [music] was just another door I close. So I'll light the fire. [music] I'll make the stars. Turn all these dreams [music and singing] into works of art. I don't need to know [singing] where the ending goes. [music] I just need a spark and the nerve to follow. [music] >> [music] >> Now, it kept generating much longer than what I inputed here. So, what I should have done instead is reduced the max duration to something that fits these lyrics a bit better. But there you go. In a nutshell, that is how you can run this textto music workflow. All right. Now, if we go back here to this preview ABC section, you can see a preview of the notation that it generated over here. So, these are basically like the notes of the vocal and the instrumental. Now, some of you are wondering if this can do instrumental only. And the answer is yes. So, for example, here's my style prompt. Epic cinematic orchestral music for a battle scene, instrumental only. And then for the lyrics, I do need to enter something. So, usually I just enter what I want the instruments to sound like within square brackets. So, for example, here I put steady buildup of staccato strings and ethnic drums. And here's my generation. [music] [music] Oh, [music] [music] heat, heat. All right. So, that covers text to music. But what if you want to input an audio clip to use as reference or to generate cover songs? Well, that's what this part is for. So, let me hold down control to drag across these purple nodes and then press Ctrl +B to unbypass these nodes. Basically, this allows you to upload a reference audio to generate the ABC notation to be plugged through the song generator. And this step basically replaces this generate ABC node over here. So, first of all, let me select this top generate ABC node plus the preview node that's connected to it. I'm going to select both of them and then press CtrlB to bypass these nodes. And what we would do instead is connect the output from this node into the generate music node. So what I'm going to do is connect the output over to here. All right. So this generate music node should now be connected to this bottom section over here. And then what we need to do is for this load audio encoder, click on this dropown and select the audio encoder that you just downloaded. And then here is where we can upload a reference audio clip. For example, let me upload this segment from a song. [music and singing] watching as you walk away [singing and music] while my heart breaks. >> All right, so that was the song. Next over here, we can choose to either use the melody of this as the reference or the full song as reference. It's a really subtle difference, but if you want to do a cover song or copy just the melody of a song over, then I would select melody only. Now, over here is where you would enter the new style that you want for this song, as well as the lyrics for the song. For this example, I'm just going to enter the same lyrics as my reference audio. And then for the style prompt, let's try jazz with casual piano, saxophone, double bass, and brushed drums. Let's press run. All right, here's our result. [singing] Watching as you walk away while my heart breaks today. [music] It's as simple as that. So that's how you can use this audio reference feature. Of course, you don't have to use the same lyrics. You can also change the lyrics to something else. For example, you can make the person sing another language. So, a super flexible tool. So, that covers all the components of this UA2 workflow in Comfy UI. This is currently the best open- source music generator you can use right now. So, definitely try this out. Finally, I also want to mention the license of this. So, the code and the documentation of U2 are under the Apache 2 license which has very minimal restrictions, but the model weights are separately licensed under this creative common non-commercial license. So, in this clause, it specifically says not primarily intended or directed towards commercial advantage or monetary compensation. So, that's something to be aware of. I'm not sure if you can like post a generation on Spotify and profit from it. Anyway, that sums up my tutorial and review of UA2. This is definitely the best open-source music generator available right now. So, definitely give it a try and let me know what you think of it. And if you run into any errors during the installation, welcome to copy and paste the exact error message in the comments below, and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.