Transcript · 29,510 chars
AI never sleeps and this week has been absolutely insane. ByteDance releases the best video model out there, Seed Dance 2.5. MiniMax also releases their best video model, MiniMax H3. And the awesome thing is this will be open source. Deep Seek releases their newest model and they pulled off the impossible again. It's as good as GLM and Opus but like a hundred times cheaper and way smaller. We have a new open source model that's trained entirely on AMD chips, not Nvidia. This AI can turn an image into transparent layers which you can edit further. We have some new video world models which you can interact with in real time. Google releases their latest robotics model and a lot more. So let's jump right in. First up, Netflix releases a really cool open source AI called IDV2V. In the simplest sense, this can basically change the style of the scene without affecting the identity or the movement of the characters. So you can just plug in an existing video and then edit one keyframe to show the new look you want and then the system will spread that style across the entire video. So you can change things like the background, lighting, clothing, or overall style while keeping the face and expressions and movements of the character the same as the original clip. Now at the top of the page, if you click on this GitHub button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Notice that this can generate videos of up to 720p and this is pretty huge. So the main model is like almost 80 gigabytes in size. So this would only fit on like really high-end consumer hardware. But hopefully there will be more compressed versions of this in the future. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new open source transcription tool which is really powerful. It's called Whisper Whisper 2 and this can basically take any audio and turn it into text like this. The cool thing is this actually has two different outputs. If you turn on the verbatim mode, this includes everything including stutters, hesitations, laughter, or other meta tags, plus repeats like this. Now, alternatively, you can also turn on the intended mode, which would get rid of all these hesitations and meta tags and give you the fully polished transcript. So, a super flexible tool. Not only that, but this also gives you precise word level timing. In fact, let's try this out. They've already released a free hugging face space for you to try this out online. Here is where you can upload any audio. So, let's upload this one. We are going to generate this in verbatim mode. >> However, due to the slow communication channels, styles in the West could lag behind by 25 to 30 years. >> All right, so that was the audio and as you can see, the transcript is indeed correct and it also gives you the start and end times for each word. Or here's another example. Let me play you the audio first. >> And so, my fellow Americans, ask not what your country can do for you, ask what you can do for your country. >> And as you can see, it's able to provide me the transcript here, plus the timing for each word. Now, this supports all these different languages as you can see here, and if you look at this self-made benchmark, then Crisper Whisper 2 is indeed even better than 11 Labs or some other leading transcription tools. And if you look at this benchmark on word level timestamps, then as you can see, Crisper Whisper also contains the least amount of errors. The awesome thing is they've released the models to this. So, if you click on this button, they've released four different models in this family. The smallest one is 0.2 billion parameters, and this is less than 500 megabytes in size, so you can easily fit this on most consumer devices. You don't even need a GPU. And then, the largest one is 2 billion parameters. This one is 3 gigabytes in size, which is still fairly tiny. You should be able to fit this in most GPUs. And then if you scroll down the Hugging Face page here, it contains all the instructions on how to download and run this locally on your computer. If you're interested in reading further, I'll link to this main page in the description below. Also this week, the Goat DeepSeek is back with another mind-blowing model. They just released the latest version of DeepSeek V4 Flash, and even though this is a flash model, it even performs as good as some of the full open-source models out there, including MiniMax M3, and it's just one point below JLM 5.2. This scores way higher than even the previous DeepSeek V4 Pro version, and as you can see it's like 10 points above the previous DeepSeek V4 Flash. What an insane upgrade. Note that here it says it has the same model structure as DeepSeek V4 Flash DSpark. In fact, this is a really important architecture breakthrough, which helped it improve efficiency and throughput by a huge amount. I did a full explainer video on DSpark, so definitely see this video if you're interested in learning more. If you look at these benchmarks on agentic coding, software engineering, and cybersecurity, you can see that it absolutely clobbers the previous V4 Pro version. And for some instances, it even beats or matches JLM 5.2 or Opus 4.8. Keep in mind this is like 70% smaller than JLM 5.2, and it's also 100 times cheaper than Claude Opus. DeepSeek V4 Flash costs around 3 cents per million tokens, making it by far the cheapest model at Frontier Intelligence. And if you look all the way on the other side, GPT 5.6 and Claude Opus and Claude Fable are all the way over here. So in terms of performance versus cost, this is by far the best option to use. It's just incredibly efficient and just way cheaper than the other frontier models. Now as expected from the Goat, they've also released the model to this already. It's already out for you to download on Hugging Face and this is only 167 GB in size. So, this can easily fit on just like one DGX Spark. Now, because this is open source, the community has acted fast and for example, Unsloth has already released GGUF versions of this. The crazy thing is the smallest one-bit version is only like 82.5 GB in size. So, you can potentially fit this on just like one or two pieces of high-end hardware. And keep in mind, this has the intelligence of GLM 5.2 or Opus 4.8. It's pretty crazy that you can now run this level of intelligence locally. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have quite a useful AI called Redesign. This basically turns a flat image into layers. So, you can edit each component or move them around. And with this, you can customize or micro-edit certain things like recoloring the elements or repositioning certain things or changing the size of elements. So, think of this as like turning a screenshot back into something closer to a Figma or Photoshop project. Now, this actually orchestrates a ton of different AI tools at once. It uses PaddleOCR. In terms of generating the layers, it uses Kwen Image Layered, which I featured on my channel before. And then for detecting and segmenting different elements, it uses Dino and SAM too. Now, according to these benchmarks, it seems like this new redesign even does better than some other image-to-layer generators like Kwen Image Layered. At the top of the page, they've released everything already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. And this is actually fairly tiny at less than 4 GB in size. However, at least according to this base code, it also does require, apparently, an OpenAI API key. But you could tweak the code further to swap this out for a local model to run it completely for free. If you're interested in reading further, I'll link to this main page in the description below. Around 2 weeks ago, Moonshot AI released Kimiko 3, which is currently the most powerful open-source model you can use right now. Now, on their release page, they said that the model weights will be released on July 27th. Well, as promised, this week, on July 27th, they released the full Kimiko 3 model on Hugging Face. So, here are some specs. This is a massive 2.8 trillion parameter mixture of experts model, so think of it as like a team of specialists working together at once. And when you use it, only 104 billion of these parameters are active. For the architecture, it uses Kimiko delta attention and attention residuals, which were also designed by them. In fact, if you want to learn more about attention residuals, which is actually a really fascinating breakthrough, see this video to learn more. And also, this has vision capabilities natively baked in. And they used this Moon VIC vision encoder to make it happen. Now, if you click on files and versions, as expected with a 2.8 trillion parameter model, this is like 1.56 terabytes in size. So, you're going to need to stack like multiple enterprise GPUs in order to run this. However, because this is open-source, the community is already really quick to help quantize and compress this further. So, for example, Unsloth released some compressed GGUF versions of Kimiko 3. And the 1-bit version is only like 594 GB. Still massive, but this is a really big drop from like 1.6 terabytes in size. Anyway, if you want to download this or check out the technical details, I will link to this Hugging Face page in the description below. Also this week, AMD releases their very own open-source model called Instella MoE. And this is quite a big deal. You see, Nvidia has basically dominated the AI space because most AI tools are built off of their CUDA platform. It's really hard to train and run AI models on non-CUDA chips. But here, AMD has trained this Instella mixture of experts model from scratch on their AMD Instinct and using the AMD Rockum software stack. First of all, here are some specs of this. So, this has 16 billion total parameters and this is a mixture of experts models. So, think of it as like a team of specialist AIs working together. So, when you use it, only 2.8 billion parameters are active, making it fairly efficient. The cool thing is this reportedly beats other similar-sized models, including Gemma 4E4B and a smaller version of Quen 3.5. The awesome thing is they're not just releasing the final model, they're actually publishing the checkpoints from pre-training, mid-training, and all these subsequent stages, along with the training recipes and the code. It uses what they call a multi-head latent attention to make the attention and memory more efficient. And then they also use something called a far skip collective, which helps overlap communication and computation across GPUs. This is a bit technical, but on this page they document how exactly they did all these stages of training to create the model. And then if you click on this GitHub repo down here, it contains all the instructions on how to download and run this, as well as the training scripts. And then if you click on their hugging face repo, note that the final version after training on reinforcement learning, this is called Think and this is roughly 32 GB in size. So, it's a medium-sized model, similar to Quen 3.6, which should be able to fit on high-end hardware. If you're interested in reading further, I'll link to this main page in the description below. Also this week, one of my favorite image generators, Audiogram, has released a new tool called object remover. And here's how it works. You can simply brush over the object that you want to remove and it'll automatically highlight that object and then remove it for you. It's as simple as that. Here are some other examples for your reference. So, we can select this bike and remove it and notice that it also removes the shadow of the bike. Or here's another really tricky example where I want to remove this plant, but there's also a lamp that's kind of occluded plus some books in front of it. But as you can see, it's able to remove the plant while also preserving the lights and the books. Plus, it also is able to remove the plant's reflection on the floor. And here's another example. Now, at least according to this removal bench, you can see that Ideogram Object Remover has the lowest error rate compared to other models like Nano Banana 2 and GPT Image 2 Medium. So, if you're looking for a free object remover which you can run online, this is a nice option for you to use. You can sign up for free and you get some daily credits to run some generations. If you're interested, I'll link to this in the description below. If you want to supercharge your content creation, definitely check out Higgsfield, the sponsor of this video. Think of it as an all-in-one AI creation platform built specifically for creators. Instead of jumping between a bunch of different tools, Higgsfield gives you access to the world's leading models in one place, including Seed Dance, Kling, and more. And they've just added the ability to generate 4K videos with Seed Dance 2.0, so you can generate ultra-sharp videos with incredible details. They're also going to roll out the best video generator out there, Seed Dance 2.5, very soon, so stay tuned for that. This supports multiple inputs, so you can combine text, images, video, and audio to control the final results more precisely. And using this in Higgsfield makes everything way easier. For example, they have Marketing Studio, which is really useful if you're making marketing content. You can paste a product link or upload a product image, and it can generate multiple ad formats in one workflow, like UGC videos, tutorials, unboxings, product reviews, and more. They also have Cinema Studio, which is built as a full end-to-end filmmaking pipeline. Instead of just typing a prompt and hoping the video looks good, Cinema Studio lets you plan scenes, control the camera, add specific characters, reduce locations, and keep everything consistent across the whole project. From idea to final output, Higgsfield gives you way more control over the whole creative process. Whether you're making ads, social videos, AI influencers, product launches, cinematic clips, or any other content, this is one of the easiest platforms to start creating with AI. Try Hicksfield today using the link in the description below. Also this week, Thinking Machines, which is an AI lab that was started by OpenAI's former CTO, they just released a new open-source model called Inkling Small. Now, last week they released the full Inkling model, which is really good. This is an omnimodal model that can understand text, audio, images, and video. Well, this week they released a smaller version of this, which is roughly a quarter of the size. It's still pretty huge though at 276 billion total parameters, and when you use it, 12 billion of these are active. And its performance is actually really impressive given the cost. So, the red line is Inkling Small. As you can see, it's able to actually achieve, in some cases, even higher performance than the full model, but using less compute. And then here's the performance of Inkling Small across these different benchmarks compared with DeepSeek V4 Flash. Note that this is the older version. They released a new version this week. And then also Gemini 3.5 Flash Lights and GPT-5.6 Luna. So, Inkling Small holds up pretty well, but as you can see, the main advantage of this is it also can take in audio, whereas the other models don't really perform so well in terms of audio capabilities. Now, on Artificial Analysis, you can see that Inkling Small is over here, just 1 point below the full Inkling model, so it is a lot more cost efficient. However, it's still behind some of the other similar-sized open-source models like the latest DeepSeek V4 Flash. So, in terms of intelligence, Inkling Small isn't the best, but I do like its multimodal capabilities, especially if you need an open-source model that can also analyze audio and images, then this would be one of the best options to use. The awesome thing is this is already out. So, if you click on this Hugging Face link, and you scroll down a bit, here it contains all the instructions on how to download this. Note that at over 200 billion parameters, this is still quite huge at 532 GB in size. So, you'll need to stack like multiple DGX sparks in order to run this. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new system for robotics called Prism. This helps robots control their bodies and react to physical contact more effectively. So, what this does is it takes an ordinary robot sensor readings, images, or other instructions and it outputs the movement actions for the robot. And if you compare Prism with other similar algorithms, you can see that Prism has a much higher success rate, whereas the other ones are more likely to fail, especially when manipulating objects. Here are some additional examples for your reference. You see, the problem is that the actions that robot takes shouldn't depend on just one metric alone. It should come from different measurements like force, velocity, contact, friction, joint angle, etc. Well, this algorithm basically looks at all of these combinations of different signals to help the robot make a more informed decision and action. And the results are surprisingly strong. So, you can see across these different metrics, this new Prism algorithm has a much higher success rate and lower error rate. Now, at the top of the page, they've already released the code to this. So, if you scroll down a bit here, it contains all the instructions on how to download and run this yourself. If you're interested in reading further, I'll link to this main page in the description below. Also this week, ByteDance releases what is probably the best video model out there, Seed Dance 2.5. By the way, their previous model, Seed Dance 2, was already the best video model out there. No other generator really comes close to beating it. Well, Seed Dance 2.5 is even better. It's incredibly good at high-action fight scenes and character consistency, as you can see from some of these demos. The nice thing about it is it's multimodal. So, you can input reference videos for it to edit. Here's an example where you can input a simple 3D scene like this and get it to generate a full video that follows the composition. Or you can upload an existing video with a green screen and get SeeDance to turn the scene into something else. It's as easy as that. You can also input a storyboard like this and easily get it to generate a full video that follows all these different shots. My favorite feature is that this can generate videos of up to 30 seconds long. This enables you to do a lot more, whereas the other models can only do around 15 to 20 seconds max. Currently, this can only generate videos of up to 720p, but they are going to release the ability to generate 1080p and 4K resolution videos in the near future. Now, this can take up to 50 reference inputs including audio, images, and video. So, it's incredibly versatile. For example, I can upload all these images and get it to incorporate everything into the video. Now, currently, this is already out for certain countries on various ByteDance platforms like Dreamina, but if you're in the US, it's not available yet. Now, this is one of the most expensive video models to use. For example, a 10-second clip would cost roughly 460 credits, and at least currently in their Illumina platform, you can top up 1,000 credits for 10 bucks. So, 460 credits is roughly $4.60. So, this is much more expensive than MiniMax H3. Also note that API access is not out yet. They're planning to roll this out next week. But, here's the thing. This is the best video model out there, especially for high-action scenes or scenes with tricky physics and motion. This is still way cheaper than filming everything yourself. Now, I'm planning to do a full review video on this very soon, probably tomorrow or the day after that. So, stay tuned. If you're interested in reading further, I'll link to this main page in the description below. Also this week, MiniMax releases their latest model MiniMax H3. By the way, this is the company behind High Law. So, I think now they're just rebranding HighLaw into the MiniMax H series. Now, this is a super powerful and flexible multimodal model, which means you can input text, images, video, and audio to use as references. And this can generate 2K resolution videos, as you can see from this really detailed example. Here's an example where we can input these two images and get it to generate a trailer, and here's the result. >> It was waiting for me. >> Or here's another example where I can input this storyboard plus this logo, and then get it to make an ad for this luxury handbag based on the storyboard and logo. And here's the complete commercial. >> [music] [music] [music] >> Here's an even more impressive example where we can input images of this bold red and blue comic strip effect, and it's able to create a very nice video from this. Now, because this can take a video, you can also input a green screen video like this plus the background which you want to use and easily add that to the scene. Or here's another example with audio. I'm going to input this audio track which was made from the open-source music generator a step >> [music] >> And then I'm going to upload this image of some random K-pop group and also a few reference panels of some text, and let's get it to make an MV from this song. Show these K-pop members from this image singing and dancing to the music. Add coarse grain, glitch effects, grunge effects. Keep the edit fast and use hard cuts only. Cuts should occur within 3 seconds based on the beat of the song. Use the typographic reference from my second image. And here's the result. >> [music] [music] >> Currently, you can try this on their online platform. It allows you to select all these different aspect ratios and up to 2K in resolution, and you can set this up to 15 seconds. Now, if I select 10 seconds, note that it costs 120 credits, and similar to ByteDance, it costs $10 to top up 1,000 credits. So, 120 credits is roughly $1.20, around like three times cheaper than SeeDance 2.5. The really awesome thing is they are actually going to open source this. They'll likely release the model sometime next week, and I'll definitely make a full review and installation tutorial on this once it's out. So, stay tuned for that. I think this is going to be the best open video model For now, if you're interested in trying this out, I will link to this main release page in the description below. Also this week, Google DeepMind releases a really exciting update for robotics. So, they just introduced Gemini Robotics 2. This is a new family of models designed to control a robot from its feet all the way to its fingertips. You see, the previous Gemini Robotics model mainly focused on upper body and table top tasks, but this version can combine walking, balancing, reaching, grasping, and reasoning all in one continuous sequence. For example, you can tell this Apollo 2 humanoid robot to do a certain action, and with this model as the brain, it can understand the instruction, locate the object, walk towards it to pick it up, and then move across the room to reach the target and place the object in the correct location. Now, they've actually released three different models. The main one is Gemini Robotics 2, and this is their vision language action model. So, this turns your natural language instructions and the robot's vision via cameras on its face into physical motor commands, basically into actions that the robot should carry out. And then they also released Gemini Robotics ER2. So, this is a higher-level reasoning system. This is designed to understand the environment, the room, plan tasks, and then it's also designed to correct the failure, and it can also coordinate with multiple robots. And then there's Gemini Robotics on device 2. This is a smaller version that can run locally on a robot without needing to connect to the internet. So, this is a completely offline model. The system also brings much better hand control compared to the previous version. As you can see, it's able to perform a ton of tasks that require a lot of dexterous finger movements like unscrewing a light bulb or tying a trash bag or sealing a Ziploc bag. Now, at the bottom here, you can actually try out Gemini Robotics 2 in AI Studio, and you can also sign up for their trusted tester program. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a new video world model called Wonder. And like other world models, this allows you to generate an interactive world which you can explore in real time. Notice that this is just video, but you can press all these keys here to navigate around the scene. As you can see, this works with a variety of different environments and characters. So, here's a more anime digital art example. The quality is not perfect. There's still some noise and inconsistencies around the edges. And then here's an example of another artistic style. And then here are some realistic examples for your reference. Now, the really cool thing is instead of just inputting an image as the starting frame, you can even input a video as you can see in these demos. This will render the same movements as the video, but it allows you to now walk around the scene as if you're viewing the video in 3D. Now, again, this is far from perfect. There's a ton of noise and artifacts with its generations, but this is a pretty cool feature. And then here's another example of a fight scene, but you can like press these keys to change the camera perspective and walk around this scene. Now, at the top of the page, it does say the code and the models are coming soon. Now, this is from Adobe, which haven't really released anything, so hopefully they will stick to their word and release the model to this. For now, if you're interested, I'll link to this main page in the description below. Also this week, Google has added a new voice feature to the Gemini app. Now, currently this is only for Mac OS. Hopefully, they will also release a Windows version in the future. This is basically an AI-powered writing and editing tool. You simply hold down the function key and speak naturally into pretty much any application on your Mac, and it will use Gemini as the transcription model. It transcribes what you say, and it also removes filler words or mistakes or repeats. It also understands when you correct yourself halfway through a sentence. It's also able to help you clean up formatting, add the correct punctuation, etc. And it inserts the text directly where your cursor is located. So, it's very similar to other AI dictation apps like Typeless or Whisper Flow. You can also enable Gemini reasoning, which gives the system permission to perform more complex tasks. For example, you can highlight some documents and ask Gemini to summarize the content. So, I think this is a really useful tool. I hope they roll it out to Windows and mobile platforms as well. For now, if you have a Mac and you're interested in learning more, I'll link to this main release page in the description below. Also this week, we have a new AI called Phi-0. This is a new video world model built around what the authors call physical language. The main idea is that instead of immediately generating the next video frames, you get the model to first reason about how everything should move and change physically. Then it goes through the video generator to actually render those frames. And from that, it's actually very good at predicting what happens next. And so this has a variety of applications. For example, you can use this to generate interactive worlds. It can predict how the scene should move if you press a certain combination of keys. Or you can also use this to create videos for autonomous driving. And the same thing applies for creating videos to train robots. And if you look at its performance on physical coherence and understanding, then you can see that on average it even outperforms other similar world models. Now, at the top of the page they've released a code button to this. And here it says the code is coming soon. So, stay tuned for that. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up-to-date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.