> The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.
Pretty cool.
But assuming you have a 16GB 3060, how long would it take to generate a 15 second clip?
The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models.
The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering.
I feel like for a good while now we'll transition into a process that uses traditional "close-up" rendering/shots + AI generated wide-shots or quick cuts.
Exciting, but also troubling. This being open-weights is a massive win for the community though.
> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.
Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as well?
Neat with native frame-to-frame generation, but wonder how easy it is to "link" together clips at the intersection, typically the models kind of lose the "momentum" across these stiches, being able to merge things with frame-to-frame between clips might help with this it feels like.
This isn't the same with lookup tables, but in Explaining Attention with Program Synthesis they were able to replace some attention heads with Python programs:
I've said it before and I'll say it again, human directors are still valuable, as they use AI video editing tools to generate the shots they want and put them together in a cohesive way. Previously they might've used film and actors but if they can just prompt the AI (or create workflows as seen with ComfyUI) then they arrange them together just like how an EDM producer doesn't actually play the instruments but instead the creativity is in the arrangement.
I suspect it'll be quite a while until AI gets a good enough aesthetic sense to do this, as even with static HTML websites humans can easily see that it's AI slop.
or option 2 we'll keep making productions with film and actors. I get it's easy to feel that it's all over with how good these video models are getting but thinking we'll all be slaves to the slop machine once it gets good enough is pretty pessimistic depressing and IMO unlikely.
(I do think it will get a foothold in the "crap people are ashamed to admit they watch" sector though, which it basically already has)
case in point, the entire SaaS multimedia space has pivoted this year to agentic workflows, as in, no more generating AI for that sensitive audience, but instead automating the human work of editing and compositing of real media
if you so happen to supply generative media it will form a cohesive edit of that too
also website slop is distinct from the AI generated sites that blend in. you only notice the ones that don’t.
Reference-to-video mode seems like all that was missing to enable completely independent cinematography as right now one couldn't stitch different scenes together properly without altering substantial portions of the scene.
I dug up a few old parody ideas I’d had back in high school and threw them at MiniMax M3 on my RTX. There’s definitely still a lot of jank once you move away from fairly normal scenarios. The moment you start to veer into weirder concepts, things tend to break down a bit especially in the game show where someone is strapped to a wheel and being spun.
Still tho, I was actually shocked by how well the text-to-video turned out overall, and how fast it ran. A 10-second, half-megapixel video gen took only a few minutes which is kinda crazy especially thinking back early WAN days.
22 comments
[ 1.5 ms ] story [ 28.4 ms ] threadThere is some debate on the license for those in the US, UK, EU, plus… no comment other than whew those samples though!
This is AGI.
Pretty cool.
But assuming you have a 16GB 3060, how long would it take to generate a 15 second clip?
The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering.
I feel like for a good while now we'll transition into a process that uses traditional "close-up" rendering/shots + AI generated wide-shots or quick cuts.
Exciting, but also troubling. This being open-weights is a massive win for the community though.
Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as well?
Neat with native frame-to-frame generation, but wonder how easy it is to "link" together clips at the intersection, typically the models kind of lose the "momentum" across these stiches, being able to merge things with frame-to-frame between clips might help with this it feels like.
https://arxiv.org/abs/2606.19317
I suspect it'll be quite a while until AI gets a good enough aesthetic sense to do this, as even with static HTML websites humans can easily see that it's AI slop.
(I do think it will get a foothold in the "crap people are ashamed to admit they watch" sector though, which it basically already has)
there are sequencing AI that will edit
case in point, the entire SaaS multimedia space has pivoted this year to agentic workflows, as in, no more generating AI for that sensitive audience, but instead automating the human work of editing and compositing of real media
if you so happen to supply generative media it will form a cohesive edit of that too
also website slop is distinct from the AI generated sites that blend in. you only notice the ones that don’t.
Still tho, I was actually shocked by how well the text-to-video turned out overall, and how fast it ran. A 10-second, half-megapixel video gen took only a few minutes which is kinda crazy especially thinking back early WAN days.
Video demos:
https://imgpb.com/rllwg