Is it just me or are these video models only actual use case is misinformation and spam? Sure they show us quirky and whimsy samples on the release page, but does anyone really believe that?
This seems extraordinarily good to me, compared to what I've seen before. Their washing machine advert example seems like it's as good as anything else on social media. I'm shocked by the quality and the coherence they're able to maintain, I assume it's really good at using those reference images they mention in their prompts. I think the only one that's noticeably bad is the concert hall, where the first few seconds show almost empty stalls and then toward the end it shows a full audience, and that audience also looks a bit off.
Yes, it is awesome. Funny we got this level of capability now and we all take it in our stride. Growing up in the 80's I loved futuristic scifi, feels like I am in it now. And even I am 'just' taking it in my stride.
Speaking of the eighties, special effects were still done without CGI and typically on the cheap. That didn't make the movies and television series any less fun to watch. In the end it's about what people do with the tools, not the quality of the tools. People commenting on how these generated videos are not perfect or are getting some details wrong, look at some blockbuster movies from last century. It's very easy to spot the special effects. And it did not matter.
But yes 80s science fiction is now science fact. Talking cars like in night rider are now definitely a thing. Some can even drive themselves. The robots in Buck Rogers, Star Wars, etc. look clumsy and fake compared to the real humanoid bots we are now starting to see. More close to I Robot, which when it came out 22 years ago was pure science fiction.
You know what, in my previous life as a filmmaker I could've only dreamed of such a thing. Filmmaking is an art form which you cannot do alone. Outside of a lot of time and money, you need cooperation of a number of people and each day of production you end up accumulating a set of compromises to your vision. Your taste is what makes you tolerate that or not, and it's exhausting. It matters if you're the author since at the end of the day it's your name on it, not the crew (as much).
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
I am not a film maker but I had some sketches on my head that I always wanted to try to make. Even Gemini scratched the itch for me, and my folder “projects” is growing with other ideas that I’d like to sketch
I don't think audio, image or video generation should exist. I haven't seen enough positive applications to justify the amount of harm these tools are being used to cause.
What are the positive applications? This all seems to be generating fake images/videos just because we can, not because it solves any real problems. It makes me wonder where the money is coming from to fund this stuff. Why would a regular person want to generate fake videos except for purposes of deception or debauchery?
It seems like the race to the bottom slowed down with audio, video and image generation models.
OpenAI has completely shifted focus, Google still is working on various toolings/models but ultimately most people are unwilling to pay for these generations, whereby code, can easily be burned (and paid for) by businesses.
image gen isn't a proven market (at least not yet).
The quality is quite high, but an observation is the direction they're taking these models correlates heavily to the usage demand of China vs the West. Specifically, they're immensely focused on t2v for action / high effect shots. There's one human reference shot in the entire release page, and that one doesn't focus on dialog at all.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
It's not about cultural difference but simply what's easier.
Action sequences are forgiving. There's lots of motion and shiny vfx to fudge things over. Rapid cuts mean each shot can be short enough that the inevitable accumulation of AI hallucinations from frame to frame doesn't get distracting. Sets can be generic: if you're generating robots attacking New York, the viewer isn't going to have time to track if the bodega on the corner is in every shot, or if the Chrysler Building switches place around town. Shots can be practically from different cities and nobody will notice.
Carrying over an actor's performance is the opposite. You want long shots and impeccable scene stability. You can't just throw some additive-blended particles on the actor's face to distract from its generation deficiencies, like you can do for the Marvel scenes.
Bytedance is a corporation looking for clout, not to help filmmakers. They're releasing a model that makes them look good, and action sequences do that.
It's a counterintuitive thing to viewers who have long been told that vfx is the most expensive kind of moviemaking. With gen AI, these somewhat convincing Marvel pastiches are trivial to produce, but a sitcom episode is utterly impossible.
I2p makes more sense and easier to do. But t2v is indeed of no use in any serious production, you must have the actor, product or setting communicated to the model visually
Google and xAI only allow 720p when you start adding references, and neither allow you to sync to audio.
This is entirely because of deepfakes.
Apparently Seedance has a version without guardrails if you are a licensed production company. I suspect Google Omni would too, but I don't know anyone who uses that model.
On carrying over actors' performance side, https://beeble.ai/ works pretty well. Friends from VFX are using a mixture of Seedance and Beeble (especially their SwitchX model)
Where can you actually get access to these models that isn't an outright scam? All the sites that promised to have Seedance 2 turned out to be scams. Does anyone know how to actually use it, and is it available to run yourself?
Seedance 2.5 looks amazing, but MiniMax H3 is going to be open weights within 24 hours: https://fal.ai/minimax-h3. According to the ComfyUI team, it should even work acceptably on mid-range consumer GPUs like the 3080.
I'd honestly take the slight quality hit for more control and lower costs.
New model day, here we go again! It’s a week of madness and all the wrong advice… and at the end of the week, the model is either dead or a massive hit. There is little middle ground.
I’m watching H3 to see if it fixes the nonsense that LTX introduced and what happens next… WAN is due, LTX says a new model is coming, Flux3 does video.
It went from zero video models to a lot of options due this year.
This. I'm very much happy to see minimax-h3 be released. I have so many projects that I have on the back burner that I could complete within minutes instead of hours and hours. "Typography, UI, and Graphics".
This will make longer (~30s) narrative add creation a lot better and more interesting at a reasonable price tag--roughly $7 best I can tell. Looking forward to trying it out.
The quality of AI videos is blowing my mind. I can still see things that seem a bit "off" but it's hard to distinguish AI videos from real videos. Even blockbuster movies are starting to be dwarfed by what AI can create. How long until we see the first full length AI movie hit the theaters? No actors, no development team, just a guy prompting AI...
It has improved a lot, but these demo reels still have all AI video issues. Flash cut salad (including the scenes that should have longer cuts), unnatural motion that looks animated, unprompted YouTube-face acting, etc. Admittedly it's all a lot less pronounced in this version.
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. But then I remember that I spend upwards of $10k on inference generating well over 50k images for storyboards, training models; and probably creating almost an hour of video (I assume). Yeah, I get that things can be economic if you don't use the latest models (ran some case studies on this), but the latest models are the most fun to work with. It doesn't scale as well as "vibe coding" stuff together on the weekend. And when things work really well its almost as if you're seeing an zoopraxiscope come to life for the first time; and you just want to keep going.
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
It is (mine)! And there is a purpose to it looking AI generated. However, the sample projects where actually fully worked on; mostly as hobby projects. I will have it go through another iteration of improvements soon. This may be the 10th version of the website, since maybe 2012-2014, I did have different domains since then. But the concept of using it both as a creative and professional outlet stays alive.
Ask someone for feedback on your website. To you it might look ok, but to me and I think many others, it is unintelligible. I have no idea what I'm looking at and it is impossible to parse.
That is the primary issue, and secondarily it is a cookie-cutter AI slopfest that will make any technical person not submerged in Kool-Aid click away.
Please just get your AI to follow some basic UI best practices or pick up a UI library.
Thanks for the feedback, appreciated! Yes its due an overhaul, but a large part of the purpose of this website is going through experiments as a digital garden and having a creative outlet, I am fully aware how it looks, but there is a part of me that needs it to go through the process of looking as sloppy as possible. It's likely going to end up going through its nth iteration very soon, and then another one after that. I do run and maintain (for clients/myself) many websites that look nothing like this.
93 comments
[ 5.3 ms ] story [ 136 ms ] threadBut yes 80s science fiction is now science fact. Talking cars like in night rider are now definitely a thing. Some can even drive themselves. The robots in Buck Rogers, Star Wars, etc. look clumsy and fake compared to the real humanoid bots we are now starting to see. More close to I Robot, which when it came out 22 years ago was pure science fiction.
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
Weird-ass times.
OpenAI has completely shifted focus, Google still is working on various toolings/models but ultimately most people are unwilling to pay for these generations, whereby code, can easily be burned (and paid for) by businesses.
image gen isn't a proven market (at least not yet).
This feels like the serious inflection point for high quality full length feature film productions using this tech.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
Action sequences are forgiving. There's lots of motion and shiny vfx to fudge things over. Rapid cuts mean each shot can be short enough that the inevitable accumulation of AI hallucinations from frame to frame doesn't get distracting. Sets can be generic: if you're generating robots attacking New York, the viewer isn't going to have time to track if the bodega on the corner is in every shot, or if the Chrysler Building switches place around town. Shots can be practically from different cities and nobody will notice.
Carrying over an actor's performance is the opposite. You want long shots and impeccable scene stability. You can't just throw some additive-blended particles on the actor's face to distract from its generation deficiencies, like you can do for the Marvel scenes.
Bytedance is a corporation looking for clout, not to help filmmakers. They're releasing a model that makes them look good, and action sequences do that.
It's a counterintuitive thing to viewers who have long been told that vfx is the most expensive kind of moviemaking. With gen AI, these somewhat convincing Marvel pastiches are trivial to produce, but a sitcom episode is utterly impossible.
China have 1.5 billion people vs US 350 ish
This is entirely because of deepfakes.
Apparently Seedance has a version without guardrails if you are a licensed production company. I suspect Google Omni would too, but I don't know anyone who uses that model.
I'd honestly take the slight quality hit for more control and lower costs.
I’m watching H3 to see if it fixes the nonsense that LTX introduced and what happens next… WAN is due, LTX says a new model is coming, Flux3 does video.
It went from zero video models to a lot of options due this year.
https://xcancel.com/CuiMao/status/2058458683781365873
I can not believe how good the AI video has gotten. In one spot the written letters were obvious AI. Other than that - I couldn't find a thing.
This is blessed and cursed timeline at the same time.
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
I’ve built it as a side project, and the cost to produce one 20-40 minute video is in the 5$ range.
My friends/family have watched some of the better videos. All one shotted, and the storylines come out surprisingly well.
Currently working on a new engine that generates video, but it’s expensivee
Open to collabing
That is the primary issue, and secondarily it is a cookie-cutter AI slopfest that will make any technical person not submerged in Kool-Aid click away.
Please just get your AI to follow some basic UI best practices or pick up a UI library.