133 comments

[ 0.28 ms ] story [ 105 ms ] thread
Seems like if you build some more scaffolding around it, it wouldn't be bad. I think AI video isn't quite there yet so you probably would want to lean into that. For example you could ask for an animated or cartoon music video so the real shots don't look weird. Also if you gave it some guidance on what a good music video is like it would probably help as well. But yeah idk may be that's not the goal here.
Unsure if it's just the way they prompted it / coded it, but the output is far too much a literal direct copy of the lyrics. The best music videos have a story arc on the theme of but often not litearlly the lyrics, and start with obscurity and reveal something (following all the literary/story mechanisms)

Consider Amber Run - Found lyrics versus the video, and the story arc of the video

https://www.youtube.com/watch?v=Yj6V_a1-EUA

Fun fact (if you care): Back in the '90s, pretty much every music video produced in Azerbaijan literally matched the lyrics.
The entire thing was cringeworthy to the core. I kind of enjoyed it though because it perfectly epitomized "AI slop" in the first 30 seconds so wonderfully. "Michelle Pfeiffer, that white gold" - show a blonde woman in a gold sequined top! "Livin' it up in the city" - show a shot of a big city!

If anything, the absurd literalism of the video contrasted so perfectly with the (IMO) brilliant clever originality of the lyrics. E.g. "Michelle Pfeiffer, that white gold" is actually a not-so-subtle reference to cocaine. Imagine if the lyrics were as stupidly unoriginal as the video ("Now we're all snorting cocaine!!").

> The best music videos have a story arc on the theme of but often not litearlly the lyrics

If the music is crazy popular, you can still do it. See Land Down Under

Sometimes you can fix this by swapping one tracks music video with another, and letting the syncopation happen naturally.
Yes, LLMs are way to literal - it is a problem.

Claude can right great code, but it’ll throw comments into the code about why we chose this approach instead of the random one I discussed with it - when the comment is of no use to a future developer.

It points at some sort of theory of mind problem in LLMs imo.

Wow. These are horrible. Sort of refreshing. I thought video was better than this now, but I guess not.
These are getting really good. Much more interesting than the average music video already.
> None of the music videos were great

Glad they acknowledge this.

Curious how much time in addition to tokens this costs. If you have to spend $25 and wait 45 minutes to get a basically unwatchable video, I'm not worried about indie film makers being replaced just yet...

This wasn’t possible even a year ago, with the speed of things changing and how much money is spent on movies is there really a doubt that someone will be able to make $100 million movie for less than $1 million in token spend?
(comment deleted)
Regular music videos (including the writing/recording) can easily go into 6 figures. I wonder what the $200,000 AI music videos looks like.
(comment deleted)
Skip the Claude Fable 5 $25 video to 1:42. The disembodied Adams' Family hand is on the job.
It is jarring to me that most of the dancing seems slightly out of sync with the music. It is like a music video uncanny valley - images look good, but the lack of sync to the sound shatters the illusion entirely.
To me nearly all of the real world dancing looks like that. The dancers could literally be olympic level and I still don't see the connection between their movements and the sounds of the music. Music videos sometimes do sync for me but only on very simple motions, and for whatever reason shuffle dance on video almost always looks great.
Well yeah, because music is not a modality of the models involved at all. It's literally just an LLM with the timestamps of each line injected into its context, and then making an API request to image / video generators. I don't really understand what the point of this project even was; stitching together dumber models with bespoke glue logic like this takes us further away from general intelligence and confuses the public.
While it's made huge improvement in just the past few years, AI still hasn't quite solved motion, especially human motion. The humans in these videos are rendered really well, but move unnaturally. Like someone did a motion capture of a real human, and then played it back with a very high quality 3d model--something hard to describe is lost in the processing. Everything seems to move just too smoothly, at too exactly constant a speed. Same for camera movement. A real life camera doesn't precisely follow some exact 3D Bézier curve at an exactly constant speed. AI doesn't get this yet.
When the line was "don't believe me just watch"

And then the clip was literally just an arm wearing a watch!

That's freaking hilarious!

It's like someone playing charades

The fable $25 version was the best.
The GPT ones are strange. The $25 fable one to me is subjectively better than the others. The $100 fable one is too literal and robotic.

The jevons paradox is you need auteurs to curate vignettes or effects and cut or mask them in etc. That's not really different philosophically when software entered art in other ways. I could see errors/glitches lowering in time but I doubt there will be much acceleration.

Wow. These are all terrible. Music video producers can breathe a sigh of relief.
These are awful. It’s like Suno music. Seems convincing if you half listen. As soon as you pay attention you notice all the cracks.
Visual effects went through this same development issues as the industry matured. What took this industry decades to advance it taking months in AI. Think the spaghetti Will Smith and now this. Another one people don't mention here but is specific to video is higgsfield ai.
(comment deleted)
Unless I'm misunderstanding the article, or your comment, the models were responsible for generating the music _videos_.

The music itself is Uptown Funk... which was a very successful song in 2014 (https://www.youtube.com/watch?v=OPf0YbXqDm0)

The videos are indeed awful though.

Lmao yeah I’d rather give $100 to a college kid to film a bunch of shit and then splice it together. Would be significantly more interesting.
The more concerning part is that a much less discerning audience will happily engage with and watch endless hours of AI slop videos. For example what happens if you give a 3 year old a tablet and youtube access to keep clicking on things.

https://www.cbc.ca/news/canada/ai-baby-slop-9.7166873

https://www.nytimes.com/2026/02/26/us/ai-videos-children-you...

Or for an "Adult" audience, I'm sure you could get an AI to create videos of "OW, My balls!" from Idiocracy.

https://substackcdn.com/image/fetch/$s_!Dh4l!,f_auto,q_auto:...

https://media.licdn.com/dms/image/v2/D4E22AQEqLntg_DW7vg/fee...

I've had some real success with Suno. There's one song I made as a joke which is a house track with lyrics about me and my friends which I've genuinely downloaded and listen to regularly.
As a clanker-apologist I gotta say this is the strawman of AI slop brought to life. Letting the machine do literally everything without even supervision and see how it turns out? No surprise at all the results are so bad.

My fav AI video is still Post-Scarcity Blues from a year ago https://www.youtube.com/watch?v=q_t3h2AZ0KY There have been others I've enjoyed since then. But, that one stands out in memory. Work warning: it is occasionally just a bit spicy.

These are pretty terrible, but for me there were at least a few moments where they genuinely became "so bad they're awesome". They actually re-enforced an idea I keep coming back to: sometime soon a real artist is going to use AI to make something amazing, not by aiming for flawless "realism" or some kind of pastiche slop, but by leaning into the weirdness that often comes out of AI.
claude’s is bad and chatgpt’s is horrific
There's definitely still room for innovation in custom use cases with AI. I like to write and draw comics, but it's very time consuming to make a finished product. Working with tools in default and you'll have a bad time. You have to really guide it and if you're using a longer story to adapt, it'll compress things and lose context during Thinking. Luma labs was the most interesting tool I've seen so far, but there's still a lot of room for growth
Embarrassing!

You can still delete this, there is time.