145 comments

[ 0.19 ms ] story [ 6.2 ms ] thread
I heard a radio spot recently and I wondered if the voice was a real person or AI. It makes we wonder how such industries are dealing with this gen-AI revolution. We spend a lot of time here thinking about how it affects software developers, but I hardly ever see any commentary on how it is affecting screen and voice actors.
Sounds like a perfect application of AI. Some jobs should be automated.
I would rather hear human voices than synthesized ones. I don't care how realistic they sound. I'm not alone, and the sentiment will certainly grow.
As long as there are enough people than there will be a market. I think there are a lot of people that prefer human but don't want the human cost.
Ok, let's go the other way round. Which jobs do you think should not be automated?
Probably the ones we collectively value as human and are willing to pay for (in either our time or money or both).

And jobs where quality is paramount and or there is greater chance of risk / bodily harm

> Probably the ones we collectively value as human and are willing to pay for (in either our time or money or both).

Society seems to be very bad at that, and we end up with high frequency traders making millions while teachers are paid peanuts.

A lot of it will get automated the same way very many industries got automated. A lot of physical labour got automated once a primitive for it was created. Similarly, we have now a primitive for automating knowledge work. In the next few years to a decade, as all the right training data and runtime environments are slowly consolidated for various fields, a lot will be automated. There is no inherent reason a voice actor must be an eternal job, the same way there was no inherent reason for a draftsman or stage musician to be an eternal job.

It has nothing to do with skill. Both were very skilled jobs. Draftsman as well despite many people going to it straight from school. But computers and CAD mean that it is now necessary for someone to do a STEM degree to be a draftsman. Recorded audio made many stage musicians redundant. It is cheaper to do it this way and gets superior results, that is all, there is no further agenda.

Now too, the next generation of voice actors and many other knowledge workers will have to go up the value chain one step and operate or potentially build these tools (in whatever form they mature to in a decades time).

The current generation of voice actors will face the same situation as many before in the performance industry - stage musicians/performers for example that were made redundant by recorded audio. The reality is that most of them just left and dispersed into the economy doing completely unrelated jobs.

For software and generally computer engineers, this new primitive happens to itself be software, so it's less of a transition and an easier upskilling path to learn to build it. And building it is one step higher in the value chain than simply using it. That is a structural advantage.

I'm really sad about how little creative control these tools have. They seem great for creating slop, but pretty useless for creating content that someone would love. I love the potential but text isn't really a great medium for describing artistic vision.

All these demos are focused on how easy it makes everything. Easy is great, but if everyone is able to make instant cute cat videos or whatever it just devalues it. I want to see turning a photo into a rigged 3d model, letting the artist animate and then generate the video. This technology could be used to increase creative expression, but instead it's being used to squeeze out creative expression

Text is not the only input. You can provide 3d block outs with rudimentary animation, annotated images with arrows etc, voice recordings of one person acting out some emotion then mapping that to a different character's voice, other uses of video to video, etc.

There could easily be at least some time period of low skilled ugly people acting in approximate but shitty ways in cheap sets just to give an input reference to a model and then describing the differences in text, yielding gorgeous people doing stuff in fancy locations in the output.

That'll change. They'll get as many knobs as e.g. photoshop for image editing. You'll also be able, if not already, to use your own voice to communicate the emotion and tone you want but then have it re-render it in the voice of your choice.

The "slop" phase is early AI, like chunky ugly very early 3D graphics where you can see the triangles.

I'm already seeing images generated by first-generation diffusion models like Stable Diffusion 1.5 (which is small enough to run on a phone) used ironically as meme generators due to the now-retro silliness of what they generate.

recently I encountered several youtubers in the uncanny valley of "is it Ai or really repetitive intonation pattern"

one youtuber used irl footage, with hands and stuff - so I know there's human behind the camera

the other was a letsplay that reacted to events just fine emotionally

and yet the uncanny valley of the sound is in full force. Maybe youtube has done something with the codecs?

I feel sad

I know that there are at least a few content farms out there that specifically pump out "faceless video" to be resold, ie. those "hands in frame only" videos. The idea being that you can then put on some AI generated voiceover over them.
It'll come quickly for audio books. I've been working on a locally hosted, fully containerized web application to narrate my sci-fi novel using a full cast of characters and a few distinct narrators:

* https://i.ibb.co/ccqKZ71L/keenlore.png

* https://i.ibb.co/1t3W0JqZ/keenlore-02.png

* https://i.ibb.co/LdBqHwKB/keenlore-03.png

This is very interesting! Imagine Silly Tavern, with dialog tagging and coloring of some sort, auto-generated talking heads. Almost a game engine.
I would pay a few actors (mostly Star Trek TNG cast) money for a license to their voice.

I don't know how the next generation of beloved actors comes about and how we don't descend into a pit of neverending photocopies of things people once loved in the 1990s/2000s.

From reddit:

* $250 per finished hour for the narrator (lowest professional rate).

* 200k to 270k words.

* 22 hours, 20 hours, and 25 hours.

* Books 1 to 3 cost $5,500, $5,000, and $6,250, respectively.

My novel has about 8 major characters, including the 3 narrators, and 25+ minor characters. The price tag is daunting.

> I don't know how the next generation of beloved actors comes about

To me it's been obvious for a while now: there won't be one

"Celebrity" is an old world concept, and does not need to exist. it's find if we don't have them
it effects software develpoers because ai has compliers , tests , ci and bought tons of data in mercor.

it always sucks at everything else.

Related; here in my city, every event, restaurant, cinema, delivery service etc have all images / posters generated, very obviously, by OpenAI. ‘Authentic Italian’ and then an AI generated menu. This must hurt companies that used to make these before; a friend was a good photographer and he says he does wedding since end last year as the work was plummeting. Not sure about stats but it’s just ‘normal’ here at least.
My ex is a professional voice actor. Someone recently gave her a negative review saying she was AI, so now even the humans are getting unfairly penalized. I also heard a promotional video recently that used her voice, that she hadn't done, so someone has already trained their voice model on her.
Relatedly, a friend told me that graphical design gigs dried up because everyone can make event fliers and posters with AI with results that are often better than those of designers.
Prompt engineering tip for Google employees: just add "P.S. Make sure the page works in Firefox too."
It is in their business interest to ignore Firefox.
Then they aren't doing a very good job of it.

>Firefox makes up about 90 percent of Mozilla’s revenue, according to Muhlheim, the finance chief for the organization’s for-profit arm — which in turn helps fund the nonprofit Mozilla Foundation. About 85 percent of that revenue comes from its deal with Google, he added.

https://www.theverge.com/news/660548/firefox-google-search-r...

What people usually say is that Google merely wants Firefox to survive for anti-competitive reasons. Presumably that does not necessitate it actually being used (or be usable).
Hot take: Google keeps Mozilla/Firefox alive through the default search engine placement (which makes Mozilla millions each year) so they don't get designated as a monopoly with their browser.
That's lukewarm at most, doubt you'll find many people disagreeing here!
And of course they're doing the same for Apple/Safari, which wouldn't survive without the $20 million default search engine placement deal.
> And of course they're doing the same for Apple/Safari, which wouldn't survive without the $20 million default search engine placement deal.

Billion with a B, as in 10 zeroes, not 7.

Oops sorry, too late to edit, but you're correct!
Apple forcing users to use their own browser/browsing engine doesn't disprove my argument IMO, virtually nobody outside the apple ecosystem uses Safari, and outside the apple ecosystem is something between 80-90% of internet users.
I wasn't saying that Apple is just as monopolistic as Google; I was saying that Google is paying Apple orders of magnitude more than it is paying Mozilla. According to your argument, that would mean that Google is trying to keep Apple/Safari alive (which of course is absurd - Apple doesn't need that money to stay alive).
Google is so internally fractured, and the factions individually are so powerful, that it's really only the chrome people who care about chrome.

If anything FF gets left out because usage is so low.

> If anything FF gets left out because usage is so low.

I've been following this for a long time. They were leaving Firefox out when its usage wasn't low.

Probably the typical backdoor executive mandate that led to death by "sprint prioritization".

Yes, we will for sure work on the Firefox compatibility bug, Dave-Open-Source-Enthusiast-Google-Dev.

But we can only pick up 10 bugfixing tickets this sprint and this one, as the entire team agrees, is priority #12.

<repeat every sprint, where during the sprint 9-10 new higher priority items magically appear just in time for the next sprint>

Death by slow asphyxiation.

I heard that Firefox is banned internally at Google (something-something-security) so Devs can't even test against firefox
Not true -- at least up to 2018 when I worked there.
What I heard was more recent than 8 years ago, more like 6-12 months ago
(comment deleted)
Firefox really struggle with demo pages of text-to-video models because of the large numbers of videos in the page in my experience, this page seems to work quite fine for me tho.
No worries, even on chromium, their webpage is just horrible. They seems to have focus only on mobile reading, on desktop its a mess.
I had to change to a desktop to read the article because the videos didn't load on my phone.
Yes, that, and P.P.S. Make sure the page works in Chrome.

The page instant scrolled suddenly back to previous video examples after I scrolled down. It was disorienting.

Several Google apps do not work for me in Chrome (Maps, Earth) but work perfectly fine in Firefox... :D
So Seedance is good primarily because of TikTok and this because of YouTube. I wonder what portion of all recorded video is privately held in hard drives at people’s homes or Apple photos. Of course there is data labeling and cleaning but is the next evolution just a question of access? Same goes for LLMs. Would people be willing to sell their data? Kind of a messed up way to make yourself obsolete. Or there is a limit to scaling?
Google does anything except launch a new version of Gemini Pro.
Just because Anthropic and OpenAI really want there to be an arms race justifying the outsized investment, doesn't mean the optimal play is to build larger, more expensive, models.

The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.

I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.

If the Chinese labs can compete on a shoestring budget with access to much less powerful hardware, Google should be able to compete as well. They're becoming almost irrelevant for agentic coding right now.
AI / LLM is about more than agentic coding. It is one of the least interesting use cases to me, thinking more broadly. HN may be over-indexed on it.
I would agree with you on AI/LLM being more than agentic coding but at the same time, I think there's more nuance.

For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.

A lot of browser automation could/is also done by agentic coding.

It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.

You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).

It might sound self-explainatory but turns out that coding can represent a wide range of problems!

I get that. I use agents a lot and LLMs often reason with code. It is valuable. I just think the floor is a lot lower for general reasoning and common tasks like that. And in 6-12 months it won’t matter. Google will publish better models. The temporal distortion of how long a Sol or a Fable has existed is real. No one is suddenly missing out on some giant competitive edge because their model is a few months behind. I feel like it’s all just going to normalize and things other than how well your model can write code will matter more and more in 12 to 24 months.
Sure I understand what you mean as well and I am not asking for SoTA models to be created by Google but more so explaining why coding is still the largest focus for many labs.

I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.

> No one is suddenly missing out on some giant competitive edge because their model is a few months behind

Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.

> things other than how well your model can write code will matter more and more in 12 to 24 months.

What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?

It's not much of a shoestring budget to be receiving regular injections of investment from state lenders along with cheap credit. I don't think the comparison holds.
IDK why you'd believe they have a shoestring budget.
And it doesn't have to be either/or. They could make larger, more expensive models, just at a slower cadence.

Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.

Google paid for 3.5 Pro training. They just didn't release it.

They never gave an official answer as to why, so I'll let you draw your own conclusions.

They did not decide it wasn't worth spending the money to train.

They absolutely spent the money.

I work there. I have zero internal knowledge about the model. Opinion my own, etc. I don't think it is worth fighting to win on a month to month time horizon. When you step back and look an inch above this market, Gemini Pro 3.1 as a product was released in February. 6 months. It feels like forever and that Google is behind, but on a 2-3 year horizon? The models are going to stay similar.

Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?

I thought of a quite a few and they are far more compelling and interesting to me.

(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)

I'm the CTO of a GCP shop with an 8 figure annual commit.

If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.

In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.

Given your position and the responsibility that comes with it; I sure hope you updated your mental model in another way than simply "they are acting irrational"... It's not clear from your comment that you did, but it sounded a bit like it.
3.5 was almost certainly a 3.1 post-train, so likely a small investment on Google's part.

They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.

Yes, I agree with you that the race all the AI companies are running doesn't make sense, but at the same time, there are rumors that Google has produced newer versions of Pro without releasing them to the public.

Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.

There is more to the cost of a model than its training. While training is a significant Capex expenditure, it has very low Operational cost after training unless it is deployed for public inference.

It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.

Google doesn't have a good coding model. This is a HUGE problem. They don't need "larger more expensive models", they need a good coding model because it's a competitive advantage.
Competitive advantage why? Will it really make them more money? They already have Google Cloud.
3.7 flash ain't bad. It isn't Opus or Sol, but it fits pretty well in the second tier.
Second tier models are like self-driving cars to the point where you question if they save any time. Sol can do a deep analysis and plan a large feature. Luna can mostly execute. There's a clear qualitative difference.
All Google has to do is build a model that works good enough for the Gemini app and for Spark. And they have it.
I have a pro subscription, I think they have just given up. Likely because when they test their new models against the other frontier models they are so bad, they just pull it back. This leads them to try and innovate in other areas where there is currently less competition so they can compete. Not a bad play.
You can tell you live in the HN/tech bubble.

In the real world out there, Google and Microsoft are absolutely dominating enterprise customers.

Every single non-tech office worker I know is writing Gemini "gems" (sort of claude prompts/skills) or prompting Copilot to help drafting board meeting notes, insurance contracts updates that reflect changes in regulations, make quick loan feasibility assessments before passing them to the relevant office, presentations, etc, etc.

I'm talking insurance, banking, consultancy, manufacturing, etc, etc.

Why? Because Google and Microsoft already were in these companies, all they had to do is "oh, you also have AI now with your plans". Procurement and data compliance are the first thing businesses have to sort out. They were already sorted out.

Google doesn't need to have the best coding model or triumph in meaningless benchmarks, it only needs their models to get better and cheaper while serving them to their existing customer base.

They are playing a different game.

And Microsoft, doesn't even need to care about models at all, they can provide whatever open or closed AI with their services and have to focus on the harness in Excel or Github/Azure Copilot or whatever.

E.g. while developers in most of my clients use whatever they prefer or the company pays for, the remaining 90% uses either Google or Microsoft products.

Not a single one has incentives into venturing into OpenAI or Anthropic or Z.Ai lands because they might be better at some benchmark that is completely irrelevant to their tasks of updating powerpoints or summarizing incoming emails.

This.

Google is an advertising company with an enterprise SaaS branch. I bet they have chosen to focus on running the most efficient "everybody" model instead of running a heavy model for coders.

When using Gemini for other tasks than coding, it is actually pretty good. It grounds well with Google search and gives mostly correct, well written answers to many niche questions.

New version of Gemini Pro has now become like GTA-6
Interesting that OpenAI abandoned Sora entirely but Google are continuing to invest heavily in their own video generation.

Maybe because they see video generation as key to developing "world models"?

This is paid API access only. They are here to make money not to position themselves for an IPO. Not a value judgement only an observation.
Google has always been committed to multimodal.

And, Veo and omni simply were better than Sora

And, Youtube is huge both as a place where video contents goes and where can be trained from. Microdramas are starting to become a real category--14 Billion USD, 90% of it made with AI.

Chinese video models can be more immediately impressive, but none of them come close to beat the value of Google's Flow. Especially when you are throwing away a lot of generations as part of the creative process. Which is what you have to do to make longer content with any video model.

OpenAI needed to be able to focus. Google can walk and chew gum, and they're not going to run out of money to buy chewing gum.

Any source on the 90% number? Last I saw (recently) one of the big players just announced they were planning to do 50% AI going forward
> Microdramas are starting to become a real category--14 Billion USD

Almost all concrete english-language info I can find about microdramas is astroturfed to hell by consultants and "independent" industry publications. Wikipedia's citations for 2025 revenue are 'Reel Reel' and 'Duanju News'. Neither cite their source, though they are likely just regurgitating predictions from Omdia, a media consultancy.

> Real Reel™ works with companies across entertainment, technology and the creator economy to build relevant industry conversations around mobile-first storytelling.

Duanju News is published by 'Studio Phocéen', who run their own microdrama production house.

The Omdia revenue estimates, which is where the projected $14 billion comes from, are completely unsubstantiated as far as I can tell. They even go so far as to make "according to new research" a link that when followed sends you to their generic "/advance-your-business/media-and-entertainment" sales pitch.

https://omdia.tech.informa.com/pr/2025/oct/microdramas-to-ge...

I don't have any special insights here, it all just smells a bit like McKinsey's "The metaverse will be worth $5 trillion by 2030, you better not miss out!!! Hire our 22 year old slide deck experts today"

YouTube. And video ads.

Previously when making video ads you'd need to actually create the video. Actors, cameramen, editors - you name it. Now a new video ads is just a prompt away, directly inside the ad-spend web UI too no doubt.

People say Google have lost and that they're having their lunch eaten by anthropic, but I am not so sure...

Google owns 15% of Anthropic, Claude trains and runs on TPUs, and Google cloud is backlogged with demand from both OAI and Anthropic.

Google is selling shovels, leasing mines, buy stakes in "competitors" and doing it's own exploration/mining. When you look at the full picture, it kinda doesn't even look like Gemini matters that much to them overall.

I don't think it does. They could probably throw a shitload of compute at Gemini and have something more competitive. Instead they seem focused on the actual innovative research which I think will pay off long term.

Gemma 4 I am way more impressed by than whatever the top ranking model is

That's a very convincing answer. I hadn't considered how many advertisers need video ads now and don't have the skill or resources to create them.
> People say Google have lost and that they're having their lunch eaten by anthropic

That's old news. People are saying the chinese models/labs are eating anthropic lunch.

They're main edge has been multimodal. I think they're still the best overall on multimodal? If I were them I would try to be the best at at least something
Sora was a social network type thing. Google sells their models on a PAYG basis - and makes money off them. Nano Banana alone has changed advertising 2D mockups and Photoshop like tasks forever. Notice how GPT image 2 is now available also on a PAYG basis.

I work in advertising and some days I spent a lot of money using these models. The amount and rapidity of prototyping using them has changed everything about advertising pre production.

its probably not even a choice given what bytedance is doing with video models
(comment deleted)
its probably not even a choice given what bytedance is doing with their video models imo
its probably not even a choice given what bytedance is doing with their video models imo. the tiktok youtube rivalry is too strong
wish some of these frontier models supported 3d.
The AI brand fragmentation at Google is not yet a problem because everyone is pretending:

x There are so-called “SOTA” or “frontier” models that are more effective than the other ones (independent of harnessing and routing)

x OpenAI and Anthropic have all the SOTA models and lead all the innovation

x Google’s moat is its search bread/butter (it’s the only reason they’re relevant)

All 3 operating assumptions are - I think - false.

What Google has done that the “cuter products” (Claude, ChatGPT) haven’t is connected relatively standard LLMs to an externally valuable live service.

As more companies realize that is where all the value is (the service) and not in the AI capability, then products (and humans) become important again.

Google should just be Google again, and Gemini should be Gemini, off to the side. Omni confuses everyone (and angers some iykyk), they should resolve “AI mode”, rename Gemma? and consolidate the brand overall so it’s clear what Google is.

Google is search.

It helps people on all sides of the market find what they’re looking for.

I don’t really see how repeatedly reinventing and rebranding the same AI chat UX is accomplishing anything toward that goal.

> Google is search.

Implicit to that is "find". Their AI integration into search has really hit its stride for me. They have that search box (or speech prompt) hard wired into people and they are finally iterating and crafting AI into that experience. They really failed hard initially.

I know others have worse experiences than me but Google knows a lot about me so maybe that affects my results. YMMV

> YMMV

It does. Very commonly it will get acronyms wrong or assume I mean something else, even when it should be in my “ad profile” or whatever it bases it off.

Often I think the search engine is working perfectly then the LLM is ruining it in delivery.

There are other UIs besides chat

You (we) type in some words. 2-8 roughly. You hit enter. Are you saying that in the website results that are presented they are skewed by the LLM taking off to somewhere you aren't intending? Steering you off - like to a place they want you to go?

You say "ruining the delivery". I believe I understand. It is not traditional www text search.

I do have a problem with most chat ai's that always ask a question at the end of their answer. This is where configurability is more important to me than absolute performance. e.g. "don't end an answer with a question unless it is directly relevant to the current conversation.

Yes there are many other UIs besides chat. The default chat UI really bothered me at first. Seemed cheesy. However the majority of human interaction is "chat".

Something that is really really cool in this tech is the vision and audio and text models that naturally cluster similar things in massive multi-dimensional arrays. I have a hard time mapping nut~squirrel~food~imageofnut~dogsaspredators~whateves. How does that condense?

We have come so far on the audio transcription dimension. In this discussion context yes it is still chat but we've moved to a different place in the brain. Not radically different (humans without visual sight come to mind in how the process goes from photons to perception (also see hank green's vid on how the eyes are part of the brain)).

Anyway hope this isn't a throwaway account. That username doesn't inspire confidence in that regard.

What the hell am I doing?

[edit: made things more clear about my un-clear thoughts! I did not shift things in the intent space]

I'm still getting major uncanny valley from any of the videos featuring humans, something about them disgusts me. I guess I should be glad I'm still able to distinguish them.
Something in their eyes. Looks very robotic / lifeless for me. And the sound-mixing is very off. Clearly feels like the voice was layered on top of whatever sound is in the background and not blended.
I don't think I can see the difference. I just have my skin crawl because I'm expecting to see something off and generally have a bad feeling about it.
You knew upfront that it was generated. I did too so the first thing I did was look at the lips of the actors convincing myself that there were flaws.

When it comes to fish swimming around I don't think I would be able to reliably tell what was real vs generated even with deep inspection.

If you're using any of these generated videos in any professional setting, I don't think I will be able to ever use your business.
Only AI code is holy, video is sinful.
Nobody says that AI code is holy, I'm pretty sure vibe coded crap by people that know nothing about writing code has the same effect.
It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.

Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.

I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).

Yes, most people using AI for creative projects spend a lot of time and attention mastering their tools, and figure out how to adapt them into their creative processes. AI can dramatically lower the cost of indie productions, while also a allowing a broader range of stories to be told. Even the most successful film makers need to bow and scrape to get their projects funded, democratizing visual media can be a good thing, even if you, personally, are no more likely to do this than you are to pick up Photoshop or record a podcast.

I enjoy making short films with AI. When my latest short screens at a festival in Ocotber, alongside traditional and AI films, hopefully the audience will like it too.

The quick "one shot" video generation might be slop to you, or I. But if someone wants to send it as birthday greeting to their aunt, and they both enjoy it, what business is it of ours?

The cheaper it is to produce, the more daring it can be. Which is a good thing.
Kids love image and video generation! The former is cheap enough to do just because it is fun.

I have a young boy, and whenever he builds an impressive "scene" from LEGO (like a diorama or whatever), I take a couple of reference pictures with my phone and make it into a "real" movie scene, cartoon, or whatever. He loves it, and this motivates him to build more and bigger things out of LEGO.

If he builds something really special, I might actually fork over the $5 to use Omni to turn his LEGO creation into a 10-second video instead of a still image. It'll blow his mind!

PS: There also are cheap and even free phone apps that make stop-motion animation trivial. We've already made a couple of videos of his toys moving around that way.

Google's main source of revenue is advertising not enriching people's lives. It's not a charity.
> As in, what work or enjoyment does someone actually get from this? Why do we look at art, watch videos/movies?

I like to generate songs from obscure poems

> As in, what work or enjoyment does someone actually get from this?

I have been pondering this and have come out with a single statement. It allows you to create the missing piece in the art you want to create. Assets for games, music for the lyrics you wrote, or just a whole song to justify a crazy dance you want to do. Whenever AI art is discussed, people tend to be so purist about art. One person cannot do everything in a project they take on to express themselves. Usually people come back with then get someone else to do it for you. I think that is an economic argument as people are trying to protect artist pay. I get that. But is it fair to just not let something they want to create because they don't posses every talent needed to do this?

I let myself get mildly excited with the last Omni release, but it turns out it (and this one) can't do the one practical thing I want - Sync generated video to provided pre-existing audio.

Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.

Yeah, that'd be a great feature.
Pretty sure I heard one of the PMs in a podcast a few weeks ago say they are intentionally not building support for it out of concerns of enabling deep-fakes.

I agree though. My issue is the cost for using AI video models is way too high for anyone not building anything serious with them, at the same time they are too restricted for actually using professionally. Prompting them with text to get something generated is cute, but then you just end up creating slop that everyone hates, ultimately devaluing the power of these things.

> Minimax H3 locally on my 12Gb 4070RTX

Minimax H3 is about 240Gb alone, how do you do? How much quantised is it, and how good are the results?

I'm running it on my 8 years old 2080ti 11 gb VRAM + 32 gb ram. After tasking fable to optimize the setup for a few nights it's pretty good for some quick funny clips (2min30 for 5s , 5mins for 10s, with reference pics to insert anyone in the clip, not great quality but acceptable for some fun on a phone). I haven't had as much fun with Gen AI since the Stable Diffusion days.
The int8 release is very capable and runs acceptably on consumer gpus
So we are just making up numbers now, huh
Raw generation quality is becoming table stakes; controllability might be the more important battleground.
Google we ain't falling for it. #neverforget3.5prowithinamonth
Draft videos more efficiently in 360p

While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.

Can you hold the seed constant?
No option I could see, their changelog suggests using their new upscaler but I've had nothing but disappointment from upscalers.

Upscale when ready: Users on paid tiers can seamlessly upgrade their favorite clips using our new 360p to 720p upscaler.

>"Draft videos more efficiently in 360p

Generate lightweight previews in 360p resolution up to 60% faster

and at a third of the cost compared to Omni 1.1’s standard 720p resolution. This is helpful for rapid prototyping, storyboard iteration, and quick rendering in developer platforms."

This is a great idea, to have a low-resolution mode for additional speed to create previews, do test runs, create rapid prototypes, etc.

My curiousity is, what's the absolute useable minimum that this could be?

That is, would/could 240p resolution work? If so, what about 144p? How about lower? Then, could those images be upscaled quickly (and is the result still usable?) with a faster image upscaling-only neural network?

The reason why knowing such lower numbers / lower bounds -- is because they could be important for additional cost/time savings and/or running derived LLM's on local resource-constrained hardware...

Anyway, great post, great idea, and we welcome Gemnini Omni 1.1 Flash to the ever-expanding list of LLM/AI's!

For me it's kind of weird right it's like no one can see the bigger picture. All technology's going to be eaten by AI so companies like Google, Microsoft, Meta ... are middle men in an industry that's disappearing. No one's really going to use a search engine when a chat bot can do a better job. No one is going to pay for software when the AI can do the whole process better at a proper command. All technology and the companies in it are essentially horses and the related roles in a time of cars. That whole slice of the economy is at best going to reduce year on year eventually to disappear. No replacement AI is it.
The chat bot will still use the search engine
Would this model be available in ComfyUI, or does Google limit these to their own tools?
> What stands out about Gemini Omni Flash is its accuracy: the details hold up under scrutiny.

Quote under the video of a short Argentinian footballer wearing no 10 with "RESSC" on his back. Can't make it up.

Is there a good example someone can show of naturally generating real videos using extending and interpolation from the last frames? (For any of the video generators).

I mean I have not tinkered that much, but trying to even get a video to 30 seconds (I just want a cartoon AI avatar to narrate tutorials) has been incredibly difficult. They drift so easily.

Many AI videos you can tell just stitch short clips together. I just want a continuous scene for like 30 seconds to 2 minutes.

Anyone else becoming numb to these updates?

I feel like I should be excited about being able to generate almost perfect videos but, I just don't care anymore.

Before AI, cool and interesting shots carried the promise that it was reality - even if it was perhaps exaggerated.

The amazing videos and photos carried the promise that I could experience that for real. They were aspirational.

Today I suspect every cool shot is made of pixels arranged on a 2D screen by an algorithm. It doesn't do it for me.

That makes me sad...

I have written how is impact the AI Industry in Nasengetu Tech News website.