I heard a radio spot recently and I wondered if the voice was a real person or AI. It makes we wonder how such industries are dealing with this gen-AI revolution. We spend a lot of time here thinking about how it affects software developers, but I hardly ever see any commentary on how it is affecting screen and voice actors.
A lot of it will get automated the same way very many industries got automated. A lot of physical labour got automated once a primitive for it was created. Similarly, we have now a primitive for automating knowledge work. In the next few years to a decade, as all the right training data and runtime environments are slowly consolidated for various fields, a lot will be automated. There is no inherent reason a voice actor must be an eternal job, the same way there was no inherent reason for a draftsman or stage musician to be an eternal job.
It has nothing to do with skill. Both were very skilled jobs. Draftsman as well despite many people going to it straight from school. But computers and CAD mean that it is now necessary for someone to do a STEM degree to be a draftsman. Recorded audio made many stage musicians redundant. It is cheaper to do it this way and gets superior results, that is all, there is no further agenda.
Now too, the next generation of voice actors and many other knowledge workers will have to go up the value chain one step and operate or potentially build these tools (in whatever form they mature to in a decades time).
The current generation of voice actors will face the same situation as many before in the performance industry - stage musicians/performers for example that were made redundant by recorded audio. The reality is that most of them just left and dispersed into the economy doing completely unrelated jobs.
For software and generally computer engineers, this new primitive happens to itself be software, so it's less of a transition and an easier upskilling path to learn to build it. And building it is one step higher in the value chain than simply using it. That is a structural advantage.
I'm really sad about how little creative control these tools have. They seem great for creating slop, but pretty useless for creating content that someone would love. I love the potential but text isn't really a great medium for describing artistic vision.
All these demos are focused on how easy it makes everything. Easy is great, but if everyone is able to make instant cute cat videos or whatever it just devalues it. I want to see turning a photo into a rigged 3d model, letting the artist animate and then generate the video. This technology could be used to increase creative expression, but instead it's being used to squeeze out creative expression
Text is not the only input. You can provide 3d block outs with rudimentary animation, annotated images with arrows etc, voice recordings of one person acting out some emotion then mapping that to a different character's voice, other uses of video to video, etc.
There could easily be at least some time period of low skilled ugly people acting in approximate but shitty ways in cheap sets just to give an input reference to a model and then describing the differences in text, yielding gorgeous people doing stuff in fancy locations in the output.
That'll change. They'll get as many knobs as e.g. photoshop for image editing. You'll also be able, if not already, to use your own voice to communicate the emotion and tone you want but then have it re-render it in the voice of your choice.
The "slop" phase is early AI, like chunky ugly very early 3D graphics where you can see the triangles.
I'm already seeing images generated by first-generation diffusion models like Stable Diffusion 1.5 (which is small enough to run on a phone) used ironically as meme generators due to the now-retro silliness of what they generate.
I know that there are at least a few content farms out there that specifically pump out "faceless video" to be resold, ie. those "hands in frame only" videos. The idea being that you can then put on some AI generated voiceover over them.
It'll come quickly for audio books. I've been working on a locally hosted, fully containerized web application to narrate my sci-fi novel using a full cast of characters and a few distinct narrators:
I would pay a few actors (mostly Star Trek TNG cast) money for a license to their voice.
I don't know how the next generation of beloved actors comes about and how we don't descend into a pit of neverending photocopies of things people once loved in the 1990s/2000s.
Related; here in my city, every event, restaurant, cinema, delivery service etc have all images / posters generated, very obviously, by OpenAI. ‘Authentic Italian’ and then an AI generated menu. This must hurt companies that used to make these before; a friend was a good photographer and he says he does wedding since end last year as the work was plummeting. Not sure about stats but it’s just ‘normal’ here at least.
My ex is a professional voice actor. Someone recently gave her a negative review saying she was AI, so now even the humans are getting unfairly penalized. I also heard a promotional video recently that used her voice, that she hadn't done, so someone has already trained their voice model on her.
Relatedly, a friend told me that graphical design gigs dried up because everyone can make event fliers and posters with AI with results that are often better than those of designers.
>Firefox makes up about 90 percent of Mozilla’s revenue, according to Muhlheim, the finance chief for the organization’s for-profit arm — which in turn helps fund the nonprofit Mozilla Foundation. About 85 percent of that revenue comes from its deal with Google, he added.
What people usually say is that Google merely wants Firefox to survive for anti-competitive reasons. Presumably that does not necessitate it actually being used (or be usable).
Hot take: Google keeps Mozilla/Firefox alive through the default search engine placement (which makes Mozilla millions each year) so they don't get designated as a monopoly with their browser.
Apple forcing users to use their own browser/browsing engine doesn't disprove my argument IMO, virtually nobody outside the apple ecosystem uses Safari, and outside the apple ecosystem is something between 80-90% of internet users.
I wasn't saying that Apple is just as monopolistic as Google; I was saying that Google is paying Apple orders of magnitude more than it is paying Mozilla. According to your argument, that would mean that Google is trying to keep Apple/Safari alive (which of course is absurd - Apple doesn't need that money to stay alive).
Firefox really struggle with demo pages of text-to-video models because of the large numbers of videos in the page in my experience, this page seems to work quite fine for me tho.
So Seedance is good primarily because of TikTok and this because of YouTube. I wonder what portion of all recorded video is privately held in hard drives at people’s homes or Apple photos. Of course there is data labeling and cleaning but is the next evolution just a question of access? Same goes for LLMs. Would people be willing to sell their data? Kind of a messed up way to make yourself obsolete. Or there is a limit to scaling?
Just because Anthropic and OpenAI really want there to be an arms race justifying the outsized investment, doesn't mean the optimal play is to build larger, more expensive, models.
The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.
If the Chinese labs can compete on a shoestring budget with access to much less powerful hardware, Google should be able to compete as well. They're becoming almost irrelevant for agentic coding right now.
I would agree with you on AI/LLM being more than agentic coding but at the same time, I think there's more nuance.
For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.
A lot of browser automation could/is also done by agentic coding.
It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.
You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).
It might sound self-explainatory but turns out that coding can represent a wide range of problems!
I get that. I use agents a lot and LLMs often reason with code. It is valuable. I just think the floor is a lot lower for general reasoning and common tasks like that. And in 6-12 months it won’t matter. Google will publish better models. The temporal distortion of how long a Sol or a Fable has existed is real. No one is suddenly missing out on some giant competitive edge because their model is a few months behind. I feel like it’s all just going to normalize and things other than how well your model can write code will matter more and more in 12 to 24 months.
Sure I understand what you mean as well and I am not asking for SoTA models to be created by Google but more so explaining why coding is still the largest focus for many labs.
I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.
> No one is suddenly missing out on some giant competitive edge because their model is a few months behind
Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.
> things other than how well your model can write code will matter more and more in 12 to 24 months.
What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?
It's not much of a shoestring budget to be receiving regular injections of investment from state lenders along with cheap credit.
I don't think the comparison holds.
And it doesn't have to be either/or. They could make larger, more expensive models, just at a slower cadence.
Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.
I work there. I have zero internal knowledge about the model. Opinion my own, etc. I don't think it is worth fighting to win on a month to month time horizon. When you step back and look an inch above this market, Gemini Pro 3.1 as a product was released in February. 6 months. It feels like forever and that Google is behind, but on a 2-3 year horizon? The models are going to stay similar.
Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?
I thought of a quite a few and they are far more compelling and interesting to me.
(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)
I'm the CTO of a GCP shop with an 8 figure annual commit.
If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.
In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.
Given your position and the responsibility that comes with it; I sure hope you updated your mental model in another way than simply "they are acting irrational"... It's not clear from your comment that you did, but it sounded a bit like it.
3.5 was almost certainly a 3.1 post-train, so likely a small investment on Google's part.
They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.
Yes, I agree with you that the race all the AI companies are running doesn't make sense, but at the same time, there are rumors that Google has produced newer versions of Pro without releasing them to the public.
Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.
There is more to the cost of a model than its training.
While training is a significant Capex expenditure, it has very low Operational cost after training unless it is deployed for public inference.
It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.
Google doesn't have a good coding model. This is a HUGE problem. They don't need "larger more expensive models", they need a good coding model because it's a competitive advantage.
Second tier models are like self-driving cars to the point where you question if they save any time. Sol can do a deep analysis and plan a large feature. Luna can mostly execute. There's a clear qualitative difference.
I have a pro subscription, I think they have just given up. Likely because when they test their new models against the other frontier models they are so bad, they just pull it back. This leads them to try and innovate in other areas where there is currently less competition so they can compete. Not a bad play.
In the real world out there, Google and Microsoft are absolutely dominating enterprise customers.
Every single non-tech office worker I know is writing Gemini "gems" (sort of claude prompts/skills) or prompting Copilot to help drafting board meeting notes, insurance contracts updates that reflect changes in regulations, make quick loan feasibility assessments before passing them to the relevant office, presentations, etc, etc.
I'm talking insurance, banking, consultancy, manufacturing, etc, etc.
Why? Because Google and Microsoft already were in these companies, all they had to do is "oh, you also have AI now with your plans". Procurement and data compliance are the first thing businesses have to sort out. They were already sorted out.
Google doesn't need to have the best coding model or triumph in meaningless benchmarks, it only needs their models to get better and cheaper while serving them to their existing customer base.
They are playing a different game.
And Microsoft, doesn't even need to care about models at all, they can provide whatever open or closed AI with their services and have to focus on the harness in Excel or Github/Azure Copilot or whatever.
E.g. while developers in most of my clients use whatever they prefer or the company pays for, the remaining 90% uses either Google or Microsoft products.
Not a single one has incentives into venturing into OpenAI or Anthropic or Z.Ai lands because they might be better at some benchmark that is completely irrelevant to their tasks of updating powerpoints or summarizing incoming emails.
Google is an advertising company with an enterprise SaaS branch. I bet they have chosen to focus on running the most efficient "everybody" model instead of running a heavy model for coders.
When using Gemini for other tasks than coding, it is actually pretty good. It grounds well with Google search and gives mostly correct, well written answers to many niche questions.
And, Youtube is huge both as a place where video contents goes and where can be trained from. Microdramas are starting to become a real category--14 Billion USD, 90% of it made with AI.
Chinese video models can be more immediately impressive, but none of them come close to beat the value of Google's Flow. Especially when you are throwing away a lot of generations as part of the creative process. Which is what you have to do to make longer content with any video model.
OpenAI needed to be able to focus. Google can walk and chew gum, and they're not going to run out of money to buy chewing gum.
> Microdramas are starting to become a real category--14 Billion USD
Almost all concrete english-language info I can find about microdramas is astroturfed to hell by consultants and "independent" industry publications. Wikipedia's citations for 2025 revenue are 'Reel Reel' and 'Duanju News'. Neither cite their source, though they are likely just regurgitating predictions from Omdia, a media consultancy.
> Real Reel™ works with companies across entertainment, technology and the creator economy to build relevant industry conversations around mobile-first storytelling.
Duanju News is published by 'Studio Phocéen', who run their own microdrama production house.
The Omdia revenue estimates, which is where the projected $14 billion comes from, are completely unsubstantiated as far as I can tell. They even go so far as to make "according to new research" a link that when followed sends you to their generic "/advance-your-business/media-and-entertainment" sales pitch.
I don't have any special insights here, it all just smells a bit like McKinsey's "The metaverse will be worth $5 trillion by 2030, you better not miss out!!! Hire our 22 year old slide deck experts today"
Previously when making video ads you'd need to actually create the video. Actors, cameramen, editors - you name it. Now a new video ads is just a prompt away, directly inside the ad-spend web UI too no doubt.
People say Google have lost and that they're having their lunch eaten by anthropic, but I am not so sure...
Google owns 15% of Anthropic, Claude trains and runs on TPUs, and Google cloud is backlogged with demand from both OAI and Anthropic.
Google is selling shovels, leasing mines, buy stakes in "competitors" and doing it's own exploration/mining. When you look at the full picture, it kinda doesn't even look like Gemini matters that much to them overall.
I don't think it does. They could probably throw a shitload of compute at Gemini and have something more competitive. Instead they seem focused on the actual innovative research which I think will pay off long term.
Gemma 4 I am way more impressed by than whatever the top ranking model is
They're main edge has been multimodal. I think they're still the best overall on multimodal? If I were them I would try to be the best at at least something
Sora was a social network type thing. Google sells their models on a PAYG basis - and makes money off them. Nano Banana alone has changed advertising 2D mockups and Photoshop like tasks forever. Notice how GPT image 2 is now available also on a PAYG basis.
I work in advertising and some days I spent a lot of money using these models. The amount and rapidity of prototyping using them has changed everything about advertising pre production.
The AI brand fragmentation at Google is not yet a problem because everyone is pretending:
x There are so-called “SOTA” or “frontier” models that are more effective than the other ones (independent of harnessing and routing)
x OpenAI and Anthropic have all the SOTA models and lead all the innovation
x Google’s moat is its search bread/butter (it’s the only reason they’re relevant)
All 3 operating assumptions are - I think - false.
What Google has done that the “cuter products” (Claude, ChatGPT) haven’t is connected relatively standard LLMs to an externally valuable live service.
As more companies realize that is where all the value is (the service) and not in the AI capability, then products (and humans) become important again.
Google should just be Google again, and Gemini should be Gemini, off to the side. Omni confuses everyone (and angers some iykyk), they should resolve “AI mode”, rename Gemma? and consolidate the brand overall so it’s clear what Google is.
Google is search.
It helps people on all sides of the market find what they’re looking for.
I don’t really see how repeatedly reinventing and rebranding the same AI chat UX is accomplishing anything toward that goal.
Implicit to that is "find". Their AI integration into search has really hit its stride for me. They have that search box (or speech prompt) hard wired into people and they are finally iterating and crafting AI into that experience. They really failed hard initially.
I know others have worse experiences than me but Google knows a lot about me so maybe that affects my results. YMMV
It does. Very commonly it will get acronyms wrong or assume I mean something else, even when it should be in my “ad profile” or whatever it bases it off.
Often I think the search engine is working perfectly then the LLM is ruining it in delivery.
You (we) type in some words. 2-8 roughly. You hit enter. Are you saying that in the website results that are presented they are skewed by the LLM taking off to somewhere you aren't intending? Steering you off - like to a place they want you to go?
You say "ruining the delivery". I believe I understand. It is not traditional www text search.
I do have a problem with most chat ai's that always ask a question at the end of their answer. This is where configurability is more important to me than absolute performance. e.g. "don't end an answer with a question unless it is directly relevant to the current conversation.
Yes there are many other UIs besides chat. The default chat UI really bothered me at first. Seemed cheesy. However the majority of human interaction is "chat".
Something that is really really cool in this tech is the vision and audio and text models that naturally cluster similar things in massive multi-dimensional arrays. I have a hard time mapping nut~squirrel~food~imageofnut~dogsaspredators~whateves. How does that condense?
We have come so far on the audio transcription dimension. In this discussion context yes it is still chat but we've moved to a different place in the brain. Not radically different (humans without visual sight come to mind in how the process goes from photons to perception (also see hank green's vid on how the eyes are part of the brain)).
Anyway hope this isn't a throwaway account. That username doesn't inspire confidence in that regard.
What the hell am I doing?
[edit: made things more clear about my un-clear thoughts! I did not shift things in the intent space]
I'm still getting major uncanny valley from any of the videos featuring humans, something about them disgusts me. I guess I should be glad I'm still able to distinguish them.
Something in their eyes. Looks very robotic / lifeless for me. And the sound-mixing is very off. Clearly feels like the voice was layered on top of whatever sound is in the background and not blended.
I don't think I can see the difference. I just have my skin crawl because I'm expecting to see something off and generally have a bad feeling about it.
It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.
Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.
I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).
Yes, most people using AI for creative projects spend a lot of time and attention mastering their tools, and figure out how to adapt them into their creative processes. AI can dramatically lower the cost of indie productions, while also a allowing a broader range of stories to be told. Even the most successful film makers need to bow and scrape to get their projects funded, democratizing visual media can be a good thing, even if you, personally, are no more likely to do this than you are to pick up Photoshop or record a podcast.
I enjoy making short films with AI. When my latest short screens at a festival in Ocotber, alongside traditional and AI films, hopefully the audience will like it too.
The quick "one shot" video generation might be slop to you, or I. But if someone wants to send it as birthday greeting to their aunt, and they both enjoy it, what business is it of ours?
Kids love image and video generation! The former is cheap enough to do just because it is fun.
I have a young boy, and whenever he builds an impressive "scene" from LEGO (like a diorama or whatever), I take a couple of reference pictures with my phone and make it into a "real" movie scene, cartoon, or whatever. He loves it, and this motivates him to build more and bigger things out of LEGO.
If he builds something really special, I might actually fork over the $5 to use Omni to turn his LEGO creation into a 10-second video instead of a still image. It'll blow his mind!
PS: There also are cheap and even free phone apps that make stop-motion animation trivial. We've already made a couple of videos of his toys moving around that way.
> As in, what work or enjoyment does someone actually get from this?
I have been pondering this and have come out with a single statement. It allows you to create the missing piece in the art you want to create. Assets for games, music for the lyrics you wrote, or just a whole song to justify a crazy dance you want to do. Whenever AI art is discussed, people tend to be so purist about art. One person cannot do everything in a project they take on to express themselves. Usually people come back with then get someone else to do it for you. I think that is an economic argument as people are trying to protect artist pay. I get that. But is it fair to just not let something they want to create because they don't posses every talent needed to do this?
I let myself get mildly excited with the last Omni release, but it turns out it (and this one) can't do the one practical thing I want - Sync generated video to provided pre-existing audio.
Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.
Pretty sure I heard one of the PMs in a podcast a few weeks ago say they are intentionally not building support for it out of concerns of enabling deep-fakes.
I agree though. My issue is the cost for using AI video models is way too high for anyone not building anything serious with them, at the same time they are too restricted for actually using professionally. Prompting them with text to get something generated is cute, but then you just end up creating slop that everyone hates, ultimately devaluing the power of these things.
I'm running it on my 8 years old 2080ti 11 gb VRAM + 32 gb ram. After tasking fable to optimize the setup for a few nights it's pretty good for some quick funny clips (2min30 for 5s , 5mins for 10s, with reference pics to insert anyone in the clip, not great quality but acceptable for some fun on a phone). I haven't had as much fun with Gen AI since the Stable Diffusion days.
While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.
Generate lightweight previews in 360p resolution up to 60% faster
and at a third of the cost compared to Omni 1.1’s standard 720p resolution. This is helpful for rapid prototyping, storyboard iteration, and quick rendering in developer platforms."
This is a great idea, to have a low-resolution mode for additional speed to create previews, do test runs, create rapid prototypes, etc.
My curiousity is, what's the absolute useable minimum that this could be?
That is, would/could 240p resolution work? If so, what about 144p? How about lower? Then, could those images be upscaled quickly (and is the result still usable?) with a faster image upscaling-only neural network?
The reason why knowing such lower numbers / lower bounds -- is because they could be important for additional cost/time savings and/or running derived LLM's on local resource-constrained hardware...
Anyway, great post, great idea, and we welcome Gemnini Omni 1.1 Flash to the ever-expanding list of LLM/AI's!
For me it's kind of weird right it's like no one can see the bigger picture. All technology's going to be eaten by AI so companies like Google, Microsoft, Meta ... are middle men in an industry that's disappearing. No one's really going to use a search engine when a chat bot can do a better job. No one is going to pay for software when the AI can do the whole process better at a proper command. All technology and the companies in it are essentially horses and the related roles in a time of cars. That whole slice of the economy is at best going to reduce year on year eventually to disappear. No replacement AI is it.
Is there a good example someone can show of naturally generating real videos using extending and interpolation from the last frames? (For any of the video generators).
I mean I have not tinkered that much, but trying to even get a video to 30 seconds (I just want a cartoon AI avatar to narrate tutorials) has been incredibly difficult. They drift so easily.
Many AI videos you can tell just stitch short clips together. I just want a continuous scene for like 30 seconds to 2 minutes.
145 comments
[ 0.19 ms ] story [ 6.2 ms ] threadAnd jobs where quality is paramount and or there is greater chance of risk / bodily harm
Society seems to be very bad at that, and we end up with high frequency traders making millions while teachers are paid peanuts.
It has nothing to do with skill. Both were very skilled jobs. Draftsman as well despite many people going to it straight from school. But computers and CAD mean that it is now necessary for someone to do a STEM degree to be a draftsman. Recorded audio made many stage musicians redundant. It is cheaper to do it this way and gets superior results, that is all, there is no further agenda.
Now too, the next generation of voice actors and many other knowledge workers will have to go up the value chain one step and operate or potentially build these tools (in whatever form they mature to in a decades time).
The current generation of voice actors will face the same situation as many before in the performance industry - stage musicians/performers for example that were made redundant by recorded audio. The reality is that most of them just left and dispersed into the economy doing completely unrelated jobs.
For software and generally computer engineers, this new primitive happens to itself be software, so it's less of a transition and an easier upskilling path to learn to build it. And building it is one step higher in the value chain than simply using it. That is a structural advantage.
All these demos are focused on how easy it makes everything. Easy is great, but if everyone is able to make instant cute cat videos or whatever it just devalues it. I want to see turning a photo into a rigged 3d model, letting the artist animate and then generate the video. This technology could be used to increase creative expression, but instead it's being used to squeeze out creative expression
There could easily be at least some time period of low skilled ugly people acting in approximate but shitty ways in cheap sets just to give an input reference to a model and then describing the differences in text, yielding gorgeous people doing stuff in fancy locations in the output.
The "slop" phase is early AI, like chunky ugly very early 3D graphics where you can see the triangles.
I'm already seeing images generated by first-generation diffusion models like Stable Diffusion 1.5 (which is small enough to run on a phone) used ironically as meme generators due to the now-retro silliness of what they generate.
one youtuber used irl footage, with hands and stuff - so I know there's human behind the camera
the other was a letsplay that reacted to events just fine emotionally
and yet the uncanny valley of the sound is in full force. Maybe youtube has done something with the codecs?
I feel sad
* https://i.ibb.co/ccqKZ71L/keenlore.png
* https://i.ibb.co/1t3W0JqZ/keenlore-02.png
* https://i.ibb.co/LdBqHwKB/keenlore-03.png
I don't know how the next generation of beloved actors comes about and how we don't descend into a pit of neverending photocopies of things people once loved in the 1990s/2000s.
* $250 per finished hour for the narrator (lowest professional rate).
* 200k to 270k words.
* 22 hours, 20 hours, and 25 hours.
* Books 1 to 3 cost $5,500, $5,000, and $6,250, respectively.
My novel has about 8 major characters, including the 3 narrators, and 25+ minor characters. The price tag is daunting.
To me it's been obvious for a while now: there won't be one
it always sucks at everything else.
Articles and discussions pop up from time here: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
>Firefox makes up about 90 percent of Mozilla’s revenue, according to Muhlheim, the finance chief for the organization’s for-profit arm — which in turn helps fund the nonprofit Mozilla Foundation. About 85 percent of that revenue comes from its deal with Google, he added.
https://www.theverge.com/news/660548/firefox-google-search-r...
Billion with a B, as in 10 zeroes, not 7.
If anything FF gets left out because usage is so low.
I've been following this for a long time. They were leaving Firefox out when its usage wasn't low.
Probably the typical backdoor executive mandate that led to death by "sprint prioritization".
Yes, we will for sure work on the Firefox compatibility bug, Dave-Open-Source-Enthusiast-Google-Dev.
But we can only pick up 10 bugfixing tickets this sprint and this one, as the entire team agrees, is priority #12.
<repeat every sprint, where during the sprint 9-10 new higher priority items magically appear just in time for the next sprint>
Death by slow asphyxiation.
The page instant scrolled suddenly back to previous video examples after I scrolled down. It was disorienting.
The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.
For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.
A lot of browser automation could/is also done by agentic coding.
It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.
You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).
It might sound self-explainatory but turns out that coding can represent a wide range of problems!
I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.
> No one is suddenly missing out on some giant competitive edge because their model is a few months behind
Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.
> things other than how well your model can write code will matter more and more in 12 to 24 months.
What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?
Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.
They never gave an official answer as to why, so I'll let you draw your own conclusions.
They did not decide it wasn't worth spending the money to train.
They absolutely spent the money.
Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?
I thought of a quite a few and they are far more compelling and interesting to me.
(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)
If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.
In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.
They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.
Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.
It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.
In the real world out there, Google and Microsoft are absolutely dominating enterprise customers.
Every single non-tech office worker I know is writing Gemini "gems" (sort of claude prompts/skills) or prompting Copilot to help drafting board meeting notes, insurance contracts updates that reflect changes in regulations, make quick loan feasibility assessments before passing them to the relevant office, presentations, etc, etc.
I'm talking insurance, banking, consultancy, manufacturing, etc, etc.
Why? Because Google and Microsoft already were in these companies, all they had to do is "oh, you also have AI now with your plans". Procurement and data compliance are the first thing businesses have to sort out. They were already sorted out.
Google doesn't need to have the best coding model or triumph in meaningless benchmarks, it only needs their models to get better and cheaper while serving them to their existing customer base.
They are playing a different game.
And Microsoft, doesn't even need to care about models at all, they can provide whatever open or closed AI with their services and have to focus on the harness in Excel or Github/Azure Copilot or whatever.
E.g. while developers in most of my clients use whatever they prefer or the company pays for, the remaining 90% uses either Google or Microsoft products.
Not a single one has incentives into venturing into OpenAI or Anthropic or Z.Ai lands because they might be better at some benchmark that is completely irrelevant to their tasks of updating powerpoints or summarizing incoming emails.
Google is an advertising company with an enterprise SaaS branch. I bet they have chosen to focus on running the most efficient "everybody" model instead of running a heavy model for coders.
When using Gemini for other tasks than coding, it is actually pretty good. It grounds well with Google search and gives mostly correct, well written answers to many niche questions.
Maybe because they see video generation as key to developing "world models"?
And, Veo and omni simply were better than Sora
And, Youtube is huge both as a place where video contents goes and where can be trained from. Microdramas are starting to become a real category--14 Billion USD, 90% of it made with AI.
Chinese video models can be more immediately impressive, but none of them come close to beat the value of Google's Flow. Especially when you are throwing away a lot of generations as part of the creative process. Which is what you have to do to make longer content with any video model.
OpenAI needed to be able to focus. Google can walk and chew gum, and they're not going to run out of money to buy chewing gum.
https://youtu.be/QaiecWzeHFM?si=UV6eF-m514nCOFe0&t=156
Almost all concrete english-language info I can find about microdramas is astroturfed to hell by consultants and "independent" industry publications. Wikipedia's citations for 2025 revenue are 'Reel Reel' and 'Duanju News'. Neither cite their source, though they are likely just regurgitating predictions from Omdia, a media consultancy.
> Real Reel™ works with companies across entertainment, technology and the creator economy to build relevant industry conversations around mobile-first storytelling.
Duanju News is published by 'Studio Phocéen', who run their own microdrama production house.
The Omdia revenue estimates, which is where the projected $14 billion comes from, are completely unsubstantiated as far as I can tell. They even go so far as to make "according to new research" a link that when followed sends you to their generic "/advance-your-business/media-and-entertainment" sales pitch.
https://omdia.tech.informa.com/pr/2025/oct/microdramas-to-ge...
I don't have any special insights here, it all just smells a bit like McKinsey's "The metaverse will be worth $5 trillion by 2030, you better not miss out!!! Hire our 22 year old slide deck experts today"
Previously when making video ads you'd need to actually create the video. Actors, cameramen, editors - you name it. Now a new video ads is just a prompt away, directly inside the ad-spend web UI too no doubt.
People say Google have lost and that they're having their lunch eaten by anthropic, but I am not so sure...
Google is selling shovels, leasing mines, buy stakes in "competitors" and doing it's own exploration/mining. When you look at the full picture, it kinda doesn't even look like Gemini matters that much to them overall.
Gemma 4 I am way more impressed by than whatever the top ranking model is
That's old news. People are saying the chinese models/labs are eating anthropic lunch.
I work in advertising and some days I spent a lot of money using these models. The amount and rapidity of prototyping using them has changed everything about advertising pre production.
x There are so-called “SOTA” or “frontier” models that are more effective than the other ones (independent of harnessing and routing)
x OpenAI and Anthropic have all the SOTA models and lead all the innovation
x Google’s moat is its search bread/butter (it’s the only reason they’re relevant)
All 3 operating assumptions are - I think - false.
What Google has done that the “cuter products” (Claude, ChatGPT) haven’t is connected relatively standard LLMs to an externally valuable live service.
As more companies realize that is where all the value is (the service) and not in the AI capability, then products (and humans) become important again.
Google should just be Google again, and Gemini should be Gemini, off to the side. Omni confuses everyone (and angers some iykyk), they should resolve “AI mode”, rename Gemma? and consolidate the brand overall so it’s clear what Google is.
Google is search.
It helps people on all sides of the market find what they’re looking for.
I don’t really see how repeatedly reinventing and rebranding the same AI chat UX is accomplishing anything toward that goal.
Implicit to that is "find". Their AI integration into search has really hit its stride for me. They have that search box (or speech prompt) hard wired into people and they are finally iterating and crafting AI into that experience. They really failed hard initially.
I know others have worse experiences than me but Google knows a lot about me so maybe that affects my results. YMMV
It does. Very commonly it will get acronyms wrong or assume I mean something else, even when it should be in my “ad profile” or whatever it bases it off.
Often I think the search engine is working perfectly then the LLM is ruining it in delivery.
There are other UIs besides chat
You say "ruining the delivery". I believe I understand. It is not traditional www text search.
I do have a problem with most chat ai's that always ask a question at the end of their answer. This is where configurability is more important to me than absolute performance. e.g. "don't end an answer with a question unless it is directly relevant to the current conversation.
Yes there are many other UIs besides chat. The default chat UI really bothered me at first. Seemed cheesy. However the majority of human interaction is "chat".
Something that is really really cool in this tech is the vision and audio and text models that naturally cluster similar things in massive multi-dimensional arrays. I have a hard time mapping nut~squirrel~food~imageofnut~dogsaspredators~whateves. How does that condense?
We have come so far on the audio transcription dimension. In this discussion context yes it is still chat but we've moved to a different place in the brain. Not radically different (humans without visual sight come to mind in how the process goes from photons to perception (also see hank green's vid on how the eyes are part of the brain)).
Anyway hope this isn't a throwaway account. That username doesn't inspire confidence in that regard.
What the hell am I doing?
[edit: made things more clear about my un-clear thoughts! I did not shift things in the intent space]
When it comes to fish swimming around I don't think I would be able to reliably tell what was real vs generated even with deep inspection.
Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.
I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).
I enjoy making short films with AI. When my latest short screens at a festival in Ocotber, alongside traditional and AI films, hopefully the audience will like it too.
The quick "one shot" video generation might be slop to you, or I. But if someone wants to send it as birthday greeting to their aunt, and they both enjoy it, what business is it of ours?
I have a young boy, and whenever he builds an impressive "scene" from LEGO (like a diorama or whatever), I take a couple of reference pictures with my phone and make it into a "real" movie scene, cartoon, or whatever. He loves it, and this motivates him to build more and bigger things out of LEGO.
If he builds something really special, I might actually fork over the $5 to use Omni to turn his LEGO creation into a 10-second video instead of a still image. It'll blow his mind!
PS: There also are cheap and even free phone apps that make stop-motion animation trivial. We've already made a couple of videos of his toys moving around that way.
I like to generate songs from obscure poems
I have been pondering this and have come out with a single statement. It allows you to create the missing piece in the art you want to create. Assets for games, music for the lyrics you wrote, or just a whole song to justify a crazy dance you want to do. Whenever AI art is discussed, people tend to be so purist about art. One person cannot do everything in a project they take on to express themselves. Usually people come back with then get someone else to do it for you. I think that is an economic argument as people are trying to protect artist pay. I get that. But is it fair to just not let something they want to create because they don't posses every talent needed to do this?
Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.
I agree though. My issue is the cost for using AI video models is way too high for anyone not building anything serious with them, at the same time they are too restricted for actually using professionally. Prompting them with text to get something generated is cute, but then you just end up creating slop that everyone hates, ultimately devaluing the power of these things.
Minimax H3 is about 240Gb alone, how do you do? How much quantised is it, and how good are the results?
Most people are running the stock release of Minimax H3 using the INT8 quant and it's about ~20GB.
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffus...
https://docs.comfy.org/tutorials/video/minimax/minimax-h3
While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.
Upscale when ready: Users on paid tiers can seamlessly upgrade their favorite clips using our new 360p to 720p upscaler.
Generate lightweight previews in 360p resolution up to 60% faster
and at a third of the cost compared to Omni 1.1’s standard 720p resolution. This is helpful for rapid prototyping, storyboard iteration, and quick rendering in developer platforms."
This is a great idea, to have a low-resolution mode for additional speed to create previews, do test runs, create rapid prototypes, etc.
My curiousity is, what's the absolute useable minimum that this could be?
That is, would/could 240p resolution work? If so, what about 144p? How about lower? Then, could those images be upscaled quickly (and is the result still usable?) with a faster image upscaling-only neural network?
The reason why knowing such lower numbers / lower bounds -- is because they could be important for additional cost/time savings and/or running derived LLM's on local resource-constrained hardware...
Anyway, great post, great idea, and we welcome Gemnini Omni 1.1 Flash to the ever-expanding list of LLM/AI's!
Quote under the video of a short Argentinian footballer wearing no 10 with "RESSC" on his back. Can't make it up.
I mean I have not tinkered that much, but trying to even get a video to 30 seconds (I just want a cartoon AI avatar to narrate tutorials) has been incredibly difficult. They drift so easily.
Many AI videos you can tell just stitch short clips together. I just want a continuous scene for like 30 seconds to 2 minutes.
I feel like I should be excited about being able to generate almost perfect videos but, I just don't care anymore.
The amazing videos and photos carried the promise that I could experience that for real. They were aspirational.
Today I suspect every cool shot is made of pixels arranged on a 2D screen by an algorithm. It doesn't do it for me.
That makes me sad...