It’s frankly embarrassing at this point. I’ve got free access through buying a Pixel phone and it’s not even worth using as it’s a waste of my time. Here’s my experience so far using it for basic sysadmin Linux type stuff.
Gemini 3.1 Pro just feels a generation behind, from when models would miss easy things and make bad assumptions. Its not actively detrimental in bad way but the opportunity cost vs using something like Opus to be productive is large.
Gemini 3.5 Flash is the most annoying model I have ever used. It loves to respond in ALL CAPS like “LOOK AT THAT” for no apparent reason. I realize it’s a flash model but I will give it a basic list of tasks and the output will simply vomit “Now I will X” “Now I will Y” “Now I will Z” over and over again filling my screen with garbage. It’s also no smarter than 3.1 Pro and consumes just as many tokens as 3.1 Pro, it’s really pointless without a newer Pro model in place.
I assumed google would lean into the efficiency stuff more and try to eat the easy 80% of workloads, winning market share on volume instead of frontier if they were not able to produce frontier level models.
They're very well equipped to be the volume discount store of inference.
I don't like how it compared itself to 5.6 Luna instead of Sol, and still losing on some metrics. You will not see me using this over even 5.6 Sol-light
This model is not for builders and engineers. DeepSWE score of 49% is behind gpt 5.4 and muse spark. It's clearly intended to be an efficient model for google gemini usage.
What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models.
Presumably 3.6 Flash is primarily meant to serve their own needs for the Gemini chat app, voice app (which Sergey Brin says he uses a lot in the car on the way to work), and for their search "AI Assistant".
Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.
I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.
I think there is another angle here - maybe, just maybe google doesn't want to release a pro/stronger model sooner because of two reasons- 1) they're afraid they'll have to get their hands dirty in a certain war?? 2) what's the benefit of being on top of this chart?
so I think they're just focussed on releasing whatever helps their bottomline (improving search?). its not like anthropic and openai are making a lot of money being on the top of charts! this might just be my crazy pills talking though.
I'm not at all an industry pundit. But I suspect there's a reason we're not seeing leading models from Google recently.
Judging from my own frustrating attempts to use Gemini for vibe-coding, it seems like Google is badly over-sold (i.e., under-provisioned).
From all those promos giving away their pro-level subscription with phones; spinning up a mid-level subscription to undercut other providers and (probably most significantly) putting AI queries into ever search response because their flagship search product had become useless; they're promising a lot more processing to customers than they can reliably deliver.
The recent iterations seem to be intended not to push the capabilities forward, but to deliver capabilities at the current level while consuming less resources. That will allow them to maintain their trajectory until (I'm expecting) they get the huge infusion of extra compute resources from Space X later this year.
If I'm right, then I expect we should see Google start pushing forward again (rather than more of this lateral stuff) by the end of the year.
> Judging from my own frustrating attempts to use Gemini for vibe-coding, it seems like Google is badly over-sold (i.e., under-provisioned).
Based on my experience lately with Codex, it seems the opposite. Gemini 3.6 Flash (high) in Antigravity CLI feels a LOT faster than GPT 5.6 Luna (high) in codex. In the order of 10-50x.
I would be fine lauding Gemini models if the only benefit of them was superior understanding of intent (read between the lines). I don't need it to code because other models are tuned for that explicitly, but I would like a model that is tuned to produce less mechanical output.
I know that Flash is obviously no Opus or Fable. That being said, when I want a very quick response, my go-to is Gemini Flash. It's just really damn fast and generally tends to be accurate enough.
27 comments
[ 0.23 ms ] story [ 24.7 ms ] threadhttps://blog.google/innovation-and-ai/models-and-research/ge...
Gemini 3.1 Pro just feels a generation behind, from when models would miss easy things and make bad assumptions. Its not actively detrimental in bad way but the opportunity cost vs using something like Opus to be productive is large.
Gemini 3.5 Flash is the most annoying model I have ever used. It loves to respond in ALL CAPS like “LOOK AT THAT” for no apparent reason. I realize it’s a flash model but I will give it a basic list of tasks and the output will simply vomit “Now I will X” “Now I will Y” “Now I will Z” over and over again filling my screen with garbage. It’s also no smarter than 3.1 Pro and consumes just as many tokens as 3.1 Pro, it’s really pointless without a newer Pro model in place.
releasing a model worse than luna is pretty bad. Its clear that internally they did not decide coding was a thing until relatively recently.
They're very well equipped to be the volume discount store of inference.
What is interesting is how this is announced before any Gemini Pro progress. From the outside it seems as though Google cannot keep up with other frontier models.
Flash 3.6 is certainly capable of many coding tasks, but clearly a model of this size of not trying to compete at the frontier as a software development tool, and not clear why Google really need to complete there other than for PR-related AI bragging rights.
I don't know how a frontier model like GPT 5.6 or Fable could have done better (I have no need that justifies paying for them), but yesterday I used the free Gemini chat app (i.e. Flash 3.6) to discuss and explain this poorly written recent AI paper to me, and honestly couldn't ask for much more.
https://alignment.openai.com/measuring-reward-seeking/
so I think they're just focussed on releasing whatever helps their bottomline (improving search?). its not like anthropic and openai are making a lot of money being on the top of charts! this might just be my crazy pills talking though.
Judging from my own frustrating attempts to use Gemini for vibe-coding, it seems like Google is badly over-sold (i.e., under-provisioned).
From all those promos giving away their pro-level subscription with phones; spinning up a mid-level subscription to undercut other providers and (probably most significantly) putting AI queries into ever search response because their flagship search product had become useless; they're promising a lot more processing to customers than they can reliably deliver.
The recent iterations seem to be intended not to push the capabilities forward, but to deliver capabilities at the current level while consuming less resources. That will allow them to maintain their trajectory until (I'm expecting) they get the huge infusion of extra compute resources from Space X later this year.
If I'm right, then I expect we should see Google start pushing forward again (rather than more of this lateral stuff) by the end of the year.
Based on my experience lately with Codex, it seems the opposite. Gemini 3.6 Flash (high) in Antigravity CLI feels a LOT faster than GPT 5.6 Luna (high) in codex. In the order of 10-50x.