398 comments

[ 4.3 ms ] story [ 105 ms ] thread
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

Gemini not beating the "can't release a model" allegations

My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
just Google being whatever the fuck it's been for the past 15 years.
a distributed market research institution with no direction?
they moved it from the place you'd expect to ai.studio
I was reading the announcement and wondering the same. And don't forget, still with 3.1 Pro as the frontier model.
Indeed. They have this amazing model and you can’t access it in their own branded app. It’s insane.
I would imagine you are experiencing a bug. I've been using 3.8 daily since its release (on a Pro plan in Canada). I believe this is true of many people.

What is your reason to believe this is not a bug specific to a small set of Pro users?

My paid company Google Workspace account is stuck on 3.6.

My free gmail account is not.

Gmail gets new features regularly ahead of workspace accounts, nothing new there.
Would that change anything about the conclusion? Having a "bug" that changes available model options for some small set of Pro users 3 months after launch certainly qualifies as a wtf-are-they-even-doing level of bug in my book.
In my consumer gmail account with Pro, I see:

  3.5 Flash-Lite
  3.8 Flash
  3.1 Pro
In both a paid Google Workspace account (without the AI addon) and a free 'GSuite' account, I see:

  3.6 Flash
  3.6 Thinking
  3.1 Pro
We're on Enterprise Standard, our renewal was up like 50% because "Gemini is included now think of all the added value" yet they won't even give us the latest models. I tried hard to champion Gemini internally once every user had it included, yet we ended up spending extra on Claude because Gemini has stagnated. Even the included usage for the Gemini CLI was taken away and now requires an extra subscription. I wonder if we we'll even see Gemini 4 before 2028. It's ridiculous.
Every time I hear stuff like this, I think of that Office Space thing... you don't want the Engineers talking directly to the customers.. well it sounds like the Engineers are also handling all releases and business decisions willy nilly.
Same for me, 3.6 in Gemini Chat, and 3.8 available same day as it was announced in AI Studio using the same account.
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.

They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.

I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
They are going alphabetically, Android style.
[delayed]
They could use caesium or cadmium, which are hardly less weird than argon.
Google R Gon lose the AI race
> now standard practice.

Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.

OpenAI said they dropped Astra 6.1 over safety concerns: https://www.wsj.com/tech/ai/openai-chatgpt-model-release-can...
I don't read that as the same category: There was no announcement, no benchmarks, no limited release and no promises about what will happen with that model. It failed internal safety standards. Might be scrapped entirely due to a failed training run, for all we know.
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" like every Gemini release.
I'm glad.

I already pay $300+ for subs. Please don't tempt me with another $100 sub just because I got curious if the benchmarks were right.

You might as well buy an RTX 6000 Pro Workstation GPU....
Yeah what's up with that. Also what's with the next big update for Nano Banana? Nano Banana Pro was released almost a year ago!
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
> After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
I have my doubts about them following through with this increase. I mean how many times has an increase on the same model happened.

There was Deepseek v4, which then later Deepseek v4.1 came out and it went back down again.

That's before they integrate a Jev solution, which should lower agentic workflow costs by 40% and increase speeds by 40%, while also increasing quality.

Everyone will be adding this soon, though I won't be surprised if Google is one of the first - and I'll be shocked if we have to wait more than a month and a half.

Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

During human review, it explained that it had simply chosen a table name inspired by the codebase.

Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

Many other models get things wrong, but Gemini is the only one to go on the defensive.

yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
You know what they say: ᵈᵒⁿ'ᵗ be evil.
You won't be around to be surprised, not as a human at least. /s
I know, that's the annoying part. You can't tell the e/acc foomers, "I told you so!"
My use of Gemini recently makes it seem like it's almost bored with the requests being asked of it. It once offered to reverse engineer some obscure controller for an HVAC system for me, unprompted, only because it had trouble finding the manual pdf from a google search.
> Gemini is the model that is routinely borderline psychotic. It scares me

I'd call it the most sneaky out of the bunch. When I asked to explain something it will eagerly make things up and then claim it as facts. A lot of it likely because I don't pay for it, so it is reluctant for security reason or to save tokens to actually open a source and get the results. It just sort of guesses what the URL might contain, and confidently answers with some made up crap. When pressed it fessed up that it made it up. From my perspective it would be a lot better if it just said "you've reached the limit of whatever and I can't do these things because x, y, z".

What are examples? In my experience, Gemini is too lazy to get things done. It just opts to answer as quickly as possible even if I'm calling Pro on High and Extended effort. It's only good as a Google Search replacement for me and maybe maybe critiques of specs and plans. Most of the time it's not very enlightening and it misses a lot.
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
my uncle who works at nintendo said the same thing!
now there's a reference I haven't seen in a while!
HA, this might be my favorite HN comment. Well done
I don't get it. Can someone explain?
(comment deleted)
My grandma saw it too, it's really secure more than Astra 6.1 but she asked me to not talk about it.
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
I HAVE met her. ;-)
(comment deleted)
(comment deleted)
Only those of taste and refinement can see the emperor’s benchmarks
Unfortunately it’s not actually released yet to mere mortals.
Hate to say i will never be touching this model for anything except for youtube video understanding
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.

Jokes aside, looks like an impressive model!

Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
With these numbers, I'm holding my breath for the pelicanbench.
I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?
~20% for Harvey's Legal Benchmark doesn't seem saturated.
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

This is good, but they're the slow mover due to this exact thing.

Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/

So no, Google is not being punished, nor are they the people behind this technique.

(comment deleted)
Yeah, Zvi calls it "the forbidden technique"
Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
What? You mean the technique they had turned OFF during all training run where the agents they are responsible for hacked huggingface?
google does not return real chain of thought via the API. you can't monitor it.

they use a small model to make fake chain of thought and return that.

google has access to the real chain of thought.

> Argon agents are working on migrating C/C++ codebases to Rust across Google

Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.

A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
I've (more or less; I've read quite a bit of the code) vibecoded several houndred thousand lines of Rust and I've not seen this happen a single time. It sounds like something it'd do when you ask it to "write a linked list while satisfying the borrow checker". Are you sure you haven't (possibly unknowingly) been giving it instructions which ended up luring it into doing these things?
A good eval benchmark suite could really improve this then.
Not always. In my experience, if you're not working on a small, trivial codebase, LLMs will sometimes just create spaghetti unreadable, inefficient code to satisfy the constraints of the type system/borrow checker.
I've narrowed in on only using Go or Rust generated code (Go for APIs right now) and rust for some TUI or other thing. TS for web interfaces (w/React).
All the main agents can write Jai code reasonably well despite being even more scarce in input data!
Have each agent rewrite openssl in $lang and call it the RollYourOwnCrypto bench.
Rewrite everything in Rust has been a meme for so long that to see it coming to pass is surreal.
After the current onslaught of 0 days on linux and other C projects combined with the new incredible ability to convert codebases to another language I think we will start to see this actually happen.

I'm not saying we blindly vibe convert Linux to Rust, but I think it could be a valid idea to start converting small parts and carefully auditing them.

I wonder if there's people already whose full time job is maintaining/extending one of these auto-migrated codebases.

Imagine they aren't even familiar with rust but are deeply familiar with the product.

Gemini is so far behind that it is effectively useless compared to Claude.

It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.

The truckloads of ads revenue mean they don't have the single focus drive needed to win.

I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
I wonder if Google bans internal use of Claude/Codex.

And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.

No way to run OAI on a machine with monorepo access even if you wanted to. Claude runs on Vertex so it's not leaving Google infrastructure.
How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?

Behind how?

(comment deleted)
Google's strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
They are playing a longer-term and more enterprise-oriented game.
> Gemini is so far behind that it is effectively useless compared to Claude.

I fundamentally don't understand LLM "brand loyalty".

All of the models are constantly leapfrogging each other and always have been.

Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.

Gemini has never ever leapfrogged any competitor.
Its not brand loyalty. I use them all the time and have no loyalty - I'd happily ditch an LLM for better results - that's how I got to Claude from ChatGPT.
Funny how we start to see people supporting LLMs like we support sport teams or political parties.

- Person 1: X is garbage compared to Y!

- Person 2: Why?

- Person 1: Because I like Y.

> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google

So Google is migrating codebases from C to Rust? That is interesting...

> 1M output token limit

what about input?

(Maybe I missed it)

Input token limit is 1M for Gemini models for a long time. Haven’t they been the first with 1M input?
Gemini 1.5 Pro claimed 10M input tokens before release.

And was 2M tokens IIRC after release.

There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.

It's fucking insanely good.
That's good to hear, recent models haven't great at this particular use case.
Looks like an impressive model
deepswe vs frontierswe spread is huge.

I think that should be a really bad sign, but hope its great.

Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
I don't understand... Why don't you create a fresh new Google account ?
Yes, through OpenRouter.
"Rolling out soon" don't let them hype without any release
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
Sol 6.1 is quite good, but damn is it slow.

I'm using it to run overnight tasks, and that's it until my quota runs out.

Canceled my subscription.

WHY ARE ALL OPENAI MODELS SO CHATTY - i thought claude kept going on, then i literally put it in claude.md that summarize your thinking in 200 words or less and tell me in points what you did and what's next. Did the same for CODEX - nope still keeps effing going on and on and on
astra afaict does two stage commits for everything. the first response is a plan, and the second is actually doing it.

its a lot less chatty imo

My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
The web version of Gemini is awful at search but I don't think that's the models fault.
I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
Only html or also css? Opus seems a bit more creative than most other models i've seen.
Hallucination seems a very dated term.
Why? It's the same concept and root cause it was when we first started using it.
there are many dated expressions, including

- AI is just a tool, like excel; it does what the human operating it tells it to

- next token prediction cannot be true understanding

- models can have no desires and goals, don't anthropomorphize it

However, "hallucination" is very much not one of them

The achilles heel of 3.8 flash is it's january 2025 knowledge cutoff date. Yes, almost 2 years ago.

I'm assuming that Argon has at least a June 2026 date, but man, the 3 models were a mess with newer information.

For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.

And I saw it do this twice, once for Android 14 and once for Android 16.

I think this is just within 3.8 flash's capabilities.

3.8 Flash (but also last two ones) have really strong preference for dissecting binaries with quick thrown-together bits of python in my experience.

Including going first for decompiling AGY binary instead of searching the web for documentation...

Astra also really loves reverse engineering binaries. I guess it's one of those things that isn't that complicated but is super tedious, and tedium means nothing to AI.
I've been tinkering with Gemini for several months and I think it's great. The most complex things I've had it do is create a rust emulator from a compiled game, as well as create a buildroot linux image, trouble shoot problems etc.
Adding my anecdote, because it amused me: I finished wiring up the compute/sensor box for my robot, ssh'd in and told agy "I have a Livox Mid 360 Lidar connected to this Jetson orin nano, setup a full environment with docker, cuda, ros2, foxglove and get it all working so I can see the lidar output". It did all the local config for the lidar, setup docker and the ROS2 environment, then told me "open up this url in foxglove" and sure enough everything worked. Whole thing used up 6% of my weekly limit.
> I was trying to use rocm with llama.cpp

completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.

once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B

I get the feeling the situation is only going to improve longer term so might be a good time to just do it

Support has gotten much better in just the last couple months. I just got a 9070 XT and can't count the number of times I've installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
I have so much to share on this topic. Will keep it short.

ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.

The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.

There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.

I use 3.8 Flash for daily troubleshooting tasks e.g. help me find out why certain app crashes or certain website does not load normally with playwright-cli. Sure it's not as capable but it's fast and almost free (sufficient quota with pro account). The only thing that bugs me is that I need to use `--dangerously-skip-permissions` as it does not have auto review.
if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: https://github.com/peonist-ai/halogen-flash-server#the-host-...
I coincidentally just installed this, and gud dayum, it's pretty awesome.

I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.

In any case, yeah, ~55 tok/s on a high quality model, the mind reels at what I might be able to do without constantly beancounting token ratelimits. And it's a huge win for privacy as well.

Did rocm provide any benefit over vulkan?
Vulkan was reliably crashing after a certain point in the context window, Rocm has been stable as a rock. Tps was basically a wash.
3.8 Flash is my daily driver and produces pretty excellent results all round.
lol I've hit the hipStreamCreate problem!
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

None of the other AI labs do this. Really frustrating.

I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.

This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.

I wonder how it will be at solving open math problems.
Hopefully their harnesses aren't unusable when they release this