34 comments

[ 4.5 ms ] story [ 53.9 ms ] thread
(comment deleted)
I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap.

Not really at gpt 5.5 tier though, and probably below glm 5.2...

But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model.

Edited: gpt-5.4-mini not the base gpt-5.4

This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
It's a very good model for this size and price. I tried it with a couple of small tasks - just an year ago this would be the level of the leading models.
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization.

DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.

I don’t like DS4 in my experiences with it I still prefer qwen locally and glm on api
I was playing with Hy3 via openrouter yesterday (and I've also been using DS4 Flash/Pro as a daily driver since I cancelled my Anthropic sub a week ago).

I've found DS4 Flash to be very temperental (via Claude Code). The speed is great, but it often builds a completely wrong mental model and charges off down the wrong path. I find myself needing to rein it in regularly (and also compact the history, which undercuts the whole cache price advantage).

Hy3 isn't as fast, but so far it seems to stay on track much more reliably than DS4 Flash. It also doesn't seem to degrade as much with longer context. I'm not sure what the real pricing is, but I feel like it's a very competitive model.

As an aside, I also nabbed a 50m token pack for LongCat 2.0 to give it a whirl. Not free, but it's so cheap they're basically giving it away. Very impressed too - seems roughly on par with Hy3. Not frontier-level intelligence, but a dependable workhorse that can navigate a codebase well and can reliably execute what you tell it to do.

Been using this and GLM 5.2 back and forth. I like the speed of Hy3. Also seems very happy to follow instructions. Still haven’t found any open models that follow instructions as good as Mimo v2 pro though
Quite interesting to see them and Meta and others release before OpenAI supposedly is to release GPT 5.6 today, would it be better to release it before or after? Calm before the storm type of thing?
A month ago I wrote a blog post about how Hy3 was topping the OpenRouter rankings despite no one talking about it: https://news.ycombinator.com/item?id=48317294

As of today, it has fallen to 8/9th on the rankings. I don't see a reason where you would use this model over competitors. However, price economics are bit confusing, as currently the effective input price of Hy3 via OpenRouter is now the same as DeepSeek-hosted DeepSeek Flash V4.

https://openrouter.ai/tencent/hy3-preview

https://openrouter.ai/deepseek/deepseek-v4-flash

Because it was free with generous limits and high availability, until it wasn't.
That UI demo page is… really quite janky.
Pelican from a few days ago: https://simonwillison.net/2026/Jul/6/hy3/ - I was using the free tier on OpenRouter, which expires on July 21st.

I tried the preview model 41 days ago and got a pelican with a "change pelican color" button: https://static.simonwillison.net/static/2026/hy3-preview-pel...

Curious why TFA calls out "Tencent in China".

  tencent/Hy3. New Apache 2.0 licensed model from Tencent in China
Is there a Tencent AI lab elsewhere (MiniMax have some association with Tencent, for example)?
I quite enjoy that the "animate wheels" button animates the sun instead.
I have been overly critical and arguing in bad faith about your writing in the past. As well as negative towards you, which in turn was breeding a bad environment. While I dont really enjoy LLMs, you did help me realize my unreasonable feelings as well as realize the occupation (and the joys I got from it) is essentially dead from it’s previous iteration and that I should let go and just join in the “I’m doing it for the money and attention” crowd. I will still just hand code my own projects and not use LLMs when I can. I think it’s cool you started the pelican meme however useful it really is even if only aesthetically.
I feel like I'm taking crazy pills with hy3, it's either benchmaxxed to hell and back or skill issue on my part but I'd rather use dense gemma. I don't think there's a single model that's wasted more of my time in recent memory.
Gemma 4 31B is underrated. It surprises me a lot.
What we really need is a breakthrough in inference or LLM architecture to allow running GLM-5.2-level models at the size of Qwen 3.6 27b or smaller on consumer devices like a 48GB Macbook Pro, and at least at 100 tokens/second. My hypothesis is that a smaller, less capable but faster model paired with a good harness can run for longer and brute force its way out to solve problems that the bigger models can one-shot.
I would never use any product that can't explain on its own front page what it is and why I should use it.
Congrats, the trial chat is QR locked. The AI companies are spoiled and get crazier every day (not only this one, US companies as well).
Got really excited for a minute that the long-standing [Hy](https://hylang.org) project had had a release, but it's just some confusingly-named LLM. Shame.
The strange names are mainly initials from Chinese pinyin. The first generation of Tencent Hunyuan was released in 23Q3, so it is already quite a veteran.
I'm sorry but what on earth is going on with that bar chart, the bars are not consistent. E.g., in the frontierscience-olympiad chart Hy3 preview scores the same as DeepSeek (70.0) but Hy3 preview's bar is visibly lower.
This is admittedly an aside from the content of the post itself, but... why do so many mobile sites insist on preventing zooming in and seem to share the same incredibly buggy image zooming? It's quite frustrating.
There are now 3 tiers of competition:

-Fable + Gpt 5.6 Sol

-Opus + Gpt 5.6 Terra + Grok 4.5 + Muse Spark 1.1

-Open Chinese models: GLM + et family

The economics is on the Fable tier people are willing to spend a lot on it and on the Open tier you have to give it away to drive usage. The bottom tiers are also getting more and more competitive.

Worse than GLM-5.2, much more expensive than deepseek v4. Tough bargain.
Interesting way to show off a model last on every benchmark. Not sure any other lab is doing this