38 comments

[ 0.22 ms ] story [ 97.0 ms ] thread
The qwen thinker/speaker architecture is really fascinating and is more in line with how I imagine human multi modality works - IE, a picture of an apple, the text a p p l e and the sound all map to the same concept without going to text first.
The multilingual example in the launch graphic has Qwen3 producing the text:

> "Bonjour, pourriez-vous me dire comment se rendreà la place Tian'anmen?"

translation: "Hello, could you tell me how to get to Tiananmen Square?"

a bold choice!

Neat. I threw a couple simple audio clips at it and it was able to at least recognize the instrumentation (piano, drums, etc). I haven't seen a lot of multimodal LLM focus around recognizing audio outside of speech, so I'd love to see a deep dive of what the SOTA is.
Next steps for AI in general:

  - additional modalities
  - Faster FPS (inferences per second)
  - Reaction time tuning (latency vs quality tradeoff) for visual and audio inputs/outputs
  - built-in planning modules in the architecture (think premotor frontal lobe)
  - time awareness during inference (towards an always inferring / always learning architecture)
You can try it out on https://chat.qwen.ai/ - sign in with Google or GitHub (signed out users can't use the voice mode) and then click on the voice icon.

It has an entertaining selection of different voices, including:

*Dylan* - A teenager who grew up in Beijing's hutongs

*Peter* - Tianjin crosstalk, professionally supporting others

*Cherry* - A sunny, positive, friendly, and natural young lady

*Ethan* - A sunny, warm, energetic, and vigorous boy

*Eric* - A Sichuan Chengdu man who stands out from the crowd

*Jada* - The fiery older sister from Shanghai

Many of these voices are especially hillarious when you switch the language.

In Russian, Ryan sounds like a westerner who started reading Russian words a month ago.

Dylan sounds somewhat authentic, while everyone else is a different degree of heavy-asian-accented Russian.

Speech input + speech output is a big deal. In theory you can talk to it using voice, and it can respond in your language, or translate for someone else, without intermediary technologies. Right now you need wakeword, speech to text, and then text to speech, in addition to your core LLM. A couple can input speech, or output speech, but not both. It looks like they have at least 3 variants in the ~32b range.

Depending on the architecture this is something you could feasibly have in your house in a couple of years or in an expensive "ai toaster"

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more.

The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. What a turn of events!

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data.

I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service.

However, most people don’t care. Paying $20 per month for a ChatGPT subscription is a bargain and they automatically get access to new versions as they come.

I think the at-home self hosting hobby is interesting, but it’s never going to be a mainstream thing.

sitting here in the US, reading that China is strongly urging the adoption of Linux and pushing for open CPU architectures like RISC-V and also self-hosted open models

are we the baddies??

(comment deleted)
>Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster

So did you run the model offline on your own computer and get realtime audio?

Can you tell me the GPU or specifications you used?

I inquired with ChatGPT:

https://chatgpt.com/share/68d23c2c-2928-800b-bdde-040d8cb40b...

It seems it needs around a $2,500 GPU, do you have one?

I tried Qwen online via its website interface a few months ago, and found it to be very good.

I've run some offline models including Deepseek-R1 70B on CPU (pretty slow, my server has 128 GB of RAM but no GPU) and I'm looking into what kind of setup I would need to run an offline model on GPU myself.

Is there a AI market for open weights? Companies like Alibaba, Tencent, Meta or Microsoft makes a lot sense. They can build on open weights, and not losing values, potentially beneficial for share prices. The only winner is application and cloud providers, I don't see how they can make money from the weights itself to be honest.
The US is probably ahead but they're so obsessed with moats, IP and safety that their lagginess is self imposed.

China has nothing to lose and everything to gain by releasing stuff openly.

Once China figures put how to make high performance FPGA chips really cheap, its game over for the US. The only power the US has is over GPU supply...and even then its pretty weak.

Not to mention NVIDIA crippling its own country with low VRAM cards. China is taking older cards, stripping the RAM and upgrading other older cards.

will be interesting how it will compare with pricing with audio modality comparing to gemini 2.0 flash once many providers offer it.

Even though gemini 2.0 flash is quite old I still like it. Very cheap (each second of audio is just 32 tokens), support even more languages, non-reasoning so very fast, big rate limits.

I recently needed to scan hundreds of low quality invoices and run them through OCR for invoice numbers and dates. I really took for granted how seamless this is in some applications, and was shocked how much work went into producing decent results.

I was obviously really naive. Either way, it gets me excited any time I see progress with OCR. I should give this a try against my (small) dataset.

I just ran Qwen against some of invoices that my gnarly algorithm really struggled with (with openai fallback) and Qwen was able to extract all relevant data without any issues. I'm pretty damn impressed to be honest.
Any insights into what "native video support" actually means? Is it just good at interpreting consecutive full frame images taken at intervals (thus missing out on fast events) or is there something more elaborate to it?
I'm running Q3-Next on my MBP and seeing ~ GPT4.1 performance from it.

Impressive what these local models are now capable of.

The real point of leverage for here is performance/size. Getting traction in the open weights space kinda forces that the models need to innovate on efficiency. This means the open weight models may get leverage that the closed weight ones don't think about.

If we had some aggregated cluster reasoning mechanisms, When would 8x 30B models running on an h100 server out perform in terms of accuracy 1 240B model on the same server.

I usually ask these models to tell me a short story, and most times the prose is stiff and the story reads like a mass market straight to KDP kids book. But wow, first shot generated something, light, mildly funny, and chill. Quite a surprise.

Pasted here for your own judgement:

*Title: The Last Lightbulb*

The power had been out for three days. Rain drummed against the windows of the old cabin, and the only light came from a flickering candle on the kitchen table.

Maggie, wrapped in a wool blanket, squinted at the last working flashlight. “We’ve got one bulb left, Jack. One.”

Jack, hunched over a board game he’d dug out of the closet, didn’t look up. “Then don’t turn it on unless you’re reading Shakespeare or delivering a baby.”

She rolled her eyes. “I need it to find the can opener. I’m not eating cold beans with my fingers again.”

Jack finally glanced up, grinning. “You did that yesterday and called it ‘rustic dining.’”

“Desperate times,” she muttered, clicking the flashlight on. The beam cut through the gloom—and immediately began to dim.

“No—!” Jack lunged, but too late. The light sputtered… then died.

Silence. Then Maggie sighed. “Well. There goes civilization.”

Jack leaned back, chuckling. “Guess we’re officially cavemen now.”

“Cavewoman,” she corrected, fumbling in the dark. “And I’m going to bed. Wake me when the grid remembers we exist.”

As she shuffled off, Jack called after her, “Hey—if you find the can opener in the dark, you’re officially magic.”

A pause. Then, from down the hall: “I found socks that match. That’s basically witchcraft.”

Jack smiled into the dark. “Goodnight, witch.”

“Goodnight, caveman.”

Outside, the rain kept falling. Inside, the dark didn’t feel so heavy anymore.

— The End —

Does anyone have good resources for learning about multimodal models? I'm not sure where to begin.
All, what's the best model right now to bring a photo to life (create a short video from a photo etc) ?
[flagged]