91 comments

[ 0.26 ms ] story [ 46.1 ms ] thread
The blog write-up style is so casual haha.
Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
After Z-Image and the original Qwen-Image - I think they've pivoted completely to closed source. From the images I've seen, Qwen-Image 3.0 is just a subpar equivalent to other proprietary models like gpt-image-2 and nb-pro.
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.

Impressive.

Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?

It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
Appears to be closed-weights entirely. No word at all on any weights release.
The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
What value does the 'keywords' meta tag even have these days? 77+ KB of crap added to the page weight. Web development is full of idiots.
Interesting too that they would include "ai friend" when China just added restrictions on AI "partners" such that many services stopped offering them rather than try to adhere to the restrictions.
I think the Qwen team is probably aware that the NSFW community is very quick to adopt any new image gen model (see: Civitai). So, it seems like a good SEO approach to try and surface their model in search results that said community is likely already checking.
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
The moat is the training data. But somehow we've collectively decided that it's not.
The red-dress woman's vestigial pinkie toes …
yellow/red tint is an extremely common problem not matter the photograph source you train on

source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle

You clearly haven't met a Chinese RedNote user.
Wow, it displays Korean properly without breaking. But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
To me it feels so weird that people are trying to push these model for online shopping like "here is how this dress/shirt/pants would look on you". But these models will always make the clothes fit your body and show you in flattering light and so on. How the actual garment fits is still as elusive as before these tools
Maybe this will lead to more business for tailors doing alterations, assuming the clothes people end up buying are expensive enough to justify it.
I have a similar use case at work for previewing construction material and such in our catalogue applied to user uploaded images.

Results are mixed, expensive, but it really feels you're few months off the next improvement to really nail it. It's already good enough.

Wonder what Qwen image will provide over nano banana.

I've noticed something similar in Facebook marketplace ads for used furniture. Most of the images are AI generated to look like a Pottery Barn catalog, then the last image will be the actual item, full of scratches and other damage, sitting in a messy garage.
If it doesn’t fit, then you must have gained weight while the item was in transit
i'd like it to notice things i would miss, like "this is ring-spun shirt, so it will sit like this on your torso" or "these pleats will require you to iron them" etc
The more honest ones are “how it looks on you” and don’t promise fit that depends on so many measurements that aren’t even visible in pics.

Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.

Me, I’m a text-learner so I don’t get it at all. But I know people who get value.

Good or bad, I’m pretty sure showing things in the best light has always been the point of “marketing”. There is a reason ads aren’t filled with ugly people with misshaped bodies, and it isn’t because the intention is to reflect reality, so what you allude to as problematic is what a lot of businesses call a feature.
I use nanobanana 2 to test changes in paint and flooring/tiling to great success. And the real pro move is taking those images to a designer to tweak the remaining 20% or so.
> But these models will always make the clothes fit your body and show you in flattering light and so on

This is something that can be fixed over time. And if this forces clothing manufacturers to stick more to their advertised "specs" (width/length), then it's a win for us.

Is it gonna be less distorted than just seeing the shirt on a model photographed by a professional in the perfect light?
A prompt went viral recently where people were sharing their pictures and asking chatgpt to visualise what their looksmatched partner would look like

The result was always someone extremely good looking

There’s going to be an entirely new class of mental disorders that will emerge from people being deluded by AI

(comment deleted)
How does this fry pan look with a fish in it? Ask Mr Bean.
The short-term goal of a tool like this is to sell products. The more ambitious long-term goal is to shift cultural norms, blurring the lines between advertising and reality until the question you're asking is no longer consciously asked. At least, not by average people, and not at the point of purchase.

I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.

I've been struggling with this question myself. But isn't this a model training/use problem (a.k.a skill issue )? Isn't there a way to make these models be faithful to how people will actually look?
But they dont have to do either of these things. Thats just what people prompt
One joy of online shopping, especially for people doing it in an impulsive and/or addictive way (it's not really rare) is the satisfaction they get from the imagination of having it.

An acquaintance of mine was buying many and not wearing most, as she did not attend that many social occasions. Still, she kept buying.

Eventually, she had to face the actual problem in her life that bothered her. She ended up dealing with it, terribly.

Yet it pushes the goal of the brand/shop forward: make the product more appealing and sell
User-targeted fashion advice is the polite goal. AI designed to generate images of people is actually racing to capture the entertainment markets. They want to be ready to replace models/actors in everything from fashion mags to porn studios. That is where the money is.
It’s clear to us here on HN but to the average person today the way these models actually work is beyond the realm of constraints and reasoning.
This will push us even more to go outside and visit actual shops. Becouse thanks to AI we will have more time? Will we?
Definitely. I think the appeal to some isn't GenAI per se but image-to-image generation like "make this item look like a talented photographer shot it".

But I can't imagine that your conversion wouldn't benefit by showing actual imperfect photos of your products, especially when you run a platform with 1k new listings per month. I'm personally turned off by something that looks like 3D, especially when all the different colour variations look the exact same.

I see the ai product ads and it's pretty funny. The product is shown but is "broken" as far as scale is concerned.

Sort of like how they tried to make gandalf big and the hobbits small in lord of the rings (which wasn't very convincing to me)

If the online retailer has a no quibble returns policy - then at there is at least a strong incentive to minimise returns, rather than oversell.
This looks impressive!
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
It will be interesting to compare this to Flux 2.
I assume Van Gogh didn't paint enough hands to train from!
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.

Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.

Seems that this will not be open weights?
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
I am curious whether the model requires a font to be installed. Does it also generate the glyphs for the text?
It's a proprietary model as a service - it doesn't require anything outside of a browser. But even when they released open-weight (the original Qwen-Image) it's a diffusion model and can't use custom font files.
The "piss filter" is still everywhere.
Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
that's a very interesting point to test new models. I speak an Indian language called "Tamil" and have always tested new models with Tamil but also have tried little bit of Arabic (quranic verses) with previous GPT image 2 and Nanobanana pro and they have nailed it. Don't know if it was because of extensive training data.
I wonder if they got permission to generate that (admittedly impressive) Berserk image.
THIS IS EXACTLY WHAT I WAS LOOKING FOR
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.

But: not open-source/open-weights, and no indication that weights/source will be released either.