Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?
Claude was clearly 'pushing back' on the coziness of the cabin. But I think it did best with the cat. Grok, I fear, is making a case for euthanasia. It's suffering and I think it would be cruel to let it continue. Someone pull the plug. .
He mentions they used quite different methods to get to the end-result. I'd love to know how much changing the prompt impacts things: like telling claude to only draw and blend, similar to sol, etc.
There's likely not training data for this and Claude seemed to spend a lot of time blending the background.
As I looked through the images I was unimpressed entirely, at first. But, then I started thinking, these look a little... "childish" to me.
Childish as in... A newish artist who is drawing a concept rather than light / forms (Which is something artists typically do as they understand drawing more and more).
The rose in the vase specifically - some models understood that there was supposed to be shading, reflections, the concept of refraction - others just drew "blue = glass" and "green = stem" and "red = rose".
Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge. I'm expecting these to get better as models improve, and perhaps the artistic progression will be there along with it...
I call this "symbol drawing". Beginning artists do this. They think "This is a head, a head is round. This is where eyes go, eyes are shaped like this", and the whole thing ends up being a collage of symbols vs a representation of the space and experience of viewing a face. When I used to be OK at drawing, it was because I forced myself to use touch instead of sight to compose images. So weird to explain, but I'd feel the 3d to get the lighting and such better.
A lot of art that someone smarter than me told me to appreciate seems to follow the pattern of hitting the space and/or experience while minimizing the use of symbols. Impressionistic paintings esp avoid symbols IMHO, while bizzaro picassos abuse symbols outright and still hit the experience they are going for.
I took the same prompts to Gemini and was stunned by the results. The are completely different from the images shown in the article (and genuinely good art pieces that were generated).
- the best image related stuff I've seen is where the harness is constantly cropping and looking closer at things (likely helps a lot for computer vision in general)
- also it would be interesting because the harness could almost have its own "palette" as if it could play with blank squares and different strokes over lapping or blending before applying to the main canvas
I'm surprised that the harness wasn't better described here. Is the draw tool really just asking LLMs to provide a set of coordinates/parameters for generating shapes? Honestly, if so, the results are shockingly good. Think of how well a typical human artist would do with such a crude method of applying brushstrokes.
GPT 5.6 Sol had the best two drawings (rose and starry nights) but even more impressive was how efficient it was RE cost/time/tokens vs Fable (3.4M vs 14.6M / $7.74 vs $161!). OpenAI has quietly innovated around inference - this is will be a growing differentiator even against open models.
I do think they are going to stretch their lead in value if Anthropic doesn't wake up and stop YOLOing tokens. Kimi is an amazing achievement, but it has the same (or worse) kitchen sink approach as Fable.
At work, even if Fable is technically better I much prefer Sol because it is so much faster and concise.
The Grok ones are amusing, almost comically bad. However, whenever I've tried to pass an image creation request to any of the Opus models, it's been far worse, like first week of using Microsoft Paint bad (while ChatGPT would create social media quality images using the same prompts)
Where do these models even have this data to learn from. There must be massive computer use datasets? I have a hard time believing it’s emergent if it’s able to do something this good.
44 comments
[ 0.23 ms ] story [ 55.3 ms ] threadhttps://www.tryai.dev/?error=server_error&error_code=unexpec...
Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?
There's likely not training data for this and Claude seemed to spend a lot of time blending the background.
Childish as in... A newish artist who is drawing a concept rather than light / forms (Which is something artists typically do as they understand drawing more and more).
The rose in the vase specifically - some models understood that there was supposed to be shading, reflections, the concept of refraction - others just drew "blue = glass" and "green = stem" and "red = rose".
Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge. I'm expecting these to get better as models improve, and perhaps the artistic progression will be there along with it...
A lot of art that someone smarter than me told me to appreciate seems to follow the pattern of hitting the space and/or experience while minimizing the use of symbols. Impressionistic paintings esp avoid symbols IMHO, while bizzaro picassos abuse symbols outright and still hit the experience they are going for.
- the best image related stuff I've seen is where the harness is constantly cropping and looking closer at things (likely helps a lot for computer vision in general) - also it would be interesting because the harness could almost have its own "palette" as if it could play with blank squares and different strokes over lapping or blending before applying to the main canvas
The rectangular smudge tool is a weird tool in the first place, but it's cute to see the models try to use it.
At work, even if Fable is technically better I much prefer Sol because it is so much faster and concise.