3 comments

[ 2.4 ms ] story [ 10.2 ms ] thread
The "photo of a yellow bird and a black motorcyle" example really highlights one of my big frustrations with using models for image generation: the words "photo" and "photograph" don't reliably constrain the results.
If you use some photography specific words, like f stop and exposure settings, and imitate the language photographers use on their sites to describe their photos, you're more likely to get plausible looking photos.
Ambiguous prompts will always have that issue. Did you want an image that contains a photo, or one that is a photo?