Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
[0] https://en.wikipedia.org/wiki/Goodhart%27s_law
All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
> An elephant typing on a typewriter
A monkey, surely?
Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others
Feels like google has a different training set than the others?