I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense.
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/
also my favorite SVG was def the google/gemini-3.6-flash
Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code.
Arguably, a royal portrait is a misinterpretation of the prompt, since it's just asking for a specific facial feature. But I guess you could look at it as a bit of artistic license.
My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.
How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.
For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
Mine is any variations on mammoths in various situations, or anthropomorphic. Since mammoths are invariably majestically going from one place to another in any of the books, models have hard time imagining anything but that.
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
For those who don’t know a Habsburg jaw also known as mandibular prognathism, it is a genetic condition characterized by a protruding lower jaw, which was notably prevalent among members of the Habsburg royal family due to their history of inbreeding. This condition often resulted in significant facial deformities and difficulties with eating and speaking.
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
That looks like something from Machinarium or Robots :)
https://playcode.io/blog/macbook-svg-benchmark
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
raninemandibularprognathism-maxxed?
That's a pretty good benchmark
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
Would've wanted to see also DS4 flash.
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.