It would be interesting to have two generations per model without cherry picking, so that the Elo estimation can include an easy-to-compute standard deviation estimation.
Honestly the first one where I would have guessed "this is a pelican riding a bicycle" if presented with just the image and 0 other context. This and the voxel tower are fairly impressive - we're seeing some semblance of visual / spatial understanding with this model.
Here's Gemini Deep Think when prompted with:
"Create a svg of a pelican riding on a bicycle"
https://www.svgviewer.dev/s/5R5iTexQ
Beat Simon Willison to it :)