What happened
The event date for this report is 2026-07-22. I generated 1,000+ SVGs across 7 frontier models to test whether AI labs are training on Simon Willison’s pelican-riding-a-bicycle benchmark. The account is based on one supplied report and should be read as a record of conditions known at that point, rather than as a statement about later developments. For the past few years, Simon Willison has tested every major LLM release with the same prompt: “Generate an SVG of a pelican riding a bicycle”.
The available figures and chronology
What began as a tongue-in-cheek benchmark has become one of the most famous informal benchmarks in AI. Simon’s pelican-on-a-bicycle results are often among the most upvoted comments on Hacker News threads announcing new releases from AI labs. These details help define the scale and immediate setting of the event without extending the evidence beyond what the supplied source states. Where a number is an estimate, provisional count, forecast or company claim, it should not be treated as a final independently audited total. The evidence packet does not support assumptions about motives, responsibility or longer-term outcomes beyond the facts described here.
The operational context
The benchmark was famous enough that there’s plenty of discussion about its usefulness and about whether AI labs might be benchmaxxing1 on it. When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn’t it be tempting to pelicanmaxx your model just a bit? These details help define the scale and immediate setting of the event without extending the evidence beyond what the supplied source states. Where a number is an estimate, provisional count, forecast or company claim, it should not be treated as a final independently audited total. The evidence packet does not support assumptions about motives, responsibility or longer-term outcomes beyond the facts described here.
What remained uncertain
I wanted to find out, so I put together a small experiment. I generated 1,008 SVGs across seven frontier models, scored them with an LLM judge, and used Claude Fable 5 for the analysis. These details help define the scale and immediate setting of the event without extending the evidence beyond what the supplied source states. Where a number is an estimate, provisional count, forecast or company claim, it should not be treated as a final independently audited total. The evidence packet does not support assumptions about motives, responsibility or longer-term outcomes beyond the facts described here.
The record at the time
Taken together, the supplied material supports the main development in the headline and the specific details above. It does not establish what happened after 2026-07-22, and no later outcome is inferred here. The report therefore preserves the event-date framing, attributes the account to the available evidence and avoids turning an early figure or stated plan into a settled result. Further official data, technical findings or independent reporting could refine the picture, but those materials were not part of this packet.


