Simon Willison used an invited keynote at the AI Engineer World’s Fair in San Francisco to argue that the LLM field has moved so quickly that even a year is too much time to cover cleanly. The talk, titled The last six months in LLMs, illustrated by pelicans on bicycles, was his third appearance at the event. He said he had originally pitched a year-long survey, but the pace of releases forced him to narrow the scope to half that window.

Willison’s framing device was intentionally odd. He has been asking models to generate an SVG of a pelican riding a bicycle, a benchmark that is playful but also revealing. His point is that text models should not be able to draw anything at all, yet they can produce code, and SVG is code. He uses the test to see whether models can maintain structure, follow instructions and produce output that is close enough to the target shape to be useful. The comments in the generated SVGs also offer a glimpse of the model’s reasoning.

The talk then walked through a rapid sequence of releases. Willison highlighted Amazon’s Nova models, Meta’s Llama 3.3 70B, DeepSeek’s v3 and R1 releases, and Mistral Small 3. He said the field’s pace had made local models worth attention again, especially as smaller systems started to approach the performance of much larger ones. DeepSeek’s R1, he noted, drew unusually broad attention, while Mistral Small 3 showed that a 24B model could be practical on a laptop without consuming every last resource.

He also argued that February and March changed expectations again. Claude 3.7 Sonnet became a favorite for many users, and it was the first Anthropic model to add reasoning. OpenAI’s GPT-4.5, by contrast, was presented as expensive and underwhelming, while o1-pro was even costlier. Google’s Gemini 2.5 Pro landed as another strong option. The effect of the sequence, in Willison’s telling, was to show that model quality, price and local usability were all moving in different directions at once.

The keynote also returned to OpenAI’s image generation launch for GPT-4o, which Willison said drew 100 million new accounts in a week and set off a burst of virality. He used that moment to raise a more cautious note about memory and context control. As a power user, he said, he wants to know exactly what goes into a prompt and what the model can see from earlier conversations. His overall message was not that one model had won. It was that the ecosystem has become too fast, too varied and too practical for old benchmark habits to capture on their own.

The keynote’s larger point is that benchmarks alone no longer define the market. Willison treated cost, local execution, latency and developer ergonomics as first-class issues, not side notes. That matters because the best model for a job is now often the one that fits the workflow, not just the one that tops a chart. He treated the last six months as the useful unit of analysis because the field changes too quickly for slower summaries to stay relevant.