It's not an argument, it's an open question: do people who express a dislike of AI know when they see a poster that was generated with assistance from AI?
HN user
simonw
JSK Fellow 2020. Creator of Datasette, co-creator of Django. Co-founder of Lanyrd, YC Winter 2011.
https://simonwillison.net/ and https://til.simonwillison.net/
Gemini have absolutely been SVGmaxxing. They've openly talked about it.
The pelican prompt is ridiculous
Yes, deliberately so.
It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.
That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.
Underlying data is available on GitHub: https://github.com/dylanjcastillo/blog/tree/main/_extras/pel...
This is fantastic
I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.
Catching a lab cheating specifically on my one dumb benchmark would be really funny.
Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.
His conclusion:
Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.
What's weird here is that I should be in total agreement with you.
I love it when AI tools are used to give people the ability to solve problems that previously they could not solve well.
For the fliers I'm talking about here (think a hardware store promoting their summer sale) I don't think the impact on the local poster design community is meaningful.
And yet... the resulting posters put me off. I'm trying to figure out why that is. I think the best explanation I have is the lack of intention behind them - knowing that none of the visual details - the illustrations, the typographic flair - had anyone thinking about them while they were constructed.
How many of the people who answer polls like that also have the ability to look at a poster and instantly clock that it was AI-generated?
I don't think it's a corner cutting thing though.
These businesses weren't paying someone to make a flier. The business owner was sweating for an hour in Microsoft Word and producing something that looked terrible.
I expect they are now investing approximately the same about of time iterating in ChatGPT and producing something that genuinely looks like a huge improvement to them.
I'm fascinated by how AI poster designs have taken over local advertising in seemingly just the last six months - presumably because ChatGPT Images and Gemini Nano Banana finally got good enough at outputting text without obvious defects in the typography.
The posters all look good - much better than their desktop publishing predecessors. And yet they also eat away at the credibility of the event or business for me.
It's hard to trust a poster when the quality of the design has zero relation to the amount of effort that went into creating it.
But is that a widespread feeling, are the vibes bad for regular people, or is this effectively just old man shouting at clouds?
Have you created other burner accounts to reply to me in the past, or is this your first one?
I'm a professional blogger now. I still also work on open source software. I'm even fine being called an "influencer" (shudder), but I take offense to accusations of unethical behavior.
I think very hard about the ethics of what I'm doing and how I can best use my "platform" (shudder again) in as constructive a way as possible.
I did hit a nerve. I don't like being accused of posting comments here for "low effort personal brand promotion" or nefarious financial motives.
Linking directly to the rendered markdown as opposed to a post on my blog is a poor way to promote my blog.
You and a few other people, but enough people still appreciate the bit that I'm going to keep doing it.
They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.
Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.)
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
to precisely describe the full 3×3 grid takes a full 3.7k tokens
It's a shame they didn't share that prompt - it would make that demo more convincing.
Yes, it generates the text without using a font. Same is true of other image models like ChatGPT Images and Gemini Nano Banana and Midjourney.
I was so excited when I saw "blaizzy" in the domain, because Prince Canuma's work around MLX has been of such uniquely high quality.
The message I get from this is that you need to treat modern frontier models as if they WILL find a way to achieve a goal if there's any available path.
So if you don't want a model to do something, make sure it's running in an environment where it cannot do that thing - including via loopholes.
What method did you use to generate that one?
I tried using their OpenAI-compatible API and got one that's a lot less fun (no animation) - https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
I managed to get myself a working Qwen Cloud account, here's a pelican. The reasoning trace is pretty fun - three details I liked were "Could add helmet? No." and "Maybe add small bell? no." and "Need maybe add small fish in basket? Not necessary."
https://tools.simonwillison.net/markdown-svg-renderer#url=ht... (scroll to bottom for pelican)
I'm amused by how the whole reasoning model thing feels like a formalization of the old "think step by step" prompting hack, which was discovered against GPT-3 two years after that model was first released.
My favorite trick for controlling the reasoning level is the hack where you look at the output token stream and spot the token for "the model has concluded reasoning"... and then replace that with the tokens for "wait, but" and force it to keep going!
Show me evidence that they had previously fired any of the people who were hired.
Your comment is exactly the misinterpretation I'm pushing back against here.
The time period covered by this story starts three years ago! What kind of LLMs do you think they were using back then?
Side question -- do you worry about being so pro-LLM when the promises of LLMs are so clearly falling short
I think my record is looking pretty good here. I was early to the "LLMs are useful for writing code" thing, especially with the code interpreter pattern (both write and then execute code in a loop) which I now realize was our first hint at coding agents.
A couple of years ago I was one of the few people talking about what a natural fit LLMs were for the commandline - https://simonwillison.net/2024/Jun/17/cli-language-models/ With hindsight maybe I should have doubled-down on that!
I've also written plenty about the weaknesses of these models - in terms of security in particular - which has aged well.
As far as I can tell they all still understand that it's a joke.
I suspect that if you prefixed every user message with a date the model might get confused and start treating the date as more relevant to the current task than it actually is.
Interesting detail from Ben Thompson's piece on Chinese models - https://stratechery.com/2026/whos-afraid-of-chinese-models/ - apparently Xi Jinping gave this speech recently http://english.scio.gov.cn/topnews/2026-07/18/content_118605... which included support for open source models:
We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing.
Or they could add the frozen initial starting date once at the start, then inject an additional "the date is now X" message any time midnight passes.
Wow, that one is really good.
Including the date in the system prompt - at the cost of a cache invalidation at midnight - is an entirely reasonable decision. Most other harnesses do the same thing.
Including the full datetime would be irresponsible, but that's not what OpenCode does.
Have you watched the video where that was said, as opposed to just reading interpretations of it from people with their own assumptions?
https://youtu.be/02YLwsCKUww?is=UsJrcLRgdSUd38Jt
Here's a transcript: https://gist.github.com/simonw/0050b4ce40439b97598d4ffdb273d...