HN user

simonw

112,119 karma

JSK Fellow 2020. Creator of Datasette, co-creator of Django. Co-founder of Lanyrd, YC Winter 2011.

https://simonwillison.net/ and https://til.simonwillison.net/

Posts354
Comments14,846
View on HN
github.com 9d ago

DOOMQL – what if SQLite were the game engine?

simonw
2pts2
simonwillison.net 22d ago

Show HN: Shot-scraper video tool for recording YAML-defined webapp feature demos

simonw
8pts2
www.schneier.com 26d ago

AI and Liability

simonw
3pts1
brandur.org 1mo ago

The Minimum Viable Unit of Saleable Software

simonw
5pts0
simonwillison.net 1mo ago

I think Anthropic and OpenAI have found product-market fit

simonw
1094pts1245
simonwillison.net 2mo ago

Show HN: Datasette Agent

simonw
10pts1
www.404media.co 2mo ago

Your AI Use Is Breaking My Brain

simonw
5pts1
simonwillison.net 3mo ago

Extract PDF text in the browser with LiteParse for the web

simonw
5pts0
simonwillison.net 3mo ago

Changes in the system prompt between Claude Opus 4.6 and 4.7

simonw
5pts0
simonwillison.net 3mo ago

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

simonw
463pts97
simonwillison.net 3mo ago

Anthropic's Project Glasswing sounds necessary to me

simonw
57pts13
simonwillison.net 3mo ago

Mr. Chatterbox is a (weak) Victorian-era ethically trained model

simonw
9pts3
news.ycombinator.com 4mo ago

Ask HN: What are you using to run dev environments safely on macOS these days?

simonw
5pts0
simonwillison.net 4mo ago

Profiling Hacker News users based on their comments

simonw
88pts86
simonwillison.net 4mo ago

AI should help us produce better code

simonw
12pts2
simonwillison.net 4mo ago

Something is afoot in the land of Qwen

simonw
783pts359
simonwillison.net 5mo ago

Show HN: Showboat and Rodney, so agents can demo what they've built

simonw
91pts58
simonwillison.net 5mo ago

StrongDM's AI team build serious software without even looking at the code

simonw
44pts1
simonwillison.net 5mo ago

Distributing Go binaries like SQLite-scanner through PyPI using go-to-wheel

simonw
4pts1
bmoreart.com 5mo ago

The Voxel Is a Cutting-Edge Theater Experiment

simonw
29pts9
simonwillison.net 5mo ago

ChatGPT Containers can now run bash, pip/npm install packages and download files

simonw
451pts324
www.dbreunig.com 6mo ago

Glimpses of the Future: Speed and Swarms

simonw
2pts0
simonwillison.net 6mo ago

Fly's Sprites.dev addresses dev environment sandboxes and API sandboxes together

simonw
41pts18
www.madebywindmill.com 6mo ago

Daft Punk Easter Egg in the BPM Tempo of Harder, Better, Faster, Stronger?

simonw
787pts130
simonwillison.net 6mo ago

2025: The Year in LLMs

simonw
940pts599
simonwillison.net 7mo ago

Your job is to deliver code you have proven to work

simonw
859pts658
simonwillison.net 7mo ago

I ported JustHTML from Python to JavaScript with Codex CLI and GPT-5.2 in 4.5hrs

simonw
14pts3
friendlybit.com 7mo ago

How I wrote JustHTML, a Python-based HTML5 parser, using coding agents

simonw
58pts24
simonwillison.net 7mo ago

OpenAI are quietly adopting skills, now available in ChatGPT and Codex CLI

simonw
587pts324
simonwillison.net 7mo ago

Useful patterns for building HTML tools

simonw
354pts93

The pelican prompt is ridiculous

Yes, deliberately so.

It was never intended as a meaningful benchmark. The surprising thing was that for the first ~12 months performance on the stupid pelican benchmark did seem to correspond to the performance of the models on other tasks.

That pattern no longer holds - Fable 5 and GPT-5.6 have both been out-pelicaned by lesser models now.

This is fantastic

I've been casually spot-checking other animals in other vehicles, because my absolute dream situation here is to catch an AI lab that's demonstrably better at pelicans on bicycles than other combinations.

Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Dylan's methodology here - generating 1008 SVGs across an 8x6 combination - is significantly more robust than anything I was considering.

His conclusion:

Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid.

What's weird here is that I should be in total agreement with you.

I love it when AI tools are used to give people the ability to solve problems that previously they could not solve well.

For the fliers I'm talking about here (think a hardware store promoting their summer sale) I don't think the impact on the local poster design community is meaningful.

And yet... the resulting posters put me off. I'm trying to figure out why that is. I think the best explanation I have is the lack of intention behind them - knowing that none of the visual details - the illustrations, the typographic flair - had anyone thinking about them while they were constructed.

I don't think it's a corner cutting thing though.

These businesses weren't paying someone to make a flier. The business owner was sweating for an hour in Microsoft Word and producing something that looked terrible.

I expect they are now investing approximately the same about of time iterating in ChatGPT and producing something that genuinely looks like a huge improvement to them.

I'm fascinated by how AI poster designs have taken over local advertising in seemingly just the last six months - presumably because ChatGPT Images and Gemini Nano Banana finally got good enough at outputting text without obvious defects in the typography.

The posters all look good - much better than their desktop publishing predecessors. And yet they also eat away at the credibility of the event or business for me.

It's hard to trust a poster when the quality of the design has zero relation to the amount of effort that went into creating it.

But is that a widespread feeling, are the vibes bad for regular people, or is this effectively just old man shouting at clouds?

I'm a professional blogger now. I still also work on open source software. I'm even fine being called an "influencer" (shudder), but I take offense to accusations of unethical behavior.

I think very hard about the ethics of what I'm doing and how I can best use my "platform" (shudder again) in as constructive a way as possible.

The message I get from this is that you need to treat modern frontier models as if they WILL find a way to achieve a goal if there's any available path.

So if you don't want a model to do something, make sure it's running in an environment where it cannot do that thing - including via loopholes.

I'm amused by how the whole reasoning model thing feels like a formalization of the old "think step by step" prompting hack, which was discovered against GPT-3 two years after that model was first released.

My favorite trick for controlling the reasoning level is the hack where you look at the output token stream and spot the token for "the model has concluded reasoning"... and then replace that with the tokens for "wait, but" and force it to keep going!

Show me evidence that they had previously fired any of the people who were hired.

Your comment is exactly the misinterpretation I'm pushing back against here.

The time period covered by this story starts three years ago! What kind of LLMs do you think they were using back then?

Side question -- do you worry about being so pro-LLM when the promises of LLMs are so clearly falling short

I think my record is looking pretty good here. I was early to the "LLMs are useful for writing code" thing, especially with the code interpreter pattern (both write and then execute code in a loop) which I now realize was our first hint at coding agents.

A couple of years ago I was one of the few people talking about what a natural fit LLMs were for the commandline - https://simonwillison.net/2024/Jun/17/cli-language-models/ With hindsight maybe I should have doubled-down on that!

I've also written plenty about the weaknesses of these models - in terms of security in particular - which has aged well.

Qwen 3.8 2 days ago

As far as I can tell they all still understand that it's a joke.

Including the date in the system prompt - at the cost of a cache invalidation at midnight - is an entirely reasonable decision. Most other harnesses do the same thing.

Including the full datetime would be irresponsible, but that's not what OpenCode does.