HN user

tezza

1,820 karma

PM: hackernews (at) terrylurie DOT com

In spare time I make:

- https://epicwin.team - browser video games. Remote Teams and Solo

- https://generative-ai.review - Reviewing AI tools from a professional angle

- LLM London organiser

Based in London. Light inventor on the side of a regular day job

Any opinions expressed are my own and not those of my employer.

Posts3
Comments656
View on HN

Yes, I see your point.

Your pelican output is thus both in the training set and yet still outside the capability of the model architecture.

And so you are tracking both the capability of the training and also the capability of the querying!

When you receive your first outstanding pelican it will track a gain of capability.

(btw I first mentioned simonw-pelican-into-training-set in May 2025 on twitter.)

My 3D-egyptology-explainer showed a massive uplift for Kimi K3 and this tracks a much improved 3D capability.

Nice qualitative test!

Just like you I am super impressed by Kimi K3.

I do a qualitative benchmark series making 3D explainers and so here's Kimi K3 vs Claude Fable:

https://generative-ai.review/2026/07/kimi-k3-rush-test-vs-cl...

I've put links to the posts on GLM5.2, Opus 4.8, Chat GPT 5.5. I grab video screencaps so you can compare in detail. The full interactive Kimi output is at the bottom of the post if you want a comprehensive 3D play around

well i’ve come across it loads, especially 1/2) during REPL style build outs and 2/2) calling libraries and frameworks you are not yet familiar with.

perl another offender… is it a hash? is it an arrayref? over time you get it right, but by trial and error and looping. json suffers this too, arrays different from strings, different from numbers etc, but opaque until checked and liable to change

Claude Fable 5 1 month ago

There are benchmarks if you want quantitative results. Mine is qualitative, and clearly billed as such. Comparison and contrast still possible.

Claude Fable 5 1 month ago

check the backlinks[1][2] in the article before you start throwing around accusations. I am not (yet) a person that has advanced notice and access to models.

Fable just got announced and I did a rush out article because people are curious. I released the post mere hours afterwards and it takes time to create the output, slice into videos, make a wordpress article on top of taking my son to basketball training and eating dinner. I’m in London and this was all happening at 1am.

If you check the links my previous articles have all the juicy stuff you are criticising me for not having with little preparation.

How is a side by side direct comparison NOT precise?

[1] first in series from 2025: https://generative-ai.review/2025/05/vibe-coding-my-way-to-e... . This has all the background you are talking about in the Appendix

.

[2] https://generative-ai.review/2026/05/vibe-coding-my-way-to-e... . Second in series 2026 has a side by side table of what changed. This is what is possible with more than a few hours advanced warning.

MidJourney public discord channel.

The amount of masterpiece level art flowing per hour was astounding.

For every one doing a ninja waifu, there were ten doing art from davinci and leonardo crossed with hockney.

it almost gave you art sickness

The John Snow Pump… where they sealed it up to stop Cholera is very small.

Also Novelty Automation (WC1R 4AX) is a tiny interactive museum/wharf end arcade which has some very intricate mechanical entertainments. Some are tiny.

How is this different to system tuning parameters in Linux /proc, FreeBsd, Windows Registry, Firefox about:config, sockopt, ioctl, postgres?

Zillions of options. Some important, some not

take part of the color space and map it uniformly to a different part of the color space

fyi Affinity Photo has recolor and hue filters that will do just that.

I used it for my video game art.

My latest game BossBattle[1] (html5, have a play) uses AI for graphics, some strobe effects, a C64 loading screen shader.

I have decided to lean in to it and I will document all the places I use AI in the game on my blog[2]. Not everything works, notably 3D assets[3] and sound effects.

There is a lot of human content… i paid for a lot out of my own pocket and have limited budget. It started in 2021 before chatgpt. LLMs cannot do everything and that’s not the purpose.

Generative AI makes me as an solo indie dev able to make the game. Without the AI the game wouldn’t exist

[1] http://epicwin.team/play/solo/BossBattle/ - (public beta) .

[2] https://generative-ai.review .

[3] https://generative-ai.review/2025/08/3d-assets-made-by-genai...

Terminals are text. Text adds features missing from gui namely:

* Ad Hoc

requirements change and terminal gives ultimate empty workbench flexibility. awesome for tasks you never new you had until that moment.

* Precision

run precisely what you want, when you want it. you are not constrained by gui UX limits.

* Pipeline

cat file.txt | perl/awk/sed/jq | tee output.result

* Equal Status

everything is text so you can combine clipboard, files, netcat output, curl output and then you can transform (above) and save. whatever you like in whatever form you like, named whatever you like.

Not sure “Digital Twin Universe” is required here. They seem rather to have rediscovered Simulators in Integration Tests from first principles? The DTU comes off as XML Databases or Information Superhighway.

Still… a really good application of agent hands-off replication.

Seems like creating a quality negative mould and then that single negative mould makes multiple positive objects en-masse.

You’re not alone. I do a small blog reviewing LLMs and have detailed comparisons that go beyond personal anecdotes. Gemini struggles in many usecases.

Everyone has to find what works for them and the switching cost and evaluation cost are very low.

I see a lot of comments generally with the same pattern “i cancelled my LEADER subscription and switched to COMPETITOR”… reminiscent of astroturf. However I scanned all the posters in this particular thread and the cancellers do seem like legit HN profiles.

this is truly bizarre.

It’s as if they’ve never heard of Maslow‘s Hierarchy of Needs before and further did they don’t know Self Actualizing is right at the very top.

Without that key stone on the top the human being is still a wanting animal. And if you somehow “mission complete” one Self Actualizing, then you immediately start to want something fresh “purpose” etc.

And obviously Self Actualizing doesn’t have to come in the form of work, although often it does.

Yeah, it needs a steady hand on the tiller. However throw together improvements of 70%, -15%, 95%, 99%, -7% across all the steps and overall you're way ahead.

SimonW's approach of having a suite of dynamic tools (agents) grind out the hallucinations is a big improvement.

In this case expressing the feeback validation and investing in the setup may help smooth these sharp edges.