A classic satirical SF short story from Stanislaw Lem, containing a satirical look at AI, sexbots, and transhumanism.
HN user
chromaton
Hacker and business owner with an interest in digital manufacturing.
I did something very similar last year, but with programming languages that were REALLY out of distribution; they were generated specifically for the benchmark. I call it TiānshūBench (天书Bench): https://jeepytea.github.io/general/introduction/2025/05/29/t...
Some models were OK at solving very simple problems, but nearly all of them would, for example, hallucinate control structures that did not exist in the target language.
Historically, the cycle has been requirements -> code -> test, but with coding becoming much faster, the bottlenecks have changed. That's one of the reasons I've been working on Spark Runner to help automate testing for web apps: https://https://github.com/simonarthur/spark-runner
I've recently found that my ability to add new features and squash bugs has outpaced my ability to do full end-to-end tests. To help with this, I created Spark Runner for automated website testing. It will create and execute a plan for tasks you give it in plain text like "add an item to the shopping cart" or you just point it at your front end code and have Spark Runner create the tests for you. It also makes nice reports telling you what's working and what's not.
New project, so feedback is welcome.
TFA seems to be big on mathematical proof of correctness, but how do you ever know you're proving the right thing?
Lisp has been around for 65 years (not 50 as in the author believes), and is one of the very first high-level programming languages. If it was as great as its advocates say, surely it would have taken over the world by now. But it hasn't, and advocates like PG and this article author don't understand why or take any lessons from that.
Moravec strikes again.
For my benchmarking suite, it turns out that it's about 1/5 the price of Claude Sonnet 4.1, with roughly comparable results.
If you're looking for free API access, Google offers access to Gemini for free, including for gemini-2.5-pro with thinking turned on. The limit is... quite high, as I'm running some benchmarking and haven't hit the limit yet.
Open weight models like DeepSeek R1 and GPT-OSS are also made available with free API access from various inference providers and hardware manufacturers.
This has been available (20b version, I'm guessing) for the past couple of days as "Horizon Alpha" on Openrouter. My benchmarking runs with TianshuBench for coding and fluid intelligence were rate limited, but the initial results show worse results that DeepSeek R1 and Kimi K2.
Current AI systems don't have a great ability to take instructions or information about the state of the world and produce new output based upon that. Benchmarks that emphasize this ability help greatly in progress toward AGI.
Yes, it would be fantastic to have more languages to test off of. I picked the base language I did (Mamba) because it was easy to modify and integrate into Python.
Generating the problems: I just thought up a few simple things that the computer might be able to do. In the future, I hope to expand to more complex problems, based upon common business situations: reading CSVs, parsing data, etc. I'll probably add new tests once I get multi-shot and reliability working correctly.
New base programming languages would be great, but what would be even better is some sort of meta-language where many features can be turned on or off, rather than just scrambling the keywords like I do now.
I did some vibe testing with a current frontier model, and it gets quite confused and keeps insisting that there's a control structure that definitely doesn't exist in the TiānshūBench language with seed=1.
I find that these books have to be read by the right person at the right time. Think and Grow Rich by Napoleon Hill did nothing for me when I first was exposed to it, but later on, helped me greatly.
BTW, the business book that helped me the most is barely known: Making Money is Killing Your Business by Chuck Blakeman.
The PDF conversions I've tried in Firefox and Chromium don't work that well.
Another on that really irritates me is the kind that presents a series of integers and asks which integer comes next. Any integer will do, you just have to fit the appropriate polynomial.
This one bugs me to no end because it's part of the standard elementary school curriculum, for example here: https://byjus.com/maths/patterns-questions/
But surely someone with a strong imagination could come up with a pattern to fit any number as the next in the sequence. I doubt most elementary educators even grasp the issue.
AutoCAD automation?
Xometry, though it's also US and EU based.
I've been slowly working my way through this book for the past couple months. It's been amazingly helpful in learning all of the deep learning terminology, and giving a good overview of the technology.
I was doing all of the examples and exercises for a while, but gave up on that at some point. My main goal, after all, was to learn about how the technology works in order to separate the wheat from the chaff, not become an AI researcher.
Slop has been around a while. I was researching a topic, and noticed that most of the top search results had the same misunderstanding of some of the definitions. The writers were clearly not familiar with the topic, and I'm sure they were just copying each other. All of the articles pre-dated GPT-3.5.
The kicker is that if you ask GPT-4 about it, it spits out the same incorrect information, meaning that GPT-4 was likely trained on this bad data. FWIW, GPT-4o gives a much more accurate response.
It can't correctly identify a DXF file in my testing. It categorizes it as plain text.
developers iso2_code population fraction
20226711 US 335893238 6.02%
15528470 EU 448387872 3.46%
13326416 IN 1392329000 0.96%
9131545 CN 1409670000 0.65%
4315369 BR 203080756 2.12%
3422806 GB 67026292 5.11%
3060711 RU 146424729 2.09%
2972917 DE 84607016 3.51%
2938205 ID 279118866 1.05%
2886350 JP 124090000 2.33%
2479673 CA 40528396 6.12%By the metrics on this page, 6% of the US population has a GitHub account. Seems a bit high to me.
"I tried to picture clusters of information as they moved through the computer. What did they look like? Ships, motorcycles? Were the circuits like freeways? I kept dreaming of a world I thought I'd never see. And then, one day, I got in..."
I met Marc Thorpe back in 1995 and saw some of the Robot Wars highlight videos he was using to promote the event. My life was never the same after that.
Please remind me why landfills are so bad again.
I ran a business successfully for many years with GnuCash. I only switched because I wanted to offload bookkeeping to a professional for a couple hundred bucks a month, and she was only familiar with QuickBooks.
Assembly Atlanta has 135 acres with indoor studios, plus outdoor streets to simulate New Orleans, New York, Tribeca, and "Europe".
It's under construction, but close to completion.
Sam Walton points out in his autobiography that big cities are changing all the time, yet somehow people expect that small towns will not.
Think and Grow Rich - Napoleon Hill
Making Money Is Killing Your Business - Chuck Blakeman