I mean, surely the main motivation for "use an LLM to rewrite a huge project in a new language" was excitement about the shiny new tech that made it possible.
HN user
MrCheeze
I did this stuff: https://mrcheeze.github.io/
TBF, they did it first with ada/babbage/curie/davinci. "Sol" is a much weaker branding, though.
Does anyone understand why LLMs have gotten so good at this? Their ability to generate accurate SVG shapes seems to greatly outshine what I would expect, given their mediocre spatial understanding in other contexts.
The Claude Plays Pokemon stream with a minimal harness is a far more significant test of model intelligence compared to the Gemini Plays Pokemon stream (which automatically maintains a map of everything that has been seen on the current map) and the GPT Plays Pokemon stream (which does that AND has an extremely detailed prompt which more or less railroads the AI into not making this mistakes it wants to make). The latter two harnesses have become too easy for the latest generations of model, enough so that they're not really testing anything anymore.
Claude Plays Pokemon is currently stuck in Victory Road, doing the Sokoban puzzles which are both the last puzzles in the game and by far the most difficult for AIs to do. Opus 4.5 made it there but was completely hopeless, 4.6 made it there and is is showing some signs of maaaaaybe being eventually bruteforce through the puzzles, but personally I think it will get stuck or undo its progress, and that Claude 4.7 or 5 will be the one to actually beat the game.
Notably 45 out of the 50 days of improvement were in two specific dungeons (Silph Co and Cinnabar Mansion) where 4.5 was entirely inadequate and was looping the same mistaken ideas with only minor variation, until eventually it stumbled by chance into the solution. Until we saw how much better it did in those spots, we weren't completely sure that 4.6 was an improvement at all!
https://docs.google.com/spreadsheets/u/0/d/e/2PACX-1vQDvsy5D...
In my experience with the models (watching Claude play Pokemon), the models are similar in intelligence, but are very different in how they approach problems: Opus 4.5 hyperfocuses on completing its original plan, far more than any older or newer version of Claude. Opus 4.6 gets bored quickly and is constantly changing its approach if it doesn't get results fast. This makes it waste more time on"easy" tasks where the first approach would have worked, but faster by an order of magnitude on "hard" tasks that require trying different approaches. For this reason, it started off slower than 4.5, but ultimately got as far in 9 days as 4.5 got in 59 days.
This writeup on the underground puzzle is worth reading, it's a pretty baffling "puzzle" design. https://pokemow.com/Gen2/ShutterPuzzle/
That said, it's definitely Gem's fault that it struggled so long, considering it ignored the NPCs that give clues.
There were no such writeups, 99% of the discussion about difficulties in Crystal were in twitch and discord chats where Google doesn't scrape. (It hadn't yet gotten the public attention that Claude and Gemini's runs of Pokemon Red and Blue have gotten.)
That said, this writeup itself will probably be scraped and influence Gemini 4.
It's hard to say for sure because Gemini 3 was only tested with this prompt. But for Gemini 2.5, which is who the prompt was originally written for, yes this does cut down on bad assumptions (a specific example: the puzzle with Farfetch'd in Ilex Forest is completely different in the DS remake of the game, and models love to hallucinate elements from the remake's puzzle if you don't emphasize the need to distinguish hypothesis from things it actually observes).
Exactly what I was going to post. Optimizations like loop unrolling slow down the N64 because keeping the code size small is the most important factor. I think even compilers of the time got this wrong, not just modern ones.
The busy beaver function is interesting precisely because you _can't_ come up with any computable function that grows faster.
Claude almost universally reacts to everything with a positive exclamation as its first sentence, regardless of whether it's good or bad. If you don't believe me, just watch https://www.twitch.tv/claudeplayspokemon for about three minutes and you'll get the idea.
Alternatively, look at the system prompt, where Anthropic attempted to get it to stop doing this: > Claude never starts its response by saying a question or idea or observation was good, great, fascinating, profound, excellent, or any other positive adjective. It skips the flattery and responds directly. https://docs.anthropic.com/en/release-notes/system-prompts#a...
This problem seems highly specific to Claude. It's not exactly sycophancy so much as it is a strong bias towards this exact type of reaction to everything.
With 2, the real problem is that approximately 0% of the OpenAI employees actually believed in the mission. Pretty much every single one of them signed the letter to the board demanding that if the company's existence ever comes into conflict with humanity's survival, the company's existence comes first.
Are you thinking of Bismuth's "Speedrunning as a gateway to scientific endeavours", perhaps?
As an n=1 data point, that was my exact situation for a while. Also a lot of the people who put out high effort stuff are college students, which works for the same reason.
More interestingly and more surprisingly, some of the people who work on exploiting games _don't_ do any sort of tech work and have no background in compsci - they're purely self educated just for the sole purpose of breaking the one game they're interested in. This was the case for some of the biggest contributors to ACE in Zelda Ocarina of Time.
I've wondered myself why there's so little overlap between these two closely related interests of mine. Some of it seems to be the "But I don't want to cure cancer. I want to turn people into dinosaurs." effect, where some of the people working on exploiting games ONLY care about what can be done in their one game of interest - it doesn't always generalize to interest in using the same techniques against everything else.
Of course there's also the fact that exploiting 20-30 year old games is just vastly easier than modern software, due to the total lack of mitigations in them. And that's on top of the fact that with popular games, you're building on decades of reverse engineering work rather than (potentially) starting from scratch. And the arguably superior toolset (savestates etc).
But I think a very big factor is the one this blogpost is trying to address - most people just don't know anything at all about the vuln research industry, which is not exactly searching for attention in the ways that speedruns broadcast to hundreds of thousands of viewers for charity are.
How long until we get to the point where models know that LLMs get this wrong, and that it is an LLM, and therefore answers wrong on purpose? Has this already happened?
(I doubt it has, but there ARE already cases where models know they are LLMs, and therefore make the plausible but wrong assumption that they are ChatGPT.)
In conversational language, "All my hats..." implies that the speaker has at least two hats, which theoretically means that the sentence could be a lie from them having exactly one hat (even a green one). However, in practice, I don't think anyone would actually call that a lie. I think we treat the "I have hats" part not as part of the sentence itself, but more as an underlying premise.
You could have a similar situation in pure logic or math - if I were to say "the largest prime number is odd", is that false? Or something else entirely? (This is what Hofstadter calls mu, from a related concept in Zen.)
Two reasons why I don't particularly believe him:
1) Altman's companies have had similar clauses before: https://news.ycombinator.com/item?id=40396787
2) The entire OpenAI board debacle started because Sam wanted Helen Toner removed from the board for publishing a paper he felt was disparaging to the company: https://thezvi.substack.com/p/openai-the-battle-of-the-board
If, like me, you are suddenly curious what would happen if you added a small fourth body: https://youtu.be/WrahPSY9pf0
I have just been informed that my above comment is false, the CLIP-L is in fact referring to OpenAI's, despite that also being the name of an OpenCLIP model.
One of the diagrams says they're using CLIP-G/14 and CLIP-L/14, which are the names of two OpenCLIP models - meaning they're not using OpenAI's CLIP.
You should never ask an LLM to answer questions about itself. The answer is guaranteed to be hallucinated unless Google specifically finetuned it on an answer of that question. The answer it gave you is meaningless. (But also, coincidentally, correct.)
So the one thing this article doesn't explicitly justify is whether the ten-byte zlib string is truly the shortest possible. You could imagine that it might be possible to hand-craft a DEFLATE block which is only three bytes instead of four, and yet still decompresses to at least two bytes (the minimum decompressed size of a PNG scanline).
But by exhaustive search of all DEFLATE blocks that are one to three bytes in length, I can confirm the article is correct. All of them decompress to (at most) one byte of data.
Which isn't a surprising result, but I wasn't actually sure of it before trying it out.
It has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971