HN user

leumon

161 karma
Posts9
Comments85
View on HN
GLM 5.2 vs. Opus 1 month ago

I've seen glm 5.2 struggle writing simple compilable c code. It might be good at web, but it's world knowledge is limited due to the small model size, making it's use quite limited in my opinion.

This kind of reminds me of these malicious captchas that get you to paste some command into cmd.exe. These kind of captchas will make this situation worse, I could also see some malicious site having a qr code that will download some virus to your phone. QR code captchas are a really bad idea in my opinion.

AlphaZero, the engine that pioneered the “neural network” approach now incorporated into Stockfish

That's simply not true. While stockfish does use a neural net, it's not using the MCTS approach like LeelaChessZero, and only uses the neural net for evaluating a position, not for suggesting moves. And it was only implemented after stockfish lost to lc0 in a computer chess tournament.

Claude Sonnet 4.6 5 months ago

Asked gemini and it said to use ground handling wheels. I think it actually makes sense to use that for this distance.

Claude Sonnet 4.6 5 months ago

My locally running nemotron-3-nano quantized to Q4_K_M gets this right. (although it used 20k thought tokens before answering the question)

For programmers or people who know computers quite well the difference to claude code is small i would say. But for "Normies" its magical that you can just ask your computer to do anything from anywhere (set timers, install stable diffusion, send you a specific doc in your download folder). You don't even have to write it, you can send it a voice message and it will install whisper or send it to the openai whisper api, etc. Obviously this is more then dangerous, but looking at what passwords people still choose today (probably also the reason why everything requires MFA nowadays), most people don't care about Security.

GPT-5.3-Codex 6 months ago

they tested it at xhigh reasoning though, which is probably double the cost of Anthropic's model.

Cost to Run Artificial Analysis Intelligence Index:

GPT-5.2 Codex (xhigh): $3244

Claude Opus 4.5-reasoning: $1485

(and probably similar values for the newer models?)

according to the age-prediction page, the changes are:

If [..] you are under 18, ChatGPT turns on extra safety settings. [...] Some topics are handled more carefully to help reduce sensitive content, such as:

- Graphic violence or gore

- Viral challenges that could push risky or harmful behavior

- Sexual, romantic, or violent role play

- Content that promotes extreme beauty standards, unhealthy dieting, or body shaming

GPT Image 1.5 7 months ago

One other test you could add is generating a chessboard from a FEN. I was surprised to see NBP able to do that (however, it seems to only work with fewer pieces, after a certain amount it makes mistakes or even generates a completely wrong image) https://files.catbox.moe/uudsyt.png