HN user

juliangoetze

34 karma
Posts6
Comments30
View on HN

I thought HN was different. And yeah, wherever I go, my timeline is full of Opus is so bad today and I will switch from Fable to 5.6 Sol, it's 1.5x better and vice versa.

Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.

This supposedly is better than KimiK2.7

How can you tell?

I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks.

Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?

I am very curious about how the "threat" of local inference and open-weights-models will change the trajectory of the model labs. The air will become pretty thin

I noticed that Claude's reasoning summaries show a phenomenon when I'm using the model in German (on claude.ai) - it mixes English grammar with German vocabulary!

"Evaluating Spülenposition gegen Wasseranschlussabstand" and "Analysierend die Platzierungskonflikte und Rohrleitungszwänge klären" are examples of generated summaries.

I wish someone could explain this - because I see no way that sentences like these would show up in training data. But maybe the RL for thinking has some quirks?