Serious question: how do we verify claims like these on the effectiveness of a harness?
HN user
richardfey
Has anyone tried Kimi K3 against gpt-5.6-sol on real projects?
Exactly; this is a no-go for me, I will wait for an independent provider to sell the service, which is possible thanks to the open weights.
I understand you're trying to be funny, but my point is that with novel technology there are novel ways to claim innocence in courts because of the legislative void.
They could use an agent to summarise the source material, and then train models on those summaries, and claim that some sort of clean-room training has happened?
Or are you referring to the claims from Meta, Google, and Adobe -- which failed to hold up under independent evaluation.
This.
However, he demonstrated a clear lack of understanding regarding what makes the bits "independent" or how to resolve the independence problem.
Yes, I didn’t want to call it out explicitly, but this is exactly the kind of thing that would have made my undergraduate statistics professor lose patience.
I can't find on their website some indication of what kind of usage I can get out it, otherwise I'd be interested.
This is a great statistical analysis and it was a pleasure to read, but I wasn't expecting the claims to be so poorly supported. There's also a reply from one of the Meta authors there, worth checking out.
I remember hearing this perspective when I first started in the software industry, and I agreed with it for quite some time. But frankly, we’ve never been further from it.
The more I read into it, the more pain memory flashbacks I got. Bravo
I'm trying it right now for a side project of mine, compared to Opus it is effectively better at following instructions and somehow has better "depth" when reasoning on complex tasks. However, if it will not be part of subscriptions, I will not use it anymore.
I don't know what I am doing right, or wrong, but I have access to claude and codex and I find myself giving the more serious work to codex recently. I tend to trust it more. I might try again Fable when it's back, but this Sonnet 5 didn't work well for my current projects.
How did you give LLMs tool use?
I could spot numerous bugs in code written recently and less recently, by me or colleagues. I was not angry but grateful and I knew there was no way back!
You mean bad because they could have used a larger memory module and thus higher resolution sound samples?
To do this, your device is shouting to the world a ton of your personal information in something called a probe packet. A probe packet contains the MAC address as well as the list of all the past Wi-fi networks that your device has tried to join before, which can reveal a lot about you!
In the 2010s, maybe. Nowadays MAC address randomisation is the norm and past WiFi networks are not broadcast anymore.
I'm going to give a try to piclaw as I want to get into the mindset of author, thanks!
I have a different take on this: how many people using a rifle do you see at composite bow tournaments? We might just move this kind of activities to sandboxes and implement more strict requirements for participation. It might make them niche or indeed disappear.
"We fired all of the QA people, now there are no QA issue reports anymore".
I'd be interested in seeing how worse the Epic Citadel demo will perform with this removal.
Some reference to Les Technopères by writer Alejandro Jodorowsky?
I have 24GB VRAM available and haven't yet found a decent model or combination. Last one I tried is Qwen with continue, I guess I need to spend more time on this.
Probably JMAP support.
I wonder how cjdns would have handled this
I feel like the process of carving out any meaning out of "QA" is complete. It's cathartic, in its twisted way...
But AI is not a software application, it does not aim to replace software applications.
It's like saying that the discovery of steel did not replace any existing weapon or tool.
Please let me know what you end up doing with this, I am curious!
They were doing this kind of optical media seek times tests/optimisations for PS1 games, like Crash Bandicoot. You certainly have more and better context than me on this console/game, I just mentioned it in case it wasn't considered.
By the way, could the nonsensical offsets be checksums instead?
Nice reverse engineering work and analysis there!
This might be an optimisation to avoid disc seeks on wildly far apart distances, which would introduce more latency.
I understand, it's risk diversification.