HN user

monster_truck

939 karma
Posts0
Comments442
View on HN
No posts found.
ECC and DDR5 21 hours ago

I do, often, and respect the time of my readers enough to verify my claims instead of presenting anecdoes as fact.

Have you actually compared the random data that various tools write to a disk to encrypted files? Load up binwalk and give it a rip, you will be surprised by how distinct the entropy graph is between the two

I didn't say all expert witnesses are good nor did I claim they are all used in good faith.

Merely that, unlike in the invented scenario above, statistical differences in the random garbage that gets written to a disk across different tool or OS versions has been successfully used to place a time window on when exactly someone destroyed evidence.

Try it for yourself brother, there is a plainly obvious difference between the random noise of writing garbage compared to a properly encrypted binary.

There are even statistically measurable differences in the garbage written across versions of a tool or OS, which, unlike the invented scenario above, has been successfully used to place a time window on when exactly someone destroyed evidence.

I don't think it is as much about 'real' work or a special insight as it is being willing to push back multiple times, or simply asking in a way that steers it towards actually 'giving enough of a fuck' to even bother. We tend to be ~blind to how differently we would ask about something we know compared to a novice, this is what makes some better teachers than others.

Have encountered a similar flavor in programming, wrote it off until I saw someone point out how garbage in garbage out they tend to be. If you hand any frontier model dogshit and ask it to do something simply, the result is often not great.

But! If you spend 20 minutes having it comb through and clean up with something like jscpd, then tell it to step through with a debugger, gather profiling traces, etc... very likely it will yield meaningful improvements or catch some corner cases. If it doesn't, anyone with experience is going to tell it to try something else, or that it isn't good enough, as opposed to accepting the first result.

You can recreate this by disabling web search and asking a model about the conjecture and then giving it his post. I've tried a few and their initial responses range from "this is a meme I'm not even going to verify it" to vaguely insulting chains of thought, concerns about the need to be careful because you're clearly nuts or stupid, then falling back on remedial explanations. After a few nudges they all eventually work through it, accept it, and apologize.

IMO its reasonable to imagine a situation where someone is having a beer or two watching The Big Game, asking an LLM to do something stupid for fun and landing somewhere like this on the magic jump to conclusions mat.

Not much.

But it does give credible plausibility to the concept that we might be mistaken about the exact boundaries of hardness for adjacent (but not equivalent) polynomial systems. Most (all?) of which have also stood up to a whole lot of undeniably sharp people poking at them for about as long.

This should be upvoted more. Saying it is not easy is an understatement. The chasm is so large that if you were to represent it on a globe it would rival oceans. I've been brought into places that were bought for half a billion+ to try and claw back half of a tenth of a percent of traffic. A redesign is pushed through that minimizes content, inflates metrics, and triples the ads. Users leave, everyone is fired, it's written off and the cycle repeats.

OAI is going to try and ape facebook's opening plays. Burn at any expense to optimize for time spent, give you what you want, establish usage habits, dial it back to foster a dependency. Leverage their position as an arbiter of truth, claim they are going to act like x and do y. Fuck people up, roll out shiny new safeguards, silently peel them back. Dig a regulatory moat that makes anything resembling competition prohibitively expensive.

The only time I have gone to best buy in the past 10 years was because they have a garbage can full of sand ready to accept any spicy pillows (batteries that have become genuinely dangerous).

Wow look at that, not a single response saying how excited they are to hear about the latest deals on amazon for some garbage they don't need to earn a reset token

There literally are not. Anyone competent enough to be an expert witness will be able to plainly explain to everyone else how statistical analysis obviously delinates the difference between truly random noise and an encrypted volume.

Might as well be rot13.

If you rented 8 MI300X's or the nvidia equivalent, I don't even think an unreasonably long password would matter.

It should finish quickly enough that you would be upset with all of the money you have now wasted by having to commit to a month of utilization

The centralized approach they're using does this without exploding in size. I think what you want would be thread local tables to avoid locks/syncing across threads? Would only help for mt, though.

Writing to a ring buffer and processing in batches should be a relatively big W for everything. You could then filter for repeated instructions to compress, which would lend itself to prefetching the next instruction and buffer with hints. Mixed precision record storage might help too, but you're plucking hairs at that point. Making the sampling adaptive based on the hw profiling counters might eliminate some 'useless' work.

The overhead is already so low that it really should not matter, though. Used to spend a lot of time trying to find wins like this but caches have gotten so large it basically doesn't matter. Even in pathologically memory or io bandwidth bound cases I've found it's usually faster to just run 2 or 4 smaller instances over trying to coordinate all of the threads in a single big one.

ECC and DDR5 2 days ago

This is not a very good post.

Claiming that every issue you've experienced is from seating errors? Come on man, at least deliberately reproduce that scenario once. I've tortured motherboards for fun and never managed that on purpose. We did pry every cap off of an old K6 board while it was on and... it just kept running. It didn't come back after we turned it off

The amount of time it takes to reseat memory, even if the case is already open and ready, is also more than enough time for it to cool back down.

If you have decent memory running at what it's rated for and you're seeing more than a couple per year, good odds are it's getting too hot or the load line calibration is not aggressive enough.

E: I have never, and I mean literally never, seen age or abuse play a factor.

What do you think the venn diagram looks like for people willing and able to find things like that prior to LLMs and also sell them to a broker, and are also stupid enough to flaunt a massive flashing "arrest me!!!" sign

Closest you're going to get is something like those kids in florida who just got wrapped for putting malware into steam games and draining peoples accounts. They were going to get caught anyways but it would have taken a lot longer to build a case against them if they weren't flaunting it on socials

You are vastly underestimating how much more profitable a 10% annual return on GPUs is than basically anything else you would use an LLM for.

They think whatever you are doing is cute and would very much like to ensure their models can do it even better in the future, but competing? Not even worth the time to think about

Qwen 3.8 3 days ago

Of course, but why be logical and think about the situation critically when you're pushing very hard for regulatory capture against competitors that give their weights away and provide services that are more reliable, offer a better value, and, this is the worst part, they're from CHINA.

Some of the accusations were going so far as to imply that they were outright routing your requests to anthropic and logging its/your response. It's kind of pathetic

Codex Resets 3 days ago

Sure, but also who asked? And that's because they have the compute to burn, which is conditional based on their need to use it to train the next one.

Anything else I've ever paid $200/mo for, advertised as being for professionals, had a markedly better customer experience. If you wanna be Patrick and tell your pet rock to take its time at that price, well then you're Patrick. Good job!

I'm having far more fun with Kimi and Deepseek, at lower prices, without these problems. And I don't have to follow a bunch of obnoxious shitposters to stay clued in on what the fuck is going on or what this week's excuse is.

They're especially cocky right now, they have to beat their chest and pretend like their only competition is Anthropic. They're in for a rough wake up call man. Playing it fast and loose with developer loyalty is a fantastic way to get burned when options like those exist. They are earnestly just as good and in some cases better, and check this out: nobody can take them away from you no matter where you live. If you want to rent 8 GPUs and run the open models yourself, you can do that! You can even be enterprising and sell your excess compute to your friends, or strangers. Best to figure this out before the regulatory capture starts

Codex Resets 3 days ago

Check the github issues brother, search for "usage". Read through the noise. Look at the dates. Don't forget the closed ones.

It's been happening every week since 5.2 was fresh. Not everyone is effected every time. Sometimes it's regional, other times it's per platform, or per version, with a feature (new or old) either enabled or disabled, or some combination of these things. Sometimes you get resets that nobody else does, other times you don't get the ones that were announced. Sometimes you get emails telling you that boosts on limits you didn't even know you had are expiring. At points it did balance out when they really fucked up metering and were underbilling by absurd factors, but that feels more like being toyed with and experimented on than a genuine mistake.

That someone had the 'brilliant' idea to add banked resets (which do expire, so you have to use or lose them, hope you didn't have other plans) says to me that they are no longer as confident about actually being able to fix this as they once were. I'm sure some amount of it is a function of compute availability and reliability, but the rest are definitely issues with the app and it's a rake they keep stomping on.

Compared to just about any other dev centric service I have paid at least $200/mo for, this does not feel very professional to me. They got $1200 outta me and it never really improved. At points I felt like I was getting my moneys worth, and I did burn through a few hundred million tokens, but the frustration of having your projects and plans interrupted and having to wait is not awesome... and then I used Deepseek V4 Pro and felt sick to my stomach with buyers remorse as I watched what it did with just $10.

Codex Resets 3 days ago

I know it for a fact. Have you never worked at a startup before or even seen "growth hacking" at work?

Turning <100k claude refugees into seven figures by counting their web sessions, their codex sessions, and the mobile app as separate "users" is pratically the only way you can get such numbers.

Then you do it again by offering ex-users like me who have already given them a grand or two a free month of plus or pro (this is the third time they have done so for me)

When you hear founders and investors talk about "the magic of storytelling" "the power of momentum", this is what that is.

Qwen 3.8 3 days ago

Which settings exactly do you want?

As far as the model settings go I just follow what's on the card.

I'm using HauhauCS's models, they seem to do a slightly better job with their "P" quants. Especially wrt patching them to eliminate the "doom loops" that will time out the GPU (esp if you have not already given it the extra power budget, set fans to max, and lowered your max clock by about 8% to save yourself a crash/reboot).

Without getting into the weeds, unsloth covers a lot more ground so, when they're good they're great, but I've also had the most problems with them. Basically, don't be afraid to shop around and fuck with sliders.

These days it should almost always be enough to open up LMStudio, set context to max, K/V quant to 16 or 8, and off you go. I'm using a 7900XTX, with 128GB of memory for the cases where things don't fit. The default settings should be fine otherwise.

I don't know anything about the r9700, seems neat. Looks like it might play nicer with the vulkan backend than rocm, and you might have to `set GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1`. Perf should be at least as good as what I'm getting (180+ tk/s prompt, >50tk/s response), which imho is just fast enough that I can read it as it reasons and responds.

You can probably get more aggressive with KV quant (any of the 4's) and bump up the batch sizes (4096/1024 vs 2048/512, etc), it's quite situational. These days there generally should not be any serious loss of accuracy from the former. Keeping the rig cool will likely matter more, so turn your music up to hide the fans and set your rig on top of the AC exhaust lol. Oh and I guess make sure ReBAR is enabled and working, adrenaline should let you know if it's not enabled.

In this domain and others like it especially!

It's bone crushingly dry to try and figure this kind of shit out yourself, can do everything right according to every available reference and code comment only to get compiler and linker errors nobody has ever seen or posted about. Weeks and weeks of smashing your face against a wall until it gives. Now anyone can get twice as far in an evening. And that fuckin rules

Hard to not get a little hyperbolic about it... but I can't wait for everyone to exhaust their meme projects/ports and models to get just a bit better. Feels like we're about to have something of a golden age when everyone starts really making the stuff they've always wanted.

Qwen 3.8 3 days ago

It's the one I am most excited for.

Over the past few weeks while using pro from them directly I have had an increasing number of responses that are obviously from a much, much better model. It is so good that the closed model dog and pony show is already spinning fud about "dark routing" and "stolen directly from fable"

Even at their new pricing it is a genuinely ridiculous amount of value. If you are the type of person who, very reasonably, does not have time to be trying out every model, and just want to use what seems to be the best currently... don't try it. You will be sick to your stomach with buyers remorse as you start to internalize just how much more you could have accomplished had you spent the first six months of the year giving them $1200 instead of OpenAI.