HN user

code_biologist

1,479 karma
Posts3
Comments486
View on HN

I do mean "grounding" as seeking the primary source of information by hand as my sibling commenter said.

If an LLM says something about history, have it give you source documents so that you can verify. If it says Earth's gravitational constant is 98 [sic] m/s^2, have it generate a prediction of a timed 1m paperclip drop test, and go do it yourself.

It's perhaps excessive in the case of gravity, but for things that matter like medical stuff I tell LLMs to expose their step by step causal reasoning / inference, and then check whether I buy it myself. SotA models' biochemistry + neurochemistry is often pretty solid. Often their answers are based on a slightly hallucinated understanding if it's a spatial anatomy question. It's fun when you get some research paper citations and then upon reading them decide you don't agree with the underlying research (eg. fMRI studies often make broad / interesting claims based on small sample sizes or sample populations that are clearly biased).

To be clear, I don't think that the LLM's self explained reasoning in a single turn is how it actually arrived at a given conclusion, but I just want to expose and bind the final conclusion to a verifiable epistemic chain.

Writing and videos produced before 2022 are the information equivalent of low-background steel (steel produced before the atom bomb era). Not all are precious, but they are important. AI influences are so pervasive at this point, even in informational writing from domain experts.

Relentless grounding is my personal solution for my AI epistemic crisis, but it's an expensive solution in terms of time and effort, and triage is hard too.

My SO is on the spectrum and likes to align cups and things along the edges of tables. I ask her not to be an edgelord in the most egregious cases, but I've also gotten better at not stressing over "it's on the edge".

That cup is still not ok lol.

This isn't enough for me as 82kg mid 30s man. I will lean out by 3-5 kilos and lose strength for lifting weights.

What fats do you put on your potatoes + asparagus + vegetables, plus cooking fats for the meats? No idea how many cals "enough olive oil to lubricate" is. I consciously use cooking fats to creep in more healthy calories to sustain me.

I agree that never having the argument take place textually is important for LLM performance and behavior. I still think we’re investing the same time and intellectual energy arguing with the model, in going back and restructuring context and prompting to head off / pre-answer a refusal.

So you take action and put in more effort to cater to the LLM to get it to do what you want, but it's not arguing because there's no record of it in the chat? Presumably you put in what you would have written in the counter-argument into the new chat, just ahead of the LLM refusal? And this isn't arguing?

I've seen exactly this behavior on claude.com with no system prompt with Opus 4.8 specifically, especially around chronic illness stuff where there's established mainstream medicine dogma and reddit / internet communities with alternate causality theories and treatment approaches (PMDD and MCAS-adjacent illness). 4.6 is happy to analyze and consider them, 4.8 really doesn't like the alternate theories and treatments.

Once it's in this loop, Opus 4.8 digs in so aggressively it's structurally incapable of conceding a provided detail as correct, even if it's conceded and agreed with everything backing that detail. Like actually, structurally incapable. I've even baited it into arguing with itself when I've "conceded" its original concern tolling hard, and then the model needs to continue to be the "voice of reason" and it will argue against its original concern because I, the user, said it.

How difficult it is to resist "someone is wrong on the internet" is a perennial joke. Turns out it doesn't really matter who/what is on the other side if they seem human-like.

Leaving Mozilla 1 month ago

What? I just want to share cat pics, video clips, and memes with my friends and respond to their stuff with not-inline emojis.

It's been interesting to see how aggressively some reasoning models like to "reason" by analogy. They love to say things like "it's like a CPU" or "it's like a highway", and then they start to make logical leaps based off that rather than just using it for user explanation. Gemini 2.5 and 3.1 Pro have been particularly bad for this type of behavior. Telling models to "speak as though you are a physiologist considering the case with an expert colleague" gets them to "reason" using a more correct linguistic substrate.

The Opus models over the last year doesn't seem as vulnerable to this type of behavior and I've noticed the "identify as expert" prompt tricks aren't as meaningful there.

Your language is ambiguous — your horror is in reference to natural gas turbine generators (used at these installations) and not gasoline generators (like in a home context)?

Why the horror? I'd prefer the gas remain in the ground, but given the gassy production of US shale oil, I guess I'd rather it be used for this than just flared. I am frustrated that pollutant emissions aren't being policed, and also that the sudden turbine demand plus supply chain issues mean using aeroderivative turbines that are quite a bit less efficient than more complex combined cycle turbines.

https://www.energy.gov/hgeo/how-gas-turbine-power-plants-wor...

I'll admit that I miss having access to the ChatGPT 4.5 "absolutely gigantic model" with enough tuning to make it sane and useful. The RLVR models are superb for actual tasks in those RLVR domains, but that fine tuned view of the world as a verifiable problem to solve makes them feel worse for touchy feely stuff. Even for medical consultation and diagnosis, RLVR model's urge to reach a conclusion often is a liability.

"If there is anything the nonconformist hates worse than a conformist, it’s another nonconformist who doesn’t conform to the prevailing standard of nonconformity." - Bill Vaughan

As a bear that's been very confused by markets failing to exhibit mean reversion in 2019 and 2022, and now with the Hormuz energy crisis, I've thought a lot about this. There's a lot of new things happening. Fed/QE intervention that has never stopped, just been more well disguised. Fiscal/government spend intervention. I think Mike Green's work on the rise of passive investing is really good, in particular explaining how it prevents mean reversion in absence of changing net cash flows into passive instruments. Passive will also induce or worsen the bust if net cash ever starts to flow out passive. Green's youtube interviews are great.

All to say, your SO's dad would have been right at any point prior to the current financial cycle. Knowing what's changed doesn't make forecasting easier though.

Anybody got tips for making custom signatures work? I'm trying to rescue Unreal Engine game save files off an EXT4 drive where the containing directory was deleted. I put the `sav 0 "GVAS"` (GVAS is the UE save magic value) in my `.photorec.sig`. `fidentify` correctly works on reference UE saves I have. Grepping the raw partition device finds many hits for GVAS, but photorec runs and doesn't recover any files...

[1] https://www.cgsecurity.org/testdisk_doc/photorec_custom_sign...

People don't like the phrase enshittification, but the process Doctrow describes is so accurate (serve the users, then serve business customers at the expense of users, then serve the platform at the expense of users and business customers) it's hard not to see it everywhere. Phone platforms fit the template exactly, sadly.

Chat, is this real? I've seen this guy pop up on youtube. I assume he's a Chinese state mouthpiece as he's a westerner in the mainland with a very pro-China spin (substack recommended the other posts below), but I'm curious how strong the factual basis for this reporting is.

China's factories are in another world - Mar 23, 2025

Chinese factories build fire trucks for under $400,000 in six weeks. In the US, it's $2 million in 4 years - Apr 19, 2025

Iran is blowing up $500 million radars. China's export bans mean they are gone forever. - Mar 16, 2026

Goodbye to Sora 4 months ago

The Occam's Razor position (Sora was the most expensive to operate, least monetizable model) seems like a simpler explanation. The legal costs/difficulty on top of "most expensive" are just the cherry on top.

I'm totally with you personally, but sometimes doing the actually hard part is fun. Type 2 fun.

Long ago I took a CPU architecture class and we implemented designs in Verilog as a final project. Apparently people who took the class in the late 90s (before my time) could actually tape-out their designs and pay a few hundred dollars to get fabbed chips as part of a multiproject wafer. I was always curious if those chips actually worked, or just looked pretty.

N=1, but I’ve been doing low carb paleo for 15+ years, from about age 20 to my current late 30s. I live off of butter, tallow, and lard. My weight has only crept up when I’ve eaten a lot of processed food. I get quite lean even with high fat if I fast more frequently or dip into ketosis. I’m trying to pack on some extra muscle with weight lifting right now and it’s not easy to get enough clean calories short of eating spoonfuls of (happy, pastured) bacon grease.

All I’m trying to say is that butter isn’t the enemy. Maybe commercial dairy production practices are the enemy, won’t argue with that.

+1. I wish Gemini 2.5/3 Pro's "personality" and long context handling wasn't so erratic, because the medical stuff in there is great. Whatever they did to produce the MedGemma models is clearly built on a strong baseline. I haven't had need to try using MedGemma on x-ray imagery, but I'd be curious to hear results — imagery diagnostics is part of what it's built for.

Opus 4.5 seems good too, though getting dumber. OpenAIs fine tuning is clearly built to toe the professional medical advice line, which can be good and bad.

I had progressively worsening pelvic floor pain issues that AI helped me with and are now in remission/repair. My decade of interaction with multiple urologists and clinicians could be characterized as repeated and consistent "pretty obvious oversight from the healthcare practitioners".