HN user

encomiast

920 karma
Posts2
Comments133
View on HN

So then you need to ask: Why did they use a deliberately faulty LLM? They could have easily used a mainstream LLM from the past 18 months and it probably would have been less work to do so. But then they would not have that headline. The answers from the LLM would have likely made the participant's answers more accurate, not 3x less accurate. But then they would not have this juicy headline.

I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice".

Here is what the study says:

"The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."

So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.

If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.

Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study that told experts to lie.

A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?".

Oh wow, that is impressive and gorgeous. I love that I can get a bespoke horsehair leather brush for my binder. Not sure I would be allowed to write in it without a matching Japanese whiskey.

I've been trying binders instead of notebooks (especially the mini-binders from folks like Maruman). Binders allow you to easily remove and add pages (and tabs, dividers, etc). This gives you the satisfaction of making something aesthetically pleasing, while also knowing you can move a 'messy' page somewhere else.

AI 2040: Plan A 12 days ago

Democracy came about because it was optimal for a power seeking government, not out of the kindness of their heart

It's not clear in this context what you actually mean by "government." You are assigning agency to something in a way that seems like a reification. While a bureaucracy can seem to have a life of its own, isn't it generally people who seek power?

I guess it's obvious, but haven't really thought about this. Back in the day we had media companies and advertisers. Magazine publishers/newspapers/TV Studios on the one hand and the Leo Burnetts/Ogilvy/Wieden+Kennedy on the other. There was always a tension between what advertisers wanted and what creatives wanted. The studio/publishers knew that bowing completely to advertisers was a fast track to making crap and the advertisers knew they needed to remind the publishers that they are only as good as their audiences.

With Meta we these have more or less merged. It's not really clear now who the customer is. They need to attract both eyeballs and money. I'm not sure what the long-term consequence of that is, but my hunch is it leads to accounting and advertising winning since they actually generate money. And the result is junk creative.

Meta seems to be completely tone-deaf when it comes to products and features. I wonder if this is a function of a single person, surrounded by people with little incentive to disagree, having all the decision making power in a company.

Show HN: 18 Words 14 days ago

Same. Stopped after two words. Gave me the same feeling as taking a timed software developer screening.

Of course you might be right. But if we look at our past ability to predict the future by extrapolating from the present and recent past, you might also say it is naive to think that _this time_ is somehow different, that this time the future is clear.

Also, I think you've set up a straw man. I haven't seen anyone arguing that this future isn't possible. What I see is resistance to the idea that we can have any certainty about the future by drawing a straight line from the past.

ChatGPT launched in November, 2022. Opus in December last year. Do you see where this is going?

Maybe this is tongue-in-cheek, but in case it's not.

No, nobody does. Consider that a person in their lifetime could have seen the Wright Brothers first flight in 1903, then the first jet engine in 1939, then commercial jet travel in the 1950s. That is an amazing 50 years for air travel. If you were living in 1950 and were to extrapolate where air travel would be by 2026, you would think we would be taking routine trips to Mars. Instead what we got was TSA, cramped seats, and better safety. But "progress" leveled off dramatically. Commercial air flight has looked pretty much the same for my entire life.

We have no idea how far transformers and LLMs will take us. It certainly isn't obvious where this is going.

If anyone is interested, here is a link where you can download the study: https://www.researchgate.net/publication/221770027_Correlati...

I few interesting bits — it does involve cursive, but it's Arabic and it's graded on a rubric that includes things like "Presenting the beauty aspects of Arabic writing'. Also, given a sample of 71 students and a p<0.001 means the correlation coefficient only needs to be around 0.40 which means handwriting and drawing may only explain about 16% of the variance of these dental skills. That's not nothing, but given the subjective nature of the test and the confounders (does this handwriting sample really measure motor skills or maybe it measures care and attention to detail, or conscientiousness), I'd be a little wary of using this to argue for education policy.

Still, glad you posted it and glad I read it. It interesting.

It's not just CORS that's hard to understand. Many (most?) developers don't really understand the threat model. And even when it's explained it hard to see why it's a big deal. Part of this is that backend developers usually have to configure CORS and it's not an access privilege protection. From the point of view of the backend it doesn't seem to matter. Bad guys can't get it. From the point of view of the front-end it's often seen as a nuisance.

The article does a nice job giving a concrete example.

Pre-2022 Books 1 month ago

Yeah, Amazon has been garbage for a while with old books, especially classics. It seems like everything is just keyed off title/author so it takes a ton of effort to make sure you are getting the edition/translation you want. It's 100x worse with Kindle where it looks like some random cheap scan into-kindle format has 2k five star reviews. And of course user reviews where they seem to mix all the reviews together for various editions.

If you are really talking about dependencies, I’m not sure you’ve really thought this all the way through. Are you inspecting every line of the Python interpreter and its dependencies before running? Are you reading the compiler that built the Python interpreter?

I think this is a useful way to look at things. We often point out that LLMs are not conscious because of x, but we tend to forget that we don't really know what consciousness is, nor do we really know what intelligence is beyond the Justice Potter Stewart definition. It's helpful to occasionally remind ourselves how much uncertainty is involved here.

Point taken. Still, isn’t an activity like learning a new library, language, or platform a fundamental part of being a software developer? Haven’t we all complained at some point about companies hiring react developers because we all know the real skill is the ability to pick up new things. And to be clear, this isn’t moral panic, it’s a concern that we may end up in a future where people don’t know how systems work anymore and we are dependent on two or three companies and their data center moats to maintain any technology.

uh...okay. The legend of Faust is the classic work where a person sells his soul to the devil for power/knowledge/pleasure. Goethe has Faust make a wager with the Mephistopheles: show me the good life (pleasure, power, knowledge, whatever) it will never be enough to make me stop striving, to make me want to linger. If you can do that my soul is yours. To me it reads a _lot_ like our contract with AI.

"so as long as I maintain my ability to reason about code…what’s the issue?"

It seems like that is the open question. The article suggests that people don't maintain this ability:

"The AI group scored 17% lower on conceptual understanding, debugging, and code reading. The largest gap was in debugging, the exact skill you need to catch what AI gets wrong. One hour of passive AI-assisted work produced measurable skill erosion."

From my own (anecdotal) experience I am seeing a lot more cases of what I call developer bullshit where developers can't even talk about the work they are vibe-coding on in a coherent way. Management doesn't notice this since it's all techno-bable to them and sounds fancy, but other developers do.