HN user

djhn

753 karma
  my hn username [at] jyl [dot] li
Posts3
Comments542
View on HN

(Apart from rent payments).

Privilege enables you to rent competence, historically by paying other people. The slop companies will now sell you a simulacrum of competence by the token.

The fact that competence can (could?) only be acquired through sustained effort over a long period of time is (was?) levelling the field.

Selling simulated competence perpetuates privilege, instead of dismantling it like you seem to claim.

Yes, and they’re bound to abuse it.

There’s a similar thing going on with emails. Dozens of services ”decide” that you need to update your email address, because ”they can’t reach you”. Many of them even stop sending you emails you explicitly subscribed to, perhaps to maintain an archive, ”because you don’t seem to open them”.

No, dear Linkedin and others, you’re reaching me just fine, and it’s none of your business whether, when and where I open them. Maybe I just read my emails offline and strip your tracking links (and avoid clicking on links in emails in general).

Inexplicably LinkedIn’s UX for changing the old email address, the one they cannot reach you at (!), to a new email address, starts with confirming your current email address (THE ONE THEY CANNOT REACH YOU AT). Brilliant.

Surely a big part of the reason is steering leverage relative to the kerb.

You are able to one-shot reverse parallel park into a much narrower gap without hitting the bumpers of the cars around you or getting onto the pavement.

I find SQL and data(bases) in general to be LLM’s Achilles’ heel. Databases are rarely under version control, so the training data only has one half of the knowledge.

My comments are more in the context of OLAP queries and other non-normalised data often queried via SQL.

I train non-LLM transformer models on (older and rarer) datasets, and automating the ingestion of sprawling datasets with hundreds of columns, often in a variety of local languages and different naming conventions adopted over decades, with quite a few duplicated columns…. The LLMs perform badly, it’s nigh impossible to test (for me as a user in prod) and it’s nearly impossible for the LLM companies to test (in training) to RLVR and RLHF this.

Bullshit Machines 2 months ago

I think I know the examples you’re talking about. They don’t show much in terms of reasoning.

The Erdős problems have turned out to be largely brute force or finding older results.

The Feb 2026 GPT-5.2 theoretical physics paper was a result of “dialogue between physicists and LLMs”, called “grad student level” by experts in the field, used a “custom harnessed” “internal OpenAI” model with “20 hours of reasoning”. Quotes from OpenAI blog.

The Matthew Schwartz physics paper with Claude this March involved “51,248 messages across 270 sessions, producing over 110 draft versions and consuming 36 million tokens”, and the actual contribution was Schwartz finding an error in Claude’s solution.

I’m afraid your numbers, all over 99%, are anchoring the conversation to an unreasonably high quality level.

I would have personally gone for 75%, 85% and 95%, which are all still best case scenario answers.

Had I taken on chatbot advice on electronics or chemistry I’d have died every couple of weeks (doing some hands-on real world R&D in my basement as a distraction from software).

Bullshit Machines 2 months ago

I’m genuinely interested in someone countering the following evidence that supports the authors.

Plane of words: broadly correct. Everything is flattened to tokens and token sequences, and the training data is dominated by text tokens.

Reasoning: CoT tokens are mostly just tokens, more appropriately called intermediate tokens, and are largely disconnected from the end result. Including them improves the end result (user satisfaction), but does not imply reasoning. See for example Turpin 2023, Mirzadeh 2024, Pournemat 2025, Palod 2025.

Synthesising evidence: You can achieve SOTA summaries with LLMs, but this involves, for example, using a harness to generate dozens of summaries with different models, separately using some kind of vector embedding model to compare results to the original, and selecting the best match. This is not how most people are using LLMs for summaries. While this is being slowly RLVR’d in post-training, a one-shot naive summary underperforms more complex methods significantly.

Ok, fair. I incorrectly assumed you meant resizing static images to create a lower resolution preview image.

Video thumbnails are a different beast altogether. And you might want to double check your assumptions about security considerations. If any of your ffmpeg, opencv, pyscenedetect code is running on your server, it might well be exploitable.

Agent Skills 3 months ago

I think I should also clarify, I work in the training of encoder-decoder transformer models. Before the ChatGPT era I worked on on encoder-only transformer models. I'm not unfamiliar with the literature and general discourse. I just do not use LLMs for programming.

Agent Skills 3 months ago

I can take on a slightly weaker form in good faith: professionally it’s a non-starter until private, open source inference can be self-hosted and the ROI is clear enough to invest in that.

And on the ROI side, trying things out regularly, I haven’t seen the positive ROI in the limited time I’ve dedicated to exploring the tools. I’ve restricted experimenting to 4 hours per month, because spending more than 2.5% of the month chasing productivity improvements that realistically seem to be 10-20%, will quickly eat into those gains. After accounting for token costs, it ends up being a wash.

Agent Skills 3 months ago

What kind of code is infrastructure in this context? Devops in a software company? Internal tooling in a software org?

At many (otherwise) world-leading facilities even just reviewing the patient history is a slog. There is rarelly any ability to keyword search the records or even filter the records by location, title and occupation of the healthcare professional making it, etc. Especially very ill people will have hundreds and hundreds of recent entries.

And stepping through those entries isn’t like browsing a modern local-first app [1], where you will just scroll through dozens of entries in milliseconds. It’s not like the slightly older and slightly slower Gmail interface. You’re clicking on each record and waiting 400ms-3s for it to load, as if instead of a 25Gb fiber connection you’re on dialup requesting the record from Epic’s headquarters in the US and proxying them via Australia.

[1] https://bugs.rocicorp.dev/p/roci

I might have too French of an attitude towards parenting for American taste, but as long as the crying and screaming isn’t based on anything real (and as long as you’ve childproofed the children’s room well enough it shouldn’t be) the child will be fine and what y’all need is sufficient distance between the bedrooms, some nice, solid brick walls in between the the rooms and some earplugs.

What could a 6-24 month old possibly do from their bed in their room, to disturb your sleep in your bed in your room? Bring a trumpet to bed and badly play Miles Davis?

What happened to lights off, door closed, do whatever you want in complete darkness in the bed that you aren’t able to climb out of?

This is a fun analogy, even if it’s just novel to me.

With any kind of creative work for hire, from architecture to advertising, from jingles to commisioned sculptures, the client’s taste and budget, more than almost anything else, determine the outcome.

Take Cannes Lions as an example of a competition and awards ceremony that essentially exists to define what ’good taste’ means within that industry. The client’s team is prominently credited alongside the creative agency. They get to climb onto the stage for the speech and they have a voice on whatever video clip is made about the project.

Partly this is to encourage more ambitious and spendy work for the industry at large. But everyone involved certainly knows, that the same creative team, with the same creative idea, could have ended up making something much worse working with a different client team.

I can’t stand AI slop, yet I think I’ve unintentionally argued in favour of people creating it, as long as it’s… good by some measure?

Do these… ”groups”, which would rather you not lump chewing tobacco together with other tobacco products… manufacture and sell chewing tobacco?

Yet another reason to doubt claims that ”software is solved”.

Anthropic did retire an interview take-home assignment involving optimising inference on exotic hardware, because Claude could one shot a solution, but that was clearly a whiteboard hypothetical instead of a real system with warts, issues and nuance.

Throughout history people have taken precautions against ceilings disintegrating. One might even say, ”strong engineering controls”.

Some of the best known laws from the ~1700BC Babylonian legal text, The Code of Hammurabi, are laws 228-233, which deal with building regulations.

229. If a builder builds a house for a man and does not make its construction firm, and the house which he has built collapses and causes the death of the owner of the house, that builder shall be put to death.

230. If it causes the death of the son of the owner of the house, they shall put to death a son of that builder.

233. If a builder constructs a house for a man but does not make it conform to specifications so that a wall then buckles, that builder shall make that wall sound using his silver (at his own expense).

That doesn’t sound like ceilings never disintegrated!

I feel compelled to concur with fwip, dpark and breezybottom. LLMs and the chatbot interfaces built for these text generating models are very good at writing fiction, including writing fictional roles and acting out those roles. Don’t get too carried away by this fiction.