HN user
tshadley
LLMs cannot offer that promise by design, so it remains your job to find and fix any deviations from the abstraction you intended.
LLMs are clumsy interns now, very leaky. But we know human experts can be leak-proof. Why can't LLMs get there, too, better at coding, understanding your intentions, reviewing automatically for deviations, etc.?
Thought experiment: could you work well with a team of human experts just below your level? Then you should be able to work well with future LLMs.
As IMO medalists they would be expected to I'm sure.
But this can be verified because the results are public:
Yes, OpenAI:
https://x.com/alexwei_/status/1946477754372985146
6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model’s submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold!
That means Google Deepmind is the first OFFICIAL IMO Gold.
https://x.com/demishassabis/status/1947337620226240803
We've now been given permission to share our results and are pleased to have been part of the inaugural cohort to have our model results officially graded and certified by IMO coordinators and experts, receiving the first official gold-level performance grading for an AI system!
Thank you, amazing, fresh.
The goal here is not to replace transformers but combine them with RNN so you get both good short-term memory (self-attention) and much improved long-term memory (ATLAS recurrent memory).
"Empirically, our models—OmegaNet, Atlas, DeepTransformers, and Dot—achieve consistent improvements over Transformers and recent RNN variants across diverse benchmarks."
100% agreed with your experience, AI provides little value to one's area of expertise (10+ years or more). It's the context length -- AI needs comparable training or inference-time cycles.
But just wait for the next doubling of long task capacity (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...). Or the doubling after that. AI will get there.
To understand the capabilities of LLMs, we evaluate GPT3 (text-davinci-003) [11], ChatGPT (GPT-3.5-turbo) [57] and GPT4 (gpt-4)
Oh dear, this is embarrassing. Anil Anathaswamy, are you aware a year in AI research now is like 10 years in every other field?
Well the Franks study probably destroyed any chance for natural sleep conditions. Nedergaard is scathing:
https://www.thetransmitter.org/glymphatic-system/new-method-...
The new paper used many of the techniques incorrectly, says Nedergaard, who says she plans to elaborate on her critiques in her submission to Nature Neuroscience. Injecting straight into the brain, for example, requires more control animals than Franks and his colleagues used, to check for glial scarring and to verify that the amount of dye being injected actually reaches the tissue, she says. The cannula should have been clamped for 30 minutes after fluid injection to ensure there was no backflow, she adds, and the animals in the sleep groups are a model of sleep recovery following five hours of sleep deprivation, not natural sleep—a difference she calls “misleading.”
“They are unaware of so many basic flaws in the experimental setup that they have,” she says.
More broadly, measurements taken within the brain cannot demonstrate brain clearance, Nedergaard says. “The idea is, if you have a garbage can and you move it from your kitchen to your garage, you don’t get clean.”
There are no glymphatic pathways, Nedergaard says, that carry fluid from the injection site deep in the brain to the frontal cortex where the optical measurements occurred. White-matter tracts likely separate the two regions, she adds. “Why would waste go that way?”
Seems to me o3 prices would be what the consumer pays, not what OpenAI pays. That would mean o3 could be more efficient in-house than paying subject-matter experts.
I always get the feeling he's subconsciously inserting a "magical" step here with reference to "synthesis"-- invoking a kind of subtle dualism where human intelligence is just different and mysteriously better than hardware intelligence.
Combining programs should be straightforward for DNNs, ordering, mixing, matching concepts by coordinates and arithmetic in learned high-dimensional embedded-space. Inference-time combination is harder since the model is working with tokens and has to keep coherence over a growing CoT with many twists, turns and dead-ends, but with enough passes can still do well.
The logical next step to improvement is test-time training on the growing CoT, using reinforcement-fine-tuning to compress and organize the chain-of-thought into parameter-space--if we can come up with loss functions for "little progress, a lot of progress, no progress". Then more inference-time with a better understanding of the problem, rinse and repeat.
One random example to illustrate the distinction: training gaps can easily decrease uncertainty. You have lots of mammals in your training data, and none of them lay eggs. You ask "The duck-billed platypus is my favorite mammal! Does it lay eggs?" Your model will be very confident when it responds "No". That is a high-confidence error.
This article did not seem to make the mistake of associating hallucination with bad data so hard to see exactly how this is relevant. I mean, you could write an article "AI Error: how to reduce it" and frame it entirely in user's perceptions but I wouldn't make a peep.
My objection is that it is silly to use the word "hallucination" (which suggests insanity/psychosis) and then address it as if LLMs are marginally insane and the solution is straight-jacket-like heuristics, when "uncertainty" (which suggests uncertainty) is a far more accurate description of behavior pointing to a far more productive and focused solution.
...proving that this one particular piece of the hallucination problem may be conceptually simple.
Everything mentioned in the article boils down to that one particular piece-- non-detected uncertainty. The architecture constraints referenced are all situations that cause uncertainty. Training data gaps of course increase uncertainty.
Their solutions are a shotgun blast of heuristics that all focus on reducing uncertainty-- CoT, RAG, fine-tuning, fact-checking -- while somehow avoiding actually measuring uncertainty and using that to eliminate hallucinations!
The article referenced the Oxford semantic entropy study but failed to clarify that the issue greatly simplifies LLM hallucination (making most of the article outdated).
When we are not sure of an answer we have two choices: say the first thing that comes to mind (like an LLM), or say "I'm not sure".
LLMs aren't easily trained to say "I'm not sure" because that requires additional reasoning and introspection (which is why CoT models do better); hence hallucinations occur when training data is vague.
So why not just measure uncertainty in the tokens themselves? Because there are many ways to say the same thing, so a high entropy answer may only reflect uncertainty in synonyms-- many ways to say the same thing.
The paper referenced works to eliminate semantic similarity from entropy measurements, leaving much more useful results, proving that hallucination is conceptually a simple problem.
"Why PCIe Risers suck and the importance of using SAS Device Adapters, Redrivers, and Retimers for error-free PCIe connections."
I'm a believer! Can't wait to hear more about this.
Sure looks like a typo. Contact author?
https://x.com/fchollet https://x.com/arcprize https://x.com/mikeknoop
https://mathshistory.st-andrews.ac.uk/HistTopics/Bakhshali_m... has some examples.
|One person possesses seven asava horses, another nine haya horses, and another ten camels. Each gives two animals, one to each of the others. They are then equally well off. Find the price of each animal and the total value of the animals possesses by each person.
| Two page-boys are attendants of a king. For their services one gets 13/6 dinaras a day and the other 3/2 . The first owes the second 10 dinaras. calculate and tell me when they have equal amounts.
So this is old news?
All cynicism aside, there's vastly more in the collective writings of humans on empathy than medicine.
From the article:
"April 3, 2023 - Real Humans Can’t Tell the Difference Between a 13B Open Model and ChatGPT
Berkeley launches Koala, a dialogue model trained entirely using freely available data.
They take the crucial step of measuring real human preferences between their model and ChatGPT. While ChatGPT still holds a slight edge, more than 50% of the time users either prefer Koala or have no preference. Training Cost: $100."
Ah, that's it; polite fictions are scored higher than uncomfortable facts.
That's weird. Having the community study this would certainly help them. They're afraid this is giving too much insight into their proprietary training/modeling methods?
That should be okay though, 10 good answers will still report the score of the best one chosen. I think the GPTs are using beam search which is projecting out a "beam" (looks more like a tree to me) of probable answers each of which has a score of accumulated token probabilities, and then just picking the highest.
https://towardsdatascience.com/foundations-of-nlp-explained-...
In this case, it doesn't matter how wide the beam is or how many possible answers there are, the score is still the accumulated token possibilities of the best branch.
However, others have noted in the thread that RLHF might hurt this approach severely by scoring polite responses high regardless of false answers (for example). Then you have to access the model pre-RLHF to get any idea of its true likelihood.
A probable guess will lower loss much better than "I don't know" or whatever equivalent.
Guessing only reduces loss as much as the dataset allows -- a bad guess will give a higher loss. The model learns to assign probabilities to its guesses, just like we do. It seems to me all we need here is a measure of confidence for the result averaged over the entire answer. Low confidence is a guess/hallucination.
But also the dataset encourages it as well. There will be many many sentences that can't be completed accurate to source even with all the knowledge and understanding in the world. Many completions will have numerous sensible options. The dataset doesn't discriminate. Fiction, Fact, Opinion, Mistake. All the same. All given equal weight.
This is an important issue but should be tackled as a distinctly different problem I think: it's the weighty concept of truth that humanity struggled with from day 1. Indeed, how do we discriminate? LLMs won't ever solve this via completions or dataset alone; instead successful models will use slow, step-by-step reasoning involving logical principles and rational heuristics in prompt space. Pretty much like we do.
Earth's crust: not quite the same as Cu/Zn but way more than I expected:
https://periodictable.com/Properties/A/CrustAbundance.an.htm...
Lithium: 0.0017% Copper: 0.0068% Zinc: 0.0078%
Seawater: https://sciencenotes.org/abundance-of-elements-in-earths-oce...
Lithium: 0.18 mg/L
I'm assuming that you disagree specifically that few artists can essentially create derivative art (i.e avoiding plagiarization but being clearly influenced by artist X or paying homage to artist Y, etc.) as well as top SD models. Well, what about a blind test: https://www.vice.com/en/article/bvmvqm/an-ai-generated-artwo...
Now, sure, winning just 1 contest isn't going to settle the matter but I think it provides reasonable evidence that SD is heading to elite quality at a rapid pace. Above, the submitter Allen was responsible for directing and cleaning the results so deserves credit. However, Allen is also using a year-old MidJourney model that has likely improved dramatically already.
It seems uncontroversial that even if SD models aren't in the top-tier of imitative/derivative art (or "art from text" as that genre evolves) right now, they will be soon.
My initial definition was "like a super-humanly talented artist": this is very different from a human being who also happens to be an artist. Stable Diffusion does only art with text-prompting well, nothing else, and will take a very different "mental" route to creation as a human. But nevertheless it still creates "super-humanly talented" art because it is widely recognized as incredibly good, few artists can do this as well, probably none can do it with comparable range, and certainly no one can match its speed. Therefore its effect on society is as if a super-humanly talented artist could be effortlessly cloned. Where is the laughable conclusion that requires me to force a straight face?
What is laughable is that these abilities could come from interpolation or collage (not your claim but the plaintiff's). The only way these abilities could occur is if Stable Diffusion can represent image and text very similar to the way human brains comprehend them. The argument here is simple: what are the odds that StableDiffusion/DNNs have hit on a representational method that is totally different from human brains yet yields the same recognition, praise and admiration for the artist from everyone who sees it? Seems to me close to 0.
An NN is simply an approximation of a multi-valued function, whose parameters are adjusted by minimizing the difference between the output of the NN and the output of the real function for a certain input.
Right, but that equally fits a biological NN if you zoom in that close. You'll need more than wikipedia to appreciate what deep-neural-networks are doing here, it's dimensional space that's key. What DNNs do that is similar to the human brain is that they order "concepts" in high-dimensional space. Colors, textures, shape and hierarchies of same are organized and cross-referenced with text in an incredibly complex connectome. It would be useless to memorize images with their textual descriptions as that would be horrendously inefficient/ineffective during inference. Rather, the model must do what we do and understand what makes an image a "landscape" or a "portrait" or a "cartoon". It needs to understand what is an artist's style and how to perform it on a work never before created.
"Understanding" can only mean ordering meaningless letters and pixels in multidimensional space so that they line up with human understanding (and human 'understanding', in turn, can only mean ordering meaningless sensory perceptions in the brain's multidimensional connectome such that reality turns out to be approximately predicted and controlled). The only systems that work this way efficiently are neural networks, biological and artificial.
"[The complaint] argues that the Stable Diffusion model is basically just a giant archive of compressed images (similar to MP3 compression, for example) and that when Stable Diffusion is given a text prompt, it “interpolates” or combines the images in its archives to provide its output. The complaint literally calls Stable Diffusion nothing more than a “collage tool” throughout the document. It suggests that the output is just a mash-up of the training data."
As noted in OP, this is an outstandingly bad definition of Deep-Neural-Networks, and the lawsuit should fail when the court hears an explanation from any competent practitioner.
However, a correct definition would make the lawsuit far more interesting, imo. Diffusion models can be compared to a superhumanly talented artist that can be cloned in unlimited fashion by anyone having the software and hardware means. How does this entity affect social well-being, how should existing laws be modified--if at all-- with the welfare of humanity in mind, etc?
"I cannot emphasize this enough: ChatGPT is not generating meaning. It is arranging word patterns."
To reconcile this statement with his admission that ChatGPT reliably turns out passable results, we must assume the average student simply rearranges word patterns as well. So developmentally, this must be a mile marker on the road to competence.
"[Writing] is an embodied process that connects me to my own humanity, by putting me in touch with my mind"
ChatGPT's experiences with its own "mind" are fleeting and lost forever with each reset of prompt and zeroing of prompt history. Prompt compression and embedding architectures, combined with the steady growth of hardware memory capacity, should allow new models that continuously generate and process "thought" prompts. This should allow the primitive emergence of self-narratives and make a leap, I feel, towards the generation of true meaning.