Thanks, interesting reference. However, their analysis doesn't tell us much about the quality of Grokipedia. Would be more interested in something like hallucination density, but I know of no way that could be measured.
HN user
13years
Software Engineer 30+ years. Fortune 100 companies.
Writes at https://www.mindprison.cc
We suck at measuring ourselves.
That is a certainty. I was once asked to calculate how much time we would save through our companies code reuse program. I read all the material on estimating savings, but then proved it was all ridiculous.
I came across a study that attempted to estimate how long it took to build libraries that had already been built. In this case, there were no unknowns, you had the entire code. Estimates were off by orders of magnitude. If we can't estimate the work when the work is already done, how could we ever estimate the work when we know less?
Not sure how you get around the contamination problems. I use these everyday and they are extremely problematic about making errors that are hard to perceive.
They are not reliable tools for any tasks that require accurate data.
A philosophical lens can sometimes help us perceive the root drivers of a set of problems. I sometimes call AI humanity's great hubris experiment.
AI's disproportionate capability to influence and capture attention versus productive output is a significant part of so many negative outcomes.
Yes, that's an excellent description.
I think it is creating a growing interest in authenticity among some. Although, it still feels like this is a minority opinion. Every content platform is being flooded with AI content. Social media floods it into all of my feeds.
I wish I could push a button and filter it all out. But that's the problem we have created. It is nearly impossible to do. If you want to consume truly human authentic content, it is nearly impossible to know. Everyone I interact with now might just be a bot.
Myself I believe technology and eventually AI were our fate once we became intelligence optimizers.
Yes, everyone talks about the Singularity, but I see the instrumental point of concern to be something prior which I've called the Event Horizon. We are optimizing, but without any understanding any longer for the outcomes.
"The point where we are now blind as to where we are going. The outcomes become increasingly unpredictable, and it becomes less likely that we can find our way back as it becomes a technology trap. Our existence becomes dependent on the very technology that is broken, fragile, unpredictable, and no longer understandable. There is just as much uncertainty in attempting to retrace our steps as there is in going forward."
AI is not inevitable fate. It is an invitation to wake up. The work is to keep dragging what is singular, poetic, and profoundly alive back into focus, despite all pressures to automate it away.
This is the struggle. The race to automate everything. Turn all of our social interactions into algorithmic digital bits. However, I don't think people are just going to wake up from calls to wake up, unfortunately.
We typically only wake up to anything once it is broken. Society has to break from the over optimization of attention and engagement. Not sure how that is going to play out, but we certainly aren't slowing down yet.
For example, take a look at the short clip I have posted here. It is an example of just how far everyone is scaling bot and content farms. It is an absolute flood of noise into all of our knowledge repositories. https://www.mindprison.cc/p/dead-internet-at-scale
that the author also tripped over
The evidence for unfaithful reasoning comes from Anthropic. It is in their system card and this Anthropic paper.
https://assets.anthropic.com/m/71876fabef0f0ed4/original/rea...
But it is not an illusion, and the answers make no sense. In some cases the models pick exactly the opposite answer. No human would do this.
Yes, outside the training patterns is the point. I have no doubt if you trained LLMs on this type of pattern with millions of examples it could get the answers reliably.
The whole point is that humans do not need data training. They understand such concepts from one example.
Take a look at this vision test - https://www.mindprison.cc/i/143785200/the-impossible-llm-vis...
It is an example that shows the difference between understanding and patterns. No model actually understands the most fundamental concept of length.
LLMs can seem to do almost anything for which there are sufficient patterns to train on. However, there aren't infinite patterns available to train on. So, edge cases are everywhere. Such as this one.
The bar you asked for was "meaningful progress". And as you state, "both are very helpful metrics", it seems the bar is met to the degree it can be.
I don't think we will see a definitive test as we can't even precisely define it. Other than heuristic signals such as stated above, the only thing left is just observing performance in the real world. But I think the current progress as measured by "benchmarks" is terribly flawed.
I think it is actually worse than that. The hype labs are still defiantly trying to convince us that somehow merely scaling statistics will lead to the emergence of true intelligence. They haven't reached the point of being "surprised" as of yet.
most people can’t reliably interpret the meaning of complex or unfamiliar text
But LLMs fail the most basic tests of understanding that don't require complexity. They have read everything that exists. What would even be considered unfamiliar in that context?
RFK Jr. is antivax because he misunderstands all the information he sees about the benefits of vaccines.
These are areas where information can be contradictory. Even this statement is questionable in its most literal interpretation. Has he made such a statement? Is that a correct interpretation of his position?
The errors we are criticizing in LLMs are not areas of conflicting information or difficult to discern truths. We are told LLMs are operating at PhD level. Yet, when asked to perform simpler everyday tasks, they often fail in ways no human normally would.
We are capable of much more, which is why we can perform tasks when no prior pattern or example has been provided.
We can understand concepts from the rules. LLMs must train on millions of examples. A human can play a game of chess from reading the instruction manual without ever witnessing a single game. This is distinctly different than pattern matching AI.
Essentially, pattern matching can outperform humans at many tasks. Just as computers and calculators can outperform humans at tasks.
So it is not that LLMs can't be better at tasks, it is that they have specific limits that are hard to discern as pattern matching on the entire world of data is kind of an opaque tool in which we can not easily perceive where the walls are and it falls completely off the rails.
Since it is not true intelligence, but a good mimic at times, we will continue to struggle with unexpected failures as it just doesn't have understanding for the task given.
or how we would measure meaningful progress in this direction.
"First, we should measure is the ratio of capability against the quantity of data and training effort. Capability rising while data and training effort are falling would be the interesting signal that we are making progress without simply brute-forcing the result.
The second signal for intelligence would be no modal collapse in a closed system. It is known that LLMs will suffer from model collapse in a closed system where they train on their own data."
They are different contexts of errors. Take any of these humans in your example, and give them an objective task, such as take any piece of literal text and reliably interpret its meaning and they can do so.
LLMs cannot do this. There are many types of human failures, but we somewhat know the parameters and context of those failures. Political/emotional/fear domains etc have their own issues, but we are aware of them.
However, LLMs cannot perform purely objective tasks like simple math reliably.
we usually say that we don't know
I think this is one of the distinguishing attributes of human failures. Human failures have some degree of predictability. We know when we aren't good at something, we then devise processes to close that gap. Which can be consultations, training, process reviews, use of tools etc.
The failures we see in LLMs are distinctly of a different nature. They often appear far more nonsensical and have more of a degree of randomness.
The LLMs as a tool would be far more useful if they could indicate what they are good at, but since they cannot self reflect over their knowledge, it is not possible. So they are equally confident in everything regardless of its correctness.
where we will have personal agents that are AGI for a huge range of use cases
We are already there for internet social media bots. I think the issue here is being able to discern the correct use cases. What is your error tolerance? For social media bots, it really doesn't matter so much.
However, mission critical business automation is another story. We need to better understand the nature of these tools. The most difficult problem is that there is no clear line for the point of failure. You don't know when you have drifted outside of the training set competency. The tool can't tell you what it is not good at. It can't tell you what it does not know.
This limits its applicability for hands-off automation tasks. If you have a task that must always succeed, there must be human review for whatever is assigned to the LLM.
GPT o3 is a better writer than most high school students at the time of graduation.
All of these claims, based on benchmarks, don't hold up in the real world on real world tasks. Which is strongly supportive of the statistical model. It will be capable of answering patterns extensively trained on. But is quickly breaks down when you step outside that distribution.
o3 is also a significant hallucinator. I spent quite a bit of time with it last weekend and found it to be probably far worse than any of the other top models. The catch is that it its hallucinations are quite sophisticated. Unless you are using it on material for which you are extremely knowledgeable, you won't know.
LLMs are probability machines. Which means they will mostly produce content that aligns to the common distribution of data. They don't analyze what is correct, but only what is probable completions for your text by common word distributions. But when scaled to incomprehensible scales of combinatorial patterns, it does create a convincing mimic of intelligence and it does have its uses.
But importantly, it diverges from the behaviors we would see in true intelligence in ways that make it inadequate for solving many of the kinds of tasks we are hoping to apply them to. The being namely the significant unpredictable behaviors. There is just no way to know what type of query/prompt will result in operating over concepts outside the training set.
Certainly random chance exists for discovery. But most revolutionary type discoveries come from deep understanding of the context.
The contribution of LLMs to knowledge is more like that of search engines. It is still the human which possesses understanding that ultimately will be the principle source of innovation. The LLM can assist with navigating and exploring existing information.
However, LLMs have significant downsides in this regard too. The hallucination problem is no joke. It can often mislead you and cause a loss of time on some tasks.
Overall, they will be somewhat useful in some manner, but substantially less so than the present hype machine suggests.
True, but OpenAI did leverage it for all that could be gained. There was no hesitation on their part to consider if it was ethical.
It is quite ironic that OpenAI loosens its rules, promoted Ghibli images, apparently in direct opposition to the viewpoints of Ghibli's founder, while also pursuing DeepSeek for using OpenAI data without permission.
FYI, some of my further elaboration on the topic:
"Everything that you loved for its uniqueness, craftsmanship, cultural significance, will be mass-produced until you despise seeing it."
https://www.mindprison.cc/p/studio-ghibli-style-ai-art-crisi...
Mention a few here, but my intent of writing this was mainly a warning of don't put much weight into any AI analysis.
If I ask sonnet what's under my bed it tells me it can't know and tells me to look under it myself.
The problem with most such questions is that these answer are likely patterns from training data. It is a typical reply.
The calculator question was interesting because the training data is unlikely to have such dialog as typical. People don't typically ask for a calculator or mention it for simple problems. Everyone has one and its use is somewhat implied.
I tried some variation of "provide accurate answers" or "accuracy is important". These did not result in the model asking for or mentioning a calculator. But as we know, results can be partially random and not always consistent especially in areas lacking strong patterns.
If I mentioned a calculator myself as part of a conversation, it would sometimes mention the need of a calculator. But every time we add more context, we are changing the probabilities for what will be generated.
We know the training data has the associations for LLM poor at math and calculator. But the references are weak. With some changes in prompting it makes the association.
For other examples of weak data and how LLMs respond, checkout these other tests I did - https://www.mindprison.cc/p/the-question-that-no-llm-can-ans...
It didn't choose to look for a calculator. LLMs that invoke tools were explicitly trained to do so. If tools are present, it will always attempt to first find a tool to satisfy the prompt.
So if tools are present, by training it will infer the intent to use the tool and not because it understands it is itself deficient in that ability.
So what we would expect to see with a LLM without tools enabled, is that it suggests that you give it access to a calculator.
If we develop real intelligence, it will be surprising. It won't just answer questions. It will tell us we are asking the wrong questions.
Sure, you can train LLMs to use tools or provide instruction in a prompt to do so.
However, it doesn't know to use the tool intuitively without explicit action to do so. It doesn't discover that is the best solution on its own. That is what is relevant. We can make them use tools. It is useful to do so.
The important takeaway is that LLMs capabilities, as useful as they may be, are not intelligence in the way they are being promoted.
Really though I think the author is highlighting that LLMs are not being used efficiently at the present moment.
Yes, that is a key point. It isn't to say they are useless tools, but that they aren't intelligent tools and that has significant meaning for what tasks we think they are appropriate for.
Unfortunately, nearly everyone has misinterpreted the intent as showing LLMs can't use tools. The point is about how LLMs work differently than most think that they do.
I'm the author.
Trying to say “this should just happen from the data” is silly, it isn’t how any of this works. It’s not how you learned things, and it’s not how LLMs-as-chatbots work.
Yes, that was the entire point of the article. We are in agreement.