HN user

Last5Digits

238 karma
Posts4
Comments81
View on HN

The fact that you refuse to engage with my points tells me otherwise.

You're drawing meaningless distinctions, anyone who has ever used Cyc will tell you that it makes massive mistakes and spits out incorrect information all the time.

But that is even true of humans, and every other system you can imagine. Facts aren't these magical things living in your brain, they're information with a high probability of accurately modeling reality.

When someone tells you x happened in y at time z. Then that only becomes a fact if the probability of the source being correct is high enough, that's it. 99% of all of your knowledge is only a fact to you because you extracted it from a source that your heuristics told you is trustworthy enough. There is never absolute certainty, it's all just probability.

Understanding comes in many forms. My uncle will not be able to model the fuel flow in his car's engine using the Navier–Stokes equations, yet he can still drive better than me. When it comes to LLMs, an understanding of the transformer architecture is wholly unnecessary to develop a good model of their capabilities and pitfalls. HN commenters tend to lack both a technical and abstract understanding of LLMs, while non-tech people tend to only lack the former.

At this point, I strongly urge you to think about what could possibly change your mind. Because if you can't think of anything, then that means that this opinion is not founded on reasoning.

The text LLMs produce is not just plausible in a "looks like human text" sense, as you'd very well know if you actually thought about it. When ChatGPT generates a fake library that looks correct, then the library must seem sensible to fool people. This can't be just a language trick anymore, it must have a similarity to the underlying structure of the problem space to look reasonable.

What's your definition of "correct" then? If a system is "accidentally correct" the majority of the time, when does it stop becoming "accidental"? You cannot trust any system in the way you want to define trust. No human, no computer, no thing in the universe is always correct. There is always a threshold.

I do research with LLMs all the time and I trust them, to a degree. Just like I trust any source and any human, to a degree. Just like I trust the output of any computer, to a degree. I don't need to verify everything they say, at all, in any way.

Genuine question, how do you think an LLM can generate "bullshit", exactly? How can it be that the system, when it doesn't know something, can output something that seems plausible? Can you explain to me how any system could do such a thing without a conception of reality and truth? Why wouldn't it just make something up that's completely removed from reality, and very obviously so, if it didn't have that?

Why do you feel the need to arbitrarily ascribe some ideology to random people based on one comment? I'm not "techno-utopian" in any sense of the word; I believe that the current AI development is highly risky and that we need to take careful measures such that society at large is prepared for the changes it may bring.

The "wide range of opinions" I see on HN are largely misinformed: they either lack the necessary technical understanding of current LLMs or are attempting to spin up some crackpot philosophical distinctions lacking in any rigor or consistency. I've never claimed that LLMs are perfect and I'd love to discuss their flaws! Believe it or not, that's why I continue to read these threads - to find genuinely informed takes contradicting my own.

Most people outside of tech tend to have no bias against or for LLMs, which gives them a leg up in finding consistent opinions about their capabilities. They tend to inform themselves with an open mind, which allows them to put things into context. Tech people have an immediate negative bias, because the implication of any system being able to write even a single line of code is an immediate intellectual threat. Therefore, things are interpreted maximally negatively.

For example, all of the talking points you mentioned are completely irrelevant unless interpreted with maximal negative bias:

- Text prediction is a general problem, being good at it requires understanding, reasoning and any other intellectual property you believe to be unique to humans.

- Every single system in existence is highly dependent on the data it uses to model the world, humans are no exception to this.

- The enormity of data required by any modern LLM is massively dwarfed by the enormity of data that was required by evolution and human civilization to get to this point.

- The energy requirements of modern LLMs are environmentally irrelevant when compared to literally any industry in either manufacturing, transportation or entertainment. We justify immensely more environmental damage for far less utility every single day.

After the giant media carousel last year, most people know what an LLM is, and the intuitive understanding they built from that reporting is way more accurate than what I have seen here. I have asked relatives and even acquaintances about just that. And as I have stated in my comment, their understanding is vastly better than that of HN.

No, the answers aren't just "plausible", they are correct the vast majority of the time. You can try this for yourself or look at any benchmark, leaderboard or even just listen to the millions of people using them every day. I fact check constantly when I use any LLM, and I can attest to you that I don't just believe that the answers I'm getting are correct, but that they actually are just that.

But they apparently actually don't get better even though every metric tells us they do, because they can't? How about making an actual argument? Why is correctness "not a property of LLMs"? Do you have a point here that I'm missing? Whether or not Kahneman thinks that there are two different systems of thinking in the human mind has absolutely no relevance here. Factualness isn't some magical circuit in the brain.

No such thing can exist.

In the same way there can exist no piece of clothing, piece of tech, piece of furniture, book, toothpick or paperclip that is environmentally friendly; yes. In any common usage, "environmentally friendly" simply means reduced impact, which is absolutely possible with LLMs, as is demonstrated by bigger models being distilled into smaller more efficient ones.

Discussing the environmental impact of LLMs has always been silly, given that we regularly blow more CO2 into the atmosphere to produce and render the newest Avengers movie or to spend one week in some marginally more comfortable climate.

The point was that for many tasks, AI has similar failure rates compared to humans while being significantly cheaper. The ability for human error rates to be reduced by spending even more money just isn't all that relevant.

Even if you had to implement checks and balances for AI systems, you'd still come away having spent way less money.

Exactly, we need a much more granular approach to evaluating intelligence and generality. Our current conception of intelligence largely works because humans share evolutionary history and partake in the same 10+ years of standardized training. As such, many dimensions of our intelligence correlate quite a bit, and you can likely infer a person's "general" proficiency or education by checking only a subset of those dimensions. If someone can't do arithmetic then it's very unlikely that they'll be able to compute integrals.

LLMs don't share that property, though. Their distribution of proficiency over various dimensions and subfields is highly variable and only slightly correlated. Therefore, it makes no sense to infer the ability or inability to perform some magically global type of reasoning or generalization from just a subset of tasks, the way we do for humans.

There is a difference between poor reasoning and no reasoning. SOTA LLMs correctly answer a significant number of these questions correctly. The likelihood of doing so without reasoning is astronomically small.

Reasoning in general is not a binary or global property. You aren't surprised when high-schoolers don't, after having learned how to draw 2D shapes, immediately go on to draw 200D hypercubes.

I honestly don't think it is a matter of opinion, though. Her voice has a few very distinct characteristics, the most significant of which being the vocal fry / huskiness, that aren't present at all in either of the Sky models.

Asking for her vocal likeness is completely in line with just wanting the association with "Her" and the big PR hit that would come along with that. They developed voice models on two different occasions and hoped twice that Johannson would allow them to make that connection. Neither time did she accept, and neither time did they release a model that sounded like her. The two day run-up isn't suspicious either, because we're talking about a general audio2audio transformer here. They could likely fine-tune it (if even that is necessary) on her voice in hours.

I don't think we're going to see this going to court. OpenAI simply has nothing to gain by fighting it. It would likely sour their relation to a bunch of media big-wigs and cause them bad press for years to come. Why bother when they can simply disable Sky until the new voice mode releases, allowing them to generate a million variations of highly-expressive female voices?

That's not the purpose though, clearly. If anything, you could make the argument that they're trading in on the association to the movie "Her", that's it. Neither Sky nor the new voice model sound particularly like ScarJo, unless you want to imply that her identity rights extend over 40% of all female voice types. People made the association because her voice was used in a movie that features a highly emotive voice assistant reminiscent of GPT-4o, which sama and others joked about.

I mean, why not actually compare the voices before forming an opinion?

https://www.youtube.com/watch?v=SamGnUqaOfU

https://www.youtube.com/watch?v=vgYi3Wr7v_g

-----

https://www.youtube.com/watch?v=iF9mrI9yoBU

https://www.youtube.com/watch?v=GV01B5kVsC0

GPT-4o 2 years ago

Apologies from me as well. I've been unnecessarily aggressive in my comments. Seeing very uninformed but smug takes on AI here over the last year has made me very wary of interactions like this, but you've been very calm in your replies and I should have been so as well.

GPT-4o 2 years ago

You consistently refuse to take the necessary reasoning steps yourself. If your next reply also requires me to lead you every single millimeter to the conclusion you should have reached on your own, then I won't reply again.

First of all, it obviously changes everything. A shortsighted person requires prescription glasses, someone that is fundamentally unable to count is incurable from our perspective. LLMs could do all of these things if we either solve tokenization or simply adapt the tokenizer to relevant tasks. This is already being done for program code, it's just that aside from gotcha arguments, nobody really cares about letter counting that much.

Secondly, the analogy was meant to convey that the intelligence of a system is not at all related to the problems at its interface. No one would say that legally blind people are less insightful or intelligent, they just require you to transform input into representations accounting for their interface problems.

Thirdly, as I thought was obvious, the tokenizer is not a uniform blur. For example, a word like "count" could be tokenized as "c|ount" or " coun|t" (note the space) or ". count" depending on the surrounding context. Each of these versions will have tokens of different lengths, and associated different letter counts. If you've been told that the cube had 10, 11 or 12 trillion constituent parts by various people depending on the random circumstances you've talked to them in, then you would absolutely start guessing through the common answers you've been given.

GPT-4o 2 years ago

No, think about it. The granularity of the interface (the tokenizer) is the problem, the actual model could count just fine.

If the legally blind person never had had good vision or corrective instruments, had never been told that their vision is compromised and had no other avenue (like touch) to disambiguate and learn, then they would tell you the same thing ChatGPT told you. "The objects blur together" implies that there is already an understanding of the objects being separate present.

You can even see this in yourself. If you did not get an education in physics and were asked to describe of how many things a steel cube is made up, you wouldn't answer that you can't tell. You would just say one, because you don't even know that atoms are a thing.

GPT-4o 2 years ago

Please try to actually understand what og_kalu is saying instead of being obtuse about something any grade-schooler intuitively grasps.

Imagine a legally blind person, they can barely see anything; just general shapes flowing into one another. In front of them is a table onto which you place a number of objects. The objects are close together and small enough such that they merge into one blurred shape for our test person.

Now when you ask the person how many objects are on the table, they won't be able to tell you! But why would that be? After all, all the information is available to them! The photons emitted from the objects hit the retina of the person, the person has a visual interface and they were given all the visual information they need!

Information lies within differentiation, and if the granularity you require is higher than the granularity of your interface, then it won't matter whether or not the information is technically present; you won't be able to access it.

Your English is absolutely fine and your answers in this thread clearly addressed the points brought up by other commenters. I have no idea what that guy is on about.

I never made the claim that there were no solutions to the problems outlined above. In fact, I implemented exactly what you've linked long ago. This doesn't, of course, change the fact that no such easy procedure exists on unrooted Android devices and that Windows, by default, has inaccurate timekeeping. Requiring extensive configuration for something as simple as accurate time is what irks me. Time sync is neither expensive nor complicated, so why is Windows so reluctant to do it often enough to keep the time?

Having to set registry values to get sane syncing behavior is just nuts when, again, my cheap Casio watch - which possesses a miniscule fraction of the features of a modern computing device - can deliver sub 500ms accuracy with absolutely zero fuss.

While I apologize for making general claims without a significantly larger sample size, I have tried this on all of my devices and those of family and friends. I only have access to phones running Android and laptops/desktops running Windows, so I cannot say whether macOS/iOS and Linux suffer from the same issue. I'm looking at ~12 devices from various manufacturers that all have similar amounts of clock drift.

All of the devices have network sync enabled, and they sync to accurate time if I manually disable and enable it again. The issue is that they don't sync regularly enough by themselves.

Thanks for this bit of sanity. I was losing my mind trying to think up an explanation as to why that comment wasn't flagged or inundated with comments pointing out the sheer stupidity and inhumane arrogance.

I don't want to be overly dramatic, but this has genuinely lowered my view of HN as a whole, and I'll think twice about reading the comments here from now on.

If I were to show this to any well-adjusted person I know, they'd probably think less of me for even being in the same community and profession as the person who wrote that comment.

Somewhat off topic, but thank you for putting into code what has been floating around in my imagination since high school. I haven't tested your project yet, but I really do hope that AI assisted roleplaying becomes a mainstay in the game development world. I got a taste when toying around with AI dungeon, and if this isn't the the next step in meaningful interactive storytelling, then I don't know what is.

Be aware that you're talking about Quora's implementation of ChatGPT here. As far as I know, the cached answers were generated with an incredibly outdated version, which is definitely not indicative of its current quality.

Even worse, I think they actually prime it with answers already posted on the thread, or even just related threads. For example, one of the answers to the first question mentions the same Altaic root as ChatGPTs answer, and I've found multiple people that are seeing their own rephrased answers in the response.

If you preprompt ChatGPT with questionable data, then the answer quality will be massively degraded. I've noticed many times now that Bing will rephrase incorrect information or construct a very shallow summary out of unrelated articles when internet searches are allowed, but is able to generate a cohesive and detailed summary when they're disabled.

Throwing random answers - some contradicting each other and some talking about subtly different aspects of the topic - into a session without further guidance just isn't a great idea.

The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably.

I think LLMs are a good start. I am certain they lack a world model, the kind you and me use.

See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear.

The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.

And then you say something like this, which is obviously wrong. No, a world model wouldn't prevent generalization failures, a perfect all-encompassing world model would. Humans experience generalization failures as well, otherwise every athlete in one sport would automatically be an expert in every other discipline or every mathematician would also be a Grand Master in chess. LLMs necessarily need a world model to generate well-formed text that isn't in their training corpus, something they are obviously capable of, it's just an imperfect world model. Ours is also imperfect, but far less so than that of LLMs.

I have linked a paper...

Except that paper is completely irrelevant to the argument you're making here. It is a useful insight into the limitations of simple metrics, but definitely does not extend to any claim of model performance, because they too use a simple metric as an replacement, even though clear qualitative differences are observed between model iterations.

Let me put it this way: Imagine I create a series of chess AIs, with each iteration better than the last. If I then show you a chart demonstrating that the ELO of my models increases linearly, would you say that my models' abilities increase linearly as well? No, obviously not, because my model needs far less strategy and complexity to go from ELO 1000 to 1100 than it needs to go from 2700 to 2800. I.e the difficulty doesn't scale linearly, and a linear increase on this nonlinear space is therefore also not really linear. Unless you believe the difficulty of accurately predicting text scales linearly, then this applies to LLMs as well.

If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.

Except that this is the entire value proposition of LLMs. They can, in the average case, actually represent concepts by the complex interplay of adjacent concepts. The entire reason why they are so impressive is that the nuances of reality are grasped and that even a noisy example of a concept can be correctly classified. Give a LLM a description that is largely incorrect and mislabeled, and chances are it gets it anyway. LLMs being unable to generalize over some concepts has as much to do with fundamental limitations as me being unable to correctly classify the shredded remains of a flower variety that I have seen once in my life has to do with me being stupid.

Look, you can argue with me or you can try it out. Push the system, see how far it can go

I have done just that for the last 6 months and have seen nothing to contradict what I've said here.

The model F=GMm/r^2, for example, has a causal and ontological semantics: F is a force, M a mass etc. these are pieces of reality. And this formula (though actual a little suspicious in many ways, GR fixes this) nevertheless says there is a force between masses that has certain properties etc.

And this model is based on the observations of Newton himself and those that came before him. There is nothing magic about observing the attraction between objects and deriving a model from that. Why are they magically "pieces of reality"? How do you know that? What differentiates mass from "funny-mass" that I just thought up and actually repels other "funny-mass"? Maybe the fact that we can test the effects described by that first model and therefore verify it as the most likely candidate?

But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).

He did derive it either from that data or his own experiences. It's true that you can construct infinite models to explain an observation, which is why the scientific method includes an Occam's razor-esque tenet to select the simplest possible model. Complex models risk contradictions with new observations, which is why you choose the one with the least assumptions. With that rule, the model to select becomes quite clear.

There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.

No, most of it told him he was right, unless you want to claim Newton was an idiot that stumbled onto the right model by accident. With "most" I obviously mean most reasonable data, people telling him he's wrong is obviously excluded from this list, if his evidence contradicted those claims.

Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.

And we compare those models by...?

What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.

You yourself just correctly made the point that we can construct endless models, well, we can create endless theories as well. And all of these theories are exactly worthless unless we test them. There is nothing "pseudo-scientific" about testing, it is literally the core of the scientific method. By your reasoning, are some crackpots coming up with the newest flat earth theory pure and unsullied by the lower demands of verification, and therefore way more scientific?

identifiable formal statistical methods entered in the early 20th C.

Formal is the important word here, statistics has been used in an informal manner from the inception of life. Formal mathematics, as in mathematics on a formal axiomatic framework, has also only been introduced in the 19th century. So what? Science owes everything to informal statistics, as does engineering and art. Rules of thumb used by engineers and creation of art that satisfies our aesthetic preferences requires sampling and approximation.

That latter system, in most cases, fails. It provides a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions

It literally doesn't and no, it doesn't need those assumptions either. The reason why normalcy is usually assumed is that it often can be assumed without significant deterioration in predictive power. That doesn't mean it needs to be assumed, in fact, it often isn't.

You constantly reference theory building, but how do you think those theories get created exactly? Through mathematical reasoning? How do you know mathematics is valid? Through logical deduction? How do you know logical deduction is valid? Through knowledge? How do you know knowledge... and so on.

Fact is, we only use these tools because they have proven their validity through being tested over and over and over again. And if you look at modern pseudoscience, it always seems to coincide with a proclivity for theory building, with very little hypothesis testing involved.

The scientific method is inherently statistical, we take a finite amount of observations and construct a model that best represents those observations. So yes, sorry, I should have said 100%.

With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples.

Also, you're describing statistics in an incredibly dismissive way. Stats is decidedly not just "taking the average". At the very least not in this brainless, first-order way you describe here.

Let's explore this with an thought experiment:

A model of some process has been confirmed across the globe, at least 5000 studies show the same result. Yet, one day, a study is published that fails to demonstrate the desired effect. Without using statistics, please tell me which action should be taken next:

A: The stray result is investigated for experimental failures.

B: The entire model of the process is immediately dismissed and we start from scratch.

By the way, you're welcome to call me uninformed, but I'd ask you to at least provide either your credentials or research that directly contradicts me.

Oh, I almost forgot. I know all of these definitions, please actually engage with what I'm saying instead of insinuating that I'm missing information.