HN user

zeknife

96 karma
Posts0
Comments52
View on HN
No posts found.

ELIZA fooled plenty of people (both originally and in the study you just linked) but i still wouldn't say Eliza passed/passes the turing test in general. It just shows that occasionally or even frequently fooling people is not a sufficient proxy for general intelligence. Ofc there isn't a standardized definition, but one thing I would personally include in a "strict" Turing test is that the human interrogee ought to be incentivized to cooperate and to make their humanity as clear as possible. And the interrogator should similarly be incentivized to find the right answer.

I didn't mean that a human driver needs to leave their vehicle to drive safely, I mean that we understand the world because we live in it. No amount of machine learning can give autonomous vehicles a complete enough world model to deal with novel situations, because you need to actually leave the road and interact with the world directly in order to understand it at that level.

How many images do you need? What are the use-cases that need a bunch of artificial yet photoreal images produced or altered without human supervision?

Thinking is subconscious when working on complex problems. Thinking is symbolic or spatial when working in relevant domains. And in my own experience, I often know what is going to come next in my internal monologues, without having to actually put words to the thoughts. That is, the thinking has already happened and the words are just narration.

OpenAI: Sora 2 years ago

Like with music generation models, the main thing that might make "open source" models better is most likely that they have no concern about excluding copyrighted material from the training data, so they actually get a good starting point instead of using a dataset consisting of youtube videos and stock footage

You said it, those tests are designed to measure human intelligence, because we know that there is a correspondence between test results and other, more general tasks - in humans. We do not know that such a correspondence exists with language models. I would actually argue that they demonstrably do not, since even an LLM that passes every IQ test you put in front of it can still trip up on trivial exceptions that wouldn't fool a child.

Anything except tasks that require having direct control of a physical body. Until fully functional androids are developed, there is a lot a human-level AI can't do.

It can't explain it's thinking after producing the answer, like you say, it'll just generate a post-hoc rationalization. But if you ask it to think step by step before making a conclusion, LLM's will somewhat reliably arrive at reasonable steps and better final answers.

Talking about sentence structure in the conventional sense may not be meaningful here, since what could be described as reasoning in LLM's happens in a more abstract space. If we're looking to understand why a small change makes a big difference, it's pretty intuitive to consider that the second instance of "her" is modified by "mother" due to attention, and ends up being a wildly different vector.

Regardless, it's reasonable to assume that certain aspects of the prompt or input structure will prime the model to be more scrutinizing. I'd be surprised to see it point out a logical inconsistency like this if it was just part of a broader context and it wasn't asked "what it thinks" or to "be logical"