HN user

MadxX79

119 karma
Posts0
Comments49
View on HN
No posts found.

Robots that replace auto industry factory workers exist; the CEO of GM didn't imagine them as part of some sort of business media induced psychotic episode.

The same is not true for the software industry execs.

Amazon lost a cumulative 2.8 billion over their first 17 quarters.

So if you're asking about time, then amazon stopped a lot faster. OpenAI is 40 quarters old.

If you are asking about money, then amazon... also stopped a lot faster. OpenAI is losing money comparable to amazon's lifetime losses every quarter.

If you want to play games like that, you could also flip it around and ask if the AI would have been eventually fired (assuming no one knew they were talking to a computer).

Not sure what that proves.

They don't want to build trust. They want to build a trust wedge between the people making the buying decisions and the people with hands on experience of the product.

When an employee says AI isn't speeding up his work, the only thing the CEO hears is "Wow, this employee is so scared of getting replaced that he's lying about how great AI is" and he will pick up the phone to Anthropic to buy more licenses.

It's sort of brilliant actually. No way to make a product grow fast enough without bypassing the employees and targeting the decision layer directly.

Yeah, I use them all the time. I just don't see any good argument that it's anything other than statistical pattern matching plus some sort of logic encoded in language. My overfitted LLM obviously didn't arrive at Harry Potter the same way JK Rowling did, so the amount of time she spent writing it is completely irrelevant to any discussion about whether or not the LLM should be able to reproduce it. discussions of AGI if it took her an hour or a decade to write it, it has seen the result, so it can reproduce it.

Yeah, what about them? As far as I read it the tasks are fixed. The AI companies should know the tasks by now, and have overfitted their models on the tests by now, in the same way I'm implying I overfitted my model to reproduce Harry Potter.

I don't know why people are so impressed by 8h.

I trained an LLM to write the whole Harry Potter series, and that took JK Rowling like 17 years.

For my next point on the graph, I'll train the LLM to write the Bible, something that took humans >1500 years.

Google is the leader, they really don't want AI to be a success, it only comes with a risk of disruption. They probably don't even really believe it's going to be that big of a deal. They are only in that game to hedge; sure they have wasted a trillion dollars if AI doesn't come through, but they will earn that back in 3-5 years. So why would they need to do deranged marketing stunts and sacrifice their credibility for that?

If OpenAI or Anthropic doesn't turn this into a trillion dollar industry FAST, they are cooked. The strategy of building up fear around your product is risky, but necessary. There is simply no way to grow the AI business fast enough if they can't talk directly to the CEOs and bypass input from the employees, and baba yaga stories are perfect for that. Every time the CEO hears an employee say that the AI isn't working great for him, he hears an employee that's scared for his job or for his life, dismisses it, and sends out a mandate that everyone needs to prompt an AI every time they as much as need to go to the toilet.

How do you propose to do a Turing test on a human (in a sense that is different from a machine simply passing the Turing test)?

Like failing to pick out all the motorcycles in a captcha, or a turing test where you have a guy chat with two people without knowing that one of them could be a computer, and the interrogator, unprompted, suggesting one of them might be a computer?

ARC-AGI-3 4 months ago

Yeah, so you are agreeing that the benchmarks are useless because they don't answer those questions.

ARC-AGI-3 4 months ago

Same question I have for all these benchmarks:

What's going to stop e.g. OpenAI from hiring a bunch of teenagers to play these games non-stop for a month and annotate the game with their logic for deriving the rules, generate a data set based on those playthroughs and fine tuning the next version of chatgpt on all those playthroughs?

Yeah, enormously. People will hedge depending on how sure they are about something. They might also have credentials in whatever you ask them, if you get legal advice from a lawyer, that can be judged to be more reliable than from a lay person.

Relationships with real people are pretty cool actually. If you talk to people that you have a longer relationship with, you might also be able to judge their areas of expertise and how prone to bullshitting they are.

I love the total lack of humility on that site. "What if the METR study turns out not to capture anything relevant? We just add a constant gap to be conservative!". But I guess these guys aren't really scientist, so it's probably a lot to ask that they relate critically to what they are doing and be honest about the limitations of their methods.

What if it turns out that the more you scale the more your LLM resembles a lobotomized human. It looks like it goes really well in the beginning, but you are just never going to get to Einstein. How does that affect everything?

What if it turned out that those AI companies were maybe having a whole bunch of humans solving the problems that are currently just below the 50% reliability threshold they set, and do fine tuning with those solutions. That will make their models perform better on the benchmark, but it's just training for the test... will the constant gap be a good approximation then?

It really dispelled the illusion for me, but it's not that easy to find those examples, but the combinatorics of possible number of guesses is untractable enough that it can't learn a good set of clues for all possible guesses.