Robots that replace auto industry factory workers exist; the CEO of GM didn't imagine them as part of some sort of business media induced psychotic episode.
The same is not true for the software industry execs.
HN user
Robots that replace auto industry factory workers exist; the CEO of GM didn't imagine them as part of some sort of business media induced psychotic episode.
The same is not true for the software industry execs.
Amazon lost a cumulative 2.8 billion over their first 17 quarters.
So if you're asking about time, then amazon stopped a lot faster. OpenAI is 40 quarters old.
If you are asking about money, then amazon... also stopped a lot faster. OpenAI is losing money comparable to amazon's lifetime losses every quarter.
If you want to play games like that, you could also flip it around and ask if the AI would have been eventually fired (assuming no one knew they were talking to a computer).
Not sure what that proves.
If you can make it 800 you can claim to be a 100x engineer!
Say what you want about nazis, but they are good at rockets.
But Google didn't go public until 2004, when they were highly profitable.
Every startup goes through a phase where they aren't profitable... For most of of them that ends when they go bankrupt.
Didn't xAI basically donate the compute for that quarter so Anthropic could get to say they turned a profit?
But then go right back to being unprofitable again afterwards, which is a little weird.
They don't want to build trust. They want to build a trust wedge between the people making the buying decisions and the people with hands on experience of the product.
When an employee says AI isn't speeding up his work, the only thing the CEO hears is "Wow, this employee is so scared of getting replaced that he's lying about how great AI is" and he will pick up the phone to Anthropic to buy more licenses.
It's sort of brilliant actually. No way to make a product grow fast enough without bypassing the employees and targeting the decision layer directly.
So, it's like if they were a pharma company that was barely profitable if you didn't take into account R&D costs?
Sounds like a healthy industry, selling tokens at 1000x below cost.
Yeah, I use them all the time. I just don't see any good argument that it's anything other than statistical pattern matching plus some sort of logic encoded in language. My overfitted LLM obviously didn't arrive at Harry Potter the same way JK Rowling did, so the amount of time she spent writing it is completely irrelevant to any discussion about whether or not the LLM should be able to reproduce it. discussions of AGI if it took her an hour or a decade to write it, it has seen the result, so it can reproduce it.
Yeah, what about them? As far as I read it the tasks are fixed. The AI companies should know the tasks by now, and have overfitted their models on the tests by now, in the same way I'm implying I overfitted my model to reproduce Harry Potter.
I don't know why people are so impressed by 8h.
I trained an LLM to write the whole Harry Potter series, and that took JK Rowling like 17 years.
For my next point on the graph, I'll train the LLM to write the Bible, something that took humans >1500 years.
Karım Kahn at the International Criminal Court would like a word about that.
Google is the leader, they really don't want AI to be a success, it only comes with a risk of disruption. They probably don't even really believe it's going to be that big of a deal. They are only in that game to hedge; sure they have wasted a trillion dollars if AI doesn't come through, but they will earn that back in 3-5 years. So why would they need to do deranged marketing stunts and sacrifice their credibility for that?
If OpenAI or Anthropic doesn't turn this into a trillion dollar industry FAST, they are cooked. The strategy of building up fear around your product is risky, but necessary. There is simply no way to grow the AI business fast enough if they can't talk directly to the CEOs and bypass input from the employees, and baba yaga stories are perfect for that. Every time the CEO hears an employee say that the AI isn't working great for him, he hears an employee that's scared for his job or for his life, dismisses it, and sends out a mandate that everyone needs to prompt an AI every time they as much as need to go to the toilet.
How do you propose to do a Turing test on a human (in a sense that is different from a machine simply passing the Turing test)?
Like failing to pick out all the motorcycles in a captcha, or a turing test where you have a guy chat with two people without knowing that one of them could be a computer, and the interrogator, unprompted, suggesting one of them might be a computer?
They won't figure it out. It's the tragedy of the commons.
Yeah, so you are agreeing that the benchmarks are useless because they don't answer those questions.
Same question I have for all these benchmarks:
What's going to stop e.g. OpenAI from hiring a bunch of teenagers to play these games non-stop for a month and annotate the game with their logic for deriving the rules, generate a data set based on those playthroughs and fine tuning the next version of chatgpt on all those playthroughs?
That pretty much describes shape up : https://basecamp.com/shapeup
I have a mixed relationship to it, but the scope cutting part of it works extremely well.
The focus it brings on focusing on the problem solved rather than on the concrete solution is also healthy I feel.
Yeah, enormously. People will hedge depending on how sure they are about something. They might also have credentials in whatever you ask them, if you get legal advice from a lawyer, that can be judged to be more reliable than from a lay person.
Relationships with real people are pretty cool actually. If you talk to people that you have a longer relationship with, you might also be able to judge their areas of expertise and how prone to bullshitting they are.
In my experience the last answer it gives is usually the right one
Great, now I have two answers and still no clue which one is the right one.
It's an interesting parallel to, especially right wingers, want project intelligence into 1 dimension so things all humans can be ordered from inferior to superior. That logic was already strained with humans, but with the introduction of AI the wheels are really coming off for that model.
Now they have agents.
People need to understand that code is a liability. LLMs hasn't changed that at all. You LLM will get every bit as confused when you have a bug somewhere in the backend and you then work around it with another line of code in the front end. line of code
Your developers were so preoccupied with whether or not they could, they didn't stop to think if they should (add 250kloc)
I get what you're saying, but I remember watching teletubbies back in the days with my nephew, and all questions of the form:
Have ____ surpassed teletubbies?
Can always be answered in the affirmative.
I love the total lack of humility on that site. "What if the METR study turns out not to capture anything relevant? We just add a constant gap to be conservative!". But I guess these guys aren't really scientist, so it's probably a lot to ask that they relate critically to what they are doing and be honest about the limitations of their methods.
What if it turns out that the more you scale the more your LLM resembles a lobotomized human. It looks like it goes really well in the beginning, but you are just never going to get to Einstein. How does that affect everything?
What if it turned out that those AI companies were maybe having a whole bunch of humans solving the problems that are currently just below the 50% reliability threshold they set, and do fine tuning with those solutions. That will make their models perform better on the benchmark, but it's just training for the test... will the constant gap be a good approximation then?
It really dispelled the illusion for me, but it's not that easy to find those examples, but the combinatorics of possible number of guesses is untractable enough that it can't learn a good set of clues for all possible guesses.