HN user

epups

1,404 karma
Posts3
Comments670
View on HN

Some important landmarks since GPT4 was first released (not in chronological order):

- Vast cost reduction (>10x)

- Performance parity of several open source models to GPT4, including some with far fewer parameters

- Much better performance, much larger context window in state-of-the-art closed source LLMs (Claude 3.5 Sonnet)

- Multimodality (audio and vision)

- Prototypes for semi-autonomous agents and chain-of-thought architectures showing promising avenues for progress

This is partially the reason why we see LLM's "plateauing" in the benchmarks. For the lmsys Arena, for example, LLM's are simply judged on whether the user liked the answer or not. Truth is a secondary part of that process, as are many other things that perhaps humans are not very good at evaluating. There is a limit to the capacity and value of having LLM's chase RLHF as a reward function. As Karpathy says here, we could even argue that it is counter productive to build a system based on human opinion, especially if we want the system to surpass us.

Large Enough 2 years ago

The graphs seem to indicate their model trades blows with Llama 3.1 405B, which has more than 3x the number of tokens and (presumably) a much bigger compute budget. It's kind of baffling if this is confirmed.

Apparently Llama 3.1 relied on artificial data, would be very curious about the type of data that Mistral uses.

This article is borderline comical. The thin list of his accomplishments focus on increasing revenue ($15B to $70B) and launching or acquiring successful initiatives (Xbox, Skype, Azure). Then the rest is a recollection of his biggest embarassments, like terrible mobile products, misguided Windows strategy and trailing AWS by 7 years.

Reporting on business issues is always muddled by a lack of proper comparisons, along with cherry picking. For example, this article makes the argument that increasing Microsoft's revenue by 4x was very impressive, even though the stock value stagnated. However, when evaluting his tenure as owner of a basketball club, he is declared successful because its value doubled. The problem is that Microsoft was eclipsed compared to its peers at the time - Google, Amazon, etc. -, and likewise the average basketball club doubled in value as well.

The author is making the point that Alphafold 3 is not so impressive - it is simply regurgitating its train set, and it's not so good for inference.

I think his central point is fair and interesting. The test train split is apparently legit, as they used structures released before 2021 for training and the rest for testing. However, there was no real check for duplicates, and the success rate might be inflated by a bunch of "me too", low hanging fruit structures that are very slight variations from what we know.

However, I'm not sure I agree with his skepticism. LLMs suffer from the exact same problems - getting it to write a Snake game in any language is trivial, but it is almost certainly regurgitating - , but can be useful as well. I mean, if for various reasons people are publishing very similar structures out there, there's certainly value in speeding up or reducing that work considerably.

The article makes it sound like it was ultimately a great decision, because now coal use is at historically low levels while renewables have been increasing. What it neglected to mention is that if nuclear had remained there, the share of coal would have directly and proportionally decreased. To achieve a 5% switch from very dirty to clean energy would be a spectacular feat, and the Germans achieved the opposite when closing their still perfectly usable nuclear plants very recently.

This is the same dynamics we have seen before. How many national supermarket chains are there? Crossing Walmart or Amazon seem very analogous here. If anything, business in the past was more feudalistic, as people were bound to a given market and they are now able to negotiate more broadly.

All of the replies here seem to point out that monopolies are bad and tech companies tend to be monopolistic. This is obvious and we all agree. However, when asked what is the difference between a mall and Apple Store, Varoufakis did not mention a monopoly or oligopoly as the issue, he mentioned rent. Specifically he mentioned that getting a percentage of profit was the biggest issue. I think this is not a strong argument for calling it a new economic model.

If it's not about percentages, then what's the difference here? Walmart and Amazon are not fundamentally different. If the Apple Store can be compared to a mall, then we are in the same economic model we've been for decades.

I like Varoufakis in general and I find this analogy interesting. However, his answer to this question was not satisfactory in my opinion, as many commercial arrangements including malls, are also based on percentages:

Q: A company like Apple might argue that instead of being a fiefdom, maybe the Apple App Store is more like a mall where companies have to rent their stores from whomever owns the building. How is technofeudalism different from the mall dynamic?

A: Well, hugely. Say you and I were going into partnership together with a fashion brand. We go to the shopping mall and we hire a shop, the rent is fixed. It is not proportional to our sales. The more money we make, the higher our price-to-rent margin. With the Apple Store, they get 30 percent of all sales. That’s not at all the same thing. That is the equivalent of the ground rent that the feudal lord used to extract from vassal capitalists.

It is not always trivial to switch. There are realistically two companies dominating phone OS, and either one comes with their poisoned pills. More importantly, I think what he is referring to here is the economical subjugation of the rest of society to tech, and to a very few tech players at that. An oligopoly might be better than a monopoly, but not much.

Sorry, could you expand on this a bit further? Are you saying that for a MoE, you want to train the exact same model, and then just finetune the feed forward networks differently for each of them? And you're saying that separately training 8 different models would not be efficient - do we have evidence for that?

A gun malfunction has a potentially lethal consequence. This bug caused a person to see a fake photo of herself with fake boobs out. Perhaps it's ok that the safety standards for the latter are lighter, no?

You also didn't address the main point of the earlier comment. Is every programmer responsible - even criminally, some suggested - for potential bugs or vulnerabilities in one of their products?

You write as if there is a finite supply of "entertainment" which comes out of thin air. If someone is consuming, then someone is supplying. You mentioned yourself that the ones supplying are the locals. So, in your analogy the locals are getting jobs in night clubs, restaurants and construction to feed the needs of this "consumption class". In economies without tourism, this is done by producing export goods.

If the manager is responsible for hiring and managing the driver and the navigator, then obviously they can also be under stress, and will be held responsible for the outcome.

Wait, so progress on Stockfish would happen regardless of Alpha Chess? I always thought they were inspired by it in the newer versions, and got much improved rating from incorporating it.

I like the article, but I feel that it could have been written 50 years ago as well. Russia being led by a despotic leader, proxy conflicts, lack of social cohesion, those are all things we have seen before in the 70's and in other points in modern history. The US even had trouble mobilizing for WW2, despite the central government trying to push for it.

The scientific consensus is that they are clearly a separate species, but there are varying ways to define these boundaries. For a primer, you can read up more here: https://www.nhm.ac.uk/discover/are-neanderthals-same-species...

Regarding your particular point about identifying "neanderthal DNA", consider that most genes are made up of thousands of base pairs. Even though there might be individual variation, it is trivial to distinguish between, say, chimpanzee and human versions of certain genes based on what we know about the populations. We also know the rate of evolution of genes and how they diverge, which means we can be pretty sure that a given version was acquired through interbreeding versus spontaneous mutation.

The main issue with populism and why people fall for it is a reliance on narratives over facts. If we discuss consolidated data instead of an anecdote, the conclusion is quite clear, and it shows that you are wrong: https://theconversation.com/hard-evidence-how-areas-with-low...

To vote a particular way in the hope of change? Why wouldn't that be rational?

There are plenty of historical examples of people voting for "change" that ended up backfiring wildly. I think you could find quite a few of those in the same ideological spectrum of people like Wilders.