HN user

impossiblefork

2,141 karma
Posts1
Comments1,375
View on HN

Mm.

Still, maybe less money could be enough.

If Euclyd's chip costs 40k per chip and we want to substitute all worldwide H200 inference capacity, so let's say 700,000 H200eds and Euclyd are right about how good their chip is, then we'd need 28,000 Euclyd cards, this is only 1.120 billion.

So maybe even 1 billion for VSORA chips this year and then two billion for VSORA and Euclyd chips next year, with the split based on how well the chips perform and some money for the lagging firm so they can keep designing things too.

I am actually not quite sure.

Here in the EU Rhea 1 was delayed and thus maybe a bit of a failure, but Rhea 2 might be good and seems to be on time. OpenChip seems to be on time too. The university work, people like Hochreiter etc. is clearly good.

What are the high-profile failures?

Yes, but how would it have been if the EU governments just piled on, subsidized whoever had reasonable qualifications who wanted to try, straight-up funded accelerator development so that the chips don't just become investment in US AI companies, straight up built the supercomputers for training, straight up offered access to groups that had previously had reasonable success in training some model.

I don't think it would even be that expensive. I think with EU chips we could do it for 20 billion EUR + 5 billion EUR per year. After all, an H200 doesn't cost 40k to manufacture, probably one-tenth, and if we are making the chips we can get the money circulating just as the Americans have.

Add in some state effort to create suitable training data-- maybe encourage academics to create some things that are AI-friendly, and let all the people trying have that. There would hardly be anyone in the game other than us if we went for this package.

We already have VSORA. Maybe Euclyd is ready soon too. If enough are ordered we have 85% of the inference need gone, so that whatever accelerators are freed can be repurposed for supercomputing and training? Maybe we could have some organized effort to make many of the VSORA and Euclyd chips so that the chips we need for training are freed up while we wait for our own training chips.

Edit: So in the end I think the thing we need to do is to order enough VSORA Jotunn8 cards and enough Euclyd cards that we flood the inference market so much that the price of H200eds etc. drops and we can buy them second hand and assemble huge supercomputers out of them. This may require some diplomatic effort from the commission, something like 20B EUR almost immediately and some coordination, but it's easy, it's doable and it requires only being a little bit decisive. Since these cards are faster less memory will be tied up in serving the models, so it'd have a positive effect on memory prices too, which in turn would have a positive effect on European supercomputer centers, who are currently of course constrained by high memory prices.

Yes, although then people came up with reasonable ways of parallelizing RNNs entirely or in part anyway. There's also the need for lookups in transformers. Queries still need to be multiplied with old keys, all of them, so RNNs don't have to be totally incompatible with parallelism. It's just that you can't have a matrix inside them like in the linear RNN h_t = Ah_t + Bx_t that you have to multiply together step after step. If A is an input-dependent scalar or something you can multiply together easily you're fine.

mLSTM layers are completely parallel (the update rule for the cell state is C_t = f_t C_{t-1} + i_t x_t where f_t and i_t are gates that can be computed from x_t alone, x_t is the input at time t and this means that you can compute F_s = \prod_{s<t} f_s as fast as a cumsum and an exponential and then compute C_t = \sum_{s<t} F_s i_t x_t, again as a cumsum, so it's as good as if though it were parallel). I think the sLSTM layers that are the other component of the xLSTM have something else like this and presumably there's some trick also to training the Kimi "delta attention" RNN.

I'm not sure whether this is hardware or optimization dependent to some degree, but I get the impression that good custom kernels are an important part of this kind of thing.

I think the Kimi thing is super cool, especially that they have so many RNN/linear attention layers (3x more than they have full attention). I haven't yet tried it though. It seems like it would be extremely reasonable for long context tasks and I guess this fits the times.

I suspect that the reason it has so many parameters is the same reason that compute optimal xLSTMs have some many parameters, and the success of this model makes me a bit unhappy that we haven't gotten an xLSTM-style model of huge size developed in Europe.

Obviously these guys are very pragmatic, they're probably not committed to anything other than what works on their internal evaluations, so they still have ordinary attention layers in the model and so on, and one can't be guaranteed that the people who come up with a good model then do the engineering in an ideal way, but I still think the success of Kimi shows what could have been if we had enough big supercomputers for LLM training and made them available to the right people-- because this is basically Hochreiter's thing. It's RNNs, or well, mostly RNNs.

I think the right idea is better patent examination.

If patent examination is correctly performed, then all patents will have real novelty and the people who have to license a patent haven't had anything taken away from them since the novelty means that they wouldn't have come up with the thing they have to license anyway.

Yes, but NPEs are necessary if patents are to be a thing also for smaller inventors.

An inventor can invent things that make a nuclear power plant cheaper or more efficient or better in some other way. He can't control whether the people and organizations who can afford nuclear plants decide to use his invention.

So NPEs are necessary for patents to work for small the inventor, and are how patents work for the small inventor. A small inventor in an expensive field is always a NPE.

It's strongly connected to the structure of the EU though, and the weak control that voters have over appointments to the commission, and every level of indirection is one at which the appointer can be influenced.

If EU institutions are used to push this sort of thing, we must treat that as what they are for. Systems do not get a pass because someone external is 'using them', but must be treated holistically.

Left-leaning politics is not at all like early 1900eds left-leaning politics.

Left-leaning politics has moved to very mild, not-even-social-democracy policies, taxation of wage income, a decreased focus on capital owners.

Left-leaning politics has thus been transformed beyond belief and has very little to do with what it used to. Most politicians have no idea about physical reality, which is the ultimate source of technology, but live sometimes in a world of administration, sometimes in a world of laws and sometimes in a world of politics only.

Left-liberals don't exist. Liberalism is a right-wing ideology: free trade, laissez-faire.

So I don't understand at all what you mean. What are the SocDems who have gone from being SocDems to not knowing what social democracy is and who now think about things like welfare and administrative stuff and living in a world of compromises attached to?

I can't see that they're attached to anything, and I think I despise them for it. At least someone who looks back to the past can look at it and critique it and see what ideas were valuable, what the real goals were, that led to different positive achievements.

I wouldn't say that making the matrix diagonal in some basis is some further step.

If we have an singular value decomposition, M=USV^*, the columns of U are linearly independent they are a basis for the space M maps things into, and the columns of V are linearly independent then it's a basis for the space it maps things from, and [M]_{BB'} = S.

That's you. but nobody In Sweden drives to work?

A smaller fraction than in the US. I think most people I know drive.

I see walking to work as an relative to each individual and their job lcoatiopna dn circumstance of where they live, not a country related thing.

Well, it isn't. It's about how walkable environments are.

GDP growth "experts" would disagree. It's the reason we don't have mandatory WFH for white collar jobs after Covid proved it's possible and salves the environment

Well, they may disagree, but the whole point is the goal of society isn't GDP, since GDP is easy to game with things like creating situation where people are effectively forced to waste energy, drive to work-- that sort of thing.

Yes, but German society is structured to require much less energy, just as Dutch society is structured to use much less land.

If you put Germans whose lives function in a US-style, even just getting to work will be a huge drag.

Misery depends on the structure of society. Here in Sweden I can walk to work. This means that I'm spending zero money on travel to work, and that my travel to work contributes $0 to Swedish GDP. But this is actually better than if Swedish GDP were higher and I was traveling by car.

This is one way in which GDP can be extremely misleading.

A lot actually. Obviously not to write these comments, but a whole lot of Claude.

With regard to the substance: sales of ASML machines are not currently connected to the EU getting chips, but to ASML getting paid for their work. For EU chips for training transformer models we'll need chip design firms, not chipmaking, and as I stated in my comment, there are some promising ones that will probably be able to design the chips we need if we order them.

Yes, but what does that affect?

If we end up with a world where only US firms can use the latest LLMs and the latest LLMs are needed to keep up in the software world, or in making prototypes, then that's a whole series of fields which are blocked from us.

So I think we need to make sure that we no only can, but build frontier LLMs from scratch-- not RLed on foreign data, but genuinely from scratch.

Even from an economic point of view, I don't think a continuous outflow of ~200 USD/month for every office job is sustainable, and that's what we'd get if the plausible scenario is borne out. An inter-EU cost of 200 USD/month for every office job though, that's survivable.

I think the way to do it is this: let EU chip design firms bid. They say what their systems can do, they give their prices, and then we choose the one that can achieve the requirements (pretrain multiple 10T+ models in a reasonable amount of time and then do RL on them) at the lowest cost.

Yes, of course people would be working there for the paycheck.

I think there's nothing special about public funding though. The field is so competitive that people will be mad if a competitive model is not achieved, making corruption more damaging to the organizations. There would also be some internal competition. There are after all several EU LLM/AI/etc. firms that would probably try to use this infrastructure.

I think the chips alone are 10B minimum. It'd be way bigger than CERN.

Provided that the systems work, they can at least be repurposed to other things. If the organizations that are to train public LLMs can't do it, we can rent the system out to Mistral or something.

So I think something like 5B, starting with 10B to get started, in public money per year, the chip firms are private, some of the LLM firms will be private, but the system is available to train European LLMs-- that's I think a realistic approach.

I think what's actually needed is two things: an EU training infrastructure that allows training of 10T+ models, and an EU inference infrastructure that is sufficient that it's possible to do RL on them.

This effectively reduces the problem to a specialized supercomputing infrastructure problem which I think is relatively easy to solve. I think the chips are coming. I think Euclyd will be able to do the inference chip and I think the training chip won't be harder. It's just a matter of accepting the need to order a huge number of them, being willing to think a little bit like the kind of people who operate corners. So we can be there next year, I think. What we then lack is a training chip-- maybe OpenChip can do it, maybe they can't, but there are reasonable but still unfinished projects. Maybe if Euclyd finishes an inference chip in 2027 we can have the state pay them to make a training version, put in fp32, put in communication tiles. If their design is real and works (which it should, since it's basically a fancier version of Groq, as it's described, and since even Groq works) I think the advantage these chips is likely to have would be enough that a training version would be NVIDIA-beating.

We probably need some solution for the data-- i.e. to allow people to do things that are against copyright law in a limited way, but I think it's a better idea to start EU firms than to try to attract Anthropic.

Because of the need for capital the hardware-software carousel is necessary. We can't pay for NVIDIA chips and then have NVIDIA feed that money into US firms. We have to feed money into EU chips that either carousel the money into EU AI firms or who just offer cheap chips.

I wish the EU were legalistic and rules based, but the commission and politicians are involved in many of these things. It's like Trump's executive stuff, just with a committee instead of a single person, and I guess, with less power.

This is not the law here in Sweden, at least.

We don't have precedent in the way that common law countries do, and the judgements in actual cases point in slightly different directions-- in one case a court felt that the failure to fire a warning shot made it not self-defence, in another fighting people trying to get into an apartment with a knife was deemed acceptable.

Generally though, if someone is breaking into your apartment while you're there, possibly trying to get at you, there's no limit, as long as you're actually trying to defend yourself (so no executing someone who you've clearly disabled, etc.).

If people are breaking into your apartment and you fire a warning shot, then proceed to shoot the attackers, no one will complain.