HN user

bluegatty

676 karma
Posts1
Comments451
View on HN

I'm not attacking you, you don't have to defend yourself.

I'm just nothing that HN rhetoric is contradictory.

But this:

"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.

If it were some true injustice, then fine, but that is clearly not the case.

There is ample room to contemplate that even copyrighted works could be considers fair use as training material.

"the corpos can get fucked as far as i'm concerned."

Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.

I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

... but Chinese SOTA foundries directly using distillation as fair game.

I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

What is more reasonable:

- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.

- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.

Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.

The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

"> but they are different.

How, and why?"

How are they even remotely the same?

They're not even used the same way.

One is raw data input, the other is training content - designed to train LLMs.

One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.

"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.

But it's debatable if that's the case.

Google stores copyrighted content and produces in in their product.

Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.

I do agree though, that we ought to draw the line somehow.

This is a misrepresentation though.

The LLM output, is not the same as the input - there is value add.

Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.

It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.

We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

Designing and creating an LLM from nothing is a monumental feat of Engineering and 'value add'.

Copying something is not.

Programming Microsoft Word is value add, copying the code is not.

Copying design ... there are some question marks there.

It's extremely easy to understand at it's core.

What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.

"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"

The training data used is part of all of this is a separate but related question.

There is value add in AI irrespective of how the data got to what it is.

Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.

'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.

I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.

We ought to identify that and integrate that into our thinking.

Making an LLM from raw data is value-add.

Distillation is just value extract.

It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.

I think we start by recognizing that ... and then try to figure it out from there.

'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.

'Libre Office' did not 'win'.

People are happy to pay $50/year per seat to have the extra features and to not have to deal with stuff.

There's an issue at the margins here:

$1000/employee is a massive cost - it has to be deeply justified. $50/employee is like ... $2 out of your pocket. It's an incremental cost. The CFO is happy to pay it if there is a lot of value.

A lot of software is in that later category.

Imagine if gasoline was 1 cent per litre - and there was 'free gas' but it was a pain to use, and you had to check a bunch of things. You may just pay the 1 cent.

AI is not quite that yet, but these dynamics will play out eventually, for a lot of things.

All of these takes are horribly 1-sided.

'China's copying / distilling strategy is working, the people getting distilled are ruining the economy!'

Or 2 days ago:

'Open Models are Communist'

Almost nothing to investigate the economic nuance of what is going on.

- Switching costs are very real, these are not perfect substitutes.

- The SOTA makers are the one's pushing the frontier, there is a kernel of truth in the fact that if they collapse, certain things will struggle to move forward.

- Nobody trusts either of those nation state, export controls are a thing, this is a very real concern.

Etc.

It's distressing that there are not sound comprehensive takes.

No fan of communism either ideologically or in practice but I hate these ignorant people who call everything they can't control 'communism'

I find it hugely hilarious that they can't realize

i) they they'd be destroying themselves by keeping US locked into high prices while the rest of the world gets access to free models

ii) no scruples at all with respect to common good

iii) the dork literally works for 'open ai' - who's original mission was 'nobody controls it'. Granted - it's hypocritical of Atman AI to flip private, but there are also legit reasons for that, as much as excuses.

a bit upside down - wealth extraction is a measure of power.

development of power can be based on all sorts of things, it depends on the framework.

a nurse who is utterly incompetent will be fired quickly.

after a certain threshold though, obviously, competition won't be about 'ability' so much - but there are baseline ability and professionalism thresholds.

"decisions after snorting a long line of social-media-psychosis and TED talks."

And yet - they are the one's paying you and everyone else somehow?

I think you might be missing something fundamental, start by considering that what is 'good software' is not an intrinsic measure, but a measure of what it does.

The only reason we really need 'intrinsically good' software, is if it's very long lived and a ton of people are going to come to depend on it.

People trying to make oak furniture when in most cases what we want is IKEA.

That said - AI or no AI - there's no excuse for not keeping a grip on things, whatever kind of 'grip' that might be.

The paper is saying 'context is context'?

That after a model has context about a project, the probes indicate a state that validates that?

Seems that the paper is highlighting the very nature of what LLMs are and what we expect them to be?

And that there is no 'thinking' here, it's just the state of the model?

You think that a fridge is about 'keeping food cold'?

It's about ease of access, price, noise, convenience, durability, features.

My folks have this fancy 2 door thing, perfectly quiet, makes the best ice you can imagine, it's hidden into the cuppboards, it's energy efficient, has these crisper things, you can see in and reach around easy, lots of space. It's a better product.

Those are announcements, not released integrated models.

There are two Tier 1 platforms today.

Meta, Google and XAi are formidable Tier 1.5 place, any one of which could rise to the fore.

My belief is that it will be Google and that probably only one of them will keep up in the long run.

There is a 'breaking point' when you start to get past 4-ish players - it really does start to introduce competitive pressures.

Your second point about Tier 2 substitution is valid, but a few things:

1) Tier 1 models are not a 'luxury good' - that has a different economic definition. They are for most applications today actually just the quality, rational choice.

2) Substitution will have different effects for different people, and you're right that AI for many tasks will be commiditized.

All of the profits in Mobile Phones go to Apple even though they are not the biggest player.

Almost all of the profits in Silicon go to the leading edge chips - even though there are a zillion fabs that make legacy chips.

Oil varies a bit but it's a commodity.

There are 3 SOTA model makers, and they have pricing power.

The 'switching costs' is not the issue so much as the inherent control over the commodity.

Think OPEC - when they acted as a cohort - they raised prices dramatically by having enough control to 'set prices'.

When OPEC lost it's pricing power ... nobody could set prices.

Fable is considerably better than GLM5 and it will have a strategic input - there is just hardly any substitute for it.

If these were cars - we'd just use whatever fuel.

But these are 'F1 races' - if you have some low grade 'dirty fuel' you will lose the race. You must have the 'top fuel'. There are 3 provides who implicitly collude and set prices.

We can try to be a big magnanimous and rise above a tiny bit of 'English on the ball', let the small stuff slide.

This is an unfair characterization.

Cohorts are not 'raising massive rounds' on selling to a couple of poor YC startups that 'don't want they product but pay for it anyhow' <- the implication there is very wrong.

They're selling to people who are using their product, but probably have bit of an open mind about things, and are willing to overlook hiccups. A bit like how buyers often have a 'national preference' in some cases.

Every startup ecosystem should do this, it's probably one of the top things any system could do to try to get things going.

Nvidia / OAI etc. are selling each other product which have a fair market value.

They're not moving money around with 'no services rendered'.

It does however mean that the situation is 'highly leveraged' - and therefor risk is more concentrated, and, they do disclose.

Nominally there's nothing wrong with investing in one's own supply chain.

It makes a lot of sense for a value chain player with huge cash position and therefore a lot of power to take a % ownership of a buyer.

Consider for a moment - what if Nvida acquired OpenAI? They would be two divisions in the same company. Would anyone consider it wrong for surpluses from one to be invested in the other? No. The moment it's 'managerial accounting' instead of 'balance sheet accounting' - nobody would care.

Imagine Google or AWS acquired Nvidia - maybe in 2017 so it would seem more realistic in terms of price - would any of this seem financially dubious? No.

Of course - those mega mergers would be bad for competition, but that's a separate question.

There's nothing inherently wrong going on here, as long as it's on the books and people understand the inherent structural risk of coupling etc..

There's nothing wrong with that, it's probably a net positive.

Small companies buying and selling each other's services is probably the #1 thing that other ecosystems should emulate.

Getting your first 3 customer is hard, you need an open and empathetic system for that - a 'cohort' is exactly that.