HN user

teravor

317 karma
Posts0
Comments144
View on HN
No posts found.

the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.

none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.

therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.

efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.

what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.

it was a mistake to respond, over 90% of the controversy is performative as it nearly always is. the ideologues will posture and claim that they are dropping mullvad while being unlikely to follow through.

most people wouldn't even know the controversy happened nor care. with the response, you now trigger various human instincts that draw attention, promote further gossiping and wondering where the fire is with all the smoke.

    > sympathetic to anti immigration sentiments
why are you rhetorically laundering illegal immigration or nebulous refugee loopholes through the reputation of legal immigration?

I expect that there will come a time where open source is not merely not winning but there are no tokens available for sale at all. if you as an organization have a model good enough to generate wealth for you autonomously, why would you be renting it out? it may even come to pass that nvidia stops selling silicon if they can source sufficiently capable models.

one of their "Curated Extensions" is Tampermonkey, a non open source userscript engine. when there are multiple open ones available.

like attracts like?

it's not surprising as LLM's will generally fail you when you yourself don't know what you want. it's surprising how many people and organizations just blunder around scribbling code without clear goals in mind - for such people LLM's are a net negative as they will just end up with even more scribbles and no hope to make sense of it all.

    > The internet is useful and we still had a dot com crash.

    "we reached a point where there will always be demand for tokens."

is it your position that the dotcom crash would have happened had there been demand for more internet at the time?

we reached a point where there will always be demand for tokens.

suppose the current frontier models are final and there is no more progress. it still wouldn't matter because the cheaper the tokens the more demand there will be which will drive the silicon demand onward.

consider what you can do if a billion tokens cost cents: you could create a scaffold that would literally burn ungodly numbers of tokens on every conceivable angle. you could create scaffolds that don't just generate the next token but perform MCTS then at the end you keep the best results.

another way to look at it, we know that there is a document in the space of all possible documents which holds an answer to every conceivable solvable problem. but the combinatorics prohibit random exhaustive search. LLM's make exhaustive search somewhat more tractable, because there will probably exist a prompt/chain of thought which can yield it (especially when you can always get experts to poke it with cheap ideas).

creating silicon fabs is probably the most reliable investment at the moment, because there will always be room around the sun. orbital launch is probably less reliable because ocean floors still have a lot of room but it will depend on nuclear regulatory innovations. amusingly, had SpaceX actually gone up it would actually draw parallels to the dotcom bubble as being too soon.

note that tokens aren't just text anymore, MCP is very effective. if nothing else with billions of cheap tokens you will be able to exhaustively create endless virtual worlds with just Unreal Engine MCP - this is the absolute floor. entertainment without end.

and this all with just the current generation models frozen in time. but suppose only the intelligence is frozen, suppose training still works. if it's cheap enough you can always just dump compute on RL on the MCP directly. and the MCP can be something like silicon design...

the intent is that no one owns those keys, your silicon should be the only entity in "possession" of those keys.

but no one uses blind signatures for attestation so it can be used to fingerprint your device's serial. they do try to make it hard. but generally you should assume that if whoever you are attesting to colludes with google they will obtain your HWID - and if it's google you are attesting to you should assume they have your HWID.

GOS uses a proxy for attestation, but it does absolutely nothing for this threat model.

PS: DRM is even worse, there is no intermediary and the APIs are open to all apps. you probably need to be a well resourced intel agency to make use of it as you need to source a valid DRM license server certificate. technically, actual license servers are in violation of their agreements with google, apple, etc if they use the license request for fingerprinting. but they do retain the ability to blacklist silicon (invalidate pirate devices from pirated media watermarks).

while it's plausible that Sam Altman could find someone to covertly exfiltrate privileged data and then somehow covertly train on it while the rest of the developers remain ignorant it would all come to naught when the companies from whom they stole data probe the model with questions it should not be able to answer.

I don't believe anyone knows how to train the model in such a way that it's guaranteed not to remember any specifics while still having the training run be worth anything.

Arc AGI are simple games, the hardness comes from the input being basically adversarial to LLM training. if you use an LLM scaffold that removes the adversarial part you are measuring something else.

the harness basically outsources the alien nature of what the LLM is asked to do to algorithms it writes. this would actually be impressive if you got it to do that for a much more complicated game than Arc.

with this harness the ARC AGI test becomes a test of whether or not the model can work out the transition rules in a very simple game.

    > simulator the model builds is comparable to the mental model of the game humans create
then they should try to use that for a more complicated game than Arc AGI. Arc games are simple by design, if you have the model simulate them they become trivial.

it looks like what they are doing is using a frontier model to write a simulator for a game and then solve using it.

it's not as impressive as it looks. the goals of Arc-AGI-like constructs is to get an IQ-like figure using raw'ish 2D measurement 'games' in the hope that it would signify something meaningful.

what this harness does is get the model to write a simulator first, it's measuring something entirely different.

a certificate that data was destroyed is absolutely worthless no matter who it comes from.

what kind of sorcery do they have to let them determine that no backups were taken before they arrived to "certify"?

    > geoengineering might lead to geopolitical conflict
I have heard of this but haven't heard a credible scenario describing it. if Europe or the US or China wanted to do it they could just be flying in circles in their own territory to deploy it safely (at the cost of some efficiency probably). only a nuclear power would likely dare to unilaterally do geoengineering.

it remains to be seen who will start the geoengineering effort, arguably Russia would benefit from global warming and so will Canada, so the US food supply should also remain secure.

I wouldn't be too worried about it this century at least. if things look like they are going to get bad we will resort to geoengineering, and once we get a taste for it we will further optimize global temperatures to our liking.

perversely, global warming earlier than expected is a good thing [for us] as it will wipe away all meaningful opposition to geoengineering.

note how according to them having 50C days is a likely outcome. no one will tolerate this. sending the sulfur planes is assured at that point. you wouldn't even need to try and convince people with harvest yields.

Precursor 9 days ago

that's actually how you do it. adversarial systems like those are prime candidates. one agent develops detection mechanisms and the other agent defeats them. progression signal is easy to get.

and you bootstrap with existing javascript detection engines.

the challenge is usually the human input data, your objective is to be clustered among the humans and for that you need to know what humans look like.

this is not an open ended arms race, it will end once the bots approximate humans to a sufficient degree - false positive rate for detection will become unacceptable even if the detection system is slightly ahead.

with bubblewrap it's better to pull a rootfs from dockerhub (eg. debian:unstable) then bootstrap it into a fully fledged distro rootfs living in its own folder. install the AI agents right into it, then create launch scripts that invoke bwrap with the distro rootfs (readonly) and a custom read-write /home/user and run whatever you want inside it - it will not see anything important outside the directory you give it. you can also run multiple agents each invisible to the others.

for bonus points you can uplift the bwrap container into an actual sandbox by invoking gvisor (`runsc ... do ...`) from inside it, or a virtual machine monitor like muvm. I'm really fond of this pattern because you can trust bwrap to set up the environment, then you just need a sandbox tool to lock it down.

bwrap by itself will probably be sufficient against most adversaries as assuming proper config it would require committing to using a linux kernel 0day to escalate privs.