HN user

itkovian_

140 karma
Posts3
Comments50
View on HN

Contrarian view; I think he’s right. Many of these ideas it’s almost shocking how many you can find sketched out in his old papers. To the point where I think it’s very wise to read all his work to see what hasn’t showed up yet but likely will. Artificial curiosity for example.

Projects like pluralis agora solve this problem. Really what you want is the model to be collectively owned and governed, not local

In many cases the GPs are not, at least nowhere near as much as you’d think. Obviously there are power laws here as well. Non partners, forget about it.

The other thing is it is maybe the most nepotistic industry out there. Which somewhat makes sense given the actual job is so relationship based.

The US isn’t what it used to be. It’s definitely not the best place in the world to live for quality of life, on basically any metric.

The requirement of being permanently obligated to pay us taxes on global income, if you have any kind of global mobility, is not worth it when you look at the situation objectively. The US is the only country that requires this, and signing up is voluntarily.

So while US immigration continues to act as though people will jump through any hoop they put up in order to be granted the extreme privilege of being able to live in the country indefinitely, it’s worth realising it’s not the 70s anymore and thats a goal many people are no longer optimizing for. In fact the opposite - the most talented people I know are all planning their lives to not settle long term in the US.

The other thing people don’t understand is exponential curves are self similar. The start of an exponential looks like an exponential. People always look at and think ‘well that’s it it’s exponential now, have missed it, can’t sustain’. Nope.

Good example of this is number of submissions to neurips/icml/iclr. In 2017 that curve was exponential.

Whether it’s actually 20% or not doesn’t matter, everyone is aware the signal of the top confs is in freefall.

There are also rings of reviewer fraud going on where groups of people in these niche areas all get assigned their own papers and recommend acceptance and in many cases the AC is part of this as well. Am not saying this is common but it is occurring.

It feels as if every layer of society is in maximum extraction mode and this is just a single example. No one is spending time to carefully and deeply review a paper because they care and they feel on principal that’s the right thing to do. People did used to do this.

I don’t think people understand the point sutton was making; he’s saying that general, simple systems that get better with scale tend to outperform hand engineered systems that don’t. It’s a kind of subtle point that’s implicitly saying hand engineering inhibits scale because it inhibits generality. He is not saying anything about the rate, doesn’t claim llms/gd are the best system, in fact I’d guess he thinks there’s likely an even more general approach that would be better. It’s comparing two classes of approaches not commenting on the merits of particular systems.

Article doesn’t say jobs aren’t about to be evicerated, says this is already happening and it’s due to capitalism, a lack of consumer protections and we require more government regulation. This never made any sense to me because we don’t have to guess how this would go - the experiment is being run in Europe right now.

Also the core of the argument is wrong, ai is clearly displacing jobs this is happening today.

The reason for this is it’s horrifying to consider that things like the Ukrainian war didn’t have to happen. It provides a huge amount of phycological relief to view these events as inevitable. I actually don’t think as humans are even able to conceptualise/internalise suffering on those scales as individuals. I can’t at least.

And then ultimately if you believe we have democracies in the west it means we are all individually culpable as well. It’s just a line of logic that becomes extremely distressing and so there’s a huge, natural and probably healthy bias away from thinking like that.

I think the better analogy is if you had someone with a superhuman, but not perfect memory read a bunch of stuff, then you were allowed to talk to the person about the things they’d read, does that violate copyright? I’d say clearly no.

Then what if their memory is so good, they repeat entire sections verbatim when asked. Does that violate it? I’d say it’s grey.

But that’s a very specific case - reproducing large chunks of owned work is something that can be quite easily detected and prevented and I’m almost certain the frontier labs are already going this.

So I think it’s just very not clear - the reality is this is a novel situation, the job of the courts is now to basically decide what’s allowed and what’s not. But the rational shouldn’t be ‘this can’t be fair use it’s just compression’. Because it’s clearly something fundamentally different and existing laws just aren’t applicable imo

Completely agree and think it’s a great summary. To summarize very succinctly; you’re chasing a moving target where the target changes based on how you move. There’s no ground truth to zero in on in value-based RL. You minimise a difference in which both sides of the equation have your APPROXIMATION in them.

I don’t think it’s hopeless though, I actually think RL is very close to working because what it lacked this whole time was a reliable world model/forward dynamics function (because then you don’t have to explore, you can plan). And now we’ve got that.

Saying we should tokenize different modalities the same would be analogous to saying that in order to be really smart, a human has to listen with its eyes. At some point there has to be SOME modality specific preprocessing. The thing is in all current sota arch.’s this modality specific preprocessing is very very shallow, almost trivially shallow. I feel this is the peice of information that may be missing for people with this view. In the multimodal models everything is moving to a shared representation very rapidly - that’s clearly already happening.

On the ‘we need to do rl loop rather than a generative model’ point - I’d say this is the consensus position today!

I don’t want to bash the guy since he’s still in his phd, but it’s written in such a confident tone for something that is so all over the place that I think it’s fair game.

Like a lot of the symbolic/embodied people, the issue is they don’t have a deep understanding of how the big models work or are trained, so they come to weird conclusions. Like things that aren’t wrong but make you go ‘ok.. but what you trying to say’.

E.g ‘Instead of pre-supposing structure in individual modalities, we should design a setting in which modality-specific processing emerges naturally.’ Seems to lack the understanding that a vision transformer is completely identical for a standard transformer except for the tokenization which is just embedding a grid of patches and adding positional embeddings. Transformers are so general, what he’s asking us to do is exactly what everyone is already doing. Everything is early fusion now too.

“The overall promise of scale maximalism is that a Frankenstein AGI can be sewed together using general models of narrow domains.” No one is suggesting this.. everyone wants to do it end to end, and also thinks that’s the most likely thing to work. Some suggestions like lecuns jepa’s do suggest to induce some structure in the arch, but still the driving force there is to allow gradients to flow everywhere.

For a lot of the other conclusions, the statements are literally almost equivalent to ‘to build agi, we need to first understand how to build agi’. Zero actionable information content.

It's an extraordinary claim. I think the reason I dismiss it as unlikely as when I look back at the steel dossier and muler investigation 1) if there was something, it's very likely they would have found it then 2) in hindsight both investigations were completely discredited and shown to be largely a institutional response to the shock which was 2016. This current re-emergence of 'trump is a Russian agent' is kinda surprising in that context. 3) I think the current behavior can be explained by a desire to end the conflict, while feeling no particular allegiance to Ukraine.

All competitive open models today share a common property; someone spent a large amount of money to train them and then released the model for free.

I don't understand why the argument continues to be we will have a rich ecosystem of base open source models; unlike opensource ai which is individuals donating time, opensource ai requires someone to donate very large amounts of capital.