And what the author goes and shows are equalities, not tautologies.
HN user
Mithriil
Currently a sofware dev.
Statistical learning, optimization, algorithms, computer sciences, and other nerdy stuff.
This is of the 26th of May.
If you hold the eraser for a second at the center, I find that it destroys the image more often than not.
It's relevant to the "thousand monkeys on a thousand typewriters".
'Always free' does not sound like an opinion.
The ratio of AI startups at YC surprised me... (slide 48). This is a clear trend.
Probabilistic analysis can carry you very, very far in doing something that looks like logical inference at the surface level, but it is nonetheless not logical inference.
A statistical approximation of logical inference (as vague as I state it) could (and will) very well pass for logical inference, at least for the common people, whose logic skills are far from perfect.
Also, humans are certainly not capable of the perfect logical inference you speak of. And I get the irony of what I'm saying with such certitude. Logic is still framed in axioms that are framed in languages, we'll never truly get there. Ah, but absoluteness gets in the way of practicality.
Yet, here we are with a tool, that is maybe not at its prime yet, that equals and beat many human beings at logical inference on some problems that are pragmatically relevant. Should I say symptoms of logical inference at that point?
As to why LLMs capacity for (apparent) logical inference is only limited to specific use cases, I don't have a clue. But I'd like to argue that, humans are like that too.
nobody understands the fundamentals
Funny statement to be found in the discussion about... research results on the fundamentals.
Asymptotics has been used to validate tons of statistical tools. This is just another tool being validated.
If you have a tool that you don't know works when data increases (n-> infinity), then you shouldn't use it.
So practicaly, I believe it has serious implications.
I don't think that this is true. You need an infinite number of dimensions for this (think Taylor's expansion, Fourier expansion, infinitely wide or deep NNs..)
As someone who worked with Nadaraya-Watson regression in the pass, the result that infinitely wide NNs converges to kernel regression baffles me.
Add the feature of doing a high five for the rare cases when it's actually good.
instantly
Shor's and Grover's still are algorithm that require a massive amount of steps...
I would expect such a law to be lobbied to death.
The Google's n-gram dataset link is outdated. You can get them here: https://storage.googleapis.com/books/ngrams/books/datasetsv3...
The half-life idea is interesting.
What's the loop behind consolidation? Random sampling and LLM to merge?
Bayesian network is a really general concept. It applies to all multidimensional probability distribution. It's a graph that encodes independence between variables. Ish.
I have not taken the time to review the paper, but if the claim stands, it means we might have another tool to our toolbox to better understand transformers.
Worry not, I came here full speed after the first paragraph to say the same thing.
Whether foreign companies pay or not for the tarrifs is clear here. However, I want to point that not receiving income from reduced trade is an impact of its own. An indirect way to pay for the tariffs, so to speak.
I think what people tend to forget when speaking of inevitability is that the scope of their statement is important.
*Existence* of a situation as inevitable isn't so bold of a claim. For example, someone will use an AI technology to cheat on an exam. Fine, it's possible. Heck, it is mathematically certain if we have a civilization that has exams and AI techs, and if that civilization runs infinitely.
*Generality* of a situation as inevitable, however, tends to go the other way.
"Mastery, even partial, is one of the few genuine avenues toward agency."
Philosophical claims have been made around this point. See, for example, "The Moral Obligation to Be Intelligent", an essay by John Erskine.
So many problems would be solved if a fraction of people would be more inclined to understand what's in front of them.
But then the pinch of resistance makes an island of likewise thinkers. And there doesn't need to be more than .05% of techies to make great products that otherwise anti-correlate with what people claim as inevitable.
We should stop with over-generalization like "The future is defined by the common man on the street." It's always much more complex than that. To every trend, there is a counter-trend (even sometimes alt-trends that are not actually opposites).
Actors that go against the current, for the sake of going against the current, exist. Always a minority, but never negligeable, I believe.
Just started going through the tutorial, and it is, indeed, mega-cool.
Btw, here's the identity matrix of size 3:
˙⊞=⇡3
(It takes the range [0,1,2] then outerproducts it with itself through equality.)
My opinion on the "Attention is all you need" paper is that its most important idea is the Positional Encoding. The transformer head itself... is just another NN block among many.
TIL that Earth crust is pushed down by glaciers, and that when glaciers subsides, the crust swells up a bit over years from the missing weight, pushing away water, slush and sliding glacier even faster.
Hard to fathom how "fluid" our ball of magma really is.
since its commands were an inscrutable jumble of ill-fitting incantations, and it has remained this way until today
What command is he talking about? When you get that git is a graph manager, it gets really easy to manage, very quickly..
I like seeing something along the line of constructive logic in the wild (i.e. not (not p) != p).
Sounds like a development à la Cory Doctorov.
(Such as in his novel Attack Surface.)
For those that argue that concepts are not orthogonal or quasi-orthogonal, then see the quasi-orthogonal case as the worst-case: "if all concepts were black and white, then how many can we fit in k dimensions". When there are nuanced concepts, then they will fit in between these quasi-orthogonals ones. What's argued here is thus a lower-bound.