If it's memorized your benchmark then your benchmark is bad, it's not cheating
HN user
aoeusnth1
What about all his other articles that had f-bombs and the predictive utility of used toilet paper?
It doesn't matter what went wrong- he is not a credible leader. If you are going to lead doctors, you should be a doctor. It doesn't matter why you didn't become a doctor.
"To all my haters who said I mispredicted the capabilities of these systems, rest assured: my statements have no predictive worth at all! I am not saying anything concrete!"
Did I miss where OpenAI plagerized the disproof of the planar unit distance problem from?
We're discussing whether they are models or not, not whether they have goals and agency. A language model does form a model of who you are and what you're thinking, because language is causally connected to those aspects of the generating distribution and modeling those aspects reduces cross-entropy.
RL provides the goals and agency. Pretraining provides the model.
When I read your reply, I’m also modeling language. Tokens are just the discretization of the model’s eyes and ears. My brain does a huge amount of work to represent what’s happening in the world based on discrete information received from the outside world, just like language models do.
The universe contains subsystems which can be described as eluded in the sense that we can take the intentional stance on these systems and describe their observable behavior as being in a state of illusion of separation.
There is nothing which makes either of them are “you.” The feeling of Self is a useful predictor which a physical subsystem uses to nagivate the world and predict observations. “I” is not a physically real label which attaches the “you”-ness to physical systems, the physical systems simply are, and are inherently first-person in character. The only real you is the global quantum wave function, or whatever the underlying real stuff is doing.
Materialism directly implies no-self and Advaita Vedanta schools of thought.
The model is the thing which is learned in order to make the probabilistic prediction with low entropy.
AI agents
I believe they were using their teeth (or lack thereof) to reference their visible meth addiction. See GP, "2) meth is neurodegenerative. heavy users end up with a permanent disability."
The article explicitly and repeatedly affirms that illegal THC vapes are dangerous because of Vitamin E Acetate, which is used as a thickener agent. TFA points out how the NYT article carefully weasels its way around admitting that the THC vaping was the cause of the teenager's lung injury - the NYT is attempting to get the audience to associate the harm with legal nicotine vapes.
Does that make more sense to you now?
Meat is optional.
And idle% is causally connected to whether you make a request or not, surely? I don't understand how your mental model works.
Babies are not conditionally created to solve a problem.
SWE-bench pro is ~20% higher than the previous .1 generation which was released 2 months ago. For their SWE benchmark, the token consumption iso-performance is down 2x from the model they released 2 months ago.
If this is a plateau I struggle to imagine what you consider fast progress.
Why is a map lookup measured in ms instead of us? Something is seriously wrong with these benchmarks.
Wobbly assumption that increasing the size of these models yields better performance.
I'm assuming you disagree that larger models are better? Can you expand on what indicates that AI will hit a wall in scaling given the evidence of the last 9 years of scaling transformers (or other models)? Where on the plot does the line go from exponential to flat?
How would you know if it wasn't an extrapolation of current knowledge? Can you point me to somethings humans have done which isn't an extrapolation?
You could reimplement it as public domain on your machine, and then edit it by hand and copyleft your own edits.
Citation needed? Tall tasks are standard practice to improve utilization and reduce hotspots by reducing load variance across tasks.
I don't think he's misleading, I think he is valuing Claude's contributions as essentially having cracked the problem open while the humans cleaned it up into something presentable.
I think your theory might be missing an extremely relevant and timely counterexample?
Unclear how much damage the designation will do to their dealmaking ability in the meantime. How long will it take for the court to reverse order?
I imagine they're also benchgooning on SVG generation
It absolutely has quasi-identity, in the sense that projecting identity on it gives better predictions about its behavior than not. Whether it has true identity is a philosophy exercise unrelated to the predictive powers of quasi-identity.
I'm an environmentalist and I agree with this framing. The solution is going to be painful and must increase prices on products and services that fossil fuels are currently the cheapest solution for. If you're not willing to personally sacrifice anything to reduce fossil fuel consumption you can see why carbon taxes are not popular, right? France's protests against them, for example, are a good example of a populist reaction against attempts to regulate the economy to have less emissions.
Amazon has 350K corporate roles, so yearly layoffs of 16K is only 5% - if you assume some modest re-hiring in lower-cost locations, this is just a relatively standard (at least lately) pivot out of high-cost US roles into other lower-cost economies.
It's easy to miss the value in something you don't do. I do fermi estimates in my head all the time and it would be exhausting to constantly pull out my phone to calculate things, to the point that I would stop attempting it as much as I do.