I think Tibo was just keeping all else fixed and it’s an illustrative example rather than a perfect real-world trajectory.
HN user
chaos_emergent
In character, in manner, in style, in all things, the supreme excellence is simplicity.
- Henry Wadsworth Longfellow
CTO @ BuildBetter.ai
The chart makes sense and is describing the cumulative cost of a trajectory. Cache tokens are created when a trajectory’s prefix is used more than once. A larger pre-compaction context window means that a greater number of cache tokens are used per turn, and a larger number of turns are completed before compaction runs. So you get a cumulative cost that grows quadratically until the compaction event.
RLHF is an increasingly small part of training though? From what I understand most of the capability gain is in RLVR
I think money’s marginal utility just changes from a vehicle for material comfort to a way to keep score.
Totally agree with you. K8s ends up being the simplest solution for a very complex problem
In addition to all of OP’s points, another reason k8s is getting popular is that LLMs have made them easier to use! It's reasonably well represented in the dataset and there are pretty strong monitoring and observability tools and verification gates to make sure that you've specified your cluster specifications correctly.
This morning, while putting in daily contacts, I realized I was down to my last few pairs. Still standing at the bathroom sink, I used the ChatGPT app on my phone to voice-command the Codex app on my computer to book an optometrist appointment for Saturday. It visited the website, figured out the API, and booked it.
No, this isn’t the same as planning a multi-day vacation. But it is plainly useful today, and it feels very close to handling more complex tasks like that.
Maybe the difference is the model and the harness. At this point, I’m starting to think some people are either gaslighting themselves about how useful these systems are, or overgeneralizing from one narrow setup. Gemini, for example, seems especially weak at agentic behavior.
The wholesale dismissal just feels strange coming from the HN community I’m used to.
We can look to other forms of automation to get a sense of what to do. For example, planes largely fly themselves and a loss of life due to manufacturing errors from the manufacturer would deem them liable for those deaths. Seems like the solution here is large penalties and generally broad disincentives for incurring harm.
Yeah I think this is more coherent than people realize. Economically relevant knowledge work is things that humans find cognitively demanding. Otherwise they wouldn't be valued in the first place.
It ties the definition to economic value, which I think is the best definition that we can conjure given that AGI is otherwise highly subjective. Economically relevant work is dictated by markets, which I think is the best proxy we have for something so ambiguous.
I wouldn’t really call it “DIY” per se, k8s has the resource API and you can create whatever scaling policies you want to with it, but I do see how that’s not obvious when it’s advertised as ‘batteries included’
I posted just that on the Twitter feed but then I realized that railroad started at the beginning of an industrial revolution where labor was a far larger portion of GDP compared to industrial production. So it kind of makes sense that the first enabling technology consumed far more GDP than current investments do, even on a marginal basis.
Thinking in counterfactuals, how would the hype around Codex would be different if it was organic and because they had built a genuinely good product? Asking as someone who genuinely loves Codex and has been in the OpenAI camp for months after buying a Claude Max plan from November to February.
Yeah I hate the title because it almost verges on clickbaity because one assumes that he's making the assertion that AI has a moral stance in the first place, versus AI being morally neutral and driven by its wielder
The reality is that none of that shit matters if you can build a product that people use and want to pay for. I would back someone who has made a dollar off of a product over someone who has built a great product that no one uses 100% of the time.
The reality is that you can make a successful business with okay engineering and great product insight. It's much more difficult to build a successful business with great engineering and poor product insight. Getting people to use and pay for what you've built gives you the product insight that you need.
All thinking does is burn output tokens for accuracy
“All that phenomenon X does is make a tradeoff of Y for Z”
It sounds like you’re indignant about it being called thinking, that’s fine, but surely you can realize that the mechanism you’re criticizing actually works really well?
An alternative but similar formulation of that statement is that Anthropic has spent more training effort in getting the model to “feel good” rather than being correct on verifiable tasks. Which more or less tracks with my experience of using the model.
Most of the value I’ve gotten out of is has been observability. Graph and DAG workflow abstractions just help OTel structure your LLM logs in a clean hierarchy of spans. I could imagine figuring out a better solution to this than the whole graph abstraction.
Other than that I’m not too sure.
Have you considered that Claude set up a crontab that does that programmatically? Every 10 mins seems awfully, idk, regular.
The writer mentioned that people's intuitions about the distribution of land and where it's most valuable are wildly off, what exactly are people's intuitions that run counter to the data presented? It seems fairly intuitive to me that property values, as you get closer to an urban center
Human-driven research is also brute-force but with a more efficient search strategy. One can think of a parameter that represents research-search-space-navigation efficiency. RL-trained agents will inevitably optimize for that parameter. I agree with your statement insomuch as the value of that efficiency parameter is lower for agents than humans today.
It's really hard to imagine that they __won't__ exceed the human value for that efficiency parameter rather soon given that 1. there are plenty of scalar value functions that can represent research efficiency, of which a subset will result in robust training, and 2. that AI labs have a massive incentive to increase their research efficiency overall, along with billions of dollars and really good human researchers working on the problem.
I feel like the dropping cost of using AI doesn't tell the full story - I feel like I'm using agents easily a hundred to two hundred times more than I used chat interfaces. The build-out seems entirely reasonable if we believe that there's going to be a similar increase in usage through the rest of the economy as user interfaces are figured out for these models.
Probably not, they’re like four years old and they’re 2500 people at the company. My guess is that there are but a handful of PMs.
100% All of the people who are floored by AI capabilities right now are software engineers, and everyone who's extremely skeptical basically has any other office job. On investigating their primary AI interaction surface, it's Microsoft Co-Pilot, which has to be the absolute shittiest implementation of any AI system so far. As a progress-driven person, it's just super disappointing to see how few people are benefiting from the productive gains of these systems.
It’s almost like the iteration loop refines itself between checks notes in Sutton search and learning
Not at all, the limitation is software to get the model on the chip and executing correctly. My bet is that they had a FDE who specializes in the chip implement Spark’s architecture on device.
I think it's a beta so they're trying to figure out pricing by deploying it.
I grew up on a rural farm in California with a dial-up connection that significantly hampered my ability to participate in the internet as a teenager. I got Starlink installed at my parents' house about five years ago, and it's resulted in me being able to spend considerably more time at home.
Even with their cheapest home plan, we're getting like 100 Mbps down and maybe 20 to 50 up. So it's just not true at all that you would have connections that are a megabit or two per second.
Calling AI an unproven market is a wild statement. My mother and every employed person around me is using AI backed by Nvidia GPUs in some way or the other on a daily basis.
When you say “in” them, are you referring to their training data, or their model weights, or the infrastructure required to run them?
Isn't the struggle of sifting through a labyrinth of physical books and learning how and where to find the right answers part of the learning process?
I would argue a machine that short-circuits the process of getting stuck in obtuse books is actually harmful long term...