HN user

marcyb5st

1,571 karma
Posts8
Comments540
View on HN

I believe part of the LLMOps (I don't like the term, but it is what it is) should be building a failover plan with proper testing that check tools trajectories and such. If you have these then you can sort the good enough models from cheaper to more expensive and have the failover you mentioned.

I saw people bulding a mapping of model->{{prompts}, {tools descriptions}, ...}, but that, to me, it feels extreme. I believe it is the model that needs to adapt to your prompts after a certain point. Models that fail to do so won't get our api requests as they will be out of the failoever roster.

It is the difference between giving an LLM an epic and say "You figure it out" and giving the single tasks' breakdown you envisioned and build incrementally on top of it.

With the latter you can, for example, say "Wait, this should be an interface because later on we need different concrete implementations". With the former, the agent doesn't do that, gets to the point where you actually need the flexibility interfaces give you and refactors everything to handle that. That is at least 2x the work/tokens. Multiply this for all the decision points you have to do to deliver a big piece of work and you have your bagillion tokens consumed.

If you like hiking and tough, but rewarding, trails, consider the one to the base camp of Everest, or the roundtrip around the Annapurna. I did both and they are truly amazing without the craziness that has become climbing those Mountains. When I reached the basecamp it really felt like being in a weird Venice or Florence, with "tourist groups" and their guides. Crazy stuff.

Especially the Annapurna one, you still climb up to 5+k meters in altitude, and seeing the actual mountain still towering over you is crazy once you realize.

High horse? Man, my country was under communist dictatorship and I didn't see a banana until I was 8 years old. However, despite that, I am aware of how lucky I got in life and I try to be mindful of those that are not as much.

Also, I don't believe one bit my life has been subsidized. Wealthy people don't throw money around if there is no profit to be made. At best they subsidize stuff to get you hooked before raising prices and raking the money in.

Additionally, I paid people to build my own house with money that I earned, and same for my car. And I now live in a country where we pay living wages and not force people to do 2-3 jobs to stay afloat. So I didn't exploit anyone as those people received months of salaries for them and their families thanks to my house.

The only thing I can concede to your argument is that things like my phone or laptop are cheap because people in the world are exploited while mining and in Foxconn like factories and for that I am trying to limit my tech footprint.

Let's go back to have open air sewers while we are at it. I mean, if the goal is to bring back stuff that lowered life expectancy and reduce 50+ years of progress, I don't see why not. I really believe that as arguments go, yours is truly terrible.

Additionally, you proved my point that the majority really doesn't care about the folks whose quality of life worsen because someone built a datacenter close to their homes. Or the fact that turbines still pollute the air with NOx emissions (which cause lung cancer and other respiratory problems).

You are basically saying that it is not so bad, but (I guess) it is because you aren't close to a datacenter not depending on the grid. And I trust more hearing about live accounts of people being affected, than a "trust me bro, it is not as bad". Are the accounts cherry-picked? Maybe, but they are also supported by videos that substantiate what they are saying.

And especially terrible for the communities that have to live with gas turbines or other local power generators as neighbors. Noise and air pollution constantly [1].

But fuck them, they are poor people so we don't care about them /s .

Additionally, people against data centers are accused of being paid by China [2]

[1] https://www.theguardian.com/environment/2026/feb/13/elon-mus... [2] https://fortune.com/2026/06/10/kevin-oleary-trump-administra...

Well put. I belong to the latter group as I feed small, granular tasks that I describe thoroughly to the LLM. I tried, however, to just give it a bigger scope task. Even best models produce sloppy code.

While the single functions/classes/structs/... can be well though out the code tends to lack cohesion, and especially maintainability. For instance, it never thinks: "I could put this logic in an interface/trait so that if the requirements change I can simply add a concrete implementation that satisfies the new requirements (and potentially use one of these for testing)".

It is not just the big cities. You go downtown even in small towns and the roads tend to be so narrow that a big SUV becomes a liability more than a vehicle. Without mentioning that I'm not even sure parking stalls fit big US cars lengthwise and they barely do widthwise

True, but they also took huge debts to build AI DCs and not sure if the DB part of the company can cushion such a fall. According to [1] their IaaS line of business brings 4.8B USD/quarter (so say 20B/year), but they have ~120B of debt (outstanding + new debt they are trying to find people to pay for).

They are justifying that on commitments (500+B USD), but 300B of those are tied to OpenAI. So, if OpenAI goes belly up or at least doesn't follow their crazy growth projections, they would have to find the same amount of consumption quickly to repay the interest on said debt and eventually the principal.

It is a lot of money for a company the size of Oracle (~500B market cap).

[1] https://finance.yahoo.com/markets/stocks/articles/oracle-500...

I think native speakers of Latin derived languages have an advantage given the proposed words in my run. The list was overly biased that way. In fact, many of the advance and grandmaster levels words are basically that. Latin derived words.

At least that was my experience as a native Italian speaker. My English vocabulary is good, but not great by any means and by reading books in English I know that there are plenty of words that are not derived from Latin

I think we (as in Switzerland) are preparing for a future in which there is not much snow melt/precipitations to fuel hydro production year round.

In fact, if the AMOC weakens/stops then there will be a drastic drop in precipitation across Europe and funnily enough maybe the temperature drop so much that the little snow there will be won't melt in big enough quantities.

Of course this is just a ban lift, meaning that there are no concrete plans to build one or more, but if there is a need to move "fast" (nuclear is not, I know) at least there is one less hurdle. I sincerily hope we invest in other technologies, especially now that Sodium batteries seem on their way to solve grid level storage, but I don't necessarily see this as a bad move per se.

And kill the savings of what remains of the middle class. Probably they will do it though, as it is a slow thing and is not felt by the average Joe like a tax hike or loss of benefits. So the policians won't trigger an outrage by doing so.

As an European, yeah, we probably are doing really good with basic science, but what about innovation when it comes to productivity? Why there is no AI lab (apart from Mistral) in EU? Why there is no European model (and hasn't been probably ever) in the pareto fronteer? Or any other really innovative company in the last while (I believe Spotify was the last European unicorn that transformed the landscape in the market they operate into).

Don't get me wrong, I rather lose the superpower race but enjoy my privacy and work benefits that folks in the US dream of. But the topic was superpower competition and I don't see the EU going anywhere in that front.

We are fragmented, among the top 4 EU economies 2 are struggling with debt (France & Italy), Germany economy is stagnating and the amount of bureaucracy hinders any attempt at innovation, ... .

I don't see that happening. The US debt will hinder any big expense that could keep it in any game long term.

Take AI for instance. The US grid is struggling to keep up with demand, while Chinese one has a lot of headway [1]. Usually, this could be solved by an increase in spending lasting a few years which would make the debt tick up, but that would've been an absolutely fine use of debt since it buys some shiny new infra that will pay dividends for the next 20ish years.

Now? Not possible. The US is already drowning in debt and the usual buyers are not showing up to buy it because of the Iran fiasco. With oil so expensive everyone was using their USD reserves to buy oil, not debt. Which mades interest rates go up considerably, and for a country with already ~130% of debt/gdp ratio these are terrible news.

So, I don't think there will be a great power race. Europe is fucked by both high debt, and lack of innovation. Russia is struggling already to finance a war of conquest they started. China is the only one that can run if it comes down to it (unless of course the numbers coming out of China are mega bogus, but for that I don't know enough to have an opinion).

[1] https://fortune.com/2025/08/14/data-centers-china-grid-us-in...

Yeah, big lol on the Recursive Self-Improvement.

I mean, firstly you get to have an "Agent" actually capable of really long horizon tasks without getting stuck in tools loops and having its context rot. Secondly, each trial (ie a model fully trained to convergence) costs millions and takes O(weeks). You can probably run 1 or 2 of those experiments in parallel even at big AI labs as the hardware is scarce and they are costly as mentioned before. Assuming this agent needs something like a hundred tries to just show some improvement, we are looking at years.

And you can't early stop training for candidates that are not promising due to "emerging capabilities". At some point you might get a big drop in loss even if the model has been plateuing for a while. And you can't really scale down models for running trials quickly either also due to these emerging capabilities. In fact you might create a model that is great at converging with in small trials (fewer params, fewer tokens), but that is uncapable of developing those unexpected traits. And this will likely happen as you created a sort of evolutionary pressure in this direction: if you are good at learning in the first few epochs you "survive" and get to pass down your traits to future trials.

All of this to say that recursive-self-improvement is waaaaaay out of our grasp as things stand right now. We need another one or two breakthroughs to get there (IMHO).

Would be very hard to demonstrate that they did that. If all employees move to some country with a slow legal justice system and strong labor laws, they also recreate the training data because that can be transferred, they can train another version in said country which is perfectly legal.

Can you demonstrate beyond any reasonable doubt that the model weights have been transferred? No. Will the EU judges move to extradite said individuals (and many are EU citizens)? Also no, especially in the face of spurious accusations. And even if they were open to, you can stonewall everything and you will probably outlast any US administration pursuing that.

GLM 5.2 Is Out 1 month ago

For personal use I already did a few months back. Dario is more competent than Sam, but even shadier (IMHO).

Anyway, switched to Openrouter through forgecode (or pi/opencode, the jury is still out on this one).

It will take a while, but I believe that also businesses will at least hedge against US companies basically being forced to geo-fence their models. For now is Fable, but they can include any model at any time.

I am getting familiar with Rust and so I have been playing around with Quoth (https://github.com/sam0x17/quoth) for now.

It is very basic and I am no DSL expert, but my idea was to build a graph from those complex documents (maintenance manuals) a that to decide what tools can be used for a given part on a given equipment in a given situation. If there is a path from A to Z it means you can use that tool given the circumstances. Basically the DSL is about pruning the graph as you specify things. I could have very well done without, but it is a fun project to try out rust, so I said, why not :)

Anecdotal, but here's my experience.

For personal stuff I use forgecode with openrouter. Firstly, forgecode is a much better harness than Cloude code (IMHO).

Anyway, regarding the models, my experience is that there is not much difference in terms of quality, but the cost difference is insane. At least for how I use agents. Yesterday's example is the following: I am developing a small DSL for search across complex technical documents. I wanted to add a small operator to it and thought that to give fable a spin. It burned through 13 USD and while it delivered the solution it wasn't objectively better than what Deepseek v4 did for 1.7 dollars (same exact task because I was curious).

For full disclosure, I ask agents for piecemeal stuff. Like in the DSL case, I designed the operators and then asked agents to implement them one by one. Probably if I asked to design the whole thing starting from these complex documents Fable would shine, but every time I try to give agents broader scope tasks they burn through millions of tokens, generate questionable code, which I have to spend time familiarize myself with.

Yeah I agree. As the saying goes: trust is built in drops and lost in buckets. A lot of time needs to pass while the US behaves as a proper ally if they want to go back to the status quo. At least IMHO as this is how I feel as an EU citizen.

Yeah, it is crazy to me. Yesterday I did the math how much it would take to fully replace "just" 1M SWEs: https://news.ycombinator.com/item?id=48382414 . It turns out you need 380GW of constant power (or 80%+ of US current production). And I conservatively assumed 0.5J / token, which was a number calculated for llama3 8B parameters. Yeah, hardware and models are more efficient now, but I expect SOTA models to be at least 10x that and I don't think there was a 10x in efficiency since llama 3.

All of this to say that the AI hype is not considering the energy portion of the equation enough. It won't automate everything not because it can't but because there is just not enough energy to go around unless there is a 100x or more efficiency gain just around the corner.

I think you are assuming cost per task will become cheaper and that there is unlimited energy supply.

While tokens costs are going down, the number of token burned is going up and up. Case in point Sam Altman is complaining about their top token users burning through 100B tokens per month [1]. So you have token prices going down but token usage going up 10x per year (if you extrapolate linearly from what Sam was ranting about). This is happening because people trust more and more LLMs and give them more autonomy and more complex tasks (IMHO).

So if you really need a true unsupervised agent that replaces SWEs you need how probably much more than that. Say 20x that number (2T tokens/month) for each SWE. I'm gonna focus on the energy part as this is more tangible. Trying with some realistic numbers:

- To replace 1M SWEs for a year you need 2T tokens/month * 12 months * 1M SWEs ( = 2.410^19 tokens)

- Assuming 0.5J per token you get 1.210^19J [2] (I took the number for an llama3 8B model, probably is much more for SOTA models IMHO).

- A year has 31M seconds

- Over a year that is 380 GW of constant power that is needed only for replacing 1M SWEs and that is around 80% of all the current US energy consumption (450GW). And apparently there are 47ish Million SWEs globally as of 2025 [3]

I don't think there is enough power capacity to deliver all of this without pivoting all of society into building data centers and power plants.

So unless there is some breakthrough in efficiency/intelligence (ie you need way fewer tokens for what you have to do) your job is gonna be safish at least.

Of course I pulled that 20x out of my ass, but I believe it is somewhat realistic for a truly autonomous agent(s) that replace SWEs.

[1] https://finance.yahoo.com/sectors/technology/articles/sam-al... [2] https://arxiv.org/html/2512.03024v1 [3] https://www.slashdata.co/post/global-developer-population-tr...