HN user

martinald

7,143 karma

Feel free to reach out: martinalderson AT gmail DOT com

meet.hn/city/gb-Cardiff

Posts94
Comments1,720
View on HN
martinalderson.com 6d ago

Winners and losers in the coming AI margin collapse

martinald
2pts0
martinalderson.com 16d ago

GLM 5.2 and the coming AI margin collapse

martinald
694pts469
martinalderson.com 29d ago

Expert-aware quantisation: near-Q4 quality at near-Q2 size?

martinald
3pts0
martinalderson.com 1mo ago

xAI is looking more like a datacentre REIT than a frontier lab

martinald
692pts552
martinalderson.com 1mo ago

Is datacentre sovereignty that important?

martinald
2pts0
martinalderson.com 2mo ago

Managed agents are the new Lambda

martinald
1pts0
twitter.com 2mo ago

New Claude Code programmatic usage restrictions

martinald
52pts42
martinalderson.com 2mo ago

Local LLM Speed Calculator

martinald
2pts0
martinalderson.com 2mo ago

Open weights are quietly closing up – and that's a problem

martinald
5pts1
martinalderson.com 2mo ago

29th August 2026: A Scenario

martinald
3pts0
limitacions.vatard.com 2mo ago

Map of track defects across the Spanish rail network

martinald
5pts0
martinalderson.com 3mo ago

Figma's woes compound with Claude Design

martinald
122pts98
renewables-map.robinhawkes.com 3mo ago

UK total wind generation record beaten today

martinald
57pts33
martinalderson.com 4mo ago

Using agents and Wine to move off Windows

martinald
2pts0
martinalderson.com 4mo ago

Why Claude's new 1M context length is a big deal

martinald
3pts0
martinalderson.com 4mo ago

Is the AI Compute Crunch Here?

martinald
2pts0
martinalderson.com 4mo ago

Why on-device agentic AI can't keep up

martinald
2pts0
martinalderson.com 4mo ago

Using OpenCode in CI/CD for AI pull request reviews

martinald
2pts0
martinalderson.com 4mo ago

Which web frameworks are most token-efficient for AI agents?

martinald
3pts0
martinalderson.com 5mo ago

Who fixes the zero-days AI finds in abandoned software?

martinald
1pts0
martinalderson.com 5mo ago

Anthropic's 500 vulns are the tip of the iceberg

martinald
2pts1
martinalderson.com 5mo ago

Attack of the SaaS Clones

martinald
2pts1
martinalderson.com 5mo ago

Self-Improving Claude.md Files

martinald
2pts0
martinalderson.com 5mo ago

Two kinds of AI users are emerging

martinald
355pts341
martinalderson.com 5mo ago

Turns out I was wrong about TDD

martinald
3pts0
martinalderson.com 5mo ago

Why sandboxing coding agents is harder than you think

martinald
2pts3
martinalderson.com 6mo ago

The coming AI compute crunch

martinald
1pts0
martinalderson.com 6mo ago

Which programming languages are most token-efficient?

martinald
6pts0
martinalderson.com 6mo ago

I ported Photoshop 1.0 to C# in 30 minutes

martinald
1pts0
martinalderson.com 6mo ago

Why I'm building my own CLIs for agents

martinald
4pts0

I don't think so. According to some very basic research there are around 8bn searches a day, or 250bn a month.

Let's assume Google serves AI overviews on every SERP (they don't) and don't cache them (they do, afiak).

And let's assume that each AI overview is 2000 tokens (blended input/output), that's 500T tokens a month.

It's rumoured that anthropic is serving somewhere close to 10Q tokens a month.

Now it may be that AI overviews uses vastly more tokens than that per search, but I doubt it based on speed to render the overview.

My very rough napkin math on this is that maybe AI overviews is consuming 100T tokens/month max (after adjusting for caching and SERPs that don't have them), which would be 1% of Anthropic token volume.

If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.

Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).

Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...

Yes agreed - I wrote this up a while back https://martinalderson.com/posts/whats-going-on-with-gemini/

My view then was they are optimising the models for inference ability on their own hardware AND use cases, which is often speed and time to first token.

They've somehow seemed to end up with terrible compute shortages, which again is surprising given how good Google is at infra deployments AND have their own hardware. From rumors out there they are turning down enterprise deals for Gemini because they don't have the compute.

The problem is they're falling further and further behind on frontier class on coding especially, and since I wrote that article it's got even worse with open weights models undercutting them on price AND intelligence.

MCP makes a lot, lot more sense when you think of it as as a auth standard and not a comparison with CLIs. It obviously does more than just auth, but having standardised auth (which CLIs definitely do not) is the real 'killer' feature.

I don't think that's inevitable with RL.

Imagine in C# you are training the model with RL loops in a harness. One uses C#12 and one uses C#15 (when released), with union types (and importantly - includes the release notes in the harness). Union types if used properly will reduce the amount of bugs/issues in theory from "forgetting" about certain conditions, because the compiler will enforce that better.

In theory, the one with union types will "win" (less errors/fewer edits required) in certain conditions, which makes it more likely to be used going forward.

Basically I think it looks less about 'ingest lots of slop' but 'how do we give our RL harnesses the best possible tools and documentation to make the best* code'. I think this is exactly what good engineering teams do.

For example, if I put 'use C#15 union types' in my CLAUDE.md/AGENTS.md on a .net11 preview project, it is very good at using them when required. It doesn't take much instruction for an agent to use new language features.

_However_ what it does do is change the language feature adoption from 'many developers' to 'eval writers and people that put features into CLAUDE.md'. This obviously changes things massively - though I sort of suspect very few developers _actually_ adopt new language features quickly.

Final thought is that I think we may see a lot of different features being adopted. Instead of what makes code readable to humans, what makes code better on evals. I sort of suspect we'll end up with some Frankenstein language in the future that is difficult for humans to write but agents can write extremely well, with esoteric language features that no (sane) human would think to use.

Micron said that they tried to tell 2 of their largest customers (one almost certainly Apple) that the prices they were demanding would result in the cancellation of a lot new construction in 2023, which wasn't in the industries best interests.

It is sort of Apple's fault. They are probably the biggest single buyer of DRAM and NAND globally and they pride themselves on their supply chain management under Cook.

It seems they over optimised this too far.

Om Malik has died 27 days ago

Really sad. I grew up reading his writing. I emailed him some thoughts on one of his blog and he immediately replied in a lovely way very recently. What a shock and a loss.

Well, the EU insists that track & train operations are separate. (ironically the UK _is_ combining passenger operations and track somewhat back together, which is only possible because of brexit).

The bigger issue tbh is the enormous cost inflation in civil engineering in general. This seems to be a problem everywhere. There's no doubt some of this is caused by material cost increases, labour shortages etc, but I'd say the huge amounts of regulation added over the years is really a core driver of this.

Yes agreed, for example, there was an interesting table on the starlink page I used to check every so often showing which countries had access to starlink as it was rolled out. Was interesting to see the expansion.

Of course, some editor decided it was 'marketing' for starlink so it got deleted despite loads of people protesting. It was the only source I could find easily for showing which country got starlink when.

A huge list of prose is still on the page (not marketing?) showing the updates in a very hard to read and not comprehensive way. Something is really quite wrong over there.

Yes 32B dense is a weird one to choose.

But in reality, 32B dense is very similar* to 32B activated on MoE in terms of inference costs. And I highly suspect eg Opus is around that level of active params.

A 284ba13b model at scale, is almost certainly cheaper to serve than a 32b dense model.

*as you can shard the model across multiple GPUs at scale. but in reality you have some loss of efficiency from GPU coordination and expert routing

If you set aside political menace, this is a huge problem with Anthropic's strategy.

You _cannot_ say that Mythos is super dangerous and can only be rolled out to certain people, but then release Fable with anything other than bulletproof cyber denials.

Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work.

So you've ended up in a situation where Anthropic are simultaneously claiming it's a incredibly dangerous model _and_ there are (minor, potentially) problems with the security "protections".

As technical people we understand that nothing can be perfect, esp in LLM world. But all my non technical friends were really confused how they had managed to make the model "safe" so quickly when it was released and the general sentiment was it shouldn't have been released - and now to an outsider I think it looks like it was never safe at all to release, so I can totally see how the current US administration have got themselves very upset with it.

_Even if_ there was no political bad will, it's a bit of a silly scenario to end up in, and really quite easily foreseen.

Keep in mind Google also rents GPUs via GCP, so they could be just reselling these to GCP customers?

Thing is though, Anthropic was really against the wall with lack of compute pre xAI deal. And tbh, Gemini reliability has been abysmal which probably points to real compute shortages.

And nearly _every_ major DC project is really up against it with massive delays, etc. Stargate UAE has been badly affected by the Iran conflict.

So maybe long term this isn't a great business, but _right now_ I'm not convinced it's all financial engineering. There is a enormous shortage of compute and xAI has a load of it _available now_.

Don't think so - the 3x is a separate cap. It actually reduces it down from market cap.

Eg say spaceX has $50bn of float at $1.5T valuation. If there wasn't _any_ cap at all, the full $1.5T would be used as the market cap. With the (new) 3x cap, it means only $150bn of the $1.5T valuation is taken into account in the index weighting.

Before this change, SpaceX wouldn't clear the 10% requirement to be listed in QQQ at all. So the 3x basically allows them to be included but _does not_ increase their market cap from $1.5T to $4.5T.

Btw, for clarity, I'm not saying there isn't questionable behaviour going on here. My main point is that even if SpaceX, openai and anthropic all went to 0 (unlikely IMO), it's not going to have a material impact on people's retirements which is what OP was proposing.

But the US has never had $1T+ IPOs before. And also a huge amount of enormous private companies that don't want to go public for various reasons.

Also, the rules have changed before. It's not the first time these rules have changed.

I see both sides of the argument (it's definitely _not_ good for 401k investors if Anthropic/OpenAI/SpaceX make huge leaps in technology that allow for far higher earnings that they aren't able to access, for example).

But my main point is that these investors regardless would "only" have 5% exposure to these. That surely cannot be considered a systemic risk that the OP is inferring.

Is it? I thought the idea was diversity of risk, not "mitigating risk". You clearly don't want 100% of your 401k in OpenAI or Anthropic. But you probably do want 1 or 2% of it in, to give you the long term growth potential?

Regardless SPY is actually a pretty "risky" index fund on some measures - it pays a (very) low dividend compared to many other intl/ETF funds and is weighted very heavily towards tech stocks (atm).

If you genuinely wanted to mitigate risk you would probably not choose SPY.

Let's get it in perspective though. The S&P500 market cap is currently $70T.

Assume that Anthropic, OpenAI and SpaceX all IPO and get included in SPY with the new fast listing rules. They are likely to be worth $3-4T combined, which means 'retail' investors are going to have perhaps 5% of their portfolio in it.

_Arugably_ that's a pretty fair allocation for retail investors to have to these "moonshot" style companies.

Also - if any one of these IPOs don't go well; I suspect the other(s) will have to postpone, further reducing exposure.