HN user

nl

33,425 karma

Formerly founder of an ML company in Adelaide and Sydney. Exited 2021.

Did some decentralized databases work and private ML work.

Formally CTO at a startup working on AI for strategy/policy work.

Always happy for emails or tweets at me.

http://twitter.com/nlothian

firstname.lastname at gmail.com

Posts223
Comments12,777
View on HN
newsletter.semianalysis.com 23d ago

AI Value Capture

nl
3pts0
desfontain.es 1mo ago

Noise infusion banned from statistical products published by Census Bureau

nl
899pts604
www.theguardian.com 1mo ago

The household battery revolution that could change energy bills and the world

nl
8pts0
techcrunch.com 2mo ago

Anthropic says it's about to have its first profitable quarter

nl
3pts1
epoch.ai 2mo ago

OpenAI Stargate: where the US sites stand

nl
4pts0
blog.kilo.ai 2mo ago

We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6

nl
1pts0
www.cnbc.com 2mo ago

US/China talks to make sure non-state actors don't get a hold of these AI models

nl
3pts1
sql-benchmark.nicklothian.com 3mo ago

Show HN: An Interactive Text to SQL Agent Benchmark

nl
1pts0
garberfiles.substack.com 4mo ago

The Pentagon's UFO Psyop

nl
3pts1
vivsha.ws 4mo ago

Stress testing Claude's language skills

nl
2pts0
chatjimmy.ai 5mo ago

Chat with Llamma 8B at 16,000 TPS

nl
3pts0
taalas.com 5mo ago

Llamma 3.1 8B in hardware, 16,000 TPS

nl
4pts0
github.com 5mo ago

DeepMind Aletheia [pdf]

nl
5pts0
arxiv.org 5mo ago

Accelerating Scientific Research with Gemini: Case Studies and Common Techniques

nl
4pts0
www.coatue.com 5mo ago

Agentic coding is accelerating app releases

nl
1pts0
bmin.ai 6mo ago

Four Ingredients for Successful Retrofitting

nl
2pts0
www.deeplearning.ai 6mo ago

In Defense of Data Centers

nl
1pts1
twitter.com 6mo ago

Erdos 281 solved with ChatGPT 5.2 Pro

nl
308pts294
www.msn.com 6mo ago

Microsoft's spending on Anthropic AI on track to reach $500M

nl
4pts1
timdettmers.com 6mo ago

Tim Dettmers: A Personal Guide to Automating Your Own Work

nl
3pts0
twitter.com 6mo ago

DHH: AI models are now good enough

nl
12pts16
newsletter.semianalysis.com 7mo ago

MI300X vs. H100 vs. H200 Benchmark Part 1: Training

nl
5pts0
www.erdosproblems.com 7mo ago

AI just proved Erdos Problem #124

nl
242pts98
github.com 8mo ago

Show HN: Vibe Prolog

nl
44pts15
lowninstitute.org 8mo ago

Healthcare corporations spent trillions of dollars on payouts to shareholders

nl
5pts1
finance.yahoo.com 8mo ago

Coinbase CEO Exposes Prediction Market Vulnerability

nl
3pts0
www.citizen.org 11mo ago

Deleting Tech Enforcement: Dropping Lawsuits Against Tech

nl
4pts0
finance.yahoo.com 1y ago

Mark Zuckerberg has rebuilt Meta around Llama

nl
1pts1
www.crowdsupply.com 2y ago

AI in a Box

nl
1pts0
www.forbes.com 3y ago

New Wildcatters Are Drilling for Limitless ‘Geologic’ Hydrogen

nl
6pts2

theoretically possible

I mean I guess, but not in a performant way if there are ever any hardware failures. And with 100K GPUs there are multiple hardware failures per day.

Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.

This seems both arrogantly dismissive ("you are holding it wrong") and incorrect.

Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.

I don't understand this sentiment at all.

Is it a claim that "breaking into Hugging Face's production infrastructure" didn't happen? That it's not actually all that severe? That it was done by hand by OpenAI employees and they fooled Hugging Face?

That the blog post exaggerates something, somehow?

What exactly do you mean?

At the moment it just reads like a thoughtless dismissal.

Deterministic seeds barely work on a single machine, small scale training run.

They just don't work at all on a many month long, 100K+ GPU cluster training run.

Great, but that seems a different concern to the auditability of a model.

You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.

But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.

I know someone who runs AI training.

They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.

They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.

There's a pretty big gap between the things we talk about here, and where the world is at.

because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?

Is this an assertion that is backed by evidence?

From the Elon/OpenAI trial:

On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”

https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...

Have you ever worked with a non-programmer and helped them setup their AI workflows?

You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.

It's a lot of work, and most people will never want to change it once they have it working.

I'm all for open models, but people seem to misunderstand what they are. They aren't the same thing as open source code!

open weights, open code and open data

Even if you have all these things you still can't replicate a model because of randomness.

You can backdoor a model with less than 1000 examples and it is impossible to detect.

China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.

Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.

The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.

But the implication is that somehow Chinese competition is illegitimate because "strategic".

having PhDs doing night shift lab tech work for pennies

I don't know why people keep bringing this up as though it is surprising.

In almost any field other than AI PhDs are underpaid on average.

There are many, many bio PhDs working as lab technicians.

No reason to think it will. Paul Allen died after donating the money to create AllenAI and I don't think there are any links.

Hopefully it somehow works out though!

Meta isn't investing in frontier big models anymore

Yes they are. Meta Muse is their attempt.

It's below frontier performance at the moment but they are spending on getting there.

If the path that was taken to commit this is full of "oops" and "fix" messages

great way to encourage people to rebase then!

I wouldn't be too fixated on the specific numbers in that post.

Anthropic was extremely capacity constrained at that point. They still are but not to that extent.

I'd note that OpenAI offers 24 hour caching. I'd be surprised if Anthropic hasn't optimised their caching for Claude code too.

SemiAnalysis recently posted that their actual Opus usage works out at $0.99 because of caching.

The principles remain though.

It's not one prompt, but here is a parametric rod connector:

  Use SCAD and design a connector for square rods.
 
  The rods are 18.2 mm square. I want to connect two end-to-end.
..
  make if the bolt holes are created optional for each side - I want to set them separately. Make them M3.5 countersunk
..
  it's the +X or -X sides I want to turn the screw holes on or off.
..
  on each of the 4 sides of the connector add additional connectors at 90 degrees. Make each optional

etc etc

I think that applies to military involvement abroad generally.

If you are dropping bombs on someone I'm unconvinced the use of AI will make them like you more or less.

I've been doing a lot of 3D design in Codex GPT 5.5 (I found Opus 4.7 wasn't as good - haven't experimented much with 4.8 or Fable).

OpenSCAD is a parametric CAD programming language, and the models know it well.

The biggest challenge is communicating words like "inside" and "above" to the model - inevitably it's idea of which direction is which is often different.

I can't say I've done anything very hard, but for things like ESP32 cases, or parametric rod connectors it is great.

You can do things like "add snap connectors" and it'll do a great job.

Their methodology isn't published.

Its widely accepted[1] that it runs the same query through the model in parallel and then has a model that either selects the best answer or synthesizes an answer from the multiple ones generated.

I believe most people think it runs 6 sub-models, but I think that is based on the pricing.

It's a pity that OpenAI doesn't publish details like this.

[1]eg https://news.ycombinator.com/item?id=48799977

The source is the GPT 5.5 System Card:

We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. As noted below, we separately evaluate GPT-5.5 Pro in certain cases because we judge that the setting could materially impact the relevant risks or appropriate safeguards posture.

https://deploymentsafety.openai.com/gpt-5-5/model-data-and-t...

There have been multiple podcasts with people from OpenAI which have confirmed this.

I like this format:

"I love Lean because <abc>. I found it failed in <xyz> case because <123>. I created a thing <blah> which handles that like this: <ahhh>.

I'd love feedback! It's open source here: "