HN user

gpjt

1,788 karma

https://www.gilesthomas.com/

Posts83
Comments256
View on HN
www.gilesthomas.com 12d ago

Building intuition about LLM parameter counts

gpjt
2pts0
www.gilesthomas.com 14d ago

Poppy the training box, part 1: the beginnings

gpjt
3pts0
www.gilesthomas.com 14d ago

From bigrams to GPT-2, one component at a time (in Jax)

gpjt
1pts0
www.gilesthomas.com 22d ago

Building a Jax training loop for an LLM training run

gpjt
2pts0
www.gilesthomas.com 28d ago

Thoughts on Role Confusion

gpjt
3pts0
www.gilesthomas.com 1mo ago

Flax debugging: making a hash of things

gpjt
2pts0
www.gilesthomas.com 1mo ago

10Gb/s Ethernet: switching to a Broadcom SFP+ module

gpjt
195pts170
www.gilesthomas.com 1mo ago

Jax: Commitment Issues

gpjt
4pts0
www.gilesthomas.com 1mo ago

Jax Back Ends and Devices

gpjt
2pts0
www.gilesthomas.com 1mo ago

Using Safetensors with Flax

gpjt
2pts0
www.gilesthomas.com 1mo ago

First Looking into Jax

gpjt
3pts0
www.gilesthomas.com 2mo ago

10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module

gpjt
3pts0
www.gilesthomas.com 2mo ago

10Gb/s Ethernet: what I did to get it working in my home

gpjt
232pts177
www.gilesthomas.com 2mo ago

10Gb Ethernet: what I had to (re)learn

gpjt
1pts1
www.gilesthomas.com 3mo ago

LLM from scratch, part 33 – what I learned from the appendices

gpjt
5pts0
www.gilesthomas.com 3mo ago

LLM from scratch (32l) – Interventions: updated instruction fine-tuning results

gpjt
1pts0
www.gilesthomas.com 3mo ago

How an LLM becomes more coherent as we train it

gpjt
3pts0
www.gilesthomas.com 3mo ago

LLM from scratch, part 32k – Interventions: gradient accumulation

gpjt
2pts0
provision.sh 3mo ago

Provision: LLM-powered server setup from Markdown

gpjt
2pts0
www.gilesthomas.com 3mo ago

LLM from scratch, part 32j – trying to train a better model in the cloud

gpjt
2pts0
www.gilesthomas.com 3mo ago

Writing an LLM from scratch, part 32i – Interventions: what is in the noise?

gpjt
1pts0
www.gilesthomas.com 3mo ago

Writing an LLM from scratch, part 32h – Interventions: full fat float32

gpjt
7pts0
www.gilesthomas.com 4mo ago

Writing an LLM from scratch, part 32g – Interventions: weight tying

gpjt
2pts0
www.gilesthomas.com 4mo ago

Writing an LLM from scratch, part 32f – Interventions: weight decay

gpjt
6pts0
www.gilesthomas.com 4mo ago

Writing an LLM from scratch, part 32e – Interventions: the learning rate

gpjt
3pts0
www.gilesthomas.com 5mo ago

Writing an LLM from scratch, part 32d – Interventions: adding attention bias

gpjt
6pts0
www.gilesthomas.com 5mo ago

Writing an LLM from scratch, part 32c – Interventions: removing dropout

gpjt
1pts0
www.gilesthomas.com 5mo ago

Writing an LLM from scratch, part 32B – Interventions: gradient clipping

gpjt
2pts0
www.gilesthomas.com 5mo ago

Writing an LLM from scratch, part 32a – Interventions: training a baseline model

gpjt
1pts0
www.gilesthomas.com 5mo ago

Getting a Custom PyTorch LLM onto the Hugging Face Hub

gpjt
1pts0

In the EU, at least one of the problems is regulatory/tax for digital services. They need to charge you the VAT rate for the country where you are based. For that, they need two pieces of evidence about your location. There are various things they can use for that -- telephone number, IP address, card billing address, and so on. If they can collect two that indicate the same country, they're safe -- but if all of them point in different directions then they could get in trouble during a tax audit.

Of course, for larger transactions you'd expect that a human in the loop could work with you to get the right info so that they would be covered. But I guess for Microsoft, their definition of "larger" might be more than a few grand...

OP here -- yes indeed. I ever do a new series of posts on upgrading my network to anything faster, the first one will probably be titled something like "25Gb/s Ethernet: how much it costs to rip out your CAT-6A and replace it with fibre".

OP here -- yes, 100%! Within my study it's DACs all the way. I only use 10GBASE-T where it's the only option: through the wall cabling, and from the connector on my ISP's crappy router to the mini-PC that acts as the real router.

Doesn't say anything about citizenship though. There are plenty of US residents who are not citizens. And a lot of people abroad appear to use US billing address credit cards -- in my last company we had hundreds of people with the same US billing address who appeared to be managing Africa-focused businesses and used IPs that matched that.

Excellent points. It's also worth noting that many people don't wind up working in the field they studied at university. CS grads have probably been the exception in recent years, because the industry has been booming, but the two most successful entrepreneurs I know studied philosophy and art history. A friend who is very senior in recruitment studied economics.

I was thinking the same. It's a simple idea, heavily over-explained. The code is similar, massively overengineered for such a simple test.

I have it running on a Proxmox VM. It basically just sends me summaries -- it reads a bunch of RSS feeds I pointed it to (news, tech, etc) and gives me a daily summary, along with an image of the day based on that. It also sends me recommendations for times to go for a run based on the weather and my Strava activity, daily recommendations for stargazing (what's visible and when, weather, etc), and a couple of daily reminders for things I tend to forget.

I'm using Claude as the model, though, so it's smart but pricey. Should configure it to use different models for different things, but it's trickier than I would have expected to do that.

It's a bit of a double edged sword. As someone who smoked and found it impossible to quit for decades, I'm very happy to have been able to switch to (reuasable) vaping. It's probably added years to my life expectancy.

OTOH the upsurge in nicotine use amongst young people feels suboptimal, and disposable vapes are a scourge.

Looked to me like it was trying to work out whether it was edible. Sensible behaviour for an animal in a world where unfamiliar edible things appear from time to time.

Not sure that every browser advertises English, but mine certainly does. However, as I'm in Portugal, many websites ignore what my browser says and send me to translated versions, I assume based on my IP. That causes problems because the translations are often quite bad, and they do it with redirects to PT URLs so I can't share links with people who don't speak the language.

Awesome, thanks! I'm still doing trains on the big machines right now (hopefully will write up over xmas) but I think once I've worked out the sweet spot for memgatokens per dollar for this model, it's time to start tweaking the other controls -- LR and cosine variation of it, as you said, and also dropout, bias, weight tying, and definitely gradient clipping (which should at least get better bang for the buck from time/$ spent). I'll leave it to Google to follow up Chinchilla with a "best batch size across a thousand trained models" paper ;-)

I think the punctuation makes it clear -- imagine "How I invented Facebook. In 2001." The full stop in the middle of the sentence breaks it and makes you realise he's speaking figuratively.

Thanks re: gradient accumulation, I'm glad to hear my intuition was right!

As part of the upcoming post I'm running the DDP train on A100s with 40 GiB and 80 GiB, H100s with 80 GiB, and B200s with 160 GiB, so I'll have at least three loss vs. batch size points to plot. So that might be interesting.

I guess a full test would be to train at various batch sizes on the 160 GiB machine and plot the resulting loss. That would be very expensive as a hobby project (the bs=64 train cost a bit more than $40 excluding overhead) so I won't do it.

But perhaps a shorter train would still be of value? That is, train for 300M tokens for a tenth of the cost and see where the loss landed? The problem with that would be if the impact of batch sizes varied with the length of the train, eg. if batch size 64 was better than 512 for short trains but weaker at longer ones.

Exactly! If I can get it down to an hour or two (seems very plausible on an 8x H200 with 160 GiB VRAM per GPU, though those are almost never available on Lambda Labs), I'll do the experiments with dropout and the other possible causes of issues, then see if I can bake that all into a new train on the RTX 3090 and confirm it repros there. Looks like I'll definitely need gradient accumulation there.

I assume the zero_grad would need to go in the same if block?

OP here -- with a 112M model you should be able to get something worth playing with using 2.24B tokens. The Chinchilla heuristic is tokens = 20 x parameters. Obviously you cam get a better result by grinding through more tokens, but it will be very slow progress. It's worth noting that Andrej Karpathy is using the 20x thing for his nanochat project.

I try to explain the Chinchilla paper in the post, but your favourite AI should be able to explain it well, and has the benefit that you can ask follow-up questions.

OK, early indicators support both you and Gemini quite strongly re: batch size. On my (somewhat ad-hoc) test dataset, I get losses like this:

  * OpenAI medium weights: 3.231
  * OpenAI small weights: 3.500
  * My locally trained model, FineWeb Chinchilla, batch size 6: 3.944
  * My locally trained model, FineWeb-Edu Chinchilla, batch size 6: 4.167
  * My locally trained model, FineWeb-Edu double Chinchilla, batch size 6: 4.135
  * My cloud trained model, FineWeb Chinchilla, batch size 13 \* 8 = 104: 3.674
That last one was trained on an 8x A100 machine with 40 GiB per GPU, with the same code as before, just converted to DDP. It certainly looks like the much larger batch size has improved the model significantly.

I'll be trying on larger machines. No gradient accumulation yet, but it's certainly looking like a valuable lever to pull for local training runs (and, I suspect, might also be useful on "small" cloud machines like the one I used -- will have to see what things look like with the bigger mini-batches I can squeeze onto 80 GiB and 160 GiB GPUs).