HN user

fnbr

2,044 karma

Finbarr.ca | @finbarrtimbers on Twitter | https://www.artfintel.com/

I love receiving emails. If you’re on HN, I would enjoy talking to you. Yes, you.

I'm an AI researcher and (small-time) angel investor.

I’ll happily talk to you about your startup, strategize about how to pitch VCs, or give advice on how to get a job in tech. You are not wasting my time.

Posts23
Comments493
View on HN
finbarr.ca 1y ago

The Bitter Lesson

fnbr
2pts0
www.artfintel.com 1y ago

How to hire ML engineers/researchers

fnbr
1pts0
www.vox.com 2y ago

OpenAI departures: Why can’t former employees talk?

fnbr
1254pts961
sander.ai 3y ago

Musings on Typicality

fnbr
2pts0
finbarr.ca 3y ago

Large language models aren't trained enough

fnbr
3pts0
karpathy.medium.com 3y ago

Yes, you should understand backprop

fnbr
3pts0
www.lesswrong.com 3y ago

Are We in an AI Overhang?

fnbr
1pts3
www.benkuhn.net 4y ago

Searching for Outliers

fnbr
23pts3
arxiv.org 9y ago

Submanifold Spare Convolutional Networks

fnbr
2pts1
finbarr.ca 9y ago

Notes on “Random Search for Hyper-Parameter Optimisation”

fnbr
2pts0
finbarr.ca 9y ago

Notes on “Training ImageNet in 1 Hour”

fnbr
3pts0
finbarr.ca 9y ago

Useful bash oneliners

fnbr
3pts2
news.ycombinator.com 9y ago

Ask HN: How do you set prices?

fnbr
584pts175
finbarr.ca 9y ago

Unit tests make you write down your assumptions

fnbr
2pts0
finbarr.ca 9y ago

How we designed our machine learning application

fnbr
2pts1
backchannel.com 9y ago

Google: Our assistant will trigger the next era of AI

fnbr
3pts3
www.bloomberg.com 9y ago

Inside Palantir’s War with the U.S. Army

fnbr
6pts1
boss.blogs.nytimes.com 11y ago

Why I do all of my recruiting through LinkedIn

fnbr
2pts0
www.kimonolabs.com 11y ago

Kimono – Turn websites into structured APIs

fnbr
2pts0
www.businessweek.com 12y ago

The Rise and Fall of BlackBerry

fnbr
2pts0
matt.might.net 12y ago

"My Ph.D. advisor rewrote himself in bash."

fnbr
7pts0
mobile.nytimes.com 13y ago

King of my castle? Yeah right.

fnbr
3pts0
cds.nyu.edu 13y ago

NYU MS in Data Science

fnbr
1pts0

MoEs have a lot of technical complexity and aren't well supported in the open source world. We plan to release a MoE soon(ish).

I do think that MoEs are clearly the future. I think we will release more MoEs moving forward once we have the tech in place to do so efficiently. For all use cases except local usage, I think that MoEs are clearly superior to dense models.

(I’m a researcher on the post-training team at Ai2.)

7B models are mostly useful for local use on consumer GPUs. 32B could be used for a lot of applications. There’s a lot of companies using fine tuned Qwen 3 models that might want to switch to Olmo now that we have released a 32B base model.

Like what? All the examples people have said are where either

1) the company has Nx preferences, for N >1, in which case the company has essentially failed to fundraise or

2) the company sells for less than they raised, which again, is a polite form of failure.

This is why I will never work somewhere with a short post termination exercise period (PTEP). If it’s not at least 5 years, ideally 10, they don’t seriously consider equity something that employees are owed.

Can you explain? In most cases, preferences won’t come into play, assuming you raise at a standard 1x preference and sell for more than you have raised. In that case, owning 0.5% should roughly translate into $5M (modulo dilution).

The rule of thumb is roughly 44gb, as most models are trained in bf16, and require 16 bits per parameter, so 2 bytes. You need a bit more for activations, so maybe 50GB?

you need enough RAM and HBM (GPU RAM) so it’s a constraint on both.

The New Inflection 2 years ago

Wow. This is surprising. Not even an acquihire. I’m very curious about what happened internally to lead to this.

1) i actually think that’s too high, i bet it’s more like 30%. My logic is that they have to have _some_ margin, but LLMs are too expensive to have typical software margins. Total speculation though.

2) It generally tracks pretty well unless the model is gaming the metric (training on the test set, overfit to the specific source of data, etc). The relative rankings will typically match in both.

3) alas, not with the mild winter North America’s having. They only stop below -5C or so. I am lucky though. The woodpecker stopped attacking my house and started attacking my neighbor’s. Even worse, it used to be a downy woodpecker,and it’s now been replaced by a pileated one (think: Woody).

It’s not clear to me what’s happening on the distillation front. I agree no one is doing it externally, but I suspect that the foundation model companies are doing it internally, performance is just too good.

There’s a bunch of recent work that quantizes the activations as well, like fp8-LM. I think that this will come. Quantization support in PyTorch is pretty experimental right now, so I think we’ll see a lot of improvements as it gets better support.

The KV cache piece is tied to the activations imo- once those start getting quantized effectively, the KV cache will follow.

Ah, well you could use a standard value network, but it’d end really slow, so you probably want to train a smaller one and rely on the implicit ensembling that MCTS does to make it better.

In my experience, PUCT does a lot better than UCT, so you want to also have a prior network.

You don’t have to train a new network, but in my experience, it works much better. I haven’t spent a ton of time using off the shelf networks with MCTS though. Maybe it works great.

very subtle bugs is the MCTS experience. Particularly once parallelism is involved.

It’s very difficult to implement, and requires training the network to use it.

I worked at DeepMind on projects that used MCTS. Even with access to the AlphaZero source code, it was very difficult to write an other implementation that got the same results as the original.