HN user

typon

3,863 karma

[ my public key: https://keybase.io/typon; my proof: https://keybase.io/typon/sigs/n3XfdbJqBTUy-IXBohwh-4TbEydnumleXWt7AsmXsIE ]

Posts10
Comments1,223
View on HN

I am using a model that runs on a 100 GB300s, near AGI, all bets are off, $10B training run on a 1GW cluster, and it can't realize that when I told it to "please implement v2 of feature X" that I mean delete v1, not support v1 and v2 together, in a weird Frankenstein's monster of the two. Sorry, but I think my job is quite secure for the forseeable future.

If there really was a "simple" solution to Fermat's Last Theorem, Andrew Wiles wouldn't have achieved the important result he did, ending up making connections across disparate fields of math.

The LLMs "sweeping up" easy, or previously missed, results seems like a net negative. It's probably better for humans to struggle and come up with new tools than to just "clean up" low hanging fruit that doesn't add much value to the field.

There is ton of room for improvement "down there".

* Software inference optimizations

* Heavy quantization

* Chips with hardcoded transformer architecture

* Much cheaper HBM

* Much sparser models - 1T total with ~1-10B active params e.g.

* Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.

GPT-5.6 13 days ago

DinoV3 paper: https://arxiv.org/pdf/2508.10104#page=36

"we use a rough estimate of a total 9M GPU hours"

From CoreWeave, at current prices (~$2.46/hr spot to ~$6.16/hr on demand) would correspond to $22M–$55M.

The dataset is really where the cost is though - they used LVD-1689M - 1.6B images of curated web data from roughly 17B instagram images. This probably cost a huge amount of hours in human annotation, compute for algorithmic filtering, etc and not to mention probably a 20-50 person team working on this model.

You might want to change assumptions about how expensive these models are.

I remember saving up for a year to buy the ATI Radeon 9600 XT (I think it was $200 MSRP) so I could play the game on high settings. Now we can play it inside a virtual machine on a crappy laptop. What a journey

Always pains me to see all this innovation and cool software work being applies towards making a machine that extracts money from retail investors and makes a few people extremely wealthy. What a depressing and colossal waste of time.

And the reason for _that_ is because of the callous way American society accepts the deaths of thousands of people who die due to the Healthcare Industrial complex (of which Brian Thompson was a key member of). Just because those deaths don't happen with guns doesn't make them any less important.

You will be forced to watch Firefly for eternity. Millenials will rule the internet for a 1000 years (a millenia).

CasNum 5 months ago

Like an oasis in a desert of LLM slop. Thanks I enjoyed this README

We will be old, not boomers. Boomers are a special generation at a special moment in world history - they made decisions based on the limited amount of knowledge they had about how the world works and while some think those decisions have doomed us forever, I remain optimistic.

ChatGPT Atlas 9 months ago

This xkcd comic doesn't apply anymore due to AI making generating automation code trivial.

OpenAI Grove 10 months ago

When you call yourself "Open"AI and then turn around and backstab the entire open community, its pretty hard to recover from that.