HN user

foobar10000

68 karma
Posts0
Comments67
View on HN
No posts found.

3T at nxfp4 (which is most of it) is only 1.5TB of vram - so 8x288GB B300 or MI355 will do it if you are careful with context - maybe dp-attn? Certainly not TP. 2 of those together can easily serve it. The new AMD MI400 are at 400GB+ each, so 8x of them will nicely fit with KV to spare.

Don't forget that you are not really seeing the thinking tokens used - so non-trivial to count them.

Yeah, if you have a fixed llm topology, you can just effectively burns 2 top layers of the chip as Rom (model weights) - which has a per area density even better than dram - so it’s just attention and kv streaming that is hbm to sram transfer.

Most big model weights will not fit a single reticle sized chip - so you’d have prob 30 different chips to split the model .

And you’d need super fast chip to chip comms for the all-reduce and similar.

So scaling to 1T models is hard - and a long lead time - but can be very power efficient.

The Coming Loop 30 days ago

Agreed - there was always a set of things I wanted to do that I knew the magic core for, but wanted a team of implementers for the curft, the 100k of actual testing harnesses, hyperparameter exploration, etc.. . I now have that team of implementers. All the problems seem research-y though - optimal binary transport systems that are zero-copy and compatible with languages, fast physical simulation optimizers, etc etc... So, things that all had a _LOT_ of busywork around the magic core.

We are there already pretty much - if I understand your point (“How the models are wielded”) refers to the harness - which is part of model training already. Fable was trained to use Claude code harnesses effectively to keep plugging at a problem with a lot of working memory and world knowledge and reasoning capability - and to keep hitting that search space intelligently. And it is not just cybersecurity. I’ll give several examples from my own recent experience :

1. Cyber - already discussed - has an issue that the bigger models can actually do a full end-to-end exploration of an exploit - go from theory to an actual deployable payload.

2. CAD/CAM/Mechanics - CalculiX (ccx) - an open-source FEA and similar mechanical solver - think Siemens or ANSYS, but open-source. A team I was helping was trying to do a design mount of a physical object that would need to reduce vibrations in a frequency band - think microphone mount basically. Usual loop would be design, analyze with Siemens, go to beginning. AI loop is have AI design, then analyze using ANSYS, then analyze result, change design, iterate. That loop did not produce anything useful for elastic materials because ANSYS would take 12 hrs to do acoustic analysis using a GPU. 1 week of autonomous work by a frontier model resulted in a modification and custom solver added to ccx that could simulate the acoustics (vibrations) _in that particular problem_ in about 20 seconds - mainly because it could try new mathematical ideas, then compare them against ANSYS reference for quality of solution, and iterate. And 1 week _after that_ the frontier model - iterating on one design per minute - came up with multiple 3-d printable ground-breaking mounts - including sending one off to xeometry for printing and getting it shipped back. Existing designs had a 20 db drop in the frequency ranges needed - this one had 60. For reference 40 DB is basically infinity :) While this was for microphones - you can imagine that vibration reduction is a big thing in engine, suspension, and weapon mounting, and well, in general things that move. 3 person team btw - unthinkable even 1 year ago.

3. Pharma. Different company - but given a known Density Function Theory or Kappa Cluster molecule simulator, one can run nice agentic loops over frontier models to do chemical or pharmaceutical research - there’s a reason Anthropic is launching Claude Chemistry. Note that then limiting factor is the multi-week runtime of Kappa Cluster and similarly molecular dynamics simulators. If one _could_ speed that up for a particular problem space or molecule type, one could very very quickly have a high-end reasoning model iterate to a good molecular design - and frontier models are getting very good at precisely automating the ML research needed to do that autonomously - after all, there’s a reference there. 5 ppl - 2 years ago would be a research institute.

4. Physical AI - robotics - same principle.

5. This is basically the bet Bezos is doing with his new company.

Please do not underestimate the effect these models will have on our ability to improve our ability to effect the world - this is just starting to hit now. I think we can all extrapolate the GDP and defense impact of this - or at least that there will be a very significant one.

Well - there is a giant push to allow non-qualified investors to invest their 401k (and roth and whatever) into the private equities market - pre-IPO companies and such.

I can't shake the feeling of a grand fleecing incoming - and honestly, most big financial companies I know are against it because the blowback of inevitably bankrupting the firefighters&nurses pension fund will be congressional hearings and piercings of corporate veils.

Seems like the feds are pushing for it - for reasons I cannot fathom.

Imagine an agent shadowing all your terminals, providing ideas and asking to run commands that will let it verify the hypotheses it comes up with, while at the same time doing research on vendor docs, etc...

Quite safe, and already a force multiplier - this would be a harness. Maybe have it be able to write to a shadow system with similar (ideally same) hardware to verify it's hypothesis on how the system works, etc...

Minor nit re[2]: for agentic workloads that are actually worth money - i.e., claude code and similar, things are either prefill-bound - which this does not help - or more importantly tps/user bound (at 150k+ context windows) - you want your big magic model to emit 200 tps/user. This is why Nvidia bought Groq (now LPU) and what Cerebras is trying to do, etc, etc. So for the stuff that makes money in the field - GPUs are not really compute bound once context lengths are large - but still memory transfer bound (may be KV-cache transfer, may be HBM->SRAM-on-chip, etc..)

I kindof agree that it is unattractive - but the regulators are perfectly happy with "EOD also introduces credit risk on the clearing house/bilateral." if it allows them to protect retail and institutional investors. They'd rather bail out the clearing houses - which they can do - then have the retail investor lose faith in the markets for a generation or 2 - ala the great crash.

So, to your point - yep, it is not final. But it is unattractive _to the market making and prime brokers and similar players_ which the regulators do not care much about.

Note that _passenger aviation_ is commercially non-competitive. The big 4 US airlines make money on credit cards, not airfare : they lose money on airfare. So, most people who are trying to make money will not use them as a model.

In general, safe businesses can only exist with government support or government prohibition of all other businesses globally - and that is a very hard bar to clear.

"who do carry liability when things go wrong" -> unless one pierces the corporate veil, it's just money. Not even their money. HIPAA - unless basically stealing data - will not generate personal liability. And even for SOX will only generate liability in limited amounts for limited people - and executives will go a long way towards avoiding the entire thing.

From what I have seen - most executives would rather shut down the business and quit than accept the possibility of personal liability - and just avoid the regions of the world in which they do have it.

We are multiple orders of magnitude away from Landauer limits - so next big thing in matmul could be photonic multipliers - there’s a bunch of them coming up in the next 3? years. So that’s a 2-4 order of magnitude improvement. Sigmoid?

I think the one thing you are not taking into account is that the investors on average fundamentally don’t care. Scale arbitrage means that small companies are fundamentally about velocity - and if they get sued due to regulations that do not pierce the corporate veil, they just fold. And the ones that did not get sued make money for the vc. And figure out later how to be hipaa etc compliant. Basically, I’ve been seeing over the last 10 years VCs are not caring about insurance or corporate liability - sink rate is so high it is irrelevant.

For big corps - this is different. But modulo hipaa - this is why they are gung ho hi about binding arbitration - they are trying to match velocity to some degree - and mostly failing…

I mean - I'd say electricity, agriculture, steam power, metallurgy, silicon computing (cmos), atomic power, the scientific method - these are _all_ very impressive - all lead to drastic changes for humanity. Not sure how I'd rank them.

I personally think AI will end up sitting in the top 3 of these - but that is an opinion. I do think it is obvious it is at least _somewhere_ in that list.

The EOD reconciliation (and corresponding inability to settle a position in milliseconds) is a feature - it allows "obvious erroneous trade" roll-back mechanisms, etc.

Very few people want the financial system to be a contractual suicide pact - they want it to be predictable, but when the unpredictable happens - they want the retail and institutional investor to be protected (the HFT players can go beat each other up - no one will really cry about them). And unpredictable can be anything from a power event taking out multiple exchanges in the NJ triangle (Sandy hurricane) to a cyber-attack (never happened yet) to a flash-crash driven by algorithms from multiple HFT driving each other nuts (happened at least once).

So, it is not EOD processes as such, but the ability to pause, assess the entire system holistically, and then correct it before it blows up the portfolios of everyone holding a 401k. So even though the exchanges _could_ got to 24/7 trading, I'd be surprised if we just went away from cyclical 24-hr based windows of settlement.

Rumor says Hudson River Trading just ordered a bunch. So, the finance AI guys definitely. And they (AI finance - DeShaw, HRT, Citadel, G Research, XTX) have deployed about 15% of total GPU capacity, so not small fries.

Note that google cloud has an itar-compatible gemini pro and google drive / docs - so, people do talk to it - and google is of course contractually obligated to not export it, nor to learn from it.

This is very different that AWS fed-gov bedrock thingie - where AWS promises that the models are running on hardware dedicated to you, with no external logging, etc.

The issue is not better - it’s better _AND_ fast enough. An agentic loop is essentially [think,verify] in a loop - i.e. [t1,v1,t2,v2,t3,v3,…] A model that does [t1,t2,t3,t4] in 40 minutes, if verify takes 10 min, will most likely do MUCH worse that a model that does t1 (decently worse) in 10 mins, v1 in 10 mins, t2 now based on t1 and v1 in 10 mins, v2 in 10 mins, etc..

So, for agentic workflows - ones where the model gets feedback from tools, etc…, fast enough is important.

1 token ahead or 2?

It's interesting - imo we'll soon have draft models specifically post-trained for denser, more complicated models. Wouldn't be surprised if diffusion models made a comeback for this - they can draft many tokens at once, and learning curves seem to top out at 90+% match for auto-regressive ones so quite interesting..

So, this especially bites if your validation step (let’s say integration tests) take 1hr plus. The harness is just waiting, prefix caching should happily resume things with just a minor new prefill chunk of output from the harness, and bam - completely new prefill.

Nice!

Should be able to push it more if

* we limit data shared to an atomic-writable size and have a sentinel - less mucking around with cached indexes - just spinning on (buffer_[rpos_]!=sentinel) (atomic style with proper sematics, etc..).

* buffer size is compile-time - then mod becomes compile-time (and if a power of 2 - just a bitmask) - and so we can just use a 64-bit uint to just count increments, not position. No branch to wrap the index to 0.

Also, I think there's a chunk of false sharing if the reader is 2 or 3 ahead of the writer - so performance will be best if reader and writer are cachline apart - but will slow down if they are sharing the same cacheline (and buffer_[12] and buffer_[13] very well may if the payload is small). Several solutions to this - disruptor patter or use a cycle from group theory - i.e. buffer[_wpos%9] for example (9 needs to be computed based on cache line size and size of payload).

I've seen these be able to pushed to about clockspeed/3 for uint64 payload writes on modern AMD chips on same CCD.

All are indeed plausible- translation is iffy due to diarization not being all there yet - but why the specific order of horribleness?

Live translation seems either better than autonude or worse, but not in the middle of the pack I’d assume? Am I missing something here?