The article does answer the question, in detail. What you're reading is just the first half is freely available. The 2nd half is on the same web page and for subscribers.
HN user
dylan522p
They use some of this too btw. Also wavelength level routing happens with breakout cables from ToR to Compute.
OPC? It'd be great to talk about your area, because I bet most don't know what you do, and maybe I can make it an entertaining read.
No TPUs are getting better architecturally
Your math is completely wrong dude.
It's 2000ms per token not for the whole query.
Hardware utilization rates and MFU are not the same thing, you forgot the latter.
You are pretending its perfectly parelelized on 1 GPU too. I use 8x GPU box throughput.
This article is based on wild speculations of how much things cost
Huh it literally uses real throughput figures
doesn’t account for ways to make things cheaper over time
It does in the subscriber section and it says it does say that in the free section.
We already have papers that suggest most big models are undertrained and smaller models can get the same accuracy.
Why assume 2023 model is the same as 2020 GPT-3 175B parameter
- Google qps is closer to 100k then 320k [1]
That number is wrong. I have a number from googler, not livestats which cannot have google internal data
- Not every query has to run on LLM. Probably only 10% would benefit from it
Agree, i have something different coming up that looks into this more, 10% may be too low. I know i used 100% which isn't right, and explicitly say that
- This means 10,000 queries per second, each needing 5 A100s to run, so 50,000 A100s are sufficient. Cost for that is $500MM, quadruple that to $2B with CPU/RAM/storage/network. That is peanuts for Google.
50k A100s networking ramp cost way more than $2B HW utilization rate
- Latency, not cost, is a bigger issue. This should be addressed soon by H100 and newer chips.
Thats discussed in the subscriber section. It's both, but yes latency is bigger issue. H100 helps but doesn't solve.
cutlass is faster actually, in most cases where cutlass supports same stuff.
Thanks! Ya, the archive link is useless, it only captures the free part, and I think I am very generous with what I keep in the free section on the ad-free website.
Tsmc 7nm doesn't use EUV and is fantastic
Author here. What do you mean die shot bleeding?
M1 2020, M2 2022?
It has noticed. Look at Apple architects, validation, layout, etc engineers moving to Nuvia + Rivos + Google + Amazon + Microsoft + Meta + Intel + Nvidia + AMD + Apple + Qualcomm.
It's there.
It's at the end of every article lol
the unsubstantiated BS that Samsung's chip manufacturing is a disaster.
They lost Qualcomm, Nvidia, and Cisco for the next generation. Are you disputing this fact?
They did not fully ramp 1Z or 1 Alpha dram nodes. Are you disputing this fact?
I have been sent C&D letters in the past by Arm, and even sued by others. I have resources.
If you don't want to believe it. Go ahead. The major claims are true, Qualcomm moving away. Nvidia moving away. DRAM node ramps being pitiful. DRAM engineering efforts come directly from a source there.
The cultural issues being the cause of these issues is the substantial claim.
MediaTek is killing it. They hired a lot of people from TSMC's process development kit teams which vastly improved their capabilities in integrating IP. They were the first to develop AV1 decode, their modems are pretty decent. They even have leading teams in wifi.
Most folks blame the culture shift on Paul actually. He was the cause of it all. Brian being a horrible symptom of what he created, and Brian definitely propagated it far.
If anything I am making the case that TSMC is reigning supreme but okay, believe your conspiracy
Good thing the article even says Samsung is a well oiled machine in many areas outside of the ones with issues discussed in this article. Almost like large companies can have varrying success and culture.
It's at the bottom of every article. I have to disclose given my relationship with various funds.
I am under no impression that my articles can move fucking Samsung's stock. That's hilarious you think I could profit off the market by writing this.
There's plenty of evidence though. Be an industry insider and you'd recognize it all.
I do know that my reports have moved smaller companies stocks by 20% in a single day, and have been verified true in the past. - https://semianalysis.com/short-report-nvidia-supplier-cut-ou...
If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after, like I did with the article I just linked.
My original article you are talking about explicitly stated heterogenous compute is the future. It talks about transistor budgets going to LLC, GPU, media, and ISP. Much of the CPU efficiency gain comes from that doubled LLC size.
Sophie's analysis contemplates a set unit volume and fixed costs related to design and verification and tape out. Mine is on pure cost/transistor based on wafer costs as a player like Apple has the units to spread fixed costs over.
Look at the in depth analysis I did of the die shot and SOC floorplan. Or even read the article I originally wrote I was right on the CPU architectural gains. Apple pushed clockspeed up slightly, but IPC was less than 5%.
https://semianalysis.substack.com/p/apple-a15-die-shot-and-a...
Why no on the BOM increase? That is what the switch to LP5 and A16 would be.
The IPC gains were what I specifically wrote about and those turned out to be true with <5% IPC gains on the Avalanche core CPU
They are allowed to sell chips outside of china with Arm IP licensed from Arm China
The last part of the article was not here last year. The last part and images are from an event they held recently. Notice the
"Before we get to the event they held and the significance of it, let’s do a recap."
He refuses to leave and he has the stamp. The 7-1 vote was even mentioned.
I would love if you could find those images from the event last year. You would need a time machine for that.
Arm does not manufacture chips. They license IP for per chip or blanket fees. The model of many of these licenses is irrevocable. I've seen a couple different Arm licensing contracts and they're all very different so hard to make blanket statements. The Chinese entity has the right to license to all Chinese semiconductor firms who have the right to sell their chips.