HN user

Patrick-STH

233 karma
Posts0
Comments47
View on HN
No posts found.

There is not an EPYC 7xx4. EPYC 8004 is Siena. We reviewed an ASRock Rack Siena platform this week actually.

We had this in the Genoa launch piece but AMD largely kept $/core constant between Milan and Genoa. Milan is still being sold since it uses cheaper PCIe Gen4 motherboards and DDR4.

On the server side it takes a few quarters for new products to start making up the majority of shipments. These days it is better to think of new servers as N, N-1, and still some N-2 generations being sold as new.

Easy one would be the P part is single socket only.

I was browsing one day looking for barebones to throw in one of our clusters and saw Newegg was selling these 1U servers cheap ($2200-2300)

Usually with AMD desktop and server CPUs buying high core counts at lower TDP yields a great perf/ W figure

Yes. We have done pieces on the BlueField-2 DPUs running Ubuntu and doing things like running ZFS and iSCSI off of the DPU's Arm cores as well. This is the BlueField-3 base, so a faster Arm core complex and more memory bandwidth.

Sometimes 2-3 tries but that is also from pacing. I usually have my laptop out to reference quickly but no prompter. People can tell when I read from one, so I have just had to get comfortable without reading.

Now that is a throwback! That old logo was created with GIMP at a table outside NetApp. STH started at a time when the SMB market (and VERY high-end homes) were using Dell, HP, and IBM rackmount gear off-lease. The original STH idea came from me learning Linux as an alternative to Windows for that segment. STH was what I used to learn the rest of the market outside of work. The first product we were sent for a review was a rackmount case and that started us on the path of reviewing new gear ~2010-2011 when it was a part time blog instead of having a team of folks working on the site. I still try to keep 15-20% of our content in the SMB/ high-end home arena.

There are a few big ones: - The CUDA license does not allow you to use GeForce in the data center. In the US it has become less popular, but if you look at our Inspur AIStation piece, that was a cluster located in China with GeForce cards. So it still happens, but less so. - The memory capacity is another big challenge. Newer models have 80GB which dwarfs the 24GB on a 4090. We just got the RTX 6000 Ada in, so that is an option for more memory. - For higher-end training, one of the big challenges is interconnect, so having NVLink and Infiniband or 100GbE+/ Infiniband NICs is important. The HGX A100 platform is designed for that with its NVSwitch and PCIe switch topology.

With all of that said, you are 100% right that many startups have used consumer cards for years. For example, Andrej Karpathy talked about how our DeepLearning11 build (8x 1080 Ti's) had a ~3 month payback period versus AWS https://twitter.com/karpathy/status/924340245478256640

This is a crazy oversimplification, but let me take a shot. The difference to others is easily a large "novel size" discussion.

Most training SoCs are focused on building something the size of an NVIDIA GPU, but designed for ML versus general purpose GPU HPC compute (FP64) plus ML. Often those accelerators today have a few types of models they are optimized for. NVIDIA is the baseline and so the competitors are looking for areas where they can get a large boost at a lower cost with something about the size of a A100/H100.

Cerebras is perhaps the biggest exception with its WSE-2, a wafer size chip. Having the wafer size chip means that Cerebras does not need to go into higher-latency and higher-power off-package interconnect as frequently because its chip is 50x larger. In turn, Cerebras drives performance and cost savings by not needing NVLink4 NVSwitches / InfiniBand.

Tesla's Dojo Tile is 25 chips roughly equivalent in size to a NVIDIA GPU in a single package with die-to-die communication facilitated by the base tile and then built for scale up units. Tesla also has focused on the interconnect and pipeline feeding the D1s and Tiles.

Ultimately, I think that it takes something beyond a "solution X saves 30% over NVIDIA in these workloads in performance/ $" to survive. NVIDIA has a massive software ecosystem and can handle more types of tasks versus some of the other AI accelerators. That goes beyond just the training and also to other parts of the data prep and movement pipeline. NVIDIA extracts high margins from this work so that is why some effectively are competing with "it costs less and on some problems can be faster" architectures but what Tesla, Cerebras, Google, and a few others have another level of differentiation.

Nothing is perfect, nor was that explanation, but just a high-level view of why the technology featured is impactful.

Totally correct. I mention that a bit in the video as well. Also - I do not view this as an AWS v. Colo. We use AWS for some services as well so it is a specific part of the workload we run (not just WP) that is in our hosting cluster.

And again, this is a fraction of what we have in data cetners due to the labs and such.

Very little, but that is probably because we cover the server industry in-depth with reviews and such. I mentioned a bit that that knowledge is key. This is our cost analysis and not going to be the same for everyone.

That may be true. We grow 20-30% Y/Y which is not as big as many other places (but is more than most others in our space.) 20-30% growth we handle with refresh cycles since we have extra capacity.

100%. We had 11 months where nobody touched the racks as an example and the visit 11 months later was the one to physically get a tool I had left there. I mentioned I even track drive time.

We include in our "Hardware Costs" having extra node and spares. For example, we keep a full spare node in the DC as well as spares.