How big is this market, self-hosting a model that requires 64 GPUs, H100 or better, with good interconnects between nodes?
I suspect the overlap of those that can afford it, and those that have the talent to manage it, is a fairly thin slice of the Venn diagram. Even the large corps are gonna be getting it from the inference vendors, or more likely Bedrock and friends.