HN user

trsohmers

1,536 karma

My website is trsohmers.com

2013 Thiel Fellow

2023-Present Positron AI (https://positron.ai)

2021-2023 Groq (https://groq.com)

2019-2020 Lambda Labs (http://lambdalabs.com); Managing hardware development and cloud/datacenter deployment

2013-2018 REX Computing(http://rexcomputing.com); Developed new processor architecture from scratch

@trsohmers

Follow me on twitter... @trsohmers

Posts22
Comments338
View on HN
www.eetimes.com 5mo ago

Positron's $230M Funding Led by Financial Trading Firms

trsohmers
1pts1
www.wsj.com 11mo ago

The New Chips Designed to Solve AI's Energy Problem

trsohmers
2pts0
substack.com 1y ago

An "Observatory" for a Shy Super AI?

trsohmers
3pts0
lambdalabs.com 2y ago

Lambda Raises $320M to Build a GPU Cloud for AI

trsohmers
3pts0
www.theregister.com 4y ago

AI hardware, PC gaming rig maker partner on powerful AI/ML laptop

trsohmers
2pts0
www.forbes.com 5y ago

AI Chip Startup Groq raises $300M

trsohmers
4pts0
hbr.org 5y ago

The Computerless Computer Company (1991)

trsohmers
3pts0
screenrant.com 5y ago

South Park Creators Launch New DeepFake News Show

trsohmers
1pts0
boingboing.net 5y ago

Sassy Justice with Fred Sassy: a terrific deepfake satire show

trsohmers
1pts1
lambdalabs.com 6y ago

Choosing the Best GPU for Deep Learning in 2020

trsohmers
5pts0
www.youtube.com 9y ago

The REX Neo Architecture: An energy efficient new processor architecture

trsohmers
3pts2
www.youtube.com 9y ago

Stanford Seminar: Beyond Floating Point: Next Generation Computer Arithmetic

trsohmers
8pts1
fortune.com 11y ago

Semiconductor startup with new processor scores Founders Fund investment

trsohmers
2pts0
www.technologyreview.com 11y ago

Startup Attempts to Reinvent the CPU to Make Computers Less Power-Hungry

trsohmers
20pts3
dreamscopeapp.com 11y ago

Dreamscope App: Fast deep learning dream image generator

trsohmers
4pts0
www.eetimes.com 11y ago

Startup to Open-Source Parallel CPU

trsohmers
76pts54
boingboing.net 12y ago

Disneyland's original prospectus (1953)

trsohmers
84pts16
www.eetimes.com 12y ago

Teenage CEO with HPC/Server startup to speak at EELive

trsohmers
2pts0
www.businessinsider.com 13y ago

Peter Thiel announces 2013 class of Fellows

trsohmers
2pts1
trsohmers.com 15y ago

IOS (almost) on the Motorola Xoom

trsohmers
2pts0
www.engadget.com 15y ago

Live time 3D Aerial Mapping using a quadcopter and Kinect

trsohmers
1pts0
www.pcworld.com 15y ago

Ubuntu on the Motorola Xoom

trsohmers
1pts0

I was put on it in 2015 after an acquaintance of mine that was previously on the list recommended me… I only heard from Forbes a few days before the list came out, they asked me for a photo and asked if I approved the 2 sentence blurb they prepared, and that was it. For years afterwards they would try to get me to come to their events, but I never had any interest, and I assume that was how they made money… but I never paid anything to be on the list or had real interest in being on it, and I don’t think it led to anything other than my technically illiterate parents thinking that it was impressive.

Based on their S1 filing and public statements, the average cost per WSE system for their (~90% of their total revenue) largest customer is ~$1.36M, and I’ve heard “retail” pricing of $2.5M per system. They are also 15U and due to power and additional support equipment take up an entire rack.

The other thing people don’t seem to be getting in this thread that just to hold the weights for 405B at FP16 requires 19 of their systems since it is SRAM only… rounding up to 20 to account for program code + KV cache for the user context would mean 20 systems/racks, so well over $20M. The full rack (including support equipment) also consumes 23kW, so we are talking nearly half a megawatt and ~$30M for them to be getting this performance on Llama 405B

Do you think that the 16k GPUs get used once and then are thrown away? Llama 405B was trained over 56 days on the 16k GPUs; if I round that up to 60 days and assume the current mainstream hourly rate of $2/H100/hour from the Neoclouds (which are obviously making margin), that comes out to a total cost of ~$47M. Obviously Meta is training a lot of models using their GPU equipment, and would expect it to be in service for at least 3 years, and their cost is obviously less than what the public pricing on clouds is.

+1 this commenter. I just visited the UK for the first time at the beginning of this month and had a fantastic ~3 hours at Bletchley Park, but felt I had to cram TNMOC and the amazing Colossus live demonstration (where I asked a million questions) and everything else in the museum in the 90 minutes I was there. If I assume other HN readers are like me, I would dedicate at least 2.5-3 hours for TNMOC to actually get a chance to actually see and play around with their extensive collection of vintage machines.

Rex Computing 2 years ago

We had a basic LLVM backend that supported a slightly modified clang frontend and a basic ABI. We tried to make it drastically easier for both the programmer and compiler to handle memory by having all memory (code+data) be part of a global flat address space across the chip, with guarantees being made to the compiler by the NoC on the latency of all memory accesses across one or multiple chips. We tested this with very small programs that could fit in the local memory of up to two chips (128KB of memory), but in theory it could have scaled up to the 64 bit address space limit. Compilation time for programs was long, but fully automated, specifically to improve upon problems faced by Cell and other scratchpad memory architectures… some of our original funding in 2015 from DARPA was actually for automated scratchpad memory management techniques on Texas Instruments DSPs and Cell (our paper: https://dl.acm.org/doi/pdf/10.1145/2818950.2818966)

This was all designed a decade ago, and REX has been in effectively hibernation since the end of 2017 after successfully taping out our 16 core test chip back in 2016, but being unable to raise additional funding to continue. I have continued to work on architectures that have leveraged scratchpad memories in different ways, including on cryptocurrency and machine learning ASICs, including at my current startup, Positron AI (https://positron.ai)

Rex Computing 2 years ago

Founder of REX Computing here; I highly recommend checking out my interview on the Microarch Club podcast linked elsewhere on the thread; will also answer questions on this thread if anyone has them.

Significantly more than that; MFN pricing for NVIDIA DGX H100 (which has been getting priority supply allocation, so many have been suckered into buying them in order to get fast delivery) is ~$309k, while a basically equivalent HGX H100 system is ~$250k, coming to a price per GPU at the full server level being ~$31.5k. With Meta’s custom OCP systems integrating the SXM baseboards from NVIDIA, my guess is that their cost per GPU would be in the ~$23-$25k range.

The quote from the linked press release is that they do training on TPUv4, while inference is running on GPUs. I have also heard this separately from people associated with Midjourney recently, and that they solely do training on TPUs.

Long story, but technically REX is still around but has not been able to continue to develop due to lack of funding and my cofounder and I needing to pay bills. We produced initial test silicon, but due to us having very little money after silicon bringup, most of our conversations turned to acquihire discussions.

There should be a podcast release (https://microarch.club/) in the near future that covers REX's history and a lot of lessons learned.

I thought that was clear through my profile, but yes, Positron AI is focused on providing the best performance per dollar while providing the best quality of service and capabilities rather than just focusing on a single metric of speed.

A guarantee to match the cheapest per token prices is sure a great way to lose a race to the bottom, but I do wish Groq (and everyone else trying to compete against NVIDIA) the greatest luck and success. I really do think that the great single batch/user performance by Groq is a great demo, but is not the best solution for a wide variety of applications, but I hope it can find its niche.

Groq states in this article [0] that they used 576 chips to achieve these results, and continuing with your analysis, you also need to factor in that for each additional user you want to have requires a separate KV cache, which can add multiple more gigabytes per user.

My professional independent observer opinion (not based on my 2 years of working at Groq) would have me assume that their COGS to achieve these performance numbers would exceed several million dollars, so depreciating that over expected usage at the theoretical prices they have posted seems impractical, so from an actual performance per dollar standpoint they don’t seem viable, but do have a very cool demo of an insane level of performance if you throw cost concerns out the window.

[0] https://www.nextplatform.com/2023/11/27/groq-says-it-can-dep...

"The current round" of AI accelerators you are referring to are things that were designed 2015-2022; There are a number of startups (including my own) that are actually designing for the real bottlenecks that differentiate Transformers (plus SSMs and other emerging architectures) from "old" CNNs, RNNs, etc.

Obviously I think my company is doing this in an unique and "correct" way, but I know of half a dozen other companies founded in the past ~18 months that are focused on the memory capacity and bandwidth bottlenecks that exist... the massive failures of the previous decade do not mean that they are going to be repeated.

The research is slightly misleading... the models they experimented all had an original pretrained context length significantly less than the fine tuned context length they tested for, e.g. they used MPT-30B-Instruct, which was pretrained for 2k sequence length and then fine tuned for 8k sequence length. A real test of if current self attention has this issue would be natively training a model with the extended sequence length.

I appreciate your feedback, and I agree I could have better worded and taken it in a more constructive direction. Thank you, and I hope myself and others will try to have more productive discourse.

Sadly I’m on my phone at the moment and can’t find the specific post, but in that PR or related discussion there was talk of only a few GB of the weights actually being used during the computation, which anyone who understands how a multi headed attention transformer works would know is impossible… your QKV matmuls need to touch all of the weights once you go through all the layers. Since that post yesterday getting 1200+ upvotes resulted in multiple conversations in my social circle that took that untrue statement as fact.

The tl;dr as I understand it is that jart had a misunderstanding of how what was actually happening and the benefits of the map optimization… the claims of actually being able to shrink the model size from 20GB > 6GB were just completely false, and while there was a model loading time improvement, actual memory required and used did not change.

A number of people saw this and said that making a breaking change to the repo that a lot of people are using and have forked for other models was a bad idea, thus this new PR.

It’s in large part due to the wire density… a 6T SRAM bit cell by itself could scale transistor density well, but the word lines and bit lines are the limiters as they typically have to go up 4 metal layers. Also, each read/write port you add to the SRAM macro increases the size exponentially rather than the linear scaling of the bit cells.