Our skippy library is a patch queue on top of llama that allows us to access internal information, such as activations, and filter tensors on model load.
To be honest, both are very tough problems we don't have a good answer for yet. If that is something that concerns you, look into building a private mesh with trusted peers.
Each stage has its own KV for the layers it hosts. You are on the money there, when one stage is waiting it's free for more parallelism. I am planning on exploiting this for more token verification through ngram spec decoding.
The lab features two Mac Studios: an Apple M3 Ultra (32 CPU cores, 80 GPU cores, 256 GB unified memory) and an Apple M1 Ultra (20 CPU cores, 48 GPU cores, 128 GB unified memory), both connected via 1Gbit Ethernet.
We use a customized Q2 quantization that preserves sensitive tensors at Q8.
To reduce compute time per layer, we are developing a custom GLM DSA Metal graph.
While we are not yet approaching MTP, we plan to port our existing MTP implementations from versions 4.7 and 5.1 to 5.2.
Since GLM's MTP acceptance rate is very high for a single predicted token, we are exploring token prediction techniques to widen the predicted tokens and utilize parallelism for verification.
This was done on my home lab simulating 5ms latency and jitter between machines. Splits work quite well if you your nodes are over WAN at metro latency’s but not super fast on global WAN.
The idea is that you could take several machines without dedicated RDMA or NVLINK fabric and use them to serve a large model on hardware you own then share it with others.
I’m currently working on GLM 5.2 on my lab environment with around 10 tok/s on the same split.
I’m one of the contributors to Mesh LLM and happy to answer any questions. I authored the skippy engine that allows you to split large models across nodes.
This is very revisionist. While they have been catching up quickly there was no master 4D chess strategy here. Google was incredibly late to this game - Sergey had to come back from retirement because most of the research team had a Sarah Connor complex and couldn’t ship. The saving grace is that AdWords picked up the tab again and founders shook the place up when it became clear the golden goose was being cooked.
Android supremacy at its finest. I would never recommend a family member buying one. The history of this kind of thing is long and keeps continuing to happen.
Honestly good luck. I left a company and removed myself from the meta account which triggered the deletion of our whole app. Impossible to recover. Meta suck.
Not sure why this is downvoted. Economic activity should be enjoyed by the commons. For example LNG being exported UNDER international value and Aussies buying it at international prices is idiotic.
There used to be a TRIM program you could install. Used one when I swapped out the super drive for an ssd in a 2012 MacBook Pro (I think at around the same time?)