Pipeline-parallel LLM inference across GPUs on separate machineshttps://github.com/leyten/shard by ngaut • 1 month ago 5 0 1 month agoGIgithub.com