How do run finetuned models in a multi-tenant/shared GPU setup?

https://news.ycombinator.com/item?id=41718501
by iamzycon • 2 years ago
1 0 2 years ago

I'm considering setting up a fine-tuning and inference platform for Llama that would allow customers to host their fine-tuned models. Would it be necessary to allocate a dedicated infrastructure for each fine-tuned model, or could a shared infrastructure work? Are there any existing solutions for this?

Related Stories

Loading related stories...

Source preview

news.ycombinator.com