A common reason is to reduce cost and latency. Larger models typically require GPUs with more memory (and hence higher costs), plus the time to serve requests is also higher (more matrix multiplications to be done).
rish-b
https://twitter.com/rish_bhargava
Posts3
Comments2