Show HN: Makes local LLMs faster and more reliable by optimizing for your device

https://www.autotunellm.com/
by tanavc • 9 hours ago
2 0 9 hours ago

Time to first token is 39% faster

Agent wall times decrease by 46%

No swaps

Tracks your resource usage in real-time and adjusts how the model runs so that it works perfectly on your device.

Implements KV cache sizing, prefix caching, live RAM pressure management, context trimming, KV quantization, and more.

Built a ton of features

10k downloads and climbing

consider using it and giving it a star :)

Related Stories

Loading related stories...

Source preview

autotunellm.com