HN user

om8

65 karma
Posts2
Comments56
View on HN

manufacturing companies making the flimsiest, cheapest, plastic crap to save 1/3 of a cent on every mop they produce. Designed to work for the least amount of time before needing replaced

We live in a world with such companies, and we can still buy quality things. If there is a demand for the purely-human generated texts, they will be around. Perhaps a lot of people around you will read ai text instead, and you'll get upset because of it, but it's their choice. You'll still have your thing

https://docs.vllm.ai/en/v0.20.0/api/vllm/model_executor/laye...

`vllm.model_executor.layers.quantization.turboquant`

The technique implemented here consists of the scalar case of the HIGGS quantization method (Malinovskii et al., "Pushing the Limits of Large Language Model Quantization via the Linearity Theorem", NAACL 2025; preprint arXiv:2411.17525): rotation + optimized grid + optional re-normalization, applied to KV cache compression. A first application of this approach to KV-cache compression is in "Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models" (Shutova et al., ICML 2025; preprint arXiv:2501.19392). Both these references pre-date the TurboQuant paper (Zandieh et al., ICLR 2026).

Claude Opus 4.6 6 months ago

Is there a way to disable it? Sometimes I value agent not having knowledge that it needs to cut corners

Ollama Turbo 12 months ago

It’s unfortunate that llama.cpp’s code is a mess. It’s impossible to make any meaningful contributions to it.

I'm currently in Russia trying to get a US visa for my CS PhD. Because I do CS, I got into a thing called administrative processing. For 95% of people, it takes days -- weeks, tops. Because of the colour of my passport, it has already lasted for 3 months. I know people who are waiting for 2 years to pass it.

Why do you think I shouldn't have access to this website in Russia?

What would you do if some of the deps started to have conflicts in them? Also, what are your plans for migration when you'll need to move from one os version to another?

Implicit solutions like yours have lower cost of entrance, but larger cost of support. uv python scripts just work if you set them up once

a local LLM is going to be the way to go

Non-technicals don't know how LLMs work, and, more importantly, don't care about their privacy.

For a technology to be widely used, by definition, you need to make it appealing to the masses, and there is almost zero demand for private LLM right now.

That's why I don't think that local llms will win. There are narrow use cases where regulations can force local llm usage (like for medical stuff), but overall I think that services will win (as they always do)