HN user

nabakin

2,183 karma
Posts14
Comments636
View on HN

Are you running qwen3.6-27b on one 3090 with your KV cache at q4? Ime there is significant long-context recall accuracy degradation at that precision. I prefer putting the KV cache at q8 and working with the 120k context

DeepSeek v4 3 months ago

I think they were mistaken or maybe they were just referring to inference because I don't see anyone making that claim and it would be quite the news.

DeepSeek v4 3 months ago

Probably because you said you used DeepSeek. People don't want to see AI in the comments and don't trust AI responses.

DeepSeek v4 3 months ago

Also, note that there's zero CUDA dependency. It runs entirely on Huawei chips.

That is a huge claim to make with no evidence.

I researched what you said, and I have found no statement to that effect in their paper[0], on huggingface[1], twitter[2], WeChat[3], or in their news release[4].

They only mention as a footnote in only the Chinese version of their news release that they plan to reduce inference costs with the Ascend 950 supernode when it releases[5]. The only mention of Huawei in their paper is that they validated a technique to lower interconnect bandwidth on Ascend NPUs and Nvidia GPUs[6].

[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main...

[1] https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro

[2] https://xcancel.com/deepseek_ai/status/2047516922263285776

[3] https://mp.weixin.qq.com/s/8bxXqS2R8Fx5-1TLDBiEDg

[4] https://api-docs.deepseek.com/news/news260424

[5] https://api-docs.deepseek.com/zh-cn/img/v4-price.png

[6] Page 16

When it comes to information transfer and processing, light can do things that electricity can’t. Photons — particles of light — are far zippier than electrons at working their way through circuits.

Electrons themselves don't move at the speed of light, but information transfer (i.e. communication) via electrons does happen close to the speed of light.

A subtle, but important, distinction that's often misunderstood and means computational performance gains would probably come from bandwidth, not latency.

If OP meant they have the fastest implementation of Gemma 4 on Blackwell at the moment, I guess that is technically true. I doubt that will hold up when TensorRT-LLM finishes their implementation though.

I know Arc AGI 2 has a private test set and they have a good amount of results[0] but it's not a conventional benchmark.

Looking around, SWE Rebench seems to have decent protection against training data leaks[1]. Kagi has one that is fully private[2]. One on HuggingFace that claims to be fully private[3]. SimpleBench[4]. HLE has a private test set apparently[5]. LiveBench[6]. Scale has some private benchmarks but not a lot of models tested[7]. vals.ai[8]. FrontierMath[9]. Terminal Bench Pro[10]. AA-Omniscience[11].

So I guess we do have some decent private benchmarks out there.

[0] https://arcprize.org/leaderboard

[1] https://swe-rebench.com/about

[2] https://help.kagi.com/kagi/ai/llm-benchmark.html

[3] https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard

[4] https://simple-bench.com/

[5] https://agi.safe.ai/

[6] https://livebench.ai/

[7] https://labs.scale.com/leaderboard

[8] https://www.vals.ai/about

[9] https://epoch.ai/frontiermath/

[10] https://github.com/alibaba/terminal-bench-pro

[11] https://artificialanalysis.ai/articles/aa-omniscience-knowle...

It's easy to game and human evaluation data has its trade-offs, but it's way easier to fake public benchmark results. I wish we had a source of high quality private benchmark results across a vast number of models like Lmarena. Having high quality human evaluation data would be a plus too.

Public benchmarks can be trivially faked. Lmarena is a bit harder to fake and is human-evaluated.

I agree it's misleading for them to hyper-focus on one metric, but public benchmarks are far from the only thing that matters. I place more weight on Lmarena scores and private benchmarks.

This is a better link from a French privacy non-profit but I can't change it now: https://mamot.fr/@LaQuadrature/115581775965025042

@dang or other mods, could you change it?

Google Translated text:

Two articles in Le Parisien yesterday, followed today by one in Le Figaro, have launched a shameful attack against GrapheneOS, a free and accessible open-source operating system for phones. At La Quadrature du Net, it's one of the tools we favor and regularly recommend for protecting against advertising tracking and spyware.

Echoing the propaganda of the Ministry of the Interior, newspapers describe GrapheneOS as a "crime-related phone solution," and a police officer adds that its use is suspicious in itself because it indicates an "intention to conceal." By portraying GrapheneOS as a technology linked to drug trafficking, this attack aims to criminalize what is actually a secure privacy-preserving tool.

In these articles, the head of the cybercrime section of the Paris prosecutor's office – who was behind the arrest of Pavel Durov – also threatens the developers of GrapheneOS. In an interview, she warns that she will "not hesitate to prosecute the publishers if links are discovered with a criminal organization and they do not cooperate with the justice system." https://archive.is/20251119110251/https://www.leparisien.fr/...

The government regularly tries to link privacy technologies, particularly encryption, to criminal behavior in order to undermine them and justify surveillance policies. This was the case in the so-called "December 8th" case, where a police narrative was constructed around the (secure) digital practices of the accused to portray a "clandestine" and "conspiratorial" group. https://www.laquadrature.net/2023/06/05/affaire-du-8-decembr...

Now, drug trafficking is being used to attack these technologies and justify the surveillance of communications. The so-called "Drug Trafficking" law was thus used as a pretext to try to legalize "backdoors" in encrypted applications like Signal or WhatsApp, without success. https://www.laquadrature.net/2025/03/18/le-gouvernement-pret...

An article in Le Monde diplomatique from November extensively examines the history of the political exploitation of drug trafficking to justify security and surveillance policies. The police attack on GrapheneOS fits perfectly within this pattern. https://www.monde-diplomatique.fr/2025/11/BONELLI/68915

In its response published yesterday, GrapheneOS points to the authoritarian tendencies of the French government, one of the most fervent supporters of the "ChatControl" regulation under discussion at the European level, one of whose goals is to put an end to end-to-end encryption. https://grapheneos.social/@GrapheneOS/115575997104456188

Additional context:

https://grapheneos.social/deck/@GrapheneOS/11557599710445618...

https://grapheneos.social/@GrapheneOS/115583866253016416

https://grapheneos.social/@LaQuadrature@mamot.fr/11558177594...

https://grapheneos.social/@GrapheneOS/115589833471347871

https://grapheneos.social/@GrapheneOS/115594002434998739

Steam Frame 8 months ago

I haven't bought a VR headset since the Oculus Rift CV1, but this might do it for me

Steam Frame 8 months ago

I don't think a lot of people realize how big of a deal this is. You used to have to choose between wireless and slow or wired and fast. Now you can have both wireless and fast. It's insane.

Steam Frame 8 months ago

And foveated streaming has a 1-2ms wireless latency on modern GPUs according to LTT. Insane.