So, this is 'open' as in 'OpenAI', not as in 'open source'. What's the benefit of this compared to something like NOW Payments for merchants (or the myriad of alternatives)? NOW Payments supports a wide range of crypto currencies as well.
HN user
Palmik
Commodity providers aren't a good indicator. They have margins too.
Remember they ~doubled the price going from GLM 5 to GLM 5.2, despite same [1] cost of inference.
[1] GLM 5.2 is actually slightly more efficient, thanks to baked in indexer cache.
By requiring various forms of identification to use social media, it will be harder to criticize your leaders anonymously without fear of retribution.
The company representative said that they report all users that use Graphene OS, without any additional qualifiers. Presumably after they've already uploaded their personal details. That's the egregious part.
DeepSeek V4's KV cache is very efficient due to its heavily compressed and sparse attention architecture.
DeepSeek V3.2 which uses DSA only (sparse attention, but without compression from HCA and CSA) is a smaller model but uses 10x more memory at 1M context window compared to DS V4 Pro.
Also, I have to say, DeepSeek's API has a very good cache hit rate. With the same workload, I see ~80% KV cache hit rate with the DS API vs ~50% with the major western inference providers for open weight models.
I really hope Huawei ramps up Ascend production and DeepSeek open sources their optimized inference engine (they already open source a lot of their kernels -- kudos to them). This could shake things up.
There are several things at play:
Inference stack efficiency: Many of these providers take off the shelf sglang / vllm / trtllm and hope for the best. Meanwhile DeepSeek team is known for pushing the boundary of optimizations.
Now, sglang and vllm are great pieces of software, but take DeepSeek's Sparse Attention (DSA). Introduced 1.5 years ago (https://arxiv.org/abs/2512.02556), used by DeepSeek 3.2, GLM 5, DeepSeek V4. Only now is it slowly strating to get optimized in the major inference engines: (https://github.com/sgl-project/sglang/issues/19380 https://github.com/sgl-project/sglang/pull/22851 etc.). Of course, DS V4 adds extra optimizations into the model architecture on top of DSA, and those will take more time to be taken full advantage of by the open source inference engines.
Privacy: Betting that people will pay extra for inference hosted outside China. This is especially true with DeepSeek, because DeepSeek is transparent about using API data for model improvements.
And few other things (scale (matters a lot for MoEs), reliability, soft enterprise lock in, etc.)
---
There is also, likely, tacit collusion at play here. Look at GLM 5 and GLM 5.1 prices. GLM 5 and 5.1 cost the same to run, but providers decided to charge much more for 5.1 because it is much better model, and because Z.AI raised their price as well.
Why was the title changed from "DeepSeek V4—almost on the frontier, a fraction of the price" to "DeepSeek V4—almost on the frontier"?
Surely art also exists in textual realm.
I don't think "friendly" and "publishing benchmarks" are at odds with each other.
Model makers (both open and closed weight) typically publish benchmarks against other models and when they do not, people rightfully call them out.
Including comparison against "other OSS engine" is just not helpful (what if it's a sandbagged baseline like HF Transformers?)
Similar article for vLLM: https://vllm-website-pdzeaspbm-inferact-inc.vercel.app/blog/...
Bechmarks from InferenceX (they do not have apples-to-apples setups to compare the different engines for whatever reason): https://inferencex.semianalysis.com/inference?i_hc=1&g_model...
I find it odd that sglang, vLLM, TRTLLM don't seem to want to publish benchmarks comparing each other. They used to, but now there seems to be some unspoken rule against it.
At least we get comparison against "other OSS engine" this time, but that could be HF's Transformers as well :)
Or there will be DSv4.1/2/3 ;)
Misleading conclusion.
This model is 8 times cheaper than Gemini for 1K images. Gemini is extremely overpriced.
1K image with Gemini is roughly $0.08 and only $0.01 with GPT Image.
Did you enable thinking for your experiment? Are you sure you were on the 2.0 rather than 1.5 version?
I do not think this is a good prompt or useful benchmark, but nonetheless, it seems to work better for me: https://chatgpt.com/share/69e88a94-ded8-8395-b5dc-abceb2f44d...
Could it be made even faster using some of the ideas from https://github.com/zerobootdev/zeroboot ?
Official announcement: https://www.thunderbolt.io/announcing-thunderbolt
My email does mention it clearly:
Again, your organization's Copilot interaction data is not included in model training under this new policy, but we are excited for you to enjoy the product improvements it will unlock.
Great work! There is maybe some bug. When you click on one of the 4 "opposing" countries (e.g. Czech Republic, Poland), it scrolls down and then shows that majority of the representatives from the country actually support it. Is that intended? Won't that make people from those countries "relax" even though they might have an impact by contacting their represenatives?
What are your thoughts on this? https://www.anthropic.com/news/where-stand-department-war
I am honestly unclear on the reasoning of people who flock from OpenAI to Anthropic, and doubly so of those who are not US citizens.
Except most of the world's population, and in fact large fraction of the engineers and scientists working on these things, are not US citizens.
From Anthropics recent blog post: https://www.anthropic.com/news/detecting-and-preventing-dist...
By examining request metadata, we were able to trace these accounts to specific researchers at the lab.
The volume, structure, and focus of the prompts were distinct from normal usage patterns
Clearly some employees of Anthropic personally looked at individual inputs and outputs of their API
Who here reads the full terms of service of every Google product they use? The fact that they disabled the whole Google account without warning is damning.
They could have easily just blocked the Gemini / Antigravity use and and/or sent a "final warning" kind of email beforehand.
That issue exists with the current proposal as well or any proposal that leaves the enforcement on the website.
I think in addition to what OP said, the browser/device should let you set hard domain-level filters which are enforced by the browser/device.
This will not be ideal for applications / sites with mixed content, but gives the parent / guardian more control.
I would pay for a solution that aggregates these alternative payment systems that aren't tied to VISA and Mastercard and their rules.
Wero, RuPay, WeChat Pay / AliPay, Crypto (through 3p like coinbase or directly), etc.
And handles them in nice and unified API, with hooks for subscriptions, etc. (taking into account that some payment methods do not support recurring payments, so there should be some hook to send an email to the customer to renew manually, etc.)
Yeah, after looking more into sqldef and alternatives I stumbled on Atlas too and I like the explicit support for migration based flow for exactly the same reasons. I want to know exactly what kind of migration will be applied to my prod database beforehand.
EU might be looking to do what China and Russia did earlier on and start cracking down on foreign social media
For some reason you forgot to mention "Like the US did with TikTok".
Actual Greek yogurt will have 8-12% of it's weight in protein.
Another option might be curd / quark (differs a lot per country).
Is there yet a low friction way to verify age of UK users that doesn't rely on third party services with questionable privacy implications and exorbitant pricing?
I wonder if any of the law makers are investors in those companies.