I've just found myself using OpenRouter if we need Google models for a project, it's worth the extra 5% just not to have to deal with the utter disaster that is their product offering.
HN user
byefruit
This is the wrong interpretation of the oxcaml project. If you look at the features and work on it, it's primarily performance or parallelism safety features.
The latter going much further than most mainstream languages.
7.3% return, not bad. As battery prices drop it will get even better.
And even when it does copy other products, it seems to be doing a terrible job of them.
Google's AI offering is a complete nightmare to use. Three different APIs, at least two different subscriptions, documentation that uses them interchangeably.
For Gemini's API it's often much simpler to actually pay OpenRouter the 5% surchargeto BYOK than deal with it all.
I still can't use my Google AI Pro account with gemini-cli..
How is this different from https://github.com/google-gemini/gemini-cli ?
Edit: it seems this is a hosted version. Would be nice if they actually joined up some of their products.
The openrouter rankings can be biased.
For example, Google's inexplicable design decisions around libraries and APIs means it's often worth the 5% premium to just use OpenRouter to access their models. In other cases it's about which models particular agents default to.
Sonnet 4 is extremely good for tool-usage agentic setups though - something I have found other models struggle to do over a long-context.
Before just accepting this at face value, New Statesman claim this is not the case:
https://www.newstatesman.com/politics/2025/07/the-british-we...
You are probably getting downvoted because you don't give any model generations or versions ('ChatGPT') which makes this not very credible.
100% this. We actually use OpenRouter (and pay their surcharge) with Gemini 2.5 Pro just because we can actually control spend via spent limit on keys (A++ feature) and prepaid credit.
Indeed, average in CA is $260/month so $5k pays off very fast in some places.
It's interesting that there's a price nearly 6x price difference between reasoning and no reasoning.
This implies it's not a hybrid model that can just skip reasoning steps if requested.
Anyone know what else they might be doing?
Reasoning means contexts will be longer (for thinking tokens) and there's an increase in cost to inference with a longer context but it's not going to be 6x.
Or is it just market pricing?
"In addition, S3 Express One Zone has reduced the per-GB charges for data uploads and retrievals by 60 percent, and these charges now apply to all bytes transferred rather than just portions of requests greater than 512 KB"
It's not clear but are there cases where this could be a significant price rise? If you exclusively had small objects (<512kb) being written and read then this could add up quickly.
Both of our models are trained on top of DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Qwen-32B.
Not to take away from their work but this shouldn't be buried at the bottom of the page - there's a gulf between completely new models and fine-tuning.
https://aws.amazon.com/snowball/pricing/ snowball seems to support getting data out of S3 though you still end up paying extortionate egress charges.
I think this is where (in the EU) a Subject Access Request could work.
This has already been proposed by the current government for wind farms: https://inews.co.uk/news/environment/new-energy-bill-discoun...
And the switch to zonal energy pricing will likely have a similar effect for other sources of generation.
I'm waiting for https://github.com/huggingface/trl/pull/2810 to land. I think this should work with the existing unsloth setup without changes.
This generally requires thousands of examples created by an expert in the field.
Or an AI model pretending to be an expert in the field... (works well in a few niche domains I have used this in)
I look forward to people applying the same standards to the OpenAI's O3 as they did Deepseek's R1 release and paper in the discussions last week.
That's not what I'm saying, they may be hiding their true compute.
I'm pointing out that nearly every thread covering Deepseek R1 so far has been like this. Compare to the O1 system card thread: https://news.ycombinator.com/item?id=42330666
Very different standards.
It's amazing how different the standards are here. Deepseek's released their weights under a real open source license and published a paper with their work which now has independent reproductions.
OpenAI literally haven't said a thing about how O1 even works.
What I don't understand is how you don't end up with a totally impractical number of vectors if you have one per token? Surely nobody is storing that many in any real system?
Do you have any evidence for this accusation?
O1's reasoning traces aren't even shown, are you suggesting they've somehow exfiltrated them?
This is pretty harsh on DeepSeek.
There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token prediction.
What would you recommend for building a strong linear algebra foundation?
As sibling comments have said, you can pay per month to use the web interface, pay for use of the API via self-signup or you can even use them via openrouter.
I've seen a couple of talks on DSPy and tried to use it for one of my projects but the structure always feels somewhat strained. It seems to be suited for tasks that are primarily show, don't tell but what do you do when you have significant prior instruction you want to tell?
e.g Tests I want applied to anything retrieved from the database. What I'd like is to optimise the prompt around those (or maybe even the tests themselves) but I can't seem to express that in DSPy signatures.
Bit weird to show $ per 1M tokens and not include the actual costs of the systems anywhere.
It would be interesting to know the outright prices for those systems as well as their hourly rental rates at the moment.
This is very interesting. Have you got any references describing this approach?
A troll so good it necessitated a change in the law: https://publications.parliament.uk/pa/bills/cbill/58-03/0154...
(Page 16, 57A)
"A company must not be registered under this Act by a name that, in the opinion of the Secretary of State, consists of or includes computer code."