Just sent you a DM on twitter!
HN user
numlocked
Chris Clark.
COO, OpenRouter (https://openrouter.ai)
Previously: Co-founder & CTO Grove (www.grove.co, NYSE:$GROV)
Creator of SQL Explorer (www.sqlexplorer.io)
[ my public key: https://keybase.io/cc; my proof: https://keybase.io/cc/sigs/mvtne8Fa_G2XaRwEkDFpAifq6DYrB5PY5rpj9-RHZ4A ]
Refund policies are clearly documented in our terms. We actually DO offer refunds within 24hrs of credit purchase, which is significantly more flexible than most companies that operate in a similar way. And we try to use good judgement when there are extenuating circumstances.
By default (and in most cases) investors and operators are aligned. When we diligence our investors, we call companies they worked with where things didn’t go well, and speak to those founders. Understanding how investors operate when it’s not all up-and-to-the-right is important when picking partners!
Yep!
Everyone wants a conspiracy, but what I originally posted is in fact the boring truth. Having a bunch of cash in the bank makes for a durable business!
We have two mechanisms whereby we retain data. Both are opt-in and off by default.
One mechanism where you get a discount and we can use the data (in theory this does mean sell it; but our intent is to use it to make efficient dynamic routing solutions. But absolutely we could one day sell it) and another where we retain it for you so you can see it in your logs. We have no rights to this data in any way. This is similar to how any tracing/logging solution works.
Both and opt-in. If you don’t opt in, we don’t retain anything and are a pass through with regards to your prompt data.
All of this is carefully documented and I encourage you to explore and chat with the docs.
We have never sold any prompt data to anyone, in any form, and have no plans to do so. Full stop.
Great investors are helpful, not harmful :) You want accountability from smart, experienced partners!
Our general theory of the case is that, in the not so distant future, inference will be the second largest opex line item for most companies (behind headcount) and that sourcing, measuring, and governing those tokens is a massive horizontal opportunity.
We will inevitably expand into adjacencies because we like building things and experimenting and we have a lot of people with great taste who are likely to ship cool things that customers want to use!
Edit: also - THANK YOU!
Interesting. Will look into it! We are releasing pass through API params soon which might hit the bid, but is a bit different than what you are describing.
What differentiation are looking for? We have good documentation of every provider and what their data retention stance is, and you can figure allow/blocklists for all providers.
Check out the Guardrails section under settings and tell me what’s missing!
Hi HN! OpenRouter co-founder and COO here. Lots of questions about why we raised!
First off: We remain founder-led and founder-controlled, and intend on being here for a long time, creating awesome products for builders all over the world. We are basically a bunch of tinkerers who like building things, and try to make stuff that we would like, when building with AI.
Since this is about the raise though, happy to share perspective on it.
We believe that strong companies should have a strong balance sheets. We touch large volumes of spend, and have large spend commits across the ecosystem; having the cash to withstand what may come is a responsible buy-down of risk, and makes the company extremely durable.
It also tells our larger customers and provider partners that we will be able to continue to serve them (and pay our bills) for a long time to come. We don't need venture dollars to continue scaling (indeed the business is healthy) but you know when you don't want to raise $100m? When you really need it!
This is also good validation to employees (current and future) that the value we are creating together is real. We also take seriously our obligation to make a return for anyone who invests; we aren't valuationmaxxing and have the privilege of getting to pick who we work with. I don't think that gets a lot of airtime in the overall start-up world, but I think it's important!
Happy to answer questions and THANK YOU to everyone here who uses OpenRouter, and to everyone who has feedback for how we can improve!
Thanks! We work really hard to make sure we are ready at launch :)
(openrouter co-founder here)
Yeah we should do something to indicate cardinality. I can share that there can often (I'm talking generally; not related to this model in particular) be e.g. a very large app that can be pushing a lot of volume. But in almost all cases that app has a large number of end users. Hypothetically, for instance, would Cursor be consider one user, or millions?
Will think about it! Thanks for the feedback.
Can you share more? I'm with OpenRouter and we would love to address this! We don't see this in our own testing, I don't believe -- but will share this feedback and dig in.
Doesn’t it seem more plausible that the marketing shots are AI (where the “generated by AI” note appears) rather than the cover designs themselves?
You are absolutely allowed to expose access to end users, as long as you continue to abide by terms of service. We have hundreds, if not thousands, of apps built on openrouter that in turn have end users of their own. We showcase many of them on our /apps ranking page!
COO of OpenRouter here. Thats right — we haven’t done it to date but we can’t have unlimited liabilities stacking up forever. At some point we will start expiring credits from accounts that have seen zero activity in over a year.
I watched the video at the expecting one thing and finding something completely different. Remarkable — [0] watch the video in its entirety. Not what I thought when I read “staples to repair porcelain”.
[0] intentional human use of an em-dash
As per its own FAQ this plugin is out of date and doesn’t actually do anything incremental re:caching:
"Hasn't Anthropic's new auto-caching feature solved this?"
Largely, yes — Anthropic's automatic caching (passing "cache_control": {"type": "ephemeral"} at the top level) handles breakpoint placement automatically now. This plugin predates that feature and originally filled that gap.
At the risk of totally misunderstanding this...it seems to be exfiltration by the app developer, who already has access to all of these data sources and the data that the customer is inputting into the AI KYC app (in this example)...right? I don't believe this exposes any end-user information to a third party. The AI app developer is already 'trusted' and could get access to this information regardless of the exfiltration. Maybe someone can explain this to me more clearly.
That is a way of doing that, but it's quite expensive computationally. There are some companies that can make it feasible [0], but it's often not a perfect process and different inference providers implement it different ways.
(I work at OpenRouter) If you send a PDF to our API we will:
1. Use native PDF parsing if the model supports it
2. Use this Mistral OCR model (we updated to this version yesterday)
3. UNLESS you override the "engine" param to use an alternate. We support a JS-based (non-LLM) parser as well [0]
So yes, in practice a lot of OCR jobs go to Mistral, but not all of them.
Would love to hear requests for other parsers if folks have them!
[0] https://openrouter.ai/docs/guides/overview/multimodal/pdfs#p...
I just learned that the site/magazine publishing this, Works in Progress, is owned by Stripe! I have no idea why, but the content is great so...thanks Stripe!
Good read! The idea that these marvels of artistry were painted like my 10th birthday at the local paint-your-own-pottery store always seemed incongruous, at best.
Why, then, are the reconstructions so ugly?
...may be that they are hampered by conservation doctrines that forbid including any feature in a reconstruction for which there is no direct archaeological evidence. Since underlayers are generally the only element of which traces survive, such doctrines lead to all-underlayer reconstructions, with the overlayers that were obviously originally present excluded for lack of evidence.
That seems plausible -- and somewhat reasonable! To the credit of academics, they seems aware of this (according to the article):
‘reconstructions can be difficult to explain to the public – that these are not exact copies, that we can never know exactly how they looked’.
We do this at openrouter and many apps use exactly that pattern!
(I work at OpenRouter) Certainly for individual developers / hobby projects that's the primary value prop; super easy access to all of the models.
But there's a lot more functionality that becomes relevant when building in production. We do automatic fallbacks, route between providers based on data policies, syndicate your data to agent observability tools / your logging platform of choice, user-level and api-key-level budget management and model allow/block lists, programmatic API key management, etc, etc. More good stuff shipping all the time!
(I work at OpenRouter) We add about 15ms of latency once the cache is warm (e.g. on subsequent requests) -- and if there are reliability problems, please let us know! OpenRouter should be more reliable as we will load balance and fall back between different Gemini endpoints.
(OpenRouter COO here) We are starting to test this and verify the deployments. More to come on that front -- but long story short is that we don't have good evidence that providers are doing weird stuff that materially affects model accuracy. If you have data points to the contrary, we would love them.
We are heavily incentivized to prioritize/make transparent high-quality inference and have no incentive to offer quantized/poorly-performing alternatives. We certainly hear plenty of anecdotal reports like this, but when we dig in we generally don't see it.
An exception is when a model is first released -- for example this terrific work by artificial analysis: https://x.com/ArtificialAnlys/status/1955102409044398415
It does take providers time to learn how to run the models in a high quality way; my expectation is that the difference in quality will be (or already is) minimal over time. The large variance in that case was because GPT OSS had only been out for a couple of weeks.
For well-established models, our (admittedly limited) testing has not revealed much variance between providers in terms of quality. There is some but it's not like we see a couple of providers 'cheating' by secretly quantizing and clearly serving less intelligence versions of the model. We're going to get more systematic about it though and perhaps will uncover some surprises.
I hear you, but the article is talking specifically about "embeddings as a product" -- not the embeddings that are within an LLM architecture. It starts:
As a quick review, embeddings are compressed numerical representations of a variety of features (text, images, audio) that we can use for machine learning tasks like search, recommendations, RAG, and classification.
Current standalone embedding models are not intrinsically connected to SotA LLM architectures (e.g. the Qwen reference) -- right? The article seems to mix the two ideas together.
I don’t quite understand. The article says things like:
“With the constant upward pressure on embedding sizes not limited by having to train models in-house, it’s not clear where we’ll slow down: Qwen-3, along with many others is already at 4096”
But aren’t embedding models separate from the LLMs? The size of attention heads in LLMs etc isn’t inherently connected to how a lab might train and release an embedding model. I don’t really understand why growth in LLM size fundamentally puts upward pressure on embedding size as they are not intrinsically connected.