HN user

numlocked

3,594 karma

Chris Clark.

COO, OpenRouter (https://openrouter.ai)

Previously: Co-founder & CTO Grove (www.grove.co, NYSE:$GROV)

Creator of SQL Explorer (www.sqlexplorer.io)

[ my public key: https://keybase.io/cc; my proof: https://keybase.io/cc/sigs/mvtne8Fa_G2XaRwEkDFpAifq6DYrB5PY5rpj9-RHZ4A ]

Posts72
Comments393
View on HN
openrouter.ai 2mo ago

Opus 4.7's New Tokenizer: What It Costs

numlocked
3pts0
openrouter.ai 7mo ago

Response Healing: Reduce JSON defects by 80%+

numlocked
51pts47
openrouter.ai 8mo ago

Is implicit caching prompt retention?

numlocked
6pts0
openrouter.ai 9mo ago

LLM Provider Variance: Introducing Exacto

numlocked
3pts0
news.ycombinator.com 1y ago

Ask HN: Do you regularly use AI agents to get work done? How and why?

numlocked
1pts2
github.com 2y ago

Show HN: SQL Explorer – Open-source reporting tool that Just Works

numlocked
215pts52
blog.untrod.com 2y ago

LLM-Powered Django Admin Fields

numlocked
2pts0
www.axios.com 2y ago

Furious Congress plows forward with TikTok bill after user revolt

numlocked
25pts30
www.nytimes.com 2y ago

' Oldest Pyramid' in Indonesia? A Study Draws Skepticism

numlocked
3pts0
blog.untrod.com 2y ago

Robot Dad

numlocked
238pts74
en.wikipedia.org 2y ago

Operation CHASE

numlocked
2pts0
blog.untrod.com 3y ago

Copy Editing a Novel with ChatGPT

numlocked
2pts0
replicate.com 3y ago

Fine Tune LLaMA to Speak Like Homer Simpson

numlocked
1pts0
arxiv.org 3y ago

Reflexion: An autonomous agent with dynamic memory and self-reflection

numlocked
4pts1
severelytheoretical.wordpress.com 3y ago

Pleas Stop Training Giant Language Models

numlocked
3pts1
news.ycombinator.com 3y ago

Tell HN: Dark Sky Is Working

numlocked
4pts2
www.aoml.noaa.gov 3y ago

Attempts to Stop a Hurricane in its Track

numlocked
4pts1
www.youtube.com 4y ago

Daniel Bernstein- the Post Quantum Internet (2016; Video)

numlocked
1pts0
en.wikipedia.org 4y ago

DWIM

numlocked
1pts0
blog.untrod.com 4y ago

Adventures in Candy Land

numlocked
1pts0
en.wikipedia.org 4y ago

List of Ciphertexts

numlocked
2pts0
upload.wikimedia.org 5y ago

Imperial Units of Length

numlocked
2pts0
en.wikipedia.org 5y ago

Etaoin Shrdlu

numlocked
1pts0
www.projectpluto.com 5y ago

Easter Dates

numlocked
1pts0
astropixels.com 5y ago

Bifrost Astronomical Observatory

numlocked
1pts0
www.nytimes.com 6y ago

Why Is a Tech Executive Installing Security Cameras Around San Francisco?

numlocked
7pts0
blog.prototypr.io 8y ago

Creating a UX research recruiting engine

numlocked
1pts0
www.nytimes.com 8y ago

LIGO Detects Fierce Collision of Neutron Stars for the First Time

numlocked
4pts0
blog.untrod.com 9y ago

A Simple Trending Products Recommendation Engine in Python

numlocked
1pts0
www.monroeworktoday.org 9y ago

Map of White Supremacy Mob Violence

numlocked
1pts0

Refund policies are clearly documented in our terms. We actually DO offer refunds within 24hrs of credit purchase, which is significantly more flexible than most companies that operate in a similar way. And we try to use good judgement when there are extenuating circumstances.

By default (and in most cases) investors and operators are aligned. When we diligence our investors, we call companies they worked with where things didn’t go well, and speak to those founders. Understanding how investors operate when it’s not all up-and-to-the-right is important when picking partners!

Yep!

Everyone wants a conspiracy, but what I originally posted is in fact the boring truth. Having a bunch of cash in the bank makes for a durable business!

We have two mechanisms whereby we retain data. Both are opt-in and off by default.

One mechanism where you get a discount and we can use the data (in theory this does mean sell it; but our intent is to use it to make efficient dynamic routing solutions. But absolutely we could one day sell it) and another where we retain it for you so you can see it in your logs. We have no rights to this data in any way. This is similar to how any tracing/logging solution works.

Both and opt-in. If you don’t opt in, we don’t retain anything and are a pass through with regards to your prompt data.

All of this is carefully documented and I encourage you to explore and chat with the docs.

Our general theory of the case is that, in the not so distant future, inference will be the second largest opex line item for most companies (behind headcount) and that sourcing, measuring, and governing those tokens is a massive horizontal opportunity.

We will inevitably expand into adjacencies because we like building things and experimenting and we have a lot of people with great taste who are likely to ship cool things that customers want to use!

Edit: also - THANK YOU!

Interesting. Will look into it! We are releasing pass through API params soon which might hit the bid, but is a bit different than what you are describing.

What differentiation are looking for? We have good documentation of every provider and what their data retention stance is, and you can figure allow/blocklists for all providers.

Check out the Guardrails section under settings and tell me what’s missing!

Hi HN! OpenRouter co-founder and COO here. Lots of questions about why we raised!

First off: We remain founder-led and founder-controlled, and intend on being here for a long time, creating awesome products for builders all over the world. We are basically a bunch of tinkerers who like building things, and try to make stuff that we would like, when building with AI.

Since this is about the raise though, happy to share perspective on it.

We believe that strong companies should have a strong balance sheets. We touch large volumes of spend, and have large spend commits across the ecosystem; having the cash to withstand what may come is a responsible buy-down of risk, and makes the company extremely durable.

It also tells our larger customers and provider partners that we will be able to continue to serve them (and pay our bills) for a long time to come. We don't need venture dollars to continue scaling (indeed the business is healthy) but you know when you don't want to raise $100m? When you really need it!

This is also good validation to employees (current and future) that the value we are creating together is real. We also take seriously our obligation to make a return for anyone who invests; we aren't valuationmaxxing and have the privilege of getting to pick who we work with. I don't think that gets a lot of airtime in the overall start-up world, but I think it's important!

Happy to answer questions and THANK YOU to everyone here who uses OpenRouter, and to everyone who has feedback for how we can improve!

(openrouter co-founder here)

Yeah we should do something to indicate cardinality. I can share that there can often (I'm talking generally; not related to this model in particular) be e.g. a very large app that can be pushing a lot of volume. But in almost all cases that app has a large number of end users. Hypothetically, for instance, would Cursor be consider one user, or millions?

Will think about it! Thanks for the feedback.

You are absolutely allowed to expose access to end users, as long as you continue to abide by terms of service. We have hundreds, if not thousands, of apps built on openrouter that in turn have end users of their own. We showcase many of them on our /apps ranking page!

I watched the video at the expecting one thing and finding something completely different. Remarkable — [0] watch the video in its entirety. Not what I thought when I read “staples to repair porcelain”.

[0] intentional human use of an em-dash

As per its own FAQ this plugin is out of date and doesn’t actually do anything incremental re:caching:

"Hasn't Anthropic's new auto-caching feature solved this?"

Largely, yes — Anthropic's automatic caching (passing "cache_control": {"type": "ephemeral"} at the top level) handles breakpoint placement automatically now. This plugin predates that feature and originally filled that gap.

At the risk of totally misunderstanding this...it seems to be exfiltration by the app developer, who already has access to all of these data sources and the data that the customer is inputting into the AI KYC app (in this example)...right? I don't believe this exposes any end-user information to a third party. The AI app developer is already 'trusted' and could get access to this information regardless of the exfiltration. Maybe someone can explain this to me more clearly.

Mistral OCR 3 7 months ago

(I work at OpenRouter) If you send a PDF to our API we will:

1. Use native PDF parsing if the model supports it

2. Use this Mistral OCR model (we updated to this version yesterday)

3. UNLESS you override the "engine" param to use an alternate. We support a JS-based (non-LLM) parser as well [0]

So yes, in practice a lot of OCR jobs go to Mistral, but not all of them.

Would love to hear requests for other parsers if folks have them!

[0] https://openrouter.ai/docs/guides/overview/multimodal/pdfs#p...

Good read! The idea that these marvels of artistry were painted like my 10th birthday at the local paint-your-own-pottery store always seemed incongruous, at best.

Why, then, are the reconstructions so ugly?

...may be that they are hampered by conservation doctrines that forbid including any feature in a reconstruction for which there is no direct archaeological evidence. Since underlayers are generally the only element of which traces survive, such doctrines lead to all-underlayer reconstructions, with the overlayers that were obviously originally present excluded for lack of evidence.

That seems plausible -- and somewhat reasonable! To the credit of academics, they seems aware of this (according to the article):

‘reconstructions can be difficult to explain to the public – that these are not exact copies, that we can never know exactly how they looked’.

(I work at OpenRouter) Certainly for individual developers / hobby projects that's the primary value prop; super easy access to all of the models.

But there's a lot more functionality that becomes relevant when building in production. We do automatic fallbacks, route between providers based on data policies, syndicate your data to agent observability tools / your logging platform of choice, user-level and api-key-level budget management and model allow/block lists, programmatic API key management, etc, etc. More good stuff shipping all the time!

(I work at OpenRouter) We add about 15ms of latency once the cache is warm (e.g. on subsequent requests) -- and if there are reliability problems, please let us know! OpenRouter should be more reliable as we will load balance and fall back between different Gemini endpoints.

(OpenRouter COO here) We are starting to test this and verify the deployments. More to come on that front -- but long story short is that we don't have good evidence that providers are doing weird stuff that materially affects model accuracy. If you have data points to the contrary, we would love them.

We are heavily incentivized to prioritize/make transparent high-quality inference and have no incentive to offer quantized/poorly-performing alternatives. We certainly hear plenty of anecdotal reports like this, but when we dig in we generally don't see it.

An exception is when a model is first released -- for example this terrific work by artificial analysis: https://x.com/ArtificialAnlys/status/1955102409044398415

It does take providers time to learn how to run the models in a high quality way; my expectation is that the difference in quality will be (or already is) minimal over time. The large variance in that case was because GPT OSS had only been out for a couple of weeks.

For well-established models, our (admittedly limited) testing has not revealed much variance between providers in terms of quality. There is some but it's not like we see a couple of providers 'cheating' by secretly quantizing and clearly serving less intelligence versions of the model. We're going to get more systematic about it though and perhaps will uncover some surprises.

I hear you, but the article is talking specifically about "embeddings as a product" -- not the embeddings that are within an LLM architecture. It starts:

As a quick review, embeddings are compressed numerical representations of a variety of features (text, images, audio) that we can use for machine learning tasks like search, recommendations, RAG, and classification.

Current standalone embedding models are not intrinsically connected to SotA LLM architectures (e.g. the Qwen reference) -- right? The article seems to mix the two ideas together.

I don’t quite understand. The article says things like:

“With the constant upward pressure on embedding sizes not limited by having to train models in-house, it’s not clear where we’ll slow down: Qwen-3, along with many others is already at 4096”

But aren’t embedding models separate from the LLMs? The size of attention heads in LLMs etc isn’t inherently connected to how a lab might train and release an embedding model. I don’t really understand why growth in LLM size fundamentally puts upward pressure on embedding size as they are not intrinsically connected.