Thank you for the detailed feedback! I shared this already with the team.
HN user
volodia
This looks like an inference glitch that we are working on fixing, thank you for flagging.
There are many ways to do it, but the simplest approach is block diffusion: https://m-arriola.com/bd3lms/
There are also more advanced approaches, for example FlexMDM, which essentially predicts length of the "canvas" as it "paints tokens" on it.
Would love to hear about your experience. Send us an email.
Not imminently, but hard to predict where the field will go
There are few: fast agents, deep research, real-time voice, coding. The other thing is that when you have a fast reasoning model, you spend more effort on thinking in the same latency budget, which pushed up quality.
We agree! In fact, there is an emerging class of models aimed at fast agentic iteration (think of Composer, the Flash versions of proprietary and open models). We position Mercury 2 as a strong model in this category.
That is also our view! We see Mercury 2 as enabling very fast iteration for agentic tasks. A single shot at a problem might be less accurate, but because the model has a shorter execution time, it enables users to iterate much more quickly.
You can think of Mercury 2 as roughly in the same intelligence tier as other speed-optimized models (e.g., Haiku 4.5, Grok Fast, GPT-Mini–class systems). The main differentiator is latency — it’s ~5× faster at comparable quality.
We’re not positioning it as competing with the largest models (Opus 4.5, etc.) on hardest-case reasoning. It’s more of a “fast agent” model (like Composer in Cursor, or Haiku 4.5 in some IDEs): strong on common coding and tool-use tasks, and providing very quick iteration loops.
Thanks for trying it and for the thoughtful feedback, really appreciate it. And we’re actively working on improving quality further as we scale the models.
Thank you for your patience. We are working to handle the surge in demand.
Just to clarify one point: Mercury (the original v1, non-reasoning model) is already used in production in mainstream IDEs like Zed: https://zed.dev/blog/edit-prediction-providers
Mercury v1 focused on autocomplete and next-edit prediction. Mercury 2 extends that into reasoning and agent-style workflows, and we have editor integrations available (docs linked from the blog). I’d encourage folks to try the models!
I’d push back a bit on the Pareto point.
On speed/quality, diffusion has actually moved the frontier. At comparable quality levels, Mercury is >5× faster than similar AR models (including the ones referenced on the AA page). So for a fixed quality target, you can get meaningfully higher throughput.
That said, I agree diffusion models today don’t yet match the very largest AR systems (Opus, Gemini Pro, etc.) on absolute intelligence. That’s not surprising: we’re starting from smaller models and gradually scaling up. The roadmap is to scale intelligence while preserving the large inference-time advantage.
Co-founder / Chief Scientist at Inception here. If helpful, I’m happy to answer technical questions about Mercury 2 or diffusion LMs more broadly.
There is also this one that was released in October: https://github.com/kuleshov/char-mdlm
the LLaDA paper is a scaled-up version of this paper; they cite it as an anonymous ICLR submission
Great question! The model can more efficiently leverage existing GPU hardware---it performs more computation per unit of memory transferred; this means that on older hardware one should be able to get similar inference speeds as one would get on recent hardware with a classical LLM. This is actually interesting commercially, since it opens new ways of reducing AI inference costs.
Yes, we plan to be releasing a tech report soon. We are not open sourcing the models at launch time, but we have a roadmap of future releases in which we hope to make some of our models accessible to the research community.
That's a good point. In this context, we've been using "commodity GPUs" to refer to standard Nvidia hardware, in contrast to specialized chips like Groq and Cerebras. While these chips also achieve fast speeds, they are not nearly as ubiquitous as Nvidia GPUs. We think that matching their performance on standard Nvidia hardware can make AI much more affordable. We also support any GPUs, not just H100's.
We're going to be releasing a tech report soon, stay tuned!
Good question! We are not open sourcing the models at launch time, but we have a roadmap of future releases in which we hope to make some of our models accessible to the research community.
The short answer is that we do more than one parallel pass over multiple tokens: we iteratively refine them over a few passes to fix incoherences. This can be seen as a generalization of diffusion algorithms that underlie systems like Midjourney or Sora.
Not today, but we will be following up with a technical report over the next week or so. In the meantime, you can take a look at some of the research papers that inspired our work: - https://arxiv.org/abs/2310.16834 - https://arxiv.org/abs/2406.07524
This is Volodymyr, co-founder at Inception---let us know if you have any questions about diffusion, language modeling, and our new Mercury models!
It won't run as fast on your CPU at it will run on a GPU. Also, it might clog most of your RAM; it's better to offload to a cheap GPU.
There are also equally easy to use open source systems, check out this one for example: https://github.com/kuleshov/minillm
Try this: it installs with a simple python command if you have an NVIDIA GPU: https://github.com/kuleshov/minillm
Afresh | Design, Product, Machine Learning, Backend | San Francisco, CA | Onsite | Visa | Full-Time
Afresh is a Series A startup focused on automating the food supply chain using AI with the ultimate goal of eliminating food waste. In the US, about 40% of all food waste occurs in supermarkets and downstream, largely due to inefficient manual ordering processes. This waste leads to >$80B in economic losses as well as 1.5 billion tons of greenhouse gas emissions, which is comparable to the emissions of Japan.
Afresh is commercializing a technology developed as part of a Stanford research project that automates the pen-and-paper processes used by supermarket operators. This technology cuts retail food waste by >50% and dramatically increases the stores' profit margins.
We are founded by a team of Computer Science PhDs, MBAs, designers, and engineers from Stanford, Berkeley, CMU. We're backed by former Google CEO Eric Schmidt's firm (Innovation Endeavors) and the first investors in Instagram, Stitchfix, SoFi, and Heroku (Steve Anderson of Baseline Ventures).
We're growing fast: we're in a partnership with 4 large regional grocers representing 500+ stores and >$10B in revenue. We're also looking for smart, enthusiastic, dependable people interested in applying cutting-edge technology to problems with significant societal impact.
Our open roles are:
- Lead UI/UX Designer - Lead Product Manager - Machine Learning Engineer - Senior Backend Engineer - Full job descriptions available at: https://jobs.lever.co/afreshtechnologies?
Website: http://afresh.ai
Feel free to reach out directly to volodymyr@afreshtechnologies.com (I'm the CTO)
Afresh | San Francisco, CA | Full-time | Onsite
Afresh is a Series A startup focused on automating the food supply chain using AI with the ultimate goal of eliminating food waste. In the US, about 40% of all food waste occurs in supermarkets and downstream, largely due to inefficient manual ordering processes. This waste leads to >$80B in economic losses as well as 1.5 billion tons of greenhouse gas emissions, which is comparable to the emissions of Japan.
Afresh is commercializing a technology developed as part of a Stanford research project that automates the pen-and-paper processes used by supermarket operators. This technology cuts retail food waste by >50% and dramatically increases the stores' profit margins.
We are founded by a team of Computer Science PhDs, MBAs, designers, and engineers from Stanford, Berkeley, CMU. We're backed by former Google CEO Eric Schmidt's firm (Innovation Endeavors) and the first investors in Instagram, Stitchfix, SoFi, and Heroku (Steve Anderson of Baseline Ventures).
We're growing fast: we're in a partnership with 4 large regional grocers representing 500+ stores and >$10B in revenue. We're also looking for smart, enthusiastic, dependable people interested in applying cutting-edge technology to problems with significant societal impact.
Our open roles are: * Machine Learning Engineer * Backend Engineer * Site Reliability / DevOps * Mobile Developer * Full-Stack Web Developer Full job descriptions available at: https://jobs.lever.co/afreshtechnologies?
Feel free to reach out directly to volodymyr@afreshtechnologies.com (I'm the CTO)
That's an interesting discussion! Having read the Vovk papers, this blog post definitely presents things much more clearly. The original papers often don't adhere to the standard definition/lemma/proof style of mathematical exposition, which makes them really hard to follow.
It's also an interesting coincidence that this story on the front page today. I'm giving a talk tomorrow at AAAI on some work that extends this theory. We show how to do uncertainty estimation (e.g. calibrated probabilities for ML classifiers) under fully adversarial assumptions (input data can be chosen by an adversary). I'll do a shameless plug and post the paper here, in case people are interested in this general topic:
Here is a paper that presents a graph version of the BWT:
http://bioinformatics.oxfordjournals.org/content/29/13/i361....