HN user

volodia

545 karma
Posts27
Comments61
View on HN
blog.kilo.ai 23d ago

Next-Edit in Kilo, Powered by Inception Diffusion LLMs

volodia
2pts0
www.inceptionlabs.ai 3mo ago

Mercury 2 on PinchBench: Diffusion LLM benchmarked on real OpenClaw agent tasks

volodia
2pts0
twitter.com 4mo ago

Mercury 2: Best-in-class speed-optimized intelligence at 1,200 tok/SEC

volodia
1pts0
arxiv.org 2y ago

Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers

volodia
2pts0
github.com 3y ago

LLMTune: 4-Bit finetuning of 65B LLAMA models on a single consumer GPU

volodia
3pts0
twitter.com 3y ago

LLMTune: 4-Bit Finetuning of LLMs on a Consumer GPI

volodia
2pts0
github.com 3y ago

Don't have a $5k MacBook to run LLAMA65B? MiniLLM runs LLMs on GPUs in <500 LOC

volodia
3pts2
twitter.com 3y ago

MiniLLM: A minimal system for running LLMs on consumer-grade Nvidia GPUs

volodia
31pts2
www.youtube.com 5y ago

Lecture Videos for Applied Machine Learning (Cornell Tech CS 5787, Fall 2020)

volodia
2pts0
kuleshov.github.io 8y ago

Audio Super Resolution with Neural Networks

volodia
3pts1
ermongroup.github.io 9y ago

Stanford Lecture Notes on Probabilistic Graphical Models

volodia
336pts32
www.prospectmagazine.co.uk 9y ago

America: The Failed State

volodia
2pts0
www.bloomberg.com 9y ago

A Moneymaking Machine Like Few Others

volodia
258pts205
www.thestar.com 9y ago

The North Pole is 20°C warmer than normal as winter descends

volodia
18pts2
thenextweb.com 10y ago

Prisma: neural-style instagram filters

volodia
28pts8
www.wsj.com 10y ago

Donald Trump's America

volodia
5pts1
www.quora.com 11y ago

Are we in a tech bubble?

volodia
1pts0
www.technologyreview.com 11y ago

Robot Journalist Finds New Work on Wall Street

volodia
21pts0
www.youtube.com 12y ago

Stephen Wolfram's Introduction to the Wolfram Language

volodia
2pts0
www.latimes.com 12y ago

Google plans move into San Francisco's Mission District

volodia
1pts0
euromaidan-pulse.herokuapp.com 12y ago

Show HN: Smart news aggregator to follow the events in Ukraine

volodia
1pts0
arstechnica.com 13y ago

Single click starts a 10,000-core CycleCloud cluster for $1060/hr

volodia
4pts0
www.nytimes.com 13y ago

New Cornell Technology School Tightly Bound to Business

volodia
1pts0
arstechnica.com 13y ago

Apple licensed design patents to Microsoft in "anti-cloning agreement"

volodia
1pts0
hbfs.wordpress.com 13y ago

Is Python slow?

volodia
1pts0
news.ycombinator.com 14y ago

Ask HN: Startup opportunities in bioinformatics

volodia
12pts5
scottaaronson.com 17y ago

Mathematician's Lament: An essay on math education and on how we view math

volodia
44pts18

We agree! In fact, there is an emerging class of models aimed at fast agentic iteration (think of Composer, the Flash versions of proprietary and open models). We position Mercury 2 as a strong model in this category.

That is also our view! We see Mercury 2 as enabling very fast iteration for agentic tasks. A single shot at a problem might be less accurate, but because the model has a shorter execution time, it enables users to iterate much more quickly.

You can think of Mercury 2 as roughly in the same intelligence tier as other speed-optimized models (e.g., Haiku 4.5, Grok Fast, GPT-Mini–class systems). The main differentiator is latency — it’s ~5× faster at comparable quality.

We’re not positioning it as competing with the largest models (Opus 4.5, etc.) on hardest-case reasoning. It’s more of a “fast agent” model (like Composer in Cursor, or Haiku 4.5 in some IDEs): strong on common coding and tool-use tasks, and providing very quick iteration loops.

I’d push back a bit on the Pareto point.

On speed/quality, diffusion has actually moved the frontier. At comparable quality levels, Mercury is >5× faster than similar AR models (including the ones referenced on the AA page). So for a fixed quality target, you can get meaningfully higher throughput.

That said, I agree diffusion models today don’t yet match the very largest AR systems (Opus, Gemini Pro, etc.) on absolute intelligence. That’s not surprising: we’re starting from smaller models and gradually scaling up. The roadmap is to scale intelligence while preserving the large inference-time advantage.

Great question! The model can more efficiently leverage existing GPU hardware---it performs more computation per unit of memory transferred; this means that on older hardware one should be able to get similar inference speeds as one would get on recent hardware with a classical LLM. This is actually interesting commercially, since it opens new ways of reducing AI inference costs.

That's a good point. In this context, we've been using "commodity GPUs" to refer to standard Nvidia hardware, in contrast to specialized chips like Groq and Cerebras. While these chips also achieve fast speeds, they are not nearly as ubiquitous as Nvidia GPUs. We think that matching their performance on standard Nvidia hardware can make AI much more affordable. We also support any GPUs, not just H100's.

We're going to be releasing a tech report soon, stay tuned!

Afresh | Design, Product, Machine Learning, Backend | San Francisco, CA | Onsite | Visa | Full-Time

Afresh is a Series A startup focused on automating the food supply chain using AI with the ultimate goal of eliminating food waste. In the US, about 40% of all food waste occurs in supermarkets and downstream, largely due to inefficient manual ordering processes. This waste leads to >$80B in economic losses as well as 1.5 billion tons of greenhouse gas emissions, which is comparable to the emissions of Japan.

Afresh is commercializing a technology developed as part of a Stanford research project that automates the pen-and-paper processes used by supermarket operators. This technology cuts retail food waste by >50% and dramatically increases the stores' profit margins.

We are founded by a team of Computer Science PhDs, MBAs, designers, and engineers from Stanford, Berkeley, CMU. We're backed by former Google CEO Eric Schmidt's firm (Innovation Endeavors) and the first investors in Instagram, Stitchfix, SoFi, and Heroku (Steve Anderson of Baseline Ventures).

We're growing fast: we're in a partnership with 4 large regional grocers representing 500+ stores and >$10B in revenue. We're also looking for smart, enthusiastic, dependable people interested in applying cutting-edge technology to problems with significant societal impact.

Our open roles are:

- Lead UI/UX Designer - Lead Product Manager - Machine Learning Engineer - Senior Backend Engineer - Full job descriptions available at: https://jobs.lever.co/afreshtechnologies?

Website: http://afresh.ai

Feel free to reach out directly to volodymyr@afreshtechnologies.com (I'm the CTO)

Afresh | San Francisco, CA | Full-time | Onsite

Afresh is a Series A startup focused on automating the food supply chain using AI with the ultimate goal of eliminating food waste. In the US, about 40% of all food waste occurs in supermarkets and downstream, largely due to inefficient manual ordering processes. This waste leads to >$80B in economic losses as well as 1.5 billion tons of greenhouse gas emissions, which is comparable to the emissions of Japan.

Afresh is commercializing a technology developed as part of a Stanford research project that automates the pen-and-paper processes used by supermarket operators. This technology cuts retail food waste by >50% and dramatically increases the stores' profit margins.

We are founded by a team of Computer Science PhDs, MBAs, designers, and engineers from Stanford, Berkeley, CMU. We're backed by former Google CEO Eric Schmidt's firm (Innovation Endeavors) and the first investors in Instagram, Stitchfix, SoFi, and Heroku (Steve Anderson of Baseline Ventures).

We're growing fast: we're in a partnership with 4 large regional grocers representing 500+ stores and >$10B in revenue. We're also looking for smart, enthusiastic, dependable people interested in applying cutting-edge technology to problems with significant societal impact.

Our open roles are: * Machine Learning Engineer * Backend Engineer * Site Reliability / DevOps * Mobile Developer * Full-Stack Web Developer Full job descriptions available at: https://jobs.lever.co/afreshtechnologies?

Feel free to reach out directly to volodymyr@afreshtechnologies.com (I'm the CTO)

That's an interesting discussion! Having read the Vovk papers, this blog post definitely presents things much more clearly. The original papers often don't adhere to the standard definition/lemma/proof style of mathematical exposition, which makes them really hard to follow.

It's also an interesting coincidence that this story on the front page today. I'm giving a talk tomorrow at AAAI on some work that extends this theory. We show how to do uncertainty estimation (e.g. calibrated probabilities for ML classifiers) under fully adversarial assumptions (input data can be chosen by an adversary). I'll do a shameless plug and post the paper here, in case people are interested in this general topic:

https://arxiv.org/abs/1607.03594