HN user

ozgune

1,855 karma

Ubicloud, Microsoft, Citus Data, Amazon

Posts49
Comments188
View on HN
www.promptarmor.com 4mo ago

Snowflake AI Escapes Sandbox and Executes Malware

ozgune
269pts82
defmacro.org 4mo ago

How to build great products (2013)

ozgune
1pts1
www.theregister.com 7mo ago

GitHub walks back plan to charge for self-hosted runners

ozgune
9pts3
z.ai 7mo ago

GLM-4.6V: Open-Source Multimodal Models with Native Tool Use

ozgune
4pts0
www.tomshardware.com 9mo ago

Bride surprises new husband with an RTX 5090 on wedding day

ozgune
5pts1
rubydramas.com 10mo ago

RubyDramas – Your guide to the last Ruby drama

ozgune
6pts0
yiyan.baidu.com 1y ago

The Open Source Release of the Ernie 4.5 Model Family

ozgune
4pts0
www.cnn.com 1y ago

Trump to change controversial Biden-era restrictions on AI chip exports

ozgune
1pts0
edition.cnn.com 1y ago

California overtakes Japan to become the world's fourth largest economy

ozgune
90pts108
www.cnn.com 1y ago

Economists say there's a math error in Trump's tariff formula [video]

ozgune
9pts2
blog.vllm.ai 1y ago

vLLM V1: A Major Upgrade to vLLM's Core Architecture

ozgune
2pts0
opensource.microsoft.com 1y ago

DocumentDB: Open-source MongoDB implementation based on PostgreSQL

ozgune
8pts1
news.ycombinator.com 1y ago

Show HN: QwQ-32B APIs – o1 like reasoning at 1% the cost

ozgune
17pts3
www.tomshardware.com 1y ago

Microsoft Azure flaunts first custom Nvidia Blackwell racks

ozgune
2pts2
techcrunch.com 1y ago

AMD to acquire infrastructure player ZT Systems for $4.9B to amp up AI ecosystem

ozgune
3pts0
artificialanalysis.ai 1y ago

Comparison of AI Models: Quality, Performance and Price Analysis

ozgune
2pts0
www.snowflake.com 1y ago

Achieve Low-Latency and High-Throughput Inference with Meta's Llama 3.1 405B

ozgune
1pts0
www.sequoiacap.com 2y ago

AI Is Now Shovel Ready

ozgune
1pts0
www.ubicloud.com 2y ago

How we enabled ARM64 VMs

ozgune
90pts17
www.lennysnewsletter.com 2y ago

How Perplexity Builds Product

ozgune
2pts0
contextual.ai 2y ago

RAG 2.0

ozgune
15pts4
www.cnn.com 2y ago

Amazon Web Services CEO to step down

ozgune
3pts0
old.reddit.com 2y ago

Hetzner Brings Back GPU Servers

ozgune
8pts1
tembo.io 2y ago

One-Click RESTful APIs in Postgres

ozgune
3pts0
www.ubicloud.com 2y ago

Authorization (ABAC) implementation in 130 lines of code

ozgune
2pts0
news.ycombinator.com 2y ago

Ask HN: Thoughts on Elastic V2, SSPL, or mixed software licenses?

ozgune
6pts3
www.nytimes.com 3y ago

You Can Have the Blue Pill or the Red Pill, and We’re Out of Blue Pills

ozgune
3pts1
dl.acm.org 4y ago

Citus: Distributed PostgreSQL for Data-Intensive Applications (SIGMOD '21)

ozgune
16pts0
dsf.berkeley.edu 4y ago

The Case for Shared Nothing (1985) [pdf]

ozgune
4pts0
techcommunity.microsoft.com 5y ago

Analyzing the Limits of Connection Scalability in Postgres

ozgune
112pts19

(Ozgun from Ubicloud)

I agree with the blog post's technical contents, but I feel we came across too strong in the title. For Ubicloud as a managed Postgres provider, we use strict memory overcommit. Our experience with operating Postgres at scale taught us that it's better to enable this than going with the defaults.

However, I can see many other scenarios, where using strict memory overcommit would have unanticipated side-effects. That's why Linux doesn't go with strict memory commit as its default.

DeepSeek v4 3 months ago

Ack. I took the benchmark results that AI Labs themselves published for their models. So the Opus 4.6 baseline would be from the time that Anthropic released the model.

DeepSeek v4 3 months ago

I reviewed how DeepSeek V4-Pro, Kimi 2.6, Opus 4.6, and Opus 4.7 across the same AI benchmarks. All results are for Max editions, except for Kimi.

Summary: Opus 4.6 forms the baseline all three are trying to beat. DeepSeek V4-Pro roughly matches it across the board, Kimi K2.6 edges it on agentic/coding benchmarks, and Opus 4.7 surpasses it on nearly everything except web search.

DeepSeek V4-Pro Max shines in competitive coding benchmarks. However, it trails both Opus models on software engineering. Kimi K2.6 is remarkably competitive as an open-weight model. Its main weakness is in pure reasoning (GPQA, HMMT) where it trails Opus.

Speculation: The DeepSeek team wanted to come out with a model that surpassed proprietary ones. However, OpenAI dropped 5.4 and 5.5 and Anthropic released Opus 4.6 and 4.7. So they chose to just release V4 and iterate on it.

Basis for speculation? (i) The original reported timeline for the model was February. (ii) Their Hugging Face model card starts with "We present a preview version of DeepSeek-V4 series". (iii) V4 isn't multimodal yet (unlike the others) and their technical report states "We are also working on incorporating multimodal capabilities to our models."

This update makes Kimi K2.6 the strongest open multimodal AI model. (No affiliation with Kimi.)

Here's the aggregated AI benchmark comparison for K2.6 vs Opus 4.6 (max effort).

- Agentic: Kimi wins 5. Opus wins 5.

- Coding: Kimi wins 5. Opus wins 1.

- Reasoning & knowledge: Kimi wins 1. Opus wins 4.

- Vision: Kimi wins 9. Opus wins 0.

Please note that the model publisher chooses their benchmarks, so there's a bias here. Most coding and reasoning & knowledge benchmarks in their list are pretty standard though.

I'll save everyone a web search. This is satire and there isn't any such German federal court ruling.

It also speaks to the world that we live in these days - I'm having a hard time separating satire from reality.

These changes are effective April 1st for existing and new customers. The price increase ratios are also different across product lines.

* Cloud (VMs): 38%

* Bare metal: 15%

* Memory add-on for bare metal: 575% (effective immediately)

It feels like memory add-on is intentionally set high to discourage customers from adding more memory.

AX102 (128 GB RAM) costs €124, AX162 (256 GB RAM) costs €244, but the 128 GB memory add-on alone costs €264. If we ignore the setup fee, it’s more cost-effective to provision additional servers instead of adding RAM to bare metal instances.

Here's the link to cloud and bare metal pricing changes: https://docs.hetzner.com/general/infrastructure-and-availabi...

I feel this analysis is unfair to PostgreSQL. PG is highly extensible, allowing you to extend write-ahead logs, transaction subsystem, foreign data wrappers (FDW), indexes, types, replication, others.

I understand that MySQL follows a specific pluggable storage architecture. I also understand that the direct equivalent in PG appears to be table access methods (TAM). However, you don't need to use TAM to build this - I'd argue FDWs are much more suitable.

Also, I think this design assumes that you'd swap PG's storage engine and replicate data to DuckDB through logical replication. The explanation then notes deficiencies in PG's logical replication.

I don't think this is the only possible design. pg_lake provides a solid open source implementation on how else you could build this solution, if you're familiar with PG: https://github.com/Snowflake-Labs/pg_lake

All up, I feel this explanation is written from a MySQL-first perspective. "We built this valuable solution for MySQL. We're very familiar with MySQL's internals and we don't think those internals hold for PostgreSQL."

I agree with the solution's value and how it integrates with MySQL. I just think someone knowledgeable about PostgreSQL would have built things in a different way.

Mistral OCR 3 7 months ago

Also, do you know if their benchmarks are available?

In their website, the benchmarks say “Multilingual (Chinese), Multilingual (East-asian), Multilingual (Eastern europe), Multilingual (English), Multilingual (Western europe), Forms, Handwritten, etc.” However, there’s no reference to the benchmark data.

This is huge!

When people ask me what’s missing in the Postgres market, I used to tell them “open source Snowflake.”

Crunchy’s Postgres extension is by far the most ahead solution in the market.

Huge congrats to Snowflake and the Crunchy team on open sourcing this.

If the benchmark doesn’t use AIO, why the performance difference between PG 17 and 18 in the blog post (sync, worker, and io_uring)?

Is it because remote storage in the cloud always introduces some variance & the benchmark just picks that up?

For reference, anarazel had a presentation at pgconf.eu yesterday about AIO. anarazel mentioned that remote cloud storage always introduced variance making the benchmark results hard to interpret. His solution was to introduce synthetic latency on local NVMes for benchmarks.

DeepSeek OCR 9 months ago

OmniAI has a benchmark that companies LLMs to cloud OCR services.

https://getomni.ai/blog/ocr-benchmark (Feb 2025)

Please note that LLMs progressed at a rapid pace since Feb. We see much better results with the Qwen3-VL family, particularly Qwen3-VL-235B-A22B-Instruct for our use-case.

(Disclaimer: Ozgun from Ubicloud)

I agree with you. I feel the challenge is that using AI coding tools is still an art, and not a science. That's why we see many qualitative studies that sometimes conflict with each other.

In this case, we found the following interesting. That's why we nudged Shikhar to blog about his experience and put a disclaimer at the top.

* Our codebase is in Ruby and follows a design pattern uncommon industry * We don't have a horse in this game * I haven't seen an evaluation that evaluates coding tools in (a) coding, (b) testing, and (c) debugging dimension

The SGLang Team has a follow-up blog post that talks about DeepSeek inference performance on GB200 NVL72: https://lmsys.org/blog/2025-06-16-gb200-part-1/

Just in case you have $3-4M lying around somewhere for some high quality inference. :)

SGLang quotes a 2.5-3.4x speedup as compared to the H100s. They also note that more optimizations are coming, but they haven't yet published a part 2 on the blog post.

I agree that you could get to high margins, but I think the modeling holds only if you're an AI lab operating at scale with a setup tuned for your model(s). I think the most open study on this one is from the DeepSeek team: https://github.com/deepseek-ai/open-infra-index/blob/main/20...

For others, I think the picture is different. When we ran benchmarks on DeepSeek-R1 on 8x H200 SXM using vLLM, we got up to 12K total tok/s (concurrency 200, input:output ratio of 6:1). If you're spiking up 100-200K tok/s, you need a lot of GPUs for that. Then, the GPUs sit idle most of the time.

I'll read the blog post in more detail, but I don't think the following assumptions hold outside of AI labs.

* 100% utilization (no spikes, balanced usage between day/night or weekdays) * Input processing is free (~$0.001 per million tokens) * DeepSeek fits into H100 cards in a way that network isn't the bottleneck

Their benchmarks are interesting. They are comparing to DeepSeek-V3's (non-reasoning) December and DeepSeek-R1's January releases. I feel that comparing to DeepSeek-R1-0528 would be more fair.

For example, R1 scores 79.8 on AIME 2024, R1-0528 performs 91.4.

R1 scores 70 on AIME 2025, R1-0528 scores 87.5. R1-0528 does similarly better for GPQA Diamond, LiveCodeBench, and Aider (about 10-15 points higher).

https://huggingface.co/deepseek-ai/DeepSeek-R1-0528

Question to author.

Are you planning to publish CH benchmarks (TPC-C and TPC-H combined)? I'd expect Aurora to perform much worse on CH than on TPC-C/H. That's because Aurora pushes the WAL logs to replicated shared storage. Since you only need quorum on a write, you get a fast ack on the write (TPC-C). The way you've run TPC-H doesn't modify the data that much, so you also get baseline Postgres performance.

However, when you're pushing writes and you have a sequential scan over the data, then Aurora needs to reconcile the WAL writes, manage locks, etc. CH benchmark exercises that path and I'd expect it to notably slow down Aurora.

(Disclaimer: ex-Citus and current Ubicloud founder)

However, a wise man once said: “[It] ain’t about how hard you hit. It’s about how hard you can get hit and keep moving forward; how much you can take and keep moving forward.” Ujiharu may have lost Oda Castle nine times, but that means he also won it back eight times, almost always with smaller armies. His refusal to accept defeat and his iron will to get up and keep fighting is why many historians reject the “weakest samurai warlord” nickname and instead refer to him as “The Phoenix.”

Love this paragraph from the article.

I had a related, but orthogonal question about multilingual LLMs.

When I ask smaller models a question in English, the model does well. When I ask the same model a question in Turkish, the answer is mediocre. When I ask the model to translate my question into English, get the answer, and translate the answer back to Turkish, the model again does well.

For example, I tried the above with Llama 3.3 70B, and asked it to plan me a 3-day trip to Istanbul. When I asked Llama to do the translations between English <> Turkish, the answer was notably better.

Anyone else observed a similar behavior?

In March, vLLM picked up some of the improvements in the DeepSeek paper. Through these, vLLM v0.7.3's DeepSeek performance jumped to about 3x+ of what it was before [1].

What's exciting is that there's still so much room for improvement. We benchmark around 5K total tokens/s with the sharegpt dataset and 12K total token/s with random 2000/100, using vLLM and under high concurrency.

DeepSeek-V3/R1 Inference System Overview [2] quotes "Each H800 node delivers an average throughput of 73.7k tokens/s input (including cache hits) during prefilling or 14.8k tokens/s output during decoding."

Yes, DeepSeek deploys a different inference architecture. But this goes onto show just how much room there is for improvement. Looking forward to more open source!

[1] https://developers.redhat.com/articles/2025/03/19/how-we-opt...

[2] https://github.com/deepseek-ai/open-infra-index/blob/main/20...

I feel the article presents the data selectively in some places. Two examples:

* The article compares Gemini 2.5 Pro Experimental to DeepSeek-R1 in accuracy benchmarks. Then, when the comparison becomes about cost, it compares Gemini 2.0 Flash to DeepSeek-R1.

* In throughput numbers, DeepSeek-R1 is quoted at 24 tok/s. There are half a dozen providers, who give you easily 100+ tok/s and at scale.

There's no doubt that Gemini 2.5 Pro Experimental is a state of the art model. I just think it's very hard to win on every AI front these days.

I agree with the blog post that using K8s + containers for GPU virtualization is a security disaster waiting to happen. Even if you configure your container right (which is extremely hard to do), you don't get seccomp-bpf.

People started using K8s for training, where you already had a network isolated cluster. Extending the K8s+container pattern to multi-tenant environments is scary at best.

I didn't understand the following part though.

Instead, we burned months trying (and ultimately failing) to get Nvidia’s host drivers working to map virtualized GPUs into Intel Cloud Hypervisor.

Why was this part so hard? Doing PCI passthrough with the Cloud Hypervisor (CH) is relatively common. Was it the transition from Firecracker to CH that was tricky?

Agreed. Here are three things that I find surreal about the s1 paper.

(1) The abstract changed how I thought about this domain (advanced reasoning models). The only other paper that did that for me was the "Memory Resource Management in VMware ESX Server". And that paper got published 23 years ago.

(2) The model, data, and code are open source at https://github.com/simplescaling/s1. With this, you can start training your own advanced reasoning models. All you need is a thousand well-curated questions with reasoning steps.

(3) More than half the references in the paper are from 2024 and Jan 2025. Just look at the paper's first page. https://arxiv.org/pdf/2501.19393 In which other field do you see this?

On our side, we first saw the Cloudflare outage. Then, Docker Hub started failing, followed by GitHub API errors.

It's amazing how much of the internet runs / depends on Cloudflare these days. Thank you for keeping the lights on. :)