HN user

PeterCorless

1,590 karma

Principal Product Marketing Manager Redpanda LinkedIn: https://www.linkedin.com/in/petercorless/ Email: peter.corless@redpanda.com Bluesky: @petercorless.bsky.social Twitter: @PeterCorless

Posts168
Comments612
View on HN
arxiv.org 1mo ago

Scalable Intra-Process Data Redistribution W Ring-Buffer Shuffle (Redpanda Oxla)

PeterCorless
2pts0
arxiv.org 1mo ago

The Importance of Out-of-Band Metadata for Safe Autonomous Agents [Redpanda]

PeterCorless
3pts0
medium.com 3mo ago

A Peer-Vetted AI Stack for Builders

PeterCorless
8pts8
www.redpanda.com 4mo ago

Redpanda pushes the envelope on Nvidia Vera

PeterCorless
2pts0
www.redpanda.com 5mo ago

Redpanda Agentic Data Plane (ADP) now in limited availability

PeterCorless
1pts0
www.redpanda.com 6mo ago

The convergence of AI and data streaming – Part 1: The coming brick walls

PeterCorless
2pts0
news.ycombinator.com 8mo ago

KubeCon attendees: check your flights

PeterCorless
1pts0
www.space.com 8mo ago

Where to find comets Lemmon, SWAN and 3I/ATLAS in the Halloween sky

PeterCorless
2pts0
www.redpanda.com 8mo ago

The Agentic Data Plane (Redpanda Data)

PeterCorless
6pts3
startree.ai 9mo ago

Data freshness (end-to-end latency) in ClickHouse and Apache Pinot

PeterCorless
6pts0
www.redpanda.com 9mo ago

Encrypting vector embeddings prior to data ingestion (Redpanda, Cyborg)

PeterCorless
1pts0
ashishjayamohan.github.io 1y ago

Systems Optimizations [Apache Pinot]

PeterCorless
2pts1
teamraft.com 1y ago

Secure and Low-Latency Queries at Scale with Raft Data Platform – [R]DP

PeterCorless
4pts2
www.pinot-connect.org 1y ago

Pinot-Connect (DB-API 2.0 for Querying Apache Pinot with Python)

PeterCorless
1pts0
substack.com 1y ago

How to Connect Your OLTP DB to Apache Pinot for Realtime Analytics

PeterCorless
1pts0
startree.ai 1y ago

Disaggregating Observability with Apache Kafka, StarTree Cloud, and Grafana

PeterCorless
1pts0
www.pcgamer.com 1y ago

Alan Emrich, the game designer and writer who coined the term '4X,' has died

PeterCorless
4pts1
startree.ai 1y ago

Apache Pinot Year in Review 2024

PeterCorless
2pts0
www.xenonstack.com 1y ago

Real-Time Observability with Apache Pinot for Failures

PeterCorless
1pts0
www.uber.com 1y ago

Serving Apache Pinot Queries with Neutrino

PeterCorless
1pts0
phys.org 1y ago

Astronomers deal a blow to theory Venus once had liquid water on its surface

PeterCorless
2pts0
www.collectspace.com 1y ago

Artemis moon suit designed by Axiom Space and Prada revealed in Milan

PeterCorless
49pts69
github.com 1y ago

Holocron: An object storage based leader election library

PeterCorless
2pts0
thenewstack.io 1y ago

Reimagining Observability: The Case for a Disaggregated Stack

PeterCorless
1pts0
www.uber.com 1y ago

Pinot for Low-Latency Offline Table Analytics

PeterCorless
27pts2
gigaom.com 1y ago

GigaOm Sonar Report for Real-Time Analytical Databases

PeterCorless
1pts0
www.macrumors.com 1y ago

Gurman: M4 MacBook Pro, Mac Mini, and iMac Coming This Year

PeterCorless
1pts4
www.space.com 1y ago

Happy 250th Anniversary, Oxygen

PeterCorless
1pts0
www.space.com 1y ago

Tungsten found in Tycho Brahe's lab; scientists not sure how it got there

PeterCorless
3pts0
phys.org 2y ago

Another intermediate-mass black hole discovery at the center of our galaxy

PeterCorless
2pts0

Agents are cardinality-hungry. They want the high-cardinality data you'd normally drop: individual trace IDs, per-request attributes, full tag sets. They are very patient. They will sift through it.

The agents themselves are not likely going to be doing the high cardinality queries or they will keel over. They have limited memory buffers. They will take many seconds to return results. They are likely going to be limited in terms of QPS.

From the blog: > Apache Iceberg, with data stored as Parquet on S3, and most of the system implemented in Go

You have just ensured that queries will have a p99 >1 second. This is kind of antithetical to having an agent be fast.

You couldn't run any sort of real-time service, where hundreds of thousands to millions of events were occurring per second, and you needed to adjust to that in milliseconds.

The terms "p99" and "QPS" do not occur anywhere in the article. Which leaves the question of scalability to a user's imagination.

I applaud the direction. I am looking for objective evidence.

Vera does what NVIDIA calls Spatial Multithreading, "physically partitioning each core’s resources rather than time slicing them, allowing the system to optimize for performance or density at runtime." A kind of static hyperthreading; you get two threads per core.

It's somewhat different from how x86 chips do simultaneous multithreading (SMT),

Fair. A good callout. And maybe the right move. However, a healthy IBM would not have needed to calve off its entire Global Technology Services business.

So much that we presume in the modern cloud wasn't a given when Apache Kafka was first released in 2011.

kevstev wrote just above about Kafka being written to run on spinning disks (HDDs), while Redpanda was written to take advantage of the latest hardware (local NVMe SSDs). He has some great insights.

As well, Apache Kafka was written in Java, back in an era when you were weren't quite sure what operating system you might be running on. For example, when Azure first launched they had a Windows NT-based system called Windows Azure. Most everyone else had already decided to roll Linux. Microsoft refused to budge on Linux until 2014, and didn't release its own Azure Linux until 2020.

Once everyone decided to roll Linux, the "write once run everywhere" promise of Java was obviated. But because you were still locked into a Java Virtual Machine (JVM) your application couldn't optimize itself to the underlying hardware and operating system you were running on.

Redpanda, for example, is written in C++ on top of the Seastar framework (seastar.io). The same framework at the heart of ScyllaDB. This engine is a thread-per-core shared-nothing architecture that allows Redpanda to optimize performance for hardware utilization in ways that a Java app can only dream of. CPU utilization, memory usage, IO throughput. It's all just better performance on Redpanda.

It means that you're actually getting better utility out of the servers you deploy. Less wasted / fallow CPU cycles — so better price-performance. Faster writes. Lower p99 latencies. It's just... better.

Now, I am biased. I work at Redpanda now. But I've been a big fan of Kafka since 2015. I am still bullish on data streaming. I just think that Apache Kafka, as a Java-based platform, needs some serious rearchitecture,

Even Confluent doesn't use vanilla Kafka. They rewrote their own engine, Kora. They claim it is 10x faster. Or 30x faster. Depending on what you're measuring.

1. https://www.confluent.io/confluent-cloud/kora/

2. https://www.confluent.io/blog/10x-apache-kafka-elasticity/

I have thought quite a bit today about the news from Confluent and IBM. I have friends and colleagues at both companies. When I was an undergrad at Carnegie Mellon University in the 1980s I used to wear a big brown and tan IBM button that said "THINK."

And here is a picture of Ben Lorica 罗瑞卡 interviewing Jay Kreps and other industry leaders at The Hive back on the evening of 25 February 2015. I believe they were talking about strategies for implementing Lambda Architecture.

All of which is to say: I have been a big fan of both companies for a long, long time. While today I am at employed at Redpanda Data, a direct competitor of Confluent, I hope to set aside any "team"-based bias to provide a sober and honest appraisal.

First, IBM has been shrinking. They were at 345,000 employees as of their 2020 Annual Report. But the COVID-19 pandemic was only one of many setbacks the company faced when Arvind Krishna took the helm as CEO. By December 2024 the employee base shrank to 270,000 — a drop of nearly 22%.

IBM revenue in 2020: $73.6B.

IBM revenue in 2024: $62.75B — a less-precipitous drop of 15%.

Revenue per employee over that period rose from $213k to $232k.

Confluent on its own? $400k.

And to compare: Amazon earns $580k per employee. Microsoft generates over $1M per. Nvidia? $4M-$5M.

And now, in November, they announced thousands of more layoffs. No one seems safe, regardless of job title. Those cut include positions in "artificial intelligence, marketing, software engineering and cloud technology."

Next, IBM has had a mixed record as a steward of acquisitions. Red Hat has doubled in revenues since their 2019 acquisition. For a while its headcount continued to grow, as much as 19,000 by 2023. But then it was forced into layoffs by parent IBM in April of that year, and then each year since, even while it remains one of the highest margin businesses in their portfolio.

SoftLayer — "IBM Cloud Classic" — also suffered significant layoffs in early 2025, with offshoring sending jobs to India.

DataStax had layoffs in 2023-2024, even before its acquisition was announced. Maybe they were "trimming the fat" to get into a shape to be acquired.

As a person with a long career in marketing, I know that many of the first roles to be jettisoned at a newly-acquired company tend to be in go-to-market organizations. Sales, Marketing, Developer Relations, Documentation, Training, Community, Customer Service. These tend to be seen as "nice to haves" by upper management. But their loss guts organizations and hollows out user-facing teams and open source communities.

My hope is that Confluent is spared as much of the pain and turmoil as possible. That, like Red Hat, it is run autonomously as much as possible.

[Crossposted from LinkedIn here, where you can see the photo mentioned: https://www.linkedin.com/feed/update/urn:li:activity:7404052...]

Jepsen: NATS 2.12.1 8 months ago

You can have DeepWiki literally scan the source code and tell you:

2. Delayed Sync Mode (Default)

In the default mode, writes are batched and marked with needSync = true for later synchronization filestore.go:7093-7097 . The actual sync happens during the next syncBlocks() execution.

However, if you read DeepWiki's conclusion, it is far more optimistic than what Aphyr uncovered in real-world testing.

Durability Guarantees

Even with delayed fsyncs, NATS provides protection against data loss through:

1. Write-Ahead Logging: Messages are written to log files before being acknowledged

2. Periodic Sync: The sync timer ensures data is eventually flushed to disk

3. State Snapshots: Full state is periodically written to index.db files filestore.go:9834-9850

4. Error Handling: If sync operations fail, NATS attempts to rebuild state from existing data filestore.go:7066-7072"

https://deepwiki.com/search/will-nats-lose-uncommitted-wri_b...

French is actually <30%. There is a well-sourced Wikipedia article about this.

• French (including Old French: 11.66%; Anglo-French: 1.88%; and French: 14.77%): 28.30%;

• Latin (including modern scientific and technical Latin): 28.24%;

• Germanic languages (including Old English, Proto-Germanic and others: 20.13%;

• Old Norse: 1.83%; Middle English: 1.53%; Dutch: 1.07%; excluding Germanic words borrowed from a Romance language): 25%;[a]

• Greek: 5.32%;

• no etymology given: 4.04%;

• derived from proper names: 3.28%; and

• all other languages: less than 1%

https://en.wikipedia.org/wiki/Foreign-language_influences_in...

Also, one could argue French itself is an agglomeration of Vulgar Latin (87%) as well as its own Frankish Germanic roots (10%), and a few of Gaulish and Breton Celtic origin.

https://en.wikipedia.org/wiki/List_of_French_words_of_German...

The only thing that might take "weeks" is procrastination. Presuming absolutely no background other than general data engineering, a decent beginner online course in Kafka (or Redpanda) will run about 1-2 hours.

You should be able to install within minutes.

Correct. Redpanda is source-available.

When you have C++ code, the number of external folks who want to — and who can effectively, actively contribute to the code — drops considerably. Our "cousins in code," ScyllaDB last year announced they were moving to source-available because of the lack of OSS contributors:

Moreover, we have been the single significant contributor of the source code. Our ecosystem tools have received a healthy amount of contributions, but not the core database. That makes sense. The ScyllaDB internal implementation is a C++, shard-per-core, future-promise code base that is extremely hard to understand and requires full-time devotion. Thus source-wise, in terms of the code, we operated as a full open-source-first project. However, in reality, we benefitted from this no more than as a source-available project.

Source: https://www.scylladb.com/2024/12/18/why-were-moving-to-a-sou...

People still want to get free utility of the source-available code. Less commonly they want be able to see the code to understand it and potentially troubleshoot it. Yet asking for active contribution is, for almost all, a bridge too far.

Yes, for Redpanda. There's a blog about that:

"The use of fsync is essential for ensuring data consistency and durability in a replicated system. The post highlights the common misconception that replication alone can eliminate the need for fsync and demonstrates that the loss of unsynchronized data on a single node still can cause global data loss in a replicated non-Byzantine system."

However, for all that said, Redpanda is still blazingly fast.

https://www.redpanda.com/blog/why-fsync-is-needed-for-data-s...

2) above is basically "Give a kid a hammer, and everything becomes a nail."

The third camp:

3) People who look at a task, then apply a tool appropriate for the task.

But in this case, it is like saying "You don't need a fuel truck. You can transport 9,000 gallons of gasoline between cities by gathering 9,000 1-gallon milk jugs and filling each, then getting 4,500 volunteers to each carry 2 gallons and walk the entire distance on foot."

In this case, you do just need a single fuel truck. That's what it was built for. Avoiding using a design-for-purpose tool to achieve the same result actually is wasteful. You don't need 288 cores to achieve 243,000 messages/second. You can do that kind of throughput with a Kafka-compatible service on a laptop.

[Disclosure: I work for Redpanda]

Exactly. Just yesterday someone posted how they can do 250k messages/second with Redpanda (Kafka-compatible implementation) on their laptop.

https://www.youtube.com/watch?v=7CdM1WcuoLc

Getting even less than that throughput on 3x c7i.24xlarge — a total of 288 vCPUs – is bafflingly wasteful.

Just because you can do something with Postgres doesn't mean you should.

1. One camp chases buzzwords.

2. The other camp chases common sense

In this case, is "Postgres" just being used as a buzzword?

[Disclosure: I work for Redpanda; we provide a Kafka-compatible service.]

Downdetector had 5,755 reports of AWS problems at 12:52 AM Pacific (3:53 AM Eastern).

That number had dropped to 1,190 by 4:22 AM Pacific (7:22 AM Eastern).

However, that number is back up with a vengeance. 9,230 reports as of 9:32 AM Pacific (12:32 Eastern).

Part of that could be explained by more people making reports as the U.S. west coast awoke. But I also have a feeling that they aren't yet on top of the problem.