HN user

bgentry

5,978 karma

Software engineer building River (riverqueue.com). Former Co-founder at Distru. Mux, Opendoor and Heroku alum.

[ my public key: https://keybase.io/bgentry; my proof: https://keybase.io/bgentry/sigs/RDwSk5TnJBGqGLpMbkVGHs1oi8xRF6bs-yXND13oxiM ]

Posts74
Comments676
View on HN
riverqueue.com 10mo ago

Dependabot and private Go proxies: how they work and why it matters

bgentry
1pts0
riverqueue.com 1y ago

SQLite for River, durable periodic jobs and dbsql for Pro, better performance

bgentry
2pts2
riverqueue.com 1y ago

River UI as a packaged Go module

bgentry
3pts0
www.unprepared.life 3y ago

Banning Gas Stoves Is a Terrible Idea

bgentry
14pts27
www.coindesk.com 4y ago

3LAU Raises $16M to Tokenize Music Royalties for Artists and Fans

bgentry
1pts0
royal.io 4y ago

Royal – Invest in artists, share their success

bgentry
2pts0
www.biorxiv.org 5y ago

Recovered deleted sequencing data sheds more light on early SARS-CoV-2 epidemic

bgentry
53pts25
info.crunchydata.com 5y ago

PostgreSQL Monitoring for Application Developers

bgentry
3pts0
stratechery.com 5y ago

Anti-Monopoly vs. Antitrust

bgentry
3pts0
www.sec.gov 5y ago

SEC Form 8-K: Opendoor going public through Social Capital SPAC $IPOB

bgentry
3pts0
www.cnbc.com 5y ago

Palihapitiya to acquire Opendoor, valuing home marketplace at $4.8B

bgentry
8pts1
www.forbes.com 5y ago

Why San Francisco is in trouble – highly compensated city employees

bgentry
64pts60
www.cnbc.com 6y ago

Supreme Court declines to hear cases over qualified immunity

bgentry
285pts214
judiciary.house.gov 6y ago

US House: Justice in Policing Act of 2020 [pdf]

bgentry
2pts0
www.businessinsider.com 6y ago

Segment lays off 10% of staff due to Covid-19 impact

bgentry
6pts0
thehill.com 6y ago

US House passes bill to protect cannabis industry access to banks, credit unions

bgentry
3pts0
brandur.org 7y ago

How to Manage Connections Efficiently in Postgres, or Any Database

bgentry
265pts32
medium.com 7y ago

How I gained commit access to Homebrew in 30 minutes

bgentry
415pts85
www.citusdata.com 8y ago

Fun with SQL: Window Functions in Postgres

bgentry
109pts8
techcrunch.com 8y ago

The 64 startups unveiled at Y Combinator W18 Demo Day 2

bgentry
11pts0
www.macrumors.com 8y ago

MacOS High Sierra App Store System Preferences Can Be Unlocked with Any Password

bgentry
17pts0
www.citusdata.com 8y ago

Citus Cloud 2, Postgres, and Scaling Without Compromise

bgentry
110pts40
www.citusdata.com 8y ago

How Citus works (a look at dynamic executors)

bgentry
7pts0
blog.agilebits.com 8y ago

1Password CLI public beta

bgentry
33pts2
www.philipotoole.com 9y ago

Why Slack Isn't Working

bgentry
5pts0
www.craigkerstiens.com 9y ago

Getting Started with JSONB in Postgres

bgentry
4pts0
blogs.windows.com 9y ago

Simpler Web Payments: Introducing the Payment Request API

bgentry
2pts0
redislabs.com 9y ago

First-Ever Redis Modules Hackathon Winners

bgentry
5pts0
theintercept.com 10y ago

Tech Companies Fight Back After Years of Being Deluged with Secret FBI Requests

bgentry
7pts0
www.techdirt.com 10y ago

Apple ordered to bypass auto-erase on San Bernadino shooter's iPhone

bgentry
685pts348

Thanks for sharing your report, it's frustrating to see things like this break in minor patch updates. Small tip for GitHub Gist: set the file format to markdown (give it a .md extension) so that the markdown will be rendered and won't require horizontal scrolling :)

The important quote from the timeline:

Mar 01 9:41 AM PST

We want to provide some additional information on the power issue in a single Availability Zone in the ME-CENTRAL-1 Region. At around 4:30 AM PST, one of our Availability Zones (mec1-az2) was impacted by objects that struck the data center, creating sparks and fire. The fire department shut off power to the facility and generators as they worked to put out the fire. We are still awaiting permission to turn the power back on, and once we have, we will ensure we restore power and connectivity safely. It will take several hours to restore connectivity to the impacted AZ. The other AZs in the region are functioning normally.

Here is something that gets lost in all the excitement about AI productivity: most software engineers became engineers because they love writing code.

I think there's a big split between those who derive meaning and enjoyment from the act of writing code or the code itself vs. those who derive it from solving problems (for which the code is often a necessary byproduct). I've worked with many across both of these groups throughout my career.

I am much more in the latter group, and the past 12mo are the most fun I've had writing software in over a decade. For those in the first group, it's easy to see how this can be an existential crisis.

An Update on Heroku 6 months ago

Yep, that team did great work. I remember having lunch at the Heroku office with the dotCloud team in 2011 or 2012 and also Solomon Hykes demoing Docker to us in our office’s basement before it launched. So much cool stuff was happening back then!

An Update on Heroku 6 months ago

As somebody whose first day working at Heroku was the day this acquisition closed, I think it’s mostly a misconception to blame Salesforce for Heroku’s stagnation and eventual irrelevance. Salesforce gave Heroku a ton of funding to build out a vision that was way ahead of its time. Docker didn’t even come out until 2013, AWS didn’t even have multiple regions when it was built. They mostly served as an investor and left us alone to do our thing, or so it seemed those first couple years.

The launch of the multi language Cedar runtime in 2011 led to incredible growth and by 2012 we were drowning in tech debt and scaling challenges. Despite more than tripling our headcount in that first year (~20 to 74) we could not keep up.

Mid 2012 was especially bad as we were severely impacted by two us-east-1 outages just 2 weeks apart. To the extent it wasn’t already, reliability and paying down tech debt became the main focus and I think we went about 18 months between major user-facing platform launches (Europe region and eventually larger sized dynos being the biggest things we eventually shipped after that drought). The organization lost its ability to ship significant changes or maybe never really had that ability at scale.

That time coincided with the founders taking a step back, leaving a loss of leadership and vision that was filled by people more concerned with process than results. I left in 2014 and at that time it already seemed clear to me that the product was basically stalled.

I’m not sure how much of this could have been done better even in hindsight. In theory Salesforce could have taken a more hands on approach early on but I don’t think that could have ended better. They were so far from profitability in late 2010 that they could not stay independent without raising more funding. The venture market in ~2010 was much smaller than a few years later—tiny rounds and low valuations. Had the company spent its pre-acquisition engineering cycles building for scalability & reliability at the expense of product velocity they probably would have never gotten successful.

Even still, it was the most amazing professional experience of my career, full of brilliant and passionate people, and it’s sad to see it end this way.

You'll need to unlock your iPhone first. Even though you're staring at the screen and just asked me to do something, and you saw the unlocked icon at the top of your screen before/while triggering me, please continue staring at this message for at least 5 seconds before I actually attempt FaceID to unlock your phone to do what you asked.

No, I don't think so. Oban does not rely on a large volume of NOTIFY in order to process a large volume of jobs. The insert notifications are simply a latency optimization for lower volume environments, and for inserts can be fully disabled such that they're mainly used for control flow (canceling jobs, pausing queues, etc) and gossip among workers.

River for example also uses LISTEN/NOTIFY for some stuff, but we definitely do not emit a NOTIFY for every single job that's inserted; instead there's a debouncing setup where each client notifies at most once per fetch period, and you don't need notifications at all in order to process with extremely high throughput.

In short, the fact that high volume NOTIFY is a bottleneck does not mean these systems cannot scale, because they do not rely on a high volume of NOTIFY or even require it at all.

This is largely because LISTEN/NOTIFY has an implementation which uses a global lock. At high volume this obviously breaks down: https://www.recall.ai/blog/postgres-listen-notify-does-not-s...

None of that means Oban or similar queues don't/can't scale—it just means a high volume of NOTIFY doesn't scale, hence the alternative notifiers and the fact that most of its job processing doesn't depend on notifications at all.

There are other reasons Oban recommends a different notifier per the doc link above:

That keeps notifications out of the db, reduces total queries, and allows larger messages, with the tradeoff that notifications from within a database transaction may be sent even if the transaction is rolled back

Yeah, River generally recommends this pattern as well (River co-author here :)

To get the benefits of transactional enqueueing you generally need to commit the jobs transactionally with other database changes. https://riverqueue.com/docs/transactional-enqueueing

It does not scale forever, and as you grow in throughput and job table size you will probably need to do some tuning to keep things running smoothly. But after the amount of time I've spent in my career tracking down those numerous distributed systems issues arising from a non-transactional queue, I've come to believe this model is the right starting point for the vast majority of applications. That's especially true given how high the performance ceiling is on newer / more modern job queues and hardware relative to where things were 10+ years ago.

If you are lucky enough to grow into the range of many thousands of jobs per second then you can start thinking about putting in all that extra work to build a robust multi-datastore queueing system, or even just move specific high-volume jobs into a dedicated system. Most apps will never hit this point, but if you do you'll have deferred a ton of complexity and pain until it's truly justified.

I get the temptation to attribute the popularity of these systems to lazy police with nothing better to do, but from personal experience there’s more to it.

I live in a medium sized residential development about 15 minutes outside Austin. A few years ago we started getting multiple incidents per month of attempted car theft where the thieves would go driveway to driveway checking for unlocked doors. Sometimes the resident footage revealed the thieves were armed while doing so. In a couple of cases they did actually steal a car.

The sheriffs couldn’t really do much about it because a) it was happening to most of the neighborhoods around us, b) the timing was unpredictable, and c) the manpower required to camp out to attempt to catch these in progress would be pretty high.

Our neighborhood installed Flock cameras at the sole entrance in response to growing resident concerns. We also put in a strict policy around access control by non law enforcement. In the ~two years since they were installed, we’ve had two or three incidents total whereas immediately prior it was at least as many each month. And in those cases the sheriffs could easily figure out which vehicles had entered or left during that time. I continue to see stories of attempted car thefts from adjacent neighborhoods several times per month.

I totally get the privacy concerns around this and am inherently suspicious of any new surveillance. I also get the reflexive dismissal of their value. In this case it has been a clear win for our community through the obvious deterrent factor and the much higher likelihood of having evidence if anything does happen.

Our Flock cameras do not show on the map here, btw.

Modern Postgres in particular can take you really far with this mindset. There are tons of use cases where you can use it for pretty much everything, including as a fairly high throughput transactional job queue, and may not outgrow that setup for years if ever. Meanwhile, features are easier to develop, ops are simpler, and you’re not going to risk wasting lots of time debugging and fixing common distributed systems edge cases from having multiple primary datastores.

If you really do outgrow it, only then do you have to pay the cost to move parts of your system to something more specialized. Hopefully by then you’ve achieved enough success & traction to justify doing so.

Should be the default mindset for any new project if you don’t have very demanding performance & throughput needs, IMO.

Definitely not! Jobs in River are enqueued and fetched by worker clients transactionally, but the jobs themselves execute outside a transaction. I’m guessing you’re aware of the risks of holding open long transactions in Postgres, and we definitely didn’t want to limit users to short-lived background jobs.

There is a super handy transactional completion API that lets you put some or all of a job in a transaction if you want to. Works great for making other database side effects atomic with the job’s completion. https://riverqueue.com/docs/transactional-job-completion

You can get pretty high job throughput while maintaining transactional integrity, but maybe not with Ruby and ActiveRecord :) https://riverqueue.com/docs/benchmarks

That River example has a MacBook Air doing about 2x the throughput as the Sidekiq benchmarks while still using Postgres via Go.

Can’t find any indication of whether those Sidekiq benchmarks used Postgres or MySQL/Maria, that may be a difference.

Ah, so that issue is specifically related to a statistics/count query used by the UI and not by River itself. I think it's something we'll build a more efficient solution for in the future because counting large quantities of records in Postgres tends to be slow no matter what, but hopefully it won't get in the way of regular usage.

Perceived only at this stage, though the kind of volume we’re looking at is 10s to 100s of millions of jobs per day.

Yeah that's a little over 100 jobs/sec sustained :) Shouldn't be much of an issue on appropriate hardware and with a little tuning, in particular to keep your jobs table from growing to more than a few million rows and to vacuum frequently. Definitely hit us up if you try it and start having any trouble!

Developer of River here ( https://riverqueue.com ). I'm curious if you ran into actual performance limitations based on specific testing and use cases, or if it's more of a hypothetical concern. Modern Postgres running on modern hardware and with well-written software can handle many thousands or tens of thousands of jobs per second (even without partitioning), albeit that depends on your workload, your tuning / autovacuum settings, and your job retention time.

You could make that argument for lots of services that have external side effects, but that’s about what happens after the service has been asked to do a thing (to send an email in this case).

However just because an action may be duplicated after the provider has been asked to do a thing, it does not eliminate the value of the provider being able to deduplicate that incoming request and avoiding multiple identical tasks on their end. Without API level idempotency, a single email on the client’s end could turn into many redundant emails at the service provider’s side, each of which could then be subject to those same subsequent duplications at the SMTP layer. And even then, providers can use the Message-Id header to provide idempotency in delivery as many do.

This is an unavoidable consequence of distributed systems where the client may not know if the server ever received or processed the request, and it may also occur due to client-side bugs or retries within their own software.

In other words, API level idempotency can help eliminate all duplication prior to the API; depending on the service, the provider may also be able to eliminate duplication afterward as well. So it’s strictly better than not having it, really not that difficult to implement, and makes it easier for integrators to build a robust integration with you.

The TV show Silicon Valley lampooned all these ideas a decade ago. The weird thing I’ve noticed is that when Bay Area tech people watch that show, they don’t seem to understand that they’re being made fun of. They think they’re being celebrated. That’s how thick the bubble is.

I've never met a person who didn't understand that this show is satire. Every single tech person I've talked to about Silicon Valley thinks it's funny because of how plausible and yet ridiculous it all is, and because of all the totally accurate details scattered throughout—from golden handcuffs / resting & vesting, down to minor things like which drinks were stocked in the show's office fridges. And I lived in the Bay Area during its entire run, so most of my network is current/former Bay Area tech people.

Ben Thompson has been covering Intel’s precarious position for over a decade (well before the market finally realized it) and the latest update is not looking good:

Intel’s is technically on pace to achieve the five nodes in four years Gelsinger promised (in truth two of those nodes were iterations), but they haven’t truly scaled any of them; the first attempt to do so, with Intel 3, destroyed their margins. This isn’t a surprise: the reason why it is hard to skip steps is not just because technology advances, but because you have to actually learn on the line how to implement new technology at scale, with sustainable yield. Go back to Intel’s 10nm failure: the company could technically make a 10nm chip, they just couldn’t do so economically; there are now open questions about Intel 3, much less next year’s promised 18A.

https://stratechery.com/2024/intel-honesty/