HN user

aphyr

9,565 karma

https://aphyr.com

Hacker, physics geek, clojurer, rubyist, aikidoist, photographer.

Posts46
Comments778
View on HN
aphyr.com 3mo ago

The future of everything is lies, I guess: Where do we go from here?

aphyr
745pts773
aphyr.com 3mo ago

The Future of Everything Is Lies, I Guess: New Jobs

aphyr
274pts179
aphyr.com 3mo ago

The future of everything is lies, I guess: Work

aphyr
290pts219
aphyr.com 3mo ago

The Future of Everything Is Lies, I Guess: Safety

aphyr
330pts181
aphyr.com 3mo ago

The Future of Everything Is Lies, I Guess: Psychological Hazards

aphyr
3pts0
aphyr.com 3mo ago

The future of everything is lies, I guess – Part 5: Annoyances

aphyr
283pts169
aphyr.com 3mo ago

The Future of Everything Is Lies, I Guess: Information Ecology

aphyr
10pts1
aphyr.com 3mo ago

The Future of Everything Is Lies, I Guess: Part 3 – Culture

aphyr
141pts108
jepsen.io 4mo ago

Jepsen: MariaDB Galera Cluster 12.1.2

aphyr
117pts16
jepsen.io 7mo ago

Jepsen: NATS 2.12.1

aphyr
432pts165
jepsen.io 11mo ago

Jepsen: Capela dda5892

aphyr
81pts14
jepsen.io 1y ago

Jepsen: TigerBeetle 0.16.11

aphyr
241pts86
jepsen.io 1y ago

Jepsen: Amazon RDS for PostgreSQL 17.4

aphyr
608pts146
jepsen.io 1y ago

Jepsen: Bufstream 0.1

aphyr
224pts85
jepsen.io 1y ago

Jepsen: Jetcd 0.8.2

aphyr
134pts18
jepsen.io 2y ago

Jepsen: Datomic Pro 1.0.7075

aphyr
401pts98
jepsen.io 2y ago

RavenDB 6.0.2 (A Jepsen Report)

aphyr
195pts69
jepsen.io 2y ago

Jepsen: MySQL 8.0.34

aphyr
364pts140
jepsen.io 4y ago

Jepsen: Redpanda 21.10.1

aphyr
193pts59
jepsen.io 4y ago

Jepsen: Radix DLT 1.0-Beta.35.1

aphyr
159pts74
jepsen.io 5y ago

Jepsen: Scylla 4.2-rc3

aphyr
172pts31
jepsen.io 6y ago

Jepsen: Redis-Raft 1b3fbf6

aphyr
195pts39
jepsen.io 6y ago

Jepsen: PostgreSQL 12.3

aphyr
769pts140
jepsen.io 6y ago

Jepsen: MongoDB 4.2.6

aphyr
604pts254
jepsen.io 6y ago

Jepsen: MongoDB 4.2.6

aphyr
159pts23
jepsen.io 6y ago

Jepsen: Dgraph 1.1.1

aphyr
203pts64
jepsen.io 6y ago

Jepsen: Etcd 3.4.3

aphyr
219pts55
jepsen.io 6y ago

Jepsen: YugaByte DB 1.3.1

aphyr
106pts23
jepsen.io 7y ago

Jepsen: TiDB 2.1.7

aphyr
171pts29
jepsen.io 7y ago

Jepsen: YugaByte DB 1.1.9

aphyr
115pts43

Welp. Glad to see Li Shen's using the last fifteen years of my work to automate away my job. :-/

-- edit --

I've seen clients and some colleagues working on things like this, and I can't seem to put into words how disheartening it is. With the exception of some private analysis work, I've shared everything I've built, with everyone, for free. Papers like Elle took years to think through, implement, test, and write. That's free. High-quality checkers, Knossos, Jepsen itself, and the analyses I've put my life into: all public, all free. I put a lot of time into docs and support; essentially all unpaid. I teach classes and give conference talks to make these techniques broadly accessible because I want other engineers to be able to make high-quality systems.

At the same time, I've got a giant pile of debt from an old house that just won't quit throwing curveballs at me, and it's gonna be a few more decades before I can retire. The fact that my clients are willing to pay for this work is why I can invest so much time in R&D and give it all away. When I see someone roll in and just tell an LLM "Go use Jepsen and Elle and figure this out", it's like... well fuck. Is this even possible any more?

Thankfully, LLMs are still really bad at my job, but I don't know if, or how long, that will last. They also don't need to be good to be useful.

And if these LLM tools work, it's good, right? They find bugs, systems get safer. I want systems to be safer. On the other hand, I'm motivated to share what I do because I really want to help people. If it's just LLMs... it feels hollow. I think about this every time I've tried to work on open-source in the last few months. When I spend hours trying to figure out how to keep naming consistent, how to preserve compatibility over a decade, how to make complex code approachable through quality documentation... I have a person in mind. Someone I'll never meet, but they'll see that work, and their life will be a little easier, and maybe they'll smile. I've been talking with my therapist about it: how the work I used to do thinking about other human beings now feels purposeless. How the effort I put into making these tools and ideas accessible will inevitably cannibalize my own employment, because someone, somewhere, is going to tell an LLM "Hey, go do that", and I work in a very, very small niche. It feels like incipient depression.

Recently I've been thinking about taking Jepsen and its supporting libraries closed-source, and changing the way I write reports--instead of teaching people how to test and what to look for, just telling people the results. I don't want to do this. It's bad for everyone, but maybe it buys me a few years of runway. Enough to pay down some of the debt and figure out what I can do next with this body.

Fuck.

I've put considerable time into this, including speaking with Ofcom directly. The guidance Ofcom issued for small site operators last year was that they did intend to target "one-man bands", and that there would be no guidance on specific numbers that constituted the "significant number" of UK visitors which triggers Part 3 and 5 provider restrictions.

I would! I grew up in a low-density suburb near Portland, then lived in small-town Minnesota, Madison, SF, Chicago, and Cincinnati. I didn't own a car until my 30s, and I currently drive about once a month---camping, Costco, lumber, that sort of thing. Pretty much all my day-to-day travel is and has been by bicycle (now an e-bike), foot, train, or bus.

Situations vary, obviously! I'm no stranger to rural life, I wound up in a car-dependent suburb with terrible bus service for a bit, and my partner is in the trades. Private vehicles are sensible and essential answers to lots of problems.

But as the Netherlands illustrates, it's not all-or-nothing: reductions in car utilization and car infrastructure have real benefits. Broadly speaking I think we can and should disincentivize private car use, increase public transit frequency, and build networks of protected infrastructure for pedestrians, cyclists, and other non-car means of getting around.

I think you can combine 'Incanters' and 'Process Engineers' into one - 'Users'

I wanted to talk about this more but couldn't quite figure out how to phrase it, so I cut a fair bit: with "incanters" I'm trying to point at a sort of ... intuitive, more informal practitioner knowledge / metis, and contrast it with a more statistically rigorous approach in "statistical/process engineers". I expect a lot of people will fuse the two, but I'm trying to stake out some tentpoles here. Users integrate a continuum of approaches, including individual intuition, folklore, formal and informal texts, scientific papers, and rigorously designed harnesses & in-house experiments. Like farming--there's deep, intuitive knowledge of local climate and landraces, but also big industrial practice, and also research plots, and those different approaches inform (and override) each other in complex ways.

I put... I'd guess around 60 hours into editing this piece, and had review from a dozen-odd friends, and I am still finding and fixing errors. I imagine that asking an LLM for a copyediting pass probably would have been helpful, but goshdarnit, I want to show that we can still write somewhat-passable prose by hand.

Thank you for this! I really wanted to go deeper on human factors, and I think there's a lot to be said about CRM and sociotechnical systems design, especially when ML gets used for decision support. Ultimately wound up truncating that section (along with more of the economic critique) because the piece was already far too long.

There's a copy of Das Kapital on the shelf behind me right now, though I don't count myself conversant enough to go super deep on class critique. Figured I'd point a few very vague fingers in that direction and let folks with more experience talk about it.

I actually wound up geoblocking the UK based on Ofcom's February 2025 presentation for small services providers--they said that they intended to target "one-man bands" who (e.g.) failed to perform a child risk assessment or age verification, but that a geoblock would be considered compliant. I don't like doing this, but as someone who visits the UK regularly (and has been regularly pushing Ofcom on this matter) I figure better safe than sorry.

https://player.vimeo.com/video/1053842235?app_id=122963

It's been weirdly uneven. Sections 1, 3, and 5 did well on HN; 2, 4, and 6 sank with essentially no trace. The distribution of views is presently:

1. Introduction: 33,088 (https://news.ycombinator.com/item?id=47689648)

2. Dynamics: 3,659 (https://news.ycombinator.com/item?id=47693678)

3. Culture: 5,914 (https://news.ycombinator.com/item?id=47703528)

4. Information Ecology: 777 (https://news.ycombinator.com/item?id=47718502)

5. Annoyances: 7,020 (https://news.ycombinator.com/item?id=47730981)

6. Psychological Hazards: 199 (https://news.ycombinator.com/item?id=47747936)

Feedback from early readers was that the work was too large to digest in a single reading, so I split it up into a series of posts. I'm not entirely sure this was the right call; the sections I thought were the most interesting seem to have gotten much less attention than the introductory preliminaries.

You're right, that is longer! I get why though; `filter` is a clojure.core function name people don't necessarily feel comfortable shadowing, and the Clojure and Spark versions make it clear what's a symbol in local scope versus a field in the dataset. I don't think it'd be hard to make a little wrapper for this sort of thing though! Here's an example which turns any symbols not in local scope into field lookups on an implicit row variable.

    (require '[clojure.walk :refer [postwalk]])

    (defmacro filter
      [ds & anaphoric-pred]
      (let [row-name (gensym 'row)
            pred     (postwalk (fn [form]
                                 (if (and (symbol? form) (nil? (resolve form)))
                                   `(get ~row-name ~(str form))
                                   form))
                       anaphoric-pred)]
      `(tc/select-rows ds (fn [~row-name] ~@pred))))
Now you can write
    (filter ds (> year 2008))
And it'll expand to the ts form:
    (pprint (macroexpand '(filter ds (> year 2008))))
    => (tc/select-rows ds (fn [row2411] (> (get row2411 "year") 2008)))

I was kind of surprised by this one--I know the MariaDB folks and have worked with some of them before. They made significant changes to fix the Repeatable Read issues we found in the last report, so I know the team cares about safety.

There wasn't much reaction on the mailing list to the lost-write problem back in January, or to the Jira tickets. I actually tried calling MariaDB on the phone to see if they'd like to talk about it, but no dice. I assume they're probably busy with other projects at the moment (hi, it's me too) and haven't had a chance to switch gears.

Jepsen: NATS 2.12.1 8 months ago

Jepsen started as a personal blog series in nights and weekends; jepsen.io is when I started doing it professionally, about ten years ago.

Jepsen: NATS 2.12.1 8 months ago

Ah, pardon me, spoke too quickly! I remembered that it fsynced by default, and offered batching, and forgot that the batch size is 0 by default. My bad!

Author here--from discussions with Capela's team, I think this sort of early testing can be remarkably helpful, because it offers a test suite that Capela's team can check their work against as they move forward.

I would suggest against this kind of integration test when the data model or API are in constant flux, because then you have to re-write or even re-design the test as the API changes. Small changes--adding fields or features, changing HTTP paths or renaming fields--are generally easy to keep up with, but if there were, say, a redesign that removed core operations, or changed the fundamental semantics, it might require extensive changes to the test suite.

If I might add on to what you and Joran are both saying, after some time working with TigerBeetle, I found it useful to think of Protocol-Aware Recovery as similar to TAPIR (https://syslab.cs.washington.edu/papers/tapir-tr14.pdf). Normally we build distributed systems on top of clean abstraction layers, like "Nodes are pure state machines that do not corrupt or forget state", or "the transaction protocol assumes each key is backed by a sequentially-consistent system like a Paxos state machine". TAPIR and PAR show a path for building a more efficient, or more capable, system, by breaking the boundaries between those layers and coupling them together.

Yeah, TigerBeetle's blog post goes into more detail here, but in short, the tests that were running in Antithesis (which were remarkably thorough) didn't happen to generate the precise combination of intersecting queries and out-of-order values that were necessary to find the index bug, whereas the Jepsen generator did hit that combination.

There are almost certainly blind spots in the Jepsen test generators too--that's part of why designing different generators is so helpful!

To build on this--this is something of a novel technique in Jepsen testing! We've done arbitrary state machine verification before, but usually that requires playing forward lots of alternate timelines: one for each possible ordering of concurrent operations. That search (see the Knossos linearizability checker) is an exponential nightmare.

In TigerBeetle, we take advantage of some special properties to make the state machine checking part linear-time. We let TigerBeetle tell us exactly which transactions happen. We can do this because it's a.) strong serializable, b.) immutable (in that we can inspect DB state to determine whether an op took place), and c.) exposes a totally ordered timestamp for every operation. Then we check that that timestamp order is consistent with real-time order, using a linear-time cycle detection approach called Elle. Having established that TigerBeetle's claims about the timestamp order are valid, we can apply those operations to a simulated version of the state machine to check semantic correctness!

I'd like to generalize this to other systems, but it's surprisingly tricky to find all three of those properties in one database. Maybe an avenue for future research!