HN user

philbo

2,492 karma

https://philbooth.me/blog

Posts159
Comments264
View on HN
www.bbc.co.uk 12d ago

Man nearly sucked out of window mid-air on Ryanair plane

philbo
15pts3
philbooth.me 1mo ago

Agentic Coding and Mental Models

philbo
2pts0
www.cnn.com 4mo ago

Top FEMA official said he once teleported to Waffle House

philbo
4pts0
www.theguardian.com 4mo ago

Can pumping chemicals into the ocean help stop global heating?

philbo
1pts1
www.playpawno.com 5mo ago

Show HN: Pawno, multiplayer chess in your terminal

philbo
1pts1
nealstephenson.substack.com 6mo ago

KdK part 2: a medical mystery from postwar Germany

philbo
2pts0
nealstephenson.substack.com 6mo ago

KdK (Kinetik der Kontinua) part 1: Introduction

philbo
1pts0
www.sciencefocus.com 7mo ago

The closer we look at time, the stranger it gets

philbo
84pts107
github.com 1y ago

Redis Historical Versions from 2009

philbo
1pts0
www.algolia.com 1y ago

Unintended double encryption crippled our search engine performance

philbo
17pts2
news.ycombinator.com 1y ago

Ask HN: Is anyone still programming the old-fashioned way (without LLMs)?

philbo
38pts44
www.bbc.co.uk 1y ago

Cities around the world are sinking at "worrying speed"

philbo
4pts0
www.youtube.com 1y ago

Tidy First? A Daily Exercise in Empirical Design • Kent Beck • Goto 2024 [video]

philbo
1pts0
iopscience.iop.org 1y ago

JWST Validates HST Distance Measurements

philbo
2pts0
philbooth.me 1y ago

Concurrency diagrams

philbo
2pts0
www.bbc.co.uk 1y ago

Apple accused of trapping and ripping off 40M iCloud customers

philbo
5pts0
en.wikipedia.org 1y ago

Lion-Man

philbo
4pts3
www.bbc.co.uk 1y ago

London through the ages inspires Civilization VII

philbo
3pts0
www.bbc.co.uk 1y ago

Biggest iceberg spins in ocean trap

philbo
3pts0
www.bbc.co.uk 2y ago

A Bugatti car, a first lady and the fake stories aimed at Americans

philbo
64pts101
www.independent.co.uk 2y ago

'world-changing' solar tech could mean the death of batteries

philbo
2pts2
philbooth.me 2y ago

Status Games

philbo
2pts0
www.bbc.co.uk 2y ago

The Speed Project Atacama: The rebel race across a desert

philbo
12pts3
philbooth.me 2y ago

The art of good code review

philbo
3pts0
www.bbc.co.uk 2y ago

Mouse filmed tidying up man's shed every night

philbo
48pts2
www.bbc.co.uk 2y ago

'Super-shoes', tumbling world records and the race for a sub two-hour marathon

philbo
2pts1
philbooth.me 2y ago

Vector Search for Dummies

philbo
1pts0
phosphoricons.com 2y ago

Phosphor is a flexible icon family for interfaces, diagrams, presentations

philbo
1pts0
philbooth.me 2y ago

Lessons learned from integrating with GPT in production

philbo
3pts0
philbooth.me 2y ago

Ways to shoot yourself in the foot with Redis

philbo
179pts80

https://gitlab.com/philbooth/opair

It's a coding harness that eschews autonomy and instead works like a pair programming partner, with distinct "driver" and "navigator" modes. I've only spent 3 weekends on it so far, so it's a long way from finished. But I am at least using opair to work on opair now, which is nice.

I didn't really want to write a harness, I just got frustrated enough that nobody else was writing the harness I actually want to use. I'll probably be the only person that uses this, but I'm fine with that.

YES!

It's still very wip, I spent a couple of weekends on it so far, but I'm working on a harness that eschews autonomy and instead aims to work as a pair programming partner. Key to that are distinct "driver" and "navigator" modes, with the capacity to flip between them rapidly.

https://gitlab.com/philbooth/opair

(not really usable yet, but after tomorrow's session I expect to be developing opair in opair, which is mildly exciting)

Yesterday I started working on an agent harness that tries to address some of the issues here.

What I'm hoping to build ultimately is something that works more like a pair-programming partner than existing harnesses do. I want the user to be an engaged part of the development process all the way through, I don't want the agent disappearing to work on its own. I even want to make it possible for users to swap into the driver role and have the LLM automatically assume the role of navigator when that happens.

There's more info in the readme (actually the readme is all that exists so far, I wanted to get the idea straight in my head first):

https://gitlab.com/philbooth/opair

Even if nobody else uses it, I hope it will be a useful tool for myself and help me find a way to work with LLMs that doesn't harm my mental models, which is what I feel current harnesses do.

If a coworker dumped a 5k-line code review on you, you'd tell them to come back when it's broken down into smaller, reviewable chunks. Large dumps of code are basically unreviewable by humans, but it seems like a lot of people have forgotten about that when it comes to LLMs.

https://app.bluefriday.uk/

The nichest of niche social network clients. It's for people in one particular country, who watch one particular TV program, on one particular day of the week.

Now that the cost of writing software is zero, I love that my focus have moved from vain attempts to generate passive income to just building whatever random shit I feel like. Wish I'd made that choice earlier in life, but no worries!

For decades, engineers understood that large code reviews are harder than small ones. Out of both politeness and a desire to receive better code reviews, we learned to break our large changes into smaller chunks. Some engineers took things even further and replaced code reviews with pair programming. But then LLMs showed up and everyone seems to have forgotten those lessons.

They can be still be applied now using coding agents, if you're willing to push back against the default setup and change your mode of thinking a little bit. Of course it doesn't help that an entire industry is dedicated to persuading us that maximizing token spend is the only way to get shit done.

I appreciate this probably seems like an extremist take, but I wrote some more about it here in case there's anybody out there who identifies with it:

https://philbooth.me/blog/agentic-coding-and-mental-models

I began this project as an exercise to learn Go before starting a new job, then continued it as an exercise to mess about with Claude Code. So development was LLM-assisted: the UI is 100% vibe-coded, the model is 100% vim-typed, the rest is a mix of both.

Online play does not require signup etc. Instead I create an ephemeral session when network games are created/joined and those sessions are deleted at the end of the game. Leaderboard state is entirely local.

Because the backend runs on a small instance, I've been quite aggressive with connection management. If the server goes a minute without hearing from a client (turn or heartbeat), it ends the game and awards the win to the opposing player.

It was a pretty fun project to work on, and I definitely wouldn't have finished it without the crutch of Claude Code to push me through some of the schlepp.

Russia was allowed to inherit the USSR seat on 3 conditions:

- It took on all the sovereign debt from the newly independent nations.

- It relinquished nukes that were left behind in Ukraine.

- The United Nations collectively agreed to it.

I don't think any of those things would happen in the UK's case. But of course it doesn't matter what you or I think. It only matters what _Iran_ thinks will happen if Scotland gains independence.

It's for the Scottish. It's in Iran's interests for Scotland to become independent because that would enforce change on the United Nations Security Council. The UK ceases to exist and loses its veto, then what happens on the UNSC after that is anyone's guess.

It sounds like there’s another failure here, which you could have documented. If the test team didn’t understand what they were meant to test, that’s a failure of communication. Simply saying “they were wrong” is not sufficient exploration of the failure so, if that’s the point your manager was making, I agree with them. Blaming a third party for misunderstanding is less useful than seeking to improve the clarity of your own communication.

I think this and other recent posts here hugely overcomplicate matters. I notice none of them provides an A/B test for each item of complexity they introduce, there's just a handwavy "this has proved to work over time".

I've found that a single CLAUDE.md does really well at guiding it how I want it to behave. For me that's making it take small steps and stop to ask me questions frequently, so it's more like we're pairing than I'm sending it off solo to work on a task. I'm sure that's not to everyone's taste but it works for me (and I say this as someone who was an agent-sceptic until quite recently).

Fwiw my ~/.claude/CLAUDE.md is 2.2K / 49 lines.

Does he though? Honestly I haven't seen it.

I've been through the last ~10 or ~15 posts on his Medium this evening, to check. Sentence-by-sentence I don't see anything that goes beyond "what if". Can you share some of the quotes you have in mind?

I think this is an interesting phenomenon, because it seems that lots of people throw personal insults at him (not saying that's you btw) without addressing the meat of whatever they're reacting to.

And lest we forget! One of the founding essays [1] of this very website discusses it: if you're slinging ad hominem attacks or personal insults around, you're by definition losing the "argument" (not that I think this qualifies as an "argument").

[1] https://paulgraham.com/disagree.html

I don't disagree with any of that. But as long as there are companies willing to pay me to write code the old-fashioned way, I'll keep doing it.

As one of the curious minority who keeps trying agentic coding but not liking it, I've been looking for explanations why my experience differs from the mainstream. I think it might lie in this nugget:

    > I believe with Claude Code, we are at the
    > “introduction of photography” period of
    > programming. Painting by hand just doesn’t
    > have the same appeal anymore when a single
    > concept can just appear and you shape it
    > into the thing you want with your code review
    > and editing skills.
The comparison seems apt and yet, still people paint, still people pay for paintings, still people paint for fun.

I like coding by hand. I dislike reviewing code (although I do it, of course). Given the choice, I'll opt for the former (and perhaps that's why I'm still an IC).

When people talk about coding agents as very enthusiastic but very junior engineering interns, it fills me with dread rather than joy.

No, I really mean more code. It's an unpopular opinion I know, but I think debt scales linearly with code, mainly because I also think bugs scale linearly with code. I recognise that readability and maintainability are important, but it doesn't change the basic equivalence of code = debt for me.

I have an extremist take on this:

All source code is technical debt. If you increase the amount of code, you increase the amount of debt. It's impossible to reduce debt with more code. The only way to reduce debt is by reducing code.

(and note that I'm not measuring code in bytes here; switching to single-character variable names would not reduce debt. I'm measuring it in statements, expressions, instructions; reducing those without reducing functionality decreases debt)

Minor nit, but you've spelt Stephen Hawking's name wrong in the clackset. It's "Stephen", not "Steven".

Surely you don't find writing boilerplate fun though?

Of course. So if I'm faced with some boilerplate, I try to refactor it away so it's less boilerplatey. Perhaps I'm lucky but mostly this seems to work, I don't often find myself writing boilerplate.

I don't know why someone would employ a dev working in such an inefficient way in 2025

Am I working inefficiently? I'm not sure. How much time does the typing part of programming actually take up? I guess it varies, but it's definitely less than 50% for me. Thinking/designing/communicating/listening take most of my time. The typing part is not a bottleneck.

We migrated to Cloud Run at work last year and there were some gotchas that other people might want to be aware of:

1. Long running TCP connections. By default, Cloud Run terminates inbound TCP connections after 5 minutes. If you're doing anything that uses a long-running connection (e.g. a websocket), you'll want to change that setting otherwise you will have weird bugs in production that nobody can reproduce in local. The upper limit on connections is 1 hour, so you will need some kind of reconnection logic on clients if you're running longer than that.

Ref: https://cloud.google.com/run/docs/configuring/request-timeou...

2. First/second generation. Cloud Run has 2 separate execution environments that come with tradeoffs. First generation emulates Linux (imperfectly) and has faster cold starts. Second generation runs on actual Linux and has faster CPU and faster network throughput. If you don't specify a choice, it defaults to first generation.

Ref: https://cloud.google.com/run/docs/about-execution-environmen...

3. Autoscaling. Cloud Run autoscales at 60% CPU and you can't change that parameter. You'll want to monitor your instance count closely to make sure you're not scaling too much or too little. For us it turned out to be more useful to restrict scaling on request count, which you can control in settings.

Ref: https://cloud.google.com/run/docs/about-instance-autoscaling