HN user

dave1010uk

7,211 karma

Dave Hulbert

https://dave.engineer

Work https://passenger.tech

http://github.com/dave1010 @dave1010 dave1010 at gmail dot com

Posts157
Comments491
View on HN
en.wikipedia.org 2mo ago

Biological Weapons Convention

dave1010uk
3pts0
dave.engineer 6mo ago

Giving coding agents situational awareness (from shell prompts to agent prompts)

dave1010uk
2pts1
schrodingerschatbot.substack.com 7mo ago

The Poison Pill in Anthropic's 'Soul Document' for Claude Opus 4.5

dave1010uk
5pts2
andrewzh112.github.io 1y ago

Absolute Zero: Reinforced Self-Play Reasoning with Zero Data

dave1010uk
7pts2
en.wikipedia.org 2y ago

Negative Temperature

dave1010uk
30pts15
tidyfirst.substack.com 2y ago

Canon TDD

dave1010uk
1pts0
medium.com 2y ago

Post-Quantum Cryptography: It's already here and it's not as scary as it sounds

dave1010uk
1pts1
www.atlasobscura.com 2y ago

The fight between cataphiles and police in the Paris catacombs

dave1010uk
125pts42
github.com 3y ago

Show HN: Pandora: let ChatGPT edit files, run commands and manage Docker

dave1010uk
4pts0
designnotes.blog.gov.uk 11y ago

Asking for a date of birth (2013)

dave1010uk
3pts0
blog.ircmaxell.com 12y ago

The Tale Of The Wrecked Fire Engine

dave1010uk
2pts0
arstechnica.com 12y ago

US to give up control of DNS root zone

dave1010uk
4pts1
fsf.org 12y ago

Mozilla working with Adobe to introduce DRM into Firefox

dave1010uk
2pts0
xkcd.com 12y ago

Heartbleed explanation

dave1010uk
27pts0
blog.inf.ed.ac.uk 12y ago

The Anaemic Domain Model is no Anti-Pattern, it’s a SOLID design

dave1010uk
1pts0
codecha.org 12y ago

Codecha

dave1010uk
2pts0
barisbaris.com 12y ago

Tracking Your Devices in Meatspace

dave1010uk
2pts0
www.ethicalco.de 12y ago

Ethical Code

dave1010uk
1pts0
paul-m-jones.com 12y ago

Estimation Methodology: 2 Workers, 1 Day Per Controller Method

dave1010uk
1pts0
www.openssl.org 12y ago

OpenSSL 1.0.0l released

dave1010uk
1pts0
pieces.stef.io 12y ago

Leading at the code-face: Balancing technical leadership with writing code

dave1010uk
1pts0
blogs.msdn.com 12y ago

What I’d like to see in IE12

dave1010uk
1pts0
copyrightuser.org 12y ago

Copyright User

dave1010uk
1pts0
nakedsecurity.sophos.com 12y ago

Million-dollar fine for sneaky Bitcoin botnet builders

dave1010uk
2pts0
github.com 12y ago

Potion

dave1010uk
2pts0
remotedebug.org 12y ago

RemoteDebug: Initiative to unify remote debugging across browsers

dave1010uk
6pts0
andre.arko.net 12y ago

Vim is the worst editor, except all the other editors

dave1010uk
55pts141
arduino.cc 12y ago

Arduino Yún includes a processor that runs Linux

dave1010uk
3pts0
www.sitepoint.com 12y ago

A PHP from the future

dave1010uk
2pts0
make.wordpress.org 12y ago

A New Frontier for WordPress Core Development

dave1010uk
2pts0

Wardley's original post on this [0] is worth a read. OP adds cowboys and city folk, and makes it a bit more personal.

5 interesting things that are worth making explicit:

1. These teams/roles are as much about appetite and attitude, as they are about someone's skills and capabilities. Some people just aren't comfortable or happy when operating in these environments.

2. Innovation isn't just at the early "0 to 1" stages. Townspeople need to innovate to scale from a million users to a billion.

3. Successful orgs have teams at all stages, working together. Each team evolves what the previous stage built. And cowboys / pioneers are only successful if they can build on industrialised components.

5. As a product evolves, it moves through the stages. Teams either need to let go, or change their way of working.

[0] https://blog.gardeviance.org/2015/03/on-pioneers-settlers-to...

Explores what AI cannot

In other words, gradient descent isn't good at combinatorial optimisation. I'm sure the research is better but the hype in the blog post leaves a bad taste.

There must be a version of Rich Sutton’s Bitter Lesson that applies to alternative computing like this, along with all the other exciting specialised hardware we've seen come and go over the years, like expert systems, optical computing, neuromorphic computing, etc.

Something like:

    General purpose commodity silicon with rapidly evolving software generally beats specialised hardware.
Software is just so much faster to iterate and improve than hardware. AI is also improving it too (eg AlphaEvolve).

Specialized hardware may give a single, significant improvement that grabs headlines but in the long term, compounding small improvements win.

1. Solve reinforcement learning.

2. solve unsupervised learning.

3. gradually tackle more complicated things.

what was the "real reason" they couldn't achieve their original goals?

I assume this is referring to why they gave up being a non-profit. The answer is that they needed more money.

It looks a lot like a CvRDT (i.e. a state-based CRDT).

They describe it as a commutative monoid, which means it has associativity and commutativity. CvRDTs also need idempotence, so they can handle duplicate data. Either they are idempotent too (which would make it semilattice-like), or the network protocol handles the deduplication outside of the data itself.

Letting the payload/application define the merge operation is clever. I assume it would mean contracts could opt in to idempotency if it doesn't already exist.

The other bit Freenet has added is doing all this with DHT routing and subscriptions, rather than a more basic peer mesh. This is very different to a blockchain and means it probably isn't suited for anything transactional.

Heck, there's others who were wealthy in scripture, even kings are they all doomed?

This is a great question. In the next verses, the disciples ask pretty much the same thing: "Who then can be saved?" and then Jesus explains to them:

    With men this is impossible, but with God all things are possible.
Whether it's a camel or a rope (and whether it's a literal needle or a small city gate, as some people argue), I think is less important (though still interesting). Either way, after the rich young ruler walks away, Jesus turns to his disciples and paints a picture something that's completely impossible without God, no matter how hard we might try by ourselves.

Genesis is a theological narrative, which is very different to most things we read these days, especially as a software engineer.

1. The general consensus is that there were more people. This is assumed in Genesis and it (annoyingly!) doesn't bother to explain it, as the audience at the time already assumed it. Also, the authors weren't interested in all the logistics and technicalities that we are today.

2. Cities referenced in Genesis were likely fortified settlements, rather than like modern cities.

The idea that people in Africa could only build simple huts is a myth that came from the colonial era. Africa had large cities, architecture and metallurgy while parts of Europe were still tribal.

If you're keen to learn more, there are some good books that explain this much better than a comment can, such as "How to Read the Bible for All Its Worth" by Fee & Stuart and "Genesis for Normal People" by Pete Enns. I haven't read it but "African Civilizations" by Graham Connah is probably the go-to book on how African cities and technologies were so much further ahead than traditional European/US narratives place them.

The best resource for these kinds of questions is probably "The Bible Project". They have a load of YouTube videos and podcasts that cover these kinds of questions.

The install is very opaque. It's not clear where these skills are installed, how to upgrade them or remove them.

Here's the `skills` package on NPM: https://www.npmjs.com/package/skills - it's MIT licensed but I can't find it on Github.

`skills` looks to be a wrapper around `add-skill`: https://github.com/vercel-labs/add-skill

From the docs, `add-skill` auto detects from 16 different potential paths to copy skills to in a repository (.claude/, .codex/, .Gemini/, etc).

`add-skill` also let's you install skills globally (~/). From the code, `skills` looks like it doesn't support global installs but under the hood it passes all args to add-skill, so you should be able to install skills globally or install multiple skills (even if the wrapper doesn't expect it).

Aside: although lots of agents have adopted SKILLS.md conventions, they're currently all using their own paths. There doesn't seem to be a consensus yet, like there is with AGENTS.md. There are even 3 generic paths: .agent/skills/, .agents/skills/ and just skills/

I know what you're thinking: these restrictions are easy to work around. But don't worry, we can just layer more restrictions on top. Eventually the children will be safe! The government just needs to...

- require proof of age (ID) to install apps from unofficial sources on your phone or PC. Probably best to block this at both the OS and also popular VPN downloading sites like github.com and debian.org.

- require proof of age (ID) to unblock DNS provider IP addresses like 8.8.8.8 and 1.1.1.1 at your ISP.

- make sure children aren't using any other "privacy" tools that might be a slippery slope to installing a VPN.

This makes it so much easier for the parents too! The internet will be so safe that they won't even need to talk to their children about internet safety.

Thanks Simon!

My tool collection [0] is inspired by yours, with a handful of differences. I'm only at 53 tools at the moment.

What I did differently:

Hosted on Cloudflare Pages. This gives you preview URLs for pull requests out the box. This might be possible with Github Pages but I haven't checked. I've used Vercel for similar projects in the past. Cloudflare seems to have the odd failed build that needs a kick from their dashboard.

Some tools can make use of Workers/Functions for backend processing and secrets. I try to keep these to a minimum but they're occasionally useful.

I have an AGENTS.md that's updated with a Github action to automatically pull in Claude-style Skills from the .skills directory. I blogged about this pattern and am still waiting for a standard to evolve [2].

I have a base stylesheet that I instruct agents to pull in. This gives a bit of consistency and also let's them use Tailwind, which they'd seem to love.

[0] https://tools.dave.engineer/

[1] https://github.com/dave1010/tools/tree/main/functions

[2] https://dave.engineer/blog/2025/11/skills-to-agents/

Claude Opus 4.5 8 months ago

Perhaps. Though if that were feasible, I'd expect it would have been exploited already.

I think this is more about the cost and time saving of being able to use cheaper models. Sub-agents are effectively the same as parallelization and temporary context compaction. (The same as with human teams, delegation and organisational structures.)

We're starting to see benchmarks include stats of low/medium/high reasoning effort and how newer models can match or beat older ones with fewer reasoning tokens. What would be interesting is seeing more benchmarks for different sub-agent reasoning combinations too. Eg does Claude perform better when Opus can use 10,000 tokens of Sonnet or 100,000 tokens of Haiku? What's the best agent response you can get for $1?

Where I think we might see gains in _some_ types of tasks is with vast quantities of tiny models. I.e many LLMs that are under 4B parameters used as sub-agents. I wonder what GPT-5.1 Pro would be like if it could orchestrate 1000 drone-like workers.

Claude Opus 4.5 8 months ago

The Claude Opus 4.5 system card [0] is much more revealing than the marketing blog post. It's a 150 page PDF, with all sorts of info, not just the usual benchmarks.

There's a big section on deception. One example is Opus is fed news about Anthropic's safety team being disbanded but then hides that info from the user.

The risks are a bit scary, especially around CBRNs. Opus is still only ASL-3 (systems that substantially increase the risk of catastrophic misuse) and not quite at ASL-4 (uplifting a second-tier state-level bioweapons programme to the sophistication and success of a first-tier one), so I think we're fine...

I've never written a blog post about a model release before but decided to this time [1]. The system card has quite a few surprises, so I've highlighted some bits that stood out to me (and Claude, ChatGPT and Gemini).

[0] https://www.anthropic.com/claude-opus-4-5-system-card

[1] https://dave.engineer/blog/2025/11/claude-opus-4.5-system-ca...

Two years ago I wrote an agent in 25 lines of PHP [0]. It was surprisingly effective, even back then before tool calling was a thing and you had to coax the LLM into returning structured output. I think it even worked with GPT-3.5 for trivial things.

In my mind LLMs are just UNIX strong manipulation tools like `sed` or `awk`: you give them an input and command and they give you an output. This is especially true if you use something like `llm` [1].

It then seems logical that you can compose calls to LLMs, loop and branch and combine them with other functions.

[0] https://github.com/dave1010/hubcap

[1] https://github.com/simonw/llm

Looks awesome!

This isn't so clear though: https://docs.innate.bot/main/software/basic/connecting-to-ba...

BASIC is accessible for free to all users of Innate robots for 300 cumulative hours - and probably more if you ask us.

Is BASIC used just to create the behaviours or to run them too? It sounds like this is an API you host that turns a behaviour like "pick up socks" into ROS2 motor commands for the robot. Are you open sourcing this too, so anyone can run the (presumably GPU heavy) backend?

Does the robot needs an internet connection to work?

Also, more importantly, what does it look like with googly eyes stuck on?

Open Social 10 months ago

That was an example of a social media company changing, with users not being able to migrate their data. Scroll a bit further and you'll see X.

Thanks for submitting this!

Author here. (If you can call me that. GPT-4 and Gemini did the bulk of the work)

This is a (slightly tongue in cheek) benchmark to test some LLMs. All open source and all the data is in the repo.

It makes use of the excellent `llm` Python package from Simon Willison.

I've only benchmarked a couple of local models but want to see what the smallest LLM is that will score above the estimated "human CEO" performance. How long before a sub-1B parameter model performs better than a tech giant CEO?

I 3D printed a replacement screw cap for something that GPT-4o designed for me with OpenSCAD a few months ago. It worked very well and the resulting code was easy to tweak.

Good to hear that newer models are getting better at this. With evals and RL feedback loops, I suspect it's the kind of thing that LLMs will get very good at.

Vision language models can also improve their 3D model generation if you give them renders of the output: "Generating CAD Code with Vision-Language Models for 3D Designs" https://arxiv.org/html/2410.05340v2

OpenSCAD is primitive. There are many libraries that may give LLMs a boost. https://openscad.org/libraries.html

This is my understanding of how it works, without knowing the actual maths behind the functions:

    # client
    r = random_blinding_factor()
    x = client_secret_input()
    x_blinded = blind(x, r)

    # Server
    y_blinded = OPRF(k, x_blinded)

    # Client
    y = unblind(y_blinded, r)
So you end up with y = OPRF(k, x). But the server never saw x and the client never saw k.

This feels like the same kind of unintuitive cryptography as homomorphic encryption.

If it knows it needs to backtrack then could it gain much by outputting something that tells the code to backtrack for it? For example, outputting something like "I've disproven the previous hypothesis, remove the details". Almost like asking to forget.

This could reduce the number of tokens it needs at inference time, saving compute. But with how attention works, it may not make any difference to the performance of the LLM.

Similarly, could there be gains by the LLM asking to work in parallel? For example "there's 3 possible approaches to this, clone the conversation so far and resolve to the one that results in the highest confidence".

This feels like it would be fairly trivial to implement.

This is a really cool idea using fundamental properties of physics but I don't think it works.

If the spacecraft is trusted by the person with the secret, it's much simpler to instruct the spacecraft to only disclose the secret after a set date. You don't seem to gain anything from the carded complexity.

If the spacecraft isn't trusted, there's nothing it stopping it disclosing all the key pairs at once.

I think the best bet at the moment is something like Rivest-Shamir-Wagner time lock puzzles, which requires a fixed number of sequential computations to perform.

I had a secondary goal in the back of my head... if you have a copy of GEB on your shelf collecting dust and you've never read more than a chapter or two, dust it off and see how it goes this time.

Dusted it off and after only a few pages, it's already a completely different read to when I read (a fraction of) it a decade or so ago.