HN user

dbuxton

751 karma

my public key: https://keybase.io/dbuxton; my proof: https://keybase.io/dbuxton/sigs/Ku70DiiRkCJSt4oYob2261mgf79UQ6kFCi8lSBr1_RY

Email: david at dbuxton dot com

Posts38
Comments175
View on HN
github.com 3mo ago

Show HN: Google Docs MCP that works

dbuxton
2pts0
www.notus.org 1y ago

The MAHA report cites studies that don't exist

dbuxton
97pts33
www.counteroffensive.news 1y ago

Ukraine can move beyond its Soviet architectural legacy

dbuxton
48pts18
github.com 1y ago

Show HN: [What I built in 30 mins with Cursor]: Archive.is browser extension

dbuxton
3pts0
www.nytimes.com 2y ago

GeoGuessr guy goes global IRL

dbuxton
2pts0
hrharriet.com 2y ago

Fine-tune a better-than-GPT-4 classifier for $30

dbuxton
4pts0
www.linkedin.com 4y ago

Improvised KYC for Ukrainian refugees coming to the UK

dbuxton
2pts0
twitter.com 4y ago

Why Russia will lose this war?

dbuxton
11pts0
www.ft.com 4y ago

In Siberia, a crypto boom made of ingenuity, defiance and DIY

dbuxton
1pts1
rescribe.xyz 4y ago

Rescribe: A high quality OCR tool for historic books

dbuxton
84pts10
www.cdc.gov 6y ago

The Discovery and Reconstruction of the 1918 Pandemic Virus

dbuxton
1pts0
www.nytimes.com 6y ago

Why Egypt Dominates Squash

dbuxton
2pts0
www.nationalgeographic.com 6y ago

A Remote Peak in Myanmar Nearly Broke an Elite Team of Climbers

dbuxton
1pts0
en.wikipedia.org 7y ago

Air New Zealand Flight 901

dbuxton
1pts1
www.newyorker.com 7y ago

The Mystery of the Havana Syndrome

dbuxton
71pts32
variety.com 7y ago

Filmstruck shutting down

dbuxton
9pts3
www.nytimes.com 8y ago

How a Ransom for Royal Falconers Reshaped the Middle East

dbuxton
1pts0
www.crowdjustice.org 9y ago

Aziz V. Trump: CrowdJustice Launches in the US

dbuxton
46pts12
en.wikipedia.org 9y ago

It Can't Happen Here

dbuxton
2pts0
athenapdf.com 10y ago

Show HN: Athena, drop-in replacement for wkhtmltopdf using Docker, Electron and Go

dbuxton
58pts21
www.nytimes.com 10y ago

History of a fake football team that fooled the NYT

dbuxton
46pts2
www.srin.co 10y ago

Earn It

dbuxton
2pts0
www.crowdjustice.co.uk 11y ago

Show HN: Crowdfund court cases for privacy, human rights and the environment

dbuxton
1pts0
www.crowdjustice.co.uk 11y ago

I was kidnapped and tortured by paramilitaries friendly to BP. Help me sue them

dbuxton
7pts0
www.bbc.com 11y ago

Google piloting modular mobile – in Puerto Rico

dbuxton
1pts0
blog.arachnys.com 11y ago

Lost in Translation: Using Google Input Tools

dbuxton
3pts0
cabotapp.com 12y ago

Show HN: Self-hosted, open-source infrastructure monitoring and alerting

dbuxton
82pts34
blog.arachnys.com 12y ago

Playing the Shell Game

dbuxton
2pts0
www.slate.com 12y ago

How will historians of the future run MS Word 97?

dbuxton
3pts0
cloudflare.com 12y ago

Cloudflare is down (but at least they use Cloudflare :)

dbuxton
2pts0

At the moment a lot of companies thinking of it like speculative capex - we can just stop doing it if we don’t get results, or we can easily optimize.

My concern is that with all the bundling that the model labs are doing the lock-in becomes harder than anticipated to unwind

It is just extremely difficult in most companies to draw a line between a specific feature and an amount of revenue. A lot of engineering has a stochastic type of impact.

(It’s often easier to see what a feature does to reduce costs though)

As well as cost-per-task I think it's worth thinking about speed, especially in non-coding contexts that benchmark less cleanly

We've started trying to do some comparison videos to capture more of the UX vs speed vs cost stuff e.g. https://www.linkedin.com/feed/update/urn:li:activity:7479891... which one of my team did for my LinkedIn account (disclaimer: marketing)

(In this particular case Deepseek was way slower than GPT 5.5 but I think that's because it installed Libreoffice half-way through the task!)

The problem with this is that you have a consistency problem if you want to take action. The only way of making your agents read-write rather than read-only in practice is to use the underlying systems rather than try and pool information in a data lake.

But that does make it more complex to build simple information retrieval use cases.

Working in Glass 1 month ago

I would encourage anyone who hasn’t to try and see a glassblowing demonstration. It’s about as different from programming as any creative process could be - real time, intuitive, working with an unstable material - but something in it really spoke to me.

The Chrysler museum in Norfolk VA has amazing free demonstrations almost every day which are amazing if anyone happens to find themselves in Southern Virginia.

We’re (harriethq.com) trying to do this by reframing it as a “provisioning” challenge - how do you get your connectors installed on non-technical desktops, how do you give some easy pre-bake recipes that wake them from their dogmatic slumber

Honestly though we are finding that a little FDE to set up pre-bake stuff that’s sufficiently specific to the customer is needed. Otherwise people are like, “I don’t need to close the books, I need to do a per-working-day profitability analysis for 10 EU countries with different public holidays”, and they get stuck there.

I had the same experience with chatbots, but we shipped a chatbot module a year ago that helps with complex config questions by reading and answering based on a Salesforce Experience site.

I was skeptical but it gets a 68 NPS from users, even if we do get the occasional "why are you investing in AI I hate it" coming through the feedback channel.

As ever, the issue is "what problem are you solving". If it's that you want more people to put their hand up and talk to you/order something, chatbots seem like a bad solution. If it's that you have a ton of complex docs that people have to read in order to implement and use your product, it's not the solution but it's probably part of a solution.

From the comments looks like lots of people looking at this problem from different angles.

We (harriethq.com) also have a somewhat similar insight, which is that setting up connectors is a drag for non-technical users, and a lot of systems don't support per-user connectivity so need an API shim.

The thing I like about this (Agent Vault) approach is that it's more extensible than what we're offering, which is a full managed service. But we've found that some features (e.g. ephemeral sandboxes to execute arbitrary e.g. uvx/npx based mcps) are just a big pain to self-deploy so it's easier for us to provide a service that just works out of the box.

Kudos to the team, this looks great and I'm looking forward to playing with it

I think the corrective to this is that many of these incumbents will fail to re-conceive their product stack from a user-centric perspective, and as a result they will be reduce to just a dumb data layer which is easily swappable.

Sure, they could do that, but the cultural change required is an order of magnitude harder than just sticking an agent on top of their source-of-truth and believing that the problem is solved.

Maybe it works for areas where the application is a relatively self-contained island of productivity. Figma is somewhere that a designer spends a lot of their day, so it's going to be less vulnerable, but most pieces of softare fit into broader workflows. So for Figma the disruptor is less likely to be "AI-powered designer" and more "AI-powered web builder" - e.g. Lovable or even Claude Code itself that just generates great designs.

I played this in my head a few times and don’t get it.

I assume we are talking

- China - maybe South Korea? - US (or is US not one of the 4?) - Russia (ok this is explicit)

I think there might be an interesting idea in here but there is some confusion that’s stopping it coming out

Can someone enlighten me?

Model Market Fit 6 months ago

The flip side of this is that if model capabilities are extremely strong such that they are able to saturate the benchmarks, the differentiation and defensibility of a wrapper solution built on top are significantly reduced.

IANAL but e.g. Claude Cowork is already good enough that it's hard to see how the legal tech startups are going to differentiate except around access controls, visual presentation of workflows, etc. And that's in a heavily enterprise/compliance-aware/security-focused context.

Don't get me wrong, that's still a big "except" - big enough for massive companies to be built. Personally the anxiety of being so close to being squashed by the foundation models would make me unhappy as an entrepreneur but looking at the market it seems like many people have a higher risk tolerance.

Of all the challenges you face as a startup, the legal entity you choose is possibly the least consequential. Just choose a jurisdiction where investors understand how the legals work (Delaware C-corp, UK Ltd is OK too) and there's a finite administrative burden and/or commoditized tooling in place to help you handle it.

Now, that may not work in all jurisdictions for reasons of local taxation etc (and you'll have to work out payroll tax, benefits etc) but that's almost never anything to do with the legal entity type!

Loved “The Winds of War” and “War and Remembrance” by Herman Wouk - middlebrow from the 70s but no less good for that.

Re-read “The Art of Not Being Governed” by James C Scott which is really mind-expanding stuff.

I have played with this but been underwhelmed. However I do think probably on the right track.

I know the ecosystem not-at-all (sum total knowledge of the CAD ecosystem is that my kids got a Bambu printer for Hanukkah) but it feels to me that current LLMs should be able to generate specs for something like https://partcad.readthedocs.io/en/latest/, which can then be sliced etc.

Curious to know what others think? I come at this from the position of zero interest in developing the fine design skills needed to master but wanting to be able to build and tweak basic functional designs.

The unification and seamless workflow at that scale is painfully hard to achieve

It does make you wonder, why not just be a lot smaller? It's not like most of these teams actually generate any revenue. It seems like a weird structural decision which maybe made sense when hoovering up available talent was its own defensive moat but now that strategy is no longer plausible should be rethought?

Manual: Spaces 8 months ago

One thing I find interesting about discussions of typography in Cyrillic is how poor the overall readability of text is in most fonts compared to Latin because of the relative scarcity of risers and descenders (e.g. pqlt etc)

One of my tutors at university claimed that she was able to read 9th century manuscript Cyrillic faster than modern printed books because the orthography was more varied and easier to scan/speed-read.

(That wasn't something I found to be true)

It’s hard to bet against the foundation models winning consumer use cases where you can reimagine the whole product as a single tool or small number that can be dynamically plugged in to the underlying model and doesn’t require access to proprietary data/custom context.

I think this is a relative succinct summary of the downside case for LLM code generation. I hear a lot of this and as someone who enjoys a well-structured codebase, I have a lot of instinctive sympathy.

However I think we should be thinking harder about how coding will change as LLMs change the economics of writing code: - If the cost of delivering a feature is ~0, what's the point in spending weeks prioritizing it? Maybe Product becomes more like an iterative QA function? - What are the risks that we currently manage through good software engineering practices and what's the actual impact of those risks materializing? For instance, if we expose customer data that's probably pretty existential, but most companies can tolerate a little unplanned downtime (even if they don't enjoy it!). As the economics change, how sustainable is the current cost/benefit equilibrium of high-quality code?

We might not like it but my guess is that in ≤ 5 years actual code is more akin to assembler where sure we might jump in and optimize but we are really just monitoring the test suites and coverage and risks rather than tuning whether or not the same library function is being evolved in a way which gives leverage across the code base.

Interesting that the infographic (which I thought was exceptionally well-designed, well done USGS) found it necessary to call out that 0 billion gallons/day goes to Mexico. Was this done by previous or this administration I wonder? I do recall reading something about disputes between US and Mexico over abstraction. (Presumably from Rio Grande or similar).

I once lived in Moscow on a compound with an adopted rescue dog. The compound had a shock collar and invisible fence setup.

Moscow’s street dogs are renowned for their intelligence. I have seen street dogs taking the escalators on the Metro. This dog worked out not just that the beeping + discomfort was worth the freedom, but also that he could wear out the battery faster by going up to the very edge of the fence - where the chirps became an uninterrupted beeeeep - and as soon as the beeping stopped, whoosh he was gone.

Great question. For me reliability is variance in performance and capability is average performance.

In practice high variance translates on the downside into failure to do basic things that a minimally competent human would basically never get wrong. In agents it's exacerbated by the compounding impact of repeated calls but even for basic workflows it can be annoying.

Fundamentally, we are at a point in time where models are already very capable, but not very reliable.

This is very interesting finding about how to improve capability.

I don't see reliability expressly addressed here, but my assumption is that these alloys will be less rather than more reliable - stronger, but more brittle, to extend the alloy metaphor.

Unfortunately for many if not most B2B use cases this reliability is the primary constraint! Would love to see similar ideas in the reliability space.

I had to do a MN1 application for my US-born daughter as both me and my partner were born abroad. As OP alludes to this is a sort of side quest called "registration" which if you don't do it before the child is 18 lapses (although I think there are still routes to obtain citizenship in these circs).

The most difficult part of the process (not dealt with in this version of Passport Application but maybe a future DLC pack?) was actually finding someone who could certify my evidence (you are meant to submit originals but they keep the docs including passports for 3-6 months which is a bit unrealistic if you are living abroad). I can't remember the exact rules but it wasn't possible to use a US notary or a normal solicitor certification process and instead I needed to go to a council office.

After calling about 5 councils all of whom disavowed any knowledge of the process or its requirements I ended up finding someone at Islington Council who was delightfully helpful. But it was one of the more frustrating UK government interactions I've had.

For me this is great for practice (I tried Russian). However the big missing piece for all these language learning apps is the lack of support for spotting and correcting errors in your pronunciation - as long as you say the word more or less right, the transcription gives you a pass.

I am very excited for the whole STT/TTS to go away and for us to have models that really "hear" exactly what you said.

Sometimes this is about accent but a lot of the time, the AI won't spot areas where you e.g. fudge a case ending or the stress on a word. Yes, you can get some of that pronunciation right by the AI repeating back with the correct stress or clear case, but you never really get the confidence that you would get from an actual human.

Another product suggestion - turn off transcription (at least for the tutor side of the conversation; I'd suggest both). Personally I find it distracting at best for languages I already speak well and a crutch for those I don't.

Finally, I find it really very hard to enjoy having a random conversation that's not very directed ("What interests you most about artificial intelligence?"). I'd suggest that there are ways of making it more goal focused without being explicitly gamified - maybe something like, here's a position and you have to persuade me (AI debate club!), or something that brings out an actual opinion or relates to a concrete experience ("what's your main goal in your job this year").

Overall though this is the first product I've seen in this space that I might actually use, so well done.

Microsoft Edit 1 year ago

I came here assuming that this was a clone of Gemini CLI (which is a clone of Claude Code). Both pleased and disappointed to be wrong.