HN user

ErrantX

14,554 karma

tom [at] errant [dot] me [dot] uk

http://www.errant.me.uk/

http://twitter.com/errantx

http://www.errant.me.uk/blog/2009/10/pro-tip-tell-us-exactly-what-your-offering/

If you want to chat or comment or whatever I always welcome emails

[ my public key: https://keybase.io/errant; my proof: https://keybase.io/errant/sigs/m7EZLmQV9GiiOriXznJq8uH1p6RQ3t54rx1yRf4FdOs ]

Posts82
Comments4,867
View on HN
www.bbc.co.uk 4y ago

Trigger of rare blood clots with AstraZeneca jab found by scientists

ErrantX
23pts6
coronavirustechhandbook.com 6y ago

Corona Virus Tech Handbook

ErrantX
3pts0
leanpub.com 11y ago

Show HN: Soft-Launching my book on copy writing

ErrantX
2pts3
phusion.github.io 12y ago

Your Docker image might be broken without you knowing it

ErrantX
2pts0
www.kickstarter.com 12y ago

Show HN: My Brother's Latest Project (Classical Music in Sheffield)

ErrantX
1pts1
www.bbc.co.uk 13y ago

Anonymous hacker group: Two jailed for cyber attacks

ErrantX
3pts0
www.freelancefrontier.com 13y ago

Show HN: Prelaunching "How Not to write an Ad" ebook

ErrantX
1pts0
www.bbc.co.uk 13y ago

France tells Internet service provider to end ads block

ErrantX
1pts0
www.freelancefrontier.com 13y ago

How to get your first freelancing clients

ErrantX
1pts0
www.errant.me.uk 13y ago

Show HN: My own year in review

ErrantX
2pts0
www.errant.me.uk 14y ago

Copy my idea, not my design (or what I learned from Obtvse/Svbtle)

ErrantX
2pts0
www.errant.me.uk 15y ago

Color is in a bubble all of its own

ErrantX
23pts8
appsumo.com 15y ago

Appsumo Lean Startup Challenge

ErrantX
23pts6
www.errant.me.uk 15y ago

The truth about passwords

ErrantX
1pts0
www.errant.me.uk 15y ago

Assange: Conspiracy or coincidence

ErrantX
2pts0
thelastpsychiatrist.com 16y ago

Why Parents Hate Parenting

ErrantX
2pts1
www.errant.me.uk 16y ago

I am my Web Host

ErrantX
14pts12
www.errant.me.uk 16y ago

Bank Simple actually personalise their email

ErrantX
61pts24
www.errant.me.uk 16y ago

The dangers of relying on 3rd party APIs

ErrantX
4pts0
www.errant.me.uk 16y ago

That Piracy thing again (debunking the arguments)

ErrantX
1pts0
www.errant.me.uk 16y ago

How to write a good CV

ErrantX
2pts0
www.errant.me.uk 16y ago

Why I think “Draw Mohammed Day” is a bad idea

ErrantX
2pts10
news.ycombinator.com 16y ago

Ask HN: who is still using Opera on their iPhone?

ErrantX
4pts1
www.errant.me.uk 16y ago

We all be hatin on Facebook, right?

ErrantX
1pts4
www.errant.me.uk 16y ago

The problem with Oauth/OpenID

ErrantX
2pts0
www.errant.me.uk 16y ago

Some simple, but commercially solid, Startup ideas

ErrantX
2pts0
www.errant.me.uk 16y ago

Harriton High, the shape of things to come?

ErrantX
1pts3
news.ycombinator.com 16y ago

Ask HN: Tips for publishing an Ebook

ErrantX
3pts6
news.ycombinator.com 16y ago

Enough April fools?

ErrantX
12pts3
www.errant.me.uk 16y ago

Ask HN: an ad blocking compromise?

ErrantX
11pts10

As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.

What sufficiently hard, but useful, problem would you ask the model for?

Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.

The problem isn't the test, its that is a public test.

Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.

Important context is; he is speaking to a generally small-c conservative.

So 3) he is well aware of his audience and is talking to them directly.

It's kinda a shame; I've noted Paul is leaning into the right wing rhetoric more and more on X, and then more in this speech.

It's a shame. Both AOC and PG are often right in their own way, and then deeply entrenched in others.

That's totally fair and things may change. For me its the history and the fact I can come back to it.

If I am honest I believe my final solution will be a combination of Open Claw, a custom knowledge wiki based on Wikmd. I just need a good all for Claw with history that is as good as gpt

Edit: and context too. It inferred my energy supplier from previously chats and so when I just asked a pertinent question it referenced their policy. Admittedly Google will have way more context if they get the product right.

Some of the advantages are second order.

For example; ChatGPT is replacing my Google searching. Not necessarily because it's better, or because it's summaries are better than Google (I find them subjectively better but it's not clear cut).

But because the app has a nice history; can ask a relatively complicated question and go do something else and then come back to it, ask a follow up. Etc.

None of that is specifically an AI benefit, but it's a workflow that really helps, well, flow.

Obviously I know nothing about your product, so completely uninformed!

But as an enterprise buyer $50/m and $10K/m is the same bucket in terms of cost. No one will blink until around 100K, depending on what it is.

(The point I am making is; as an enterprise buyer I absolutely know how annoying it is for me to turn up and go "this random regulation, we're interpreting it in this highly specific and unique way, and we want it asap". Hence willingness to pay down that inconvenience)

Your getting that interest because it looks like a steal. Ultimately those businesses couldn't care less about $50/m (except to chance it) but they want - or even need - the enterprise terms.

They will pay $50 for your product... And probably $950 for the terms.

(Not saying that would have been the right thing for you but my advice to folks who find themselves in this position is always 20x or 40x the price - if that is enough to make it worth your bother, then go for it. Good chance theyll pay)

Doctors make errors all the time though, so the real argument is about the error percentage. If AIs is lower then it's safer (but it's hard to have that convo, I recognise).

Besides; this article was about diagnosis not prescribing. It's pretty obvious, I think, that diagnosis is one area where AI will perform extremely well in the long run.

I think there are two metrics; the first is outright misdiagnosis, which studies put between 5 and 8% in US/Europe. That's a meaningful number to tackle.

Secondly; overdiagnosis. Where a Dr says on balance it could be X on a difficult to diagnose but dangerous problem (usually cancer). The impact of overdiagnosis is significant in terms of resources, mental health, cost etc.

What matters ultimately is the system achieves your goals. The clearer you can be about that the less the implementation detail actually matters.

For example; do you care if the UI has a purple theme or a blue one? Or if it's React or Vur. If you do that's part of your goals, if not it doesn't entirely matter if V1 is Blue and React, but V4 ends up Purple and Vue.

So, rollback and try again with the insight.

AI makes it cheap to implement complex first drafts and iterations.

I'm building a CRM system for my business; first time it took about 2 weeks to get a working prototype. V4 from scratch took about 5 hours.

I meant that frame very deliberately. Use of the word AI is misleading people that LLMs are intelligent.

They model what looks like intelligence but with very hard limits. The two advantages they have over human brains are perfect recall and data storage. They are also faster.

But the brain is vastly more intelligent:

- It can learn concepts (e.g. language) with an order of magnitude less information

- It responds in parallel to multiple formats of stimuli (e.g. sight/sound)

- LLMs lack the ability to generalise

- The brain interprets and understands what it experienced

That's just the tip of the iceberg. Don't get me wrong: I use AI, it is by far some of the most impressive tech we have built so far, and it has potential to advance society significantly.

But it is definitely, vastly, less intelligent than us.

I just feel this is a great example of someone falling into the common trap of treating an LLM like a human.

They are vastly less intelligent than a human and logical leaps that make sense to you make no sense to Claude. It has no concept of aesthetics or of course any vision.

All that said; it got pretty close even with those impediments! (It got worse because the writer tried to force it to act more like a human would)

I think a better approach would be to write a tool to compare screenshots, identity misplaced items and output that as a text finding/failure state. claude will work much better because your dodging the bits that are too interpretive (that humans rock at and LLMs don't)

If a manager is handling (almost) all disputes of all sorts, then they will fundamentally lack authority to enforce an outcome on a real dispute. They simply are too involved because resolution requires you to take some sort of side.

If my children won't speak to each other I will refuse to be the go between because I become a proxy for one to the other. If one then punches the other they won't respect my perspective that this was wrong because I've set myself up as the proxy for the others feelings.

If you need a manger to resolve the above example, the org is broken and the engineers are poor engineers.

You may not mean it but I do think sometimes framing it this way implies leading and managing is something that requires less ability (it's a skill in its own right).

What I think is true is people cap out their technical competency, and look to shift their skillset and, globally, we are bad at a) training them to be good managers (because there is a wrong assumption it's an innate skill) and b) weeding out the many who also lack the ability to be a manager.

I suspect this is written by someone who stepped into managing a team and no further.

My reflection overall is; he's probably heard of servant leadership but not understood it? It's not about sweeping away problems but more a mindset that your role is to empower. I feel strongly that all new managers should embrace and get good at this because it instills the mindset that the best leaders ultimately only succeed through their team.

A servant leader who becomes overworked is either not doing their job well (delegation isn't contrary to the mindset!) or, more likely, has a poor leader themselvesw.

I actually love the concept of transparent leadership but sadly I can't see it come through in his points. They are all things a good leader, a good servant leader, should also do.

For me transparent leadership becomes more critical as you move up the stack. Once you get to multiple teams or teams of teams leaders must pivot strongly to strategy setting, and in this your servant leadership comes in painting a clear destination for everyone to get to.

At this point I believe the best leaders are genuinely transparent and the worst keep secrets. One of my most respected mentors framed it as deliberately over-sharing. Which I love, even if I get into trouble for it constantly!

(I do like the writers anarchic streak; the best leaders are radicals)

I'm going to go out on a limb here and say NextJs with Auth.js is pretty boring technology.

I'm struggling to see what you'd choose to do differently here?

Edit: actually I'll go further and say I'm guiding against accidental complexity. For example Auth.js is really boring technology, but I am annoyed they've deprecated in favour of better Auth - it's not better and it is definitely not boring technology!

I wouldn't call that accidental complexity? It's just a set of preferences.

Your last point; feels a bit idealistic. The point of code is to achieve a goal, there are ways to achieve with optimal efficiency in construction but a lot of people call that gold plating.

The setup these prompts leave you with is boring, standard, and something surely I can do in a couple of hours. You might even skeleton it right? The thing is the AI can do it both faster in elapsed time but also, reduces my time to writing two prompts (<2 minutes) and some review 10-15 perhaps?

Also remember this was a simple example; once we get to real business logic efficiencies grow.

Yeh I think you are right and I am also finding larger apps built using SDD steadily get harder to extend.

For large existing codebases, SDD is mostly unusable.

I don't really agree with the overall blog post (my view is all of these approaches have value, and we are still to early on to fnd the One True Way) but that point is very true.

I did this first too. The trick is realising that the "spec" isn't a full system spec, per se, but a detailed description of what you want to do.

System specs are non trivial for current AI agents. Hand prompting every step is time consuming.

I think (and I am still learning!) SDD sits as a fix for that. I can give it two fairly simple prompts & get a reasonably complex result. It's not a full system but it's more than I could get with two prompts previously.

The verbose "spec" stuff is just feeding the LLMs love of context, and more importantly what I think we all know is you have to tell an agent over and over how to get the right answer or it will deviate.

Early on with speckit I found I was clarifying a lot but I've discovered that was just me being not so good at writing specs!

Example prompts for speckit;

(Specify) I want to build a simple admin interface. First I want to be able to access the interface, and I want to be able to log in with my Google Workspaces account (and you should restrict logins to my workspaces domain). I will be the global superadmin, but I also want a simple RBAC where I can apply a set of roles to any user account. For simplicity let's make a record user accounts when they first log in. The first roles I want are Admin, Editor and Viewer.

(Plan) I want to implement this as a NextJS app using the latest version of Next. Please also use Mantine for styling instead of Tailwind. I want to use DynamoDB as my database for this project, so you'll also need to use Auth.js over Better Auth. It's critical that when we implement you write tests first before writing code; forget UI tests, focus on unit and integration tests. All API endpoints should have a documented contract which is tested. I also need to be able to run the dev environment locally so make sure to localise things like the database.

Yes... That's much the point I was making.

But there is a lot more complexity than, I think, you are glossing over. For example, you also likely have at least one technical services partner in the flows, probably two.

Additionally, money often doesn't move in real time, especially when credit cards are involved. The process is, intentionally, split.

Your point on that is fair, but remember, many credit providers are also not banks, and the money is in a bank account owned by a third party. So, as a trivial example, I can't just assume money coming to me from Bank A is related to transactions from Bank A's cards.

A lot of people don't realise that the main way all of this works is through very large batch files with lists of transactions in moving back and forth between various parties behind the scenes.

(We are on semantic points, though, but I just wanted to clarify the complexity behind the scenes that most people don't see or understand)

Yep you are completely correct; people don't realise how complex the chip is - it has what you'd legitimately recognise as an operating system! It can also be reprogrammed over the wire, if your chip and pin is taking a bit toooo long that might be what's happening.

Your correct on the risk spread. I wasn't confident last night (I'm not totally versed on the terminals) but looked it up. As I understand if you choose to accept offline only payments then you accept the risk of the transaction failing. If it's the issuers choice they own the risk.

There are a few differences for sure. All entirely technical in how the money moves or clears. The most obvious point here is debit card moves your money from your account, credit moves the issuers money from their account.

But to your wider point; from a transaction fee point of view you are dead right. Of course a credit card has other attractions; for example it's credit :D but also things like section 75 protection.

Your card doesn't know the balance, it doesn't work like that.

Offline transactions mostly died off when the limit in the UK for contactless was raised to £100. At £20/30 (the original limits) issuers/merchants risk accept some payments not being valid (and the total limit before you had to chip and pin was fairly low top).

And worth saying, the merchant has some control on the terminal but mostly the decision of offline/online is down to the issuer and configured on the card.