You’re getting a lot of pushback, but I’m genuinely interested in your experiences. What differences have you seen between average and top leaders? Not in results, but in terms of how they approached the job?
HN user
jdlshore
James Shore.
Book: The Art of Agile Development (now in 2nd edition): https://www.jamesshore.com/s/aoad2
Blog: https://www.jamesshore.com/v2/new
Other stuff: https://www.jamesshore.com/v2/best
Contact: jshore at jamesshore.
Mastodon: @jamesshore@mastodon.online
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
To be clear, the issue is not that the books were used to train Claude, but that they were pirated.
There’s two types of quality: external quality, which is how well something works, and internal quality, which is how well something is built.
Nonetheless, some people have decided that the risk isn’t worth it, and have decided to geo-block the EU.
You can choose differently, of course, and I’m not making value judgments here. Just explaining why a small local news site would geo-block the EU.
It’s a letter to the editor.
The GDPR states that websites that serve EU citizens are subject to it, no matter where they are. Many US news organizations chose to geo-block the EU rather than going to the trouble of figuring out compliance.
A Reddit spam tool makes $90K per year ($60K ARR, $30K one-time payments), split 50% with “a larger partner.” The founder’s prior salary was $84K, so with the revenue split and cost of operations, the business is not yet recouping its costs.
This is pretty low-grade content marketing. If his spam tool was good (to whatever degree spam tools can be good) he’d be making a lot more money.
Quality is more about maintainability, IMO, which comes down to consistency and clarity of design, sensible abstractions, cohesion, decoupling, etc. I don’t think a rework metric tells you that.
If you can narrow the problem down, then you could design a much better interface for it than a text box and free form text (unless that's the better solution).
Yes, I agree, in that the chatbot we built probably would have worked just as well with a traditional UI, and would have been done a lot faster. But it would have been a lot less sexy (actually important for the bottom line!) and there are future directions that could take advantage of the conversational interface that’s potentially better than a traditional UI.
On the down side, good chatbots are really frikkin difficult to write. These things (LLMs) are not reliable at scale. The basic functionality came together in weeks. Getting it to behave consistently and obey guardrails took months, and even then we had to accept a low level of failed conversations.
As @simonw said, the Ford example isn’t a good one.
As for AI-assisted engineering going well, I think the jury is still out. Here on HN and with the engineers I know, you see people claiming multiples of productivity on coding tasks. But you also see people complaining about drowning in slop PRs.
I think there’s a lot of confounding factors to these reports. The type of work matters a lot: bug fixing good, prototyping good, big legacy codebases not so much, but maybe good for increased understanding. The type of automation matters: aggressive autocomplete good, vibe coding bad, dark factory (vibe coding with fancy harnesses and auto-“correcting” eval loops) questionable.
And then finally, the perennial mistake our industry makes, which is to value speed of creation over maintenance costs. Personally, I think this is where AI-assisted engineering is going to fall down really hard, but the jury’s still out on that one.
Anyway, there’s a really big spread in experiences with AI, that I think chalk up more to all this context rather than religion and belief. OP didn’t address it at all, which I think is a big gap in their essay, but I do think think they describe the executive-level mania pretty well.
Their company does data projects. That plus context makes me think they’re talking about internal work process automation type of work, although it also seems like they’re talking about conversational interfaces (chatbots).
I completely buy the “emperor’s new clothes” argument for work process automation. I’m surprised they don’t address AI-assisted engineering, which seems to be going positively for a lot of folks (although I have doubts about its sustainability). I disagree about the success of chatbots, if the problem is narrowly-defined and chosen properly. My previous company built a conversational interface to a vector database and saw good results. (Although, arguably, the vector database was the real magic, and a traditional UI would have been faster and more accurate.)
In general, I think OP is more right than wrong, though, particularly about the AI mania and unrealistic expectations sweeping the C-suite.
Bell curves are probability distributions. This is a time series, so it can’t be a bell curve. It just has the same shape.
It’s Digital Transformation all over again, which was just the latest name of a proud tradition of IT consulting going back decades.
They all have the same promise: unlock business results in your company by deploying work process automation.
They tend to have mediocre results because the low-hanging fruit was plucked long ago, and automating work processes runs into the reality that work processes are not well defined or globally understood. They’re fragmented, with a long tail of special cases, and not amenable to automation. Attempts to unify them causes huge political battles, and the special cases are too numerous to effectively automate.
Maybe things will be different this time, but I wouldn’t bet on it.
Microservices don’t reduce complexity, they just move it to the interactions between services. You have the same fundamental design problem.
In other words, if you can’t design a modular monolith, you can’t design a set of microservices.
Incredibly hard problem, but METR had a good method. They had people estimate how long a task would take (before knowing whether AI would be used), and then randomly assigned each task to “with AI” or “without AI.” When the data was in, they compared actual/estimate ratios of the two populations.
(Presumably, they used a t-test that only compared people against themselves.)
Interestingly, for that study (released in 2025), participants self-rated themselves as 20% more productive, but were measured as being 19% less productive.
Oobleck (corn starch and water) will do this too. But presumably they already knew that. The article describes it as being known to happen in “complex fluids,” but that it was news that it happens in “simple fluids.” Presumably silly putty and oobleck are “complex fluids?”
I think your underlying point is correct, but "buy" is also "buy+maintain." There's a real cost to keeping up with dependency upgrades, especially for big frameworks that like to change their fundamental public-facing API every few years.
Yep. Very slop-like. I found it grating and had to stop reading, even though I thought the subject was interesting.
Agreed. The article is very sloppy, but that doesn’t mean the central thesis is wrong. I’m not conversant enough with financial theory to say one way or another. Anybody care to critique this?
If it’s a headline, it’s coming from a press office, not researchers. The “why” is the same as for any other journalism: to get clicks.
Ramp is an expense reporting tool with minimal LLM functionality and reasonably decent ML for reading receipts and categorizing expenses. This is bog-standard content marketing, not some conspiracy to sell you AI.
Age. Hip replacement surgery is common among the elderly.
Jesus, this is the sloppiest of slop I’ve ever seen. I can’t extract any meaning from the noise.
Thank you, @dabluck, for sharing what failed. I think stories of failure are incredibly valuable, and more useful than stories of success, which are often post-hoc rationalizations.
I’m sorry all the airchair geniuses in this thread feel compelled to express how they’re so much smarter than you and would never fail… or at least, never admit it.
I think it’s incredibly poor form for somebody with the reach and clout of Jon Gruber to be naming and shaming an individual engineer like this. At the very least he could have tried to get the other guy’s side.
I think it’s more that they can’t recognize the downsides of AI. They work with AI and it’s so smart! And magical! They don’t have the expertise to recognize the problems in its output, and they’re confused by complaints. It looks like stubborn resistance to them, so they turn to evangelism to try to get people to “see the light.” They don’t engage with legitimate criticism because they don’t understand it.
(That and the normal herd of grifters who pile on to every fad.)
Unfortunately I think we're entering (have entered?) a period of insanity.
It’s been true for a long time. One of the hardest things as a senior leader in software is dealing with people demanding “accountability” (by which they mean making long-term plans with impossibly precise forecasts) and focusing on costs, all while ignoring value and refusing to engage in prioritization. (I swear, if I hear “it’s all important” again…)
People are just… shallow. They operate on feelings and vibes. They follow the herd without thinking critically. Then they get angry when their dreams clash with reality, and they blame the messenger when those dreams turn out to be fantasies.
But you’re right. AI is bringing out the worst in these tendencies. I think it’s because it’s so convincing when you don’t dig deeply, or aren’t an expert in the subject being discussed. On the plus side, it raises the floor, but I think we’re in for some difficult times before the lessons are learned. I don’t think it will take long, though: I suspect that naive use of AI is going to massively speed up the technical debt curve, and where it used to take 5-9 years to destroy a codebase, it will now take closer to 1-2.
This really needs an editor.
It starts off great, with a compelling story about responding to a roadside emergency, but then it veers off the interstate and meanders through the countryside. The paragraphs — they compound. The em-dashes — they multiply. Words upon words and I still don’t know what the author is trying to say. Something about a system for responding to corporate pseudo-emergencies. I guess.
I’m a writer. I get what a preface is for. And I get why an executive coach writes a book (and it ain’t the royalties). But please, please: hire a development editor. Preferably one not named Claude.
(PS: emdashes are supposed to be used—as a copyeditor once told me—without spaces on either side. It’s one of the ways you can tell the true lover of the emdash from an LLM.)
As VPEng, I didn’t use metrics to assess individuals. Too prone to metric gaming.
Instead, I had a career ladder with a detailed rubric describing the skills an engineer at each level was expected to have. (Including communication and peer-leadership skills.)
Managers performed qualitative assessment of employees, using the career ladder as a guide. They relied on tech leads and Staff engineers to help them understand people’s skills, and provided 1:1 feedback and coaching.
We did use impact-based metrics to assess the results of important initiatives. We solved the attribution and lagging indicator problems by estimating impact rather than measuring it, and using a series of proxy measurements (activation, usage, retention, etc.) as a feedback mechanism for revising those estimates.