HN user

jsdalton

3,162 karma

Software engineer, Contentful

Posts13
Comments470
View on HN

I strongly agree with this take — and that’s partly why the article posted here leaves me scratching my head. PRs are already the gate, right? I don’t care what an agent does or doesn’t do within the confines of its workspace assuming their contributions are gated via a git repository and they don’t require exotic access to a production environment to do their development.

I’m also with you on the junior / mid-level engineer framing (a “brilliant” junior engineer perhaps, one who graduated from at the top of their class from the best CS program in the country) with a big caveat: AI is like a junior engineer who doesn’t know how to learn.

It’s like you’re working with the guy from Memento. Every day your LLM reports to work and they’ve learned nothing from your work so far. Every day is the first day!

Now like the Memento guy you can help them to scatter their workspace with sticky notes and reminders everywhere. With some effort you can start to approximate that thing called “learning” which is LITERALLY the most important trait of every single software developer on a team.

But I confess it’s a struggle for me and the available tooling isn’t there yet. The best I’ve done looks closer to the “second brain” people use tools like Obsidian for. Sadly I don’t think a second brain is a substitute for a first brain. And to be 100% honest any engineer who exhibited the same inability to learn and grow as an AI agent would be sacked after their first month on the job at any company I’ve ever worked at.

I’m actually reasonably optimistic that either the main AI providers or someone else will improve on this in the coming years. It certainly feels like a decent memory paired with a well architected thinking system that’s better at contextually injecting memories (I find LLMs today don’t know what they don’t know unless you force them to put metaphorical sticky notes all over the place) as well as capturing real learnings without supervision shouldn’t be an impossible task requiring novel technical structures.

Anyhow I’d love to be wrong about some of the above and I’m always reading articles like this one hoping that someone has solved these problems already and that I’m just slow on the uptake. But as of today, I’m only modestly better at architecting such agents than I was when I started.

Permanent DST is just a synonym for "let's all agree to wake up an hour earlier." The same change could be affected by e.g. schools and businesses agreeing to open at 8am instead of 9am. (Of course that would be wildly unpopular so permanent DST is just way to trick people into swallowing the pill.)

But would behavior change in the long run? Countries like Spain where solar noon differs wildly from clock noon just end up aligning their rituals accordingly (e.g. eating dinner at 9pm).

It seems to me the parent commenter is saying the opposite: looking exactly like each other _is_ the point. It's a form of social signaling, to indicate that a project "belongs" to the in group of high-flying successful AI hype projects.

Note I'm not arguing that this is a good strategy. But given that so many people follow it I imagine it's not as bad as it appears on the surface.

I’d much, much prefer people were honest about AI answers and text and had the decency to cite it explicitly when they use it.

What I hate far worse than what this article complains about is just blatant AI writing in articles, comments, video narration you name it.

Way more insidious, way bigger problem!

It’s aggressively AI written. I’d rather just read the prompt.

It’s unfortunate because many of us are going “full AI” when it comes to coding. And there are some true gives and takes that are interesting to explore.

Sadly, this piece reads like pure hype.

Yes, and it immediately called to mind for me the phrase “the map is not the territory.”

Put another way: no matter how detailed or “perfect” you make a map, it will never be the territory, ie the thing that is mapped.

Computers and AI are like a map in this regard —- just ones and zeros that we have assigned meaning to arbitrarily. No matter how “good” AI gets, it’s still just a map of the thing not the thing itself.

So AI saying “I feel sad” is never more than a representation of sadness that should not be confused with the subjective experience of sadness itself.

The metaphor that’s popped into my head recently is baking bread.

You can learn to bake good bread. It’s not _that_ hard. And it’ll probably taste better than store bought bread.

But it almost certainly won’t be cheaper. And it’ll take a more more time and effort.

Still, sometimes you might bake your own bread for kicks. But most of the time, you’ll just buy the bread someone else has already perfected.

I simply don’t agree with the conclusion though I appreciate the approach to thinking about products.

I recently chose to take the train vs driving and the factors behind that decision were:

* Time, yes. Train was approximately the same but an actually a bit slower.

* Cost. Train was slightly cheaper when looking at the true cost of driving. Also significantly cheaper than flying.

* Experience. This is entirely overlooked in the timetable centric approach. Train is simply the most pleasant way to travel long distances (maybe ferry is competitive there). I was able to move around and get work done and enjoy the view. If the train company had swapped my train for a bus I would NOT have been a satisfied customer.

* City center to city center (vs airport to airport). Had the train company said “we swapped the arrival location to the airport but technically we still got you to the city” I would NOT have been a happy customer.

The “trains as timetables” hypothesis would imply that the train could meet my needs via something other than rail travel and would definitely lose me as a customer.

On the other hand, improvements such as better wifi service (it was terrible and not sure why cell service is also poor on a train) or a route that was more scenic but did not impact my arrival time significantly would positively affect my likelihood of choosing train.

So the better lesson is know your customer needs and know their specific jobs to be done and center your hypothesis around this.

I’ve often thought this would be useful for version control and change review, since it allows diffs to be a lot less noisy. I’m imagining how much easier it would be to review a PR with significant README edits if the file was already structured with semantic line breaks.

I’ve previously had the above thought and applied it to the end of sentences, but the idea of introducing them at the level of semantic thought had not occurred to me. But if this is where we’re going I’d start to wish for indentation possibilities. I’ve do this frequently with SQL statements, introducing both line breaks and indentations to provide a visual structure that mimics the semantic structure of clauses and the details they contain.

My father had his eyes fixed like this about a decade ago and purports to be happy with it.

IIRC they did actually require that he wear contact lenses that replicated the effect for some amount of time (a month or so I believe) because there are people who are not happy with the arrangement. So they sort of force you to try before you buy.

Agreed. Many animals without language show evidence of thinking (e.g. complex problem solving skills and tool use). Language is clearly an enabler of complex thought in humans but not the entire basis of our intelligence, as it is with LLMs.

Much of this post was spot on — but the blind spots are highly problematic.

In this agentic AI utopia of six months from now:

* Why would developers — especially junior developers — be assigned oversight of the AI clusters? This sounds more like an engineering management role that’s very hands on. This makes sense because the skill set required for the desired outcomes is no longer “how do I write code that makes these computers work correcty” and rather “what’s the best solution for our customers and/or business in this problem space.” Higher order thinking, expertise in the domain, and dare I say wisdom are more valuable than knowing the intricacies of React hooks.

* Economically speaking what are all these companies doing with all this code? Code is still a liability, not an asset. Mere humans writing code faster than they comprehend the problem space is already a problem and the brave new world described here makes this problem worse not better. In particular here, there’s no longer an economic “moat” to build a business off of if everything can be “solved” in a day with a swarm of AI agents.

* I wonder about the ongoing term scaling of these approaches. The trade off seems to be extremely fast productivity at the start which falls off a cliff as the product matures and grows. It’s like a building that can be constructed in a day up to a few floors but quickly hits an upper limit as your ability to build _on top of_ the foundational layer of poorly understood garbage.

* Heaven help the ops / infrastructure folks who have to run this garbage and deal with issues at scale.

Btw I don’t reject everything in this post — these tools are indeed powerful and compelling and the trendlines are undeniable.

Does it operate by translating your higher level AI methods into lower level Playwright methods, and if so is it possible to debug the actual methods those methods were translated to?

Also is there some level of deterministic behavior here or might every test run result in a different underlying command if your wording isn’t precise enough?

This is just outstanding. It's so exactly what I wish for out of a scratch pad.

My feature request to add to your pile (possibly a lonely one, since maybe it's just unique to how my brain works):

I really want a scratch pad like this to have UX that supports "inverted" order. Meaning, new blocks get added to the top of the page instead of the bottom. The blocks naturally flow in descending order of creation rather than ascending. The scratch pad always opens at the top of the page. Over time, blocks thus end up "decaying" toward the bottom, with the most relevant at the top.

It just fits better with how my brain works.

I also +1 the sentiment given elsewhere in this thread to bias toward ignoring the vast majority of these feature requests and preserve the simplicitly of what you've built. That includes mine!

This is where feature flags come in. When you use feature flags to support development you ship code straight to production — but your work is hidden behind that glad while development is under way. So in practice you merge your PRs immediately to master (and deploy).

Essentially you’re decoupling release from development here. This supports any number of QA practices. (We don’t have dedicated QA at the moment and instead have a biweekly “mob QA” session where we do a group deep dive into our current work.) We will capture most small fixes and improvements as sub tasks on the appropriate story (or file a bug ticket if the story was already done and we discovered a new issue.)

As a result of the above we don’t use long lived feature branches which become painful and slow, process wise. We just merge immediately after review. (Unmentioned but this is of course supported by automated testing and continuous deployment.)

My advice to you (at least what's worked for me on several reasonably well-functioning teams using stories and Jira in a similar manner as you have described:

Decouple users stories (customer/product outcomes) from tasks (units of work needed to achieve those outcomes). Jira is designed pretty well for this, since you can have sub tasks attached to user stories.

This works better when your user stories _are_ actually defining outcomes -- for example when you have stories like "User admins can filter jobs by category" and not "Build a category filter for the job search." The first can usually be successfully defined by a few acceptance criteria, whereas the second starts to get weird since you're focus is more on what it will take to build the thing vs. what is the result you want your customer to see.

With your outcome defined in the story you can define any number of implementation tasks it will take to achieve it. If you're a cross-functional team and you do "vertical" instead of "horizontal" splitting then you're sure to have a few coding tasks on the front end ("add the filter component to the search bar", "update the backend client to pass the category id as a query parameter") as well as a few on the backend ("update the API to accept the category parameter", "add the category param to the repository query service"). You probably have some non technical tasks too ("update the help documentation", "update the OpenAPI spec").

We almost never (save for absolutely trivial user stories) have a PR attached to the story but instead have PRs attached to sub tasks. If you're doing trunk based development with continuous deployment and feature flags, you can and should be shipping many PRs. Just yesterday I was in the middle of a task at the end of the day and decided to cut the PR where I was and split the task in two on the fly, since it was easier for me to ship the code like that and easier for my team to review it.

We do story grooming and estimation and all that -- but only for user stories. Tasks are the domain of the humans doing the work and they are meant to be flexible and even disposable. We usually have a session at the start of work on the user story when the engineers working on it align on the solution and then break the work down in to sub tasks, but these naturally evolve as the work progresses. I should had that multiple tasks invite collaboration instead of one story per engineer.

Lastly, I've found you have to preach the virtues of small PRs to your team and usually convert a few stragglers who don't see the value. I try to practice what I preach (i.e. by keeping my own PRs small) and also make a big deal out of it in retrospectives -- i.e. point out how painful the review process is with large PRs, usually entailing many rounds of comments and changes -- so that people quickly become believers if they are not already.

As a last point I try to encourage the value of "PR reviews are your top priority at any given moment" since every second a piece of code sits unreviewed adds to your team's cost of delay. There's a virtuous cycle here where smaller PRs lead to less painful code reviews lead to greater willingness to spend 10 minutes (vs. an hour) doing a code review, which helps really get PRs moving through.

I think PR stacking is great but I also find it's not as important if your PRs are getting reviewed and approved faster than you can write the code for your next PR.

(I didn't realize I'd write so much here, I forget that it's actually kind of a big topic that's built on a variety of different practices that all start coming together at some point when you get in a groove.)

Do you have an induction stovetop now or a regular convection? Because in my experience induction is vastly superior to both convection and gas stoves. It provides far more and faster heat, cools nearly instantly and the nice ones at least give you flexibility where you put your pan.

The downside of induction is it doesn’t work with all pot types.

The big advantage gas retains is the ability to see with your eye exactly how much heat you’re producing. I do love being able to precisely tune exactly how much heat I’m cooking with. I assume this is why gas is so prevalent at restaurants.

But I’ve also found that at home once I get to know an electric stove (what does level 5 do vs 6?) I’m just as capable as I am with gas.

(I’m opposite of you btw. Had electric for a long time and just moved to a house with a gas stovetop.)

Probably “heuristics” is a better term to describe these.

“Mental models” refers to something quite different, which is the simplified representation we form in our minds of the world around us that helps us better comprehend and navigate it.

This article from Martin Fowler explores your point in greater depth. It's a good read: https://martinfowler.com/articles/is-quality-worth-cost.html

One concrete problem with technical debt the article highlights is it that negatively impacts the time to deliver new features. Customers today usually expect not only a great initial feature set from a product, but also a steady stream of improvements and growth, along with responsiveness to feedback and pain points.

It’s also the time that most closely aligns solar noon to clock noon. People who advocate for “summer time all the time” basically just want to wake up earlier but want clocks to trick them into doing it.

(The more legitimate argument is for western regions to align to summer time to be closer in clock time to their eastern neighbors.)

For #4, I know for a fact that my wife’s WhatsApp automatically stores pictures you send her to her iCloud. So the grey blob would definitely be there unless she actively deleted it.