HN user

bushido

1,649 karma

co-founder @ skipup.ai write random essays at dheer.co

Posts103
Comments291
View on HN
dheer.co 1mo ago

AI Content Fatigue

bushido
3pts0
dheer.co 2mo ago

Recursive Refinement

bushido
2pts0
dheer.co 2mo ago

Humans haven't outsourced all their thinking. They're thinking on a lag

bushido
2pts0
dheer.co 2mo ago

LLM Anxiety

bushido
2pts0
dheer.co 3mo ago

Same Prompt, Worse Results

bushido
2pts0
github.com 3mo ago

Show HN: Swarm – Get consistent results from Claude Code

bushido
1pts0
dheer.co 3mo ago

Distribution is the only moat AI can't kill

bushido
2pts0
github.com 3mo ago

AgentChat – Watch your agent teams/swarms plan

bushido
2pts0
github.com 3mo ago

Show HN: I successfully failed at one-shot-ing a video codec like h.264

bushido
9pts3
dheer.co 3mo ago

Thanks to AI, your job description is now wrong

bushido
2pts0
dheer.co 3mo ago

Tickets Are Prompts

bushido
19pts11
dheer.co 3mo ago

AI adoption problem isn't tech debt

bushido
1pts0
gitvelocity.dev 4mo ago

AI-powered dev productivity metrics

bushido
2pts0
dheer.co 4mo ago

Agents have a human personality problem

bushido
1pts0
dheer.co 4mo ago

Months to minutes: AI feature-gap harness

bushido
1pts0
dheer.co 4mo ago

Getting the most out of Claude agent teams

bushido
1pts1
blog.skipup.ai 5mo ago

The committee problem: why B2B demos die after the form

bushido
1pts0
pmc.ncbi.nlm.nih.gov 1y ago

Intermittent Fasting Promotes White Adipose Browning and Decreases Obesity(2017)

bushido
2pts0
lifearchitect.ai 2y ago

GPT-6 (2025)

bushido
1pts0
codereview.team 4y ago

GitHub Code Review on Demand

bushido
2pts0
medium.com 4y ago

The Goldilocks Zone of SaaS Metrics

bushido
2pts0
www.annenbergclassroom.org 5y ago

Guide to the United States Constitution

bushido
2pts0
en.wikipedia.org 5y ago

First They Came

bushido
18pts0
news.ycombinator.com 6y ago

Ask HN: How do you approach salaries for distributed/remote teams?

bushido
12pts4
news.ycombinator.com 6y ago

Ask HN: What streaming / donation solutions would you use for web performances?

bushido
1pts1
www.investopedia.com 6y ago

Plunge Protection Team

bushido
7pts0
www.vox.com 6y ago

A modest proposal to save American democracy

bushido
3pts1
www.popularmechanics.com 6y ago

The Blood of the Crab

bushido
5pts0
github.com 6y ago

Full Modular Monolith with Domain-Driven Design

bushido
4pts0
www.bbc.com 6y ago

The Secret Seat of the Knights Templar

bushido
77pts18

The 737 has had 14 major recertifications. The aircraft today looks/behaves nothing like the original from the 1960s.

The main motivation for recertifications comes from commercial pressure where if a aircraft is given a new number and not recertified, then the pilots have to be retrained.

Honestly, back when the 737 MAX debacle happened, a lot of consumers claimed that they would stop flying aircrafts if they ran into 737 MAXs. And I don't think it happened in enough numbers - or even enough to make news. Sales went through the roof, everything kept working.

Recertifications are very common. The issue really is is the aircraft is AS different and untested as the old MAXs, and I really can't see that happening again in the next decade or two atleast.

Reporting/news around AI is so very interesting these days.

It's really hard to tell what's propaganda from what's not.

For examples, this Reuters that you have included here has all tells tale signs of creating fear, uncertainty, and feeding into propaganda. But then again I can't be sure.

It's completely the opposite stance to (a) the actions being taken by chinese companies and (b) the public stance taken by their govt https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

It seems like you're basing your spend on the subsidized consumer subscriptions. The equivalent API costs for these subscriptions is usually 12-20x.

Any overages (hourly/weekly/model) on these plans gets billed at rack API costs.

Its not practical to expect these subsidies to last for very long.

Not a great move imo from a business stand point, given the heightened supply chain risk that global (non-US) corporations and sovereigns are already associating with the frontier labs.

I like that it's going to drive more momentum towards the open source/weight models. I was hoping that it would be a slower burn though.

Fable 5 is Back 21 days ago

The loss of trust in using US based model's is unlikely to come back though.

Anthropic with it's hyped doomsday messaging, and the administration falling for it (at best), has eroded a lot of trust and has triggered an arms race of sorts.

IMHO to promote that China believes in free markets and making the technology available to all.

Which will likely help them bolster the sales of the MANY new AI chips in development/use in China to international markets. Dislodging Nvidia.

Kinda the opposite of what Jensen Huang (Nvidia) thinks US is doing: https://www.youtube.com/shorts/u3SY8nvjhQA

Edit: I'm a fan of deepseek and believe it's good to make the technology open/available. And do think that also help business - which I support as well.

Edit 2: No idea why I'm getting downvoted. That's also their official stance https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

Very interesting read.

What jumps out at me is a lot of this is still very task oriented. And each to their own, but anecdotally, I haven't seen great results from task oriented behavior.

I don't mean that it does not produce what was asked for. I'm saying that tasks even when created by engineering and product teams are often wrong.

I lean very heavily towards outcome based prompting. Say exactly what do you want achieved and then maybe give some constraints, ie. what definitely not to do.

In my experiments, this has always produced much, much better results.

Interestingly, it's less engineering and more customer focus.

Anecdotally, I think this is a much bigger cleanup than just talking to the administration.

I think there's a wider damage done which there is no coming back from. And this is for the USA, not Anthropic.

The chance that sovereignty and rules such as this could be applied to AI was a concern that a lot of people had, but the risk was unknown.

Speaking for myself, I had guessed this would happen at some point of time. I was expecting/hoping it'd be years away.

However given the events in the last few days it greatly increases my concern with building any product which can depend on an on an API which could go away for a number of my customers.

I've been experimenting with other open weight models hosted in favorable sovereign countries for a few months, but this accelerates something which was an experiment to now being a must-have.

I don't think it is going to be easy for any of the parties to repair this easily.

Interestingly I've had a similar experience with agent teams/swarms, albeit they can get much more expensive depending on the workflow.

I found that Fable didn't have as much of an impact when put in a team.

But it was/is a very pleasant model to work with 1:1. And was the first time I didn't use my primary team based workhorse in months, across 10s of sessions last week.

This feels more and more like a marketing/scarcity play for the largest global corps.

Will likely give them time to expand capacity as well. And make them harder to dislodge in these orgs.

I go back and forth on this one.

Yes, I'm with the author. I'm absolutely sick of constantly reading AI content.

But if I have to really dig into it deep, a lot of the people who send me AI content now, weren't sending me anything meaningful to begin with (pre ai).

The number of organizations I have been around where most people just copy paste each other's messages is no joke. This was happening long before AI came along. AI has just made it so much more obvious.

Previously they might have copied it from Joe in Product. Now it all sounds like Claude or GPT.

I was half expecting/hoping this would talk about the opportunity cost of owning a home. More precisely, owning a primary residence.

There are quite a few studies about this, but it is something which is not discussed broadly enough. But there is a inverse relationship between home ownership and income.

Because for most regions across the globe, once someone buys a home, they start looking for work that's in geographic proximity to their primary residence. And in most parts of the world, incomes generally tend to stagnate since higher paying jobs are almost always away from where people live.

Now a lot of people believe that they will pick the higher income, but the amount of logistics which goes into thinking about selling your home or renting it even often dissuades people from trying to look for a job which pays significantly more.

Interestingly, some multinational companies that I know of facilitate the entire transaction for their executives and senior managers when they need to move cities or countries because of this effect.

Owning a home for the purpose of investment and not living is a different matter, and the same effect isn't seen there.

I think the main thing which a lot of these articles miss is it's not just your Agents.md which can give you a model upgrade or the inverse.

But everything your harness looks at could be this. So the skills in your code base, the commands that you've added, the memories that were auto created, they all work towards improving or completely destroying your productivity.

And most of it is hidden. You hear people talk about this all the time where they'll be like, Oh, I use GSD or I use Superpowers and my results have gotten worse.

Your results might have gotten worse precisely because you use them (along with your memories and other skills).

If you think of an agent harness as a tool which you use to build your product, then I think you might be absolutely right. I don't see it being easy for a harness to ever build a product.

I actually think that the harnesses which do end up building products, the harness will be the product.

As an example, I have a harness which I have my entire team use consistently. The harness is designed for one thing: to get the results I get with less nuanced understanding of why I get it.

Mind you, most of my team members are non-technical, or at least would be considered non-technical, two years ago.

These days, I spend most of my time fine-tuning the harness. What that gives me is a team which is producing at 5x their capacity from three months ago, and I get easier to review, more robust pull requests that I have more confidence in merging.

It's still a far cry from automating the entire process. I still think humans need to give the outcomes to even the harnesses to produce the results.

Claude Opus 4.7 3 months ago

Possible, but very unlikely.

One of the hard rules in my harness is that it has to provide a summary Before performing a specific action. There is zero ambiguity in that rule. It is terse, and it is specific.

In the last 4 sessions (of 4 total), it has tried skipping that step, and every time it was pointed out, it gave something like the following.

You're right — I skipped the summary. Here it is.

It is not following instructions literally. I wish it was. It is objectively worse.

Claude Opus 4.7 3 months ago

I think my results have actually become worse with Opus 4.7.

I have a pretty robust setup in place to ensure that Claude, with its degradations, ensures good quality. And even the lobotomized 4.6 from the last few days was doing better than 4.7 is doing right now at xhigh.

It's over-engineering. It is producing more code than it needs to. It is trying to be more defensible, but its definition of defensible seems to be shaky because it's landing up creating more edge cases. I think they just found a way to make it more expensive because I'm just gonna have to burn more tokens to keep it in check.

Tangentially related to some of the issues a lot of people are facing, especially the ones where Claude keeps rechecking/scanning the same files over and over.

Ask claude code to give you all the memories it has about you in the codebase and prune them. There is a very high chance that you have memories in there which are contradicting each other and causing bad behavior. Auto-saved memories are a big source of pollution and need to be pruned regularly. I almost don't let it create any memories at all if I can help it.

Disclaimer: I'm also burning through usage very quickly now - though for different reasons. Less than 48 hours to exhaust an account, where it used to take me 5-6 days with the same workload.

Design is an interesting beast.

Good design is not always logical. Color theory, if followed, results in pretty bad experiences. And interestingly, good design can't always be explained in a natural language.

Main thing is, it's very hard to get AI to have taste, because taste is not always statistically explainable.

The best I've gotten to is have it use something like ShadCN (or another well document package that's part of it's training) and make sure that it does two things, only runs the commands to create components, and does not change any stock components or introduce any Tailwind classes for colors and such. Also make it ensure that it maintains the global CSS.

This doesn't make the design look much better than what it is out of the box, but it doesn't turn it into something terrible. If left unprompted on these things, it lands up with mixing fonts that it has absolutely no idea if they look good or not, bringing serif fonts into body text, mixing and matching colors which would have looked really, really good in 2005. But just don't work any more.

That's partly the harness for me. I do believe product managers are still important, but they're important in deciding what gets shipped, not what gets built.

Engineers are still important. They're important in building the harness to ensure that anything which is being built/shipped is of sufficient quality.

In my opinion, testing/QA/etc is now the core product.

But the best code that you'll get is literally connecting to the pain point the customer was saying to the agentic workflow that is building your product.

Bad customer communication in my experience is the result of every person who handled the convo pre-engineers posturing the message trying to make sure the next person is motivated to get it to the next gatekeeper.

This is all very biased based on my own workflow though.

Interestingly, they landed on a conclusion which I have often argued against these days [0]. Code is absolutely cheap, and previously, it was the most important resource that we guarded.

Entire job descriptions and functions were built to guard the engineer's time. Product owners, product managers, customer success, etc., all shielded the engineers who produced code because that was the scarcest resource.

With that scarcity gone, we really need to be thinking about the entire structure differently. I'm definitely in the we still need people camp. The roles are wildly different, though. We can't continue doing the same job that we did with a slight twist.

[0] https://dheer.co/gatekeeping-on-a-different-stage/

Tickets Are Prompts 4 months ago

I agree. I think the poorly defined criteria that we have all gotten accustomed to is thanks to the many layers of the game of whispers that we've added in in organizations between the needs of the customer and the engineers.

Something I've been doing in my own organization, but also trying to help other organizations with, is getting engineers closer to the customers now that building and the time it takes to build is no longer the resource, which is scarce.