HN user

Eridrus

6,080 karma

eridrus@gmail.com

Posts1
Comments2,782
View on HN

This isn't a perfect survey, but I am a Codex user and most of my employees are pretty stuck in their workflows and were not really interested in trying even when I was saying good things (even pre-Opus 4.8).

When we work trial people, 100% of people ask for Claude rather than Codex or anything else.

When I talk to people at non-AI tech events everyone basically says they use Claude and have not tried an alternative.

Developers writ large are actually not that interested in trying multiple tools, they like customizing their chosen tool and tweaking it forever.

I think developers are as susceptible to brand marketing as everyone else. It's why almost everyone has a Macbook.

I don't think it's that early tbh, agentic coding has ~90% adoption in the US.

Claude Code has largely won individual developer mindshare and has been on top ever since it came out. The benchmarks change, but almost nobody opts to use anything other than Claude IME when I ask them. Enterprise is more competitive since they care about costs and other things, but developers leaning towards Claude puts a thumb on the scales there.

The product doesn't have much lock in, so it is possible to dislodge Claude, and Anthropic could (and some may argue is likely to) just shoot themselves in the foot again and again and again, but Google has never been particularly good at enterprise sales, and they have never actually been at the frontier of intelligence.

I think Google's incentives have mostly about building models for their products, which makes them focus more on the cheap end, and while they need that, it feels like the Innovator's Dilemma is biting them here.

I own a lot of Google stock from working there in the past and have been quite happy about their trajectory up until the last 6 months, but I am getting pretty antsy about their AI story these days.

I think artists are too attached to solo work and should hire people to crank out variants on their successful vision or license their ideas so that they can continue having new ideas.

I strongly disagree. This is just the Politician's Fallacy.

I keep harping on about the "virtual staging" that real estate agents have been doing for a decade that is equally deceptive and annoying and already gets labeled, and the labels don't actually help because you're still left trying to decipher what is real yourself.

If they wanted to actually do something useful, they'd get together with the legislature and pass a law saying that real estate listings need to come with floor plans that are accurate within X% under the penalty of some sort of fine with a private right of action. But passing laws is hard and faces opposition.

I am so certain of this because I was not born yesterday and this is not my first time paying attention to (NYC) politics.

Politician makes a grand statement they do not have the authority to meaningfully act on to get headlines, DCWP issues a weak sauce disclosure rule and the news cycle moves on because this is not actually anybody's priority.

I am honestly so surprised that everyone on HN is so naive that they take political statements like this at face value.

Politicians routinely say they will do things they do not have the authority to do, and it's often very important to understanding what will actually happen to have some understanding of what authorities are available to them, or at the very least ask Google/LLMs about it.

Right, which is why all we're going to get it a label saying that AI was used (or maybe landlords will try to fly through on the label of "virtually staged" that they've been using).

Existing law doesn't have the authority to ban all AI images as inherently deceptive, and DCWP isn't going to be spending a bunch of time prosecuting individual images.

I agree with Mamdani that these images are often deceptive and misleading and sifting through the bullshit is annoying (and was annoying with virtually staged images too). It's just not going to go anywhere. The energy would be better spent on zoning and building code reform.

You should read up on what a risk retention group is and how it works. To me, it's even worse than you think.

I did some basic reading but don't really see anything particularly wrong with them.

AFAICT, the argument being advanced against Corgi is that insured customers might be doing risky things assuming their insurance will bail them out. This just doesn't ring true to me because I think most startup founders are just willing to accept more risk and accept that sometimes that includes legal risk.

When you look at Corgi's marketing, e.g. https://www.corgi.insure/ai what you'll see in the common risk triggers is basically compliance: AI Safety Audits, VC due diligence, EU Regulation. It's basically all about showing other people that you're "doing something", not because you think you need or want insurance.

I think the comparison to Delve is actually quite apt: startups generally do not care about SOC 2, they just need the checkbox that their customers are asking for. And startup's customers often themselves don't really care, they are just doing it to satisfy their own SOC 2 requirements, ad infinitum.

I think the main people that are being potentially deceived here are not Corgi's customers, those customers' customers, but I don't think they truly care either and are also checking a box.

I think that's a reasonable concern, but I feel like there's no meat to this accusation in in the article.

They found a way to sidestep regulations in a non-traditional way, they're using AI for underwriting, but like I said, there's no actual evidence the underwriting is wrong.

Is a startup gets insurance for something they couldn't get insurance for elsewhere and then Corgi goes belly up, the startup is our their premiums but otherwise in the same place.

For all we know, there are multiple risk groups under the hood for different risk types/profiles to insulate mispricing of different policy types.

Honestly, I feel like startups don't buy insurance at all unless customers ask, there's just nothing meaningful there to insure. If you fuck up that badly you're probably just going to go out of business even if the insurance check comes through.

I agree that insurers definitely faces the urge to underprice risk because the shoe will drop later, but there's no actual evidence here that they're mispricing risk of that people buying it really think it's going to save them if they do something risky.

Thanks for the article, I assume you are the author.

I think the main question about Corgi is: are they underpricing risk so severely that they go bust? And honestly, we have no idea.

For all we know startups are buying overpriced insurance from Corgi because they have a better brand and are easier to deal with than Berkshire's army of underwriters.

Though it's also worth noting that the main reasons startups buy insurance is not because they want insurance, but because enterprise customers demand insurance. Which is to say, it's not out of the realm of possibility that funded startups are not actually that price sensitive, because they just want to get the deal signed and move on.

We got our insurance elsewhere because we're a little older, so I have no actual opinion of Corgi, but there's a lot of stuff that enterprise customers demand that is driven by some compliance checklist. Delve took this to an extreme, but directionally, they were providing the service customers wanted, and at least in the insurance market, you can just pay more to paper over your problems rather than addressing the core risks in a way where there is no fraud. We pay for random shit we don't need that delivers no value for enterprise customers to tick boxes, for all I know Corgi fills the same need.

98% isn't much 15 days ago

He is saying the most inane things ever, I don't know why I need to be charitable here.

98% isn't much 15 days ago

Yeah, but the guy writing the article seems to be bad at math and thinking.

Can I imagine a venue kicking out 2% of their former clients on some criteria? Absolutely yes.

Kicking out 2% of website visitors may still be totally reasonable if the cost to serve them is meaningful, or if they are less than 2% of revenue.

His defense for 98% being bad is that some CSS thing people were arguing about only had 70% coverage on his website.

Our b2b dashboard didn't support Safari for a while at all and it was entirely not an issue because everyone had a simple workaround to just use Chrome and the dashboard wasn't really the main product.

USPS is also not paying for the pensions that they already incurred, and those are actually even larger ($20bn/yr).

Anyway, Congress did already rescind the requirement that USPS prefund the pensions (a 57bn debt forgiveness), but this is all just hiding the ball, the 9bn in pension liabilities being incurred now are real. Even if you don't prefund them now, an honest accounting says that these are part of your costs today.

A privatized USPS would likely get rid of pensions entirely.

I think the problem being solved here is largely one of waste.

USPS hides 9bn of unfunded pension obligations every year and underserves urban areas to subsidize rural areas.

Mail volume is also generally falling as everything moves to email, so it is getting both less profitable and less critical.

The US is a rich country, we can afford to waste a lot of money and not notice, and of course one person's waste is another person's easier job or subsidized service, but given the ongoing decline in the importance of mail (vs package) delivery, it's not clear that this is a particularly important utility for the government to maintain any more.

The biggest problem with obviously AI driven projects right now is that nobody knows how much expertise you have in this area (seems low given you don't know about systems that do very similar things), nor how carefully you thought about the problem and the implementation.

If I look at a project like Verus I know that experts thought about how to structure the system and it's concrete guarantees and semantics as well as the actual implementation.

So the result is that I trust Verus and I don't trust this.

You description of "design, then use goal and loop prompts to RALPH a feature" makes me even more convinced that this is slop because it sounds like you haven't thought deeply about every line of code and have delegated it to an AI that we all know makes mistakes. Thinking about the design is not really a substitute for thinking about the implementation.

People keep making AI OSS projects and expecting the same reaction as people had to OSS projects before AI, but pre-AI OSS came with a bunch of implicit promises about quality and effort and care that AI projects do not have.

I've interviewed several "I just read design docs and delegate all the actual thinking/analysis to the AI" engineers recently, and they are not grounded enough in reality to know when the AI is telling them something true or false and so confidently show me huge piles of code that purports to do things that I know are just nonsense. At some point I know these folks were good engineers, but they just drank too much of the kool aid.

To be clear, I have some of my own 100% slop projects where I have only vague ideas of what is going on beyond the high level design. But I have that currently put in a containment zone where I can easily verify the outputs or don't care about reliability (it's mostly UIs) and the scale of the damage they can do is minimal. Everything else, while basically still 100% AI generated at this point, is still reviewed and often repeatedly re-prompted to do precisely what I want the code to do, because if you let the coding agents make decisions for you, they still make bad decisions.

See, this goes back to the, all software engineers besides me are wrong, because I see this list and do not think it is anywhere close to a sufficient list for good quality software. The thing about all these criteria is that sometimes they are important, sometimes they are not.

This "standard" exists for the sake of code analysis vendors to be able to have some sort of shared taxonomy, but also provide a fig leaf of standardization to their products.

Most engineers are wrong (I obviously am the true arbiter of taste), but that doesn't mean there isn't better and worse code.

"Does it work" glosses over a bunch of things: is it fast, cheap, secure, reliable, easy to understand, easy to modify? And that's just for server software where you've nailed down all the functional requirements. Determining what the functional requirements is it's own question.

And all these other non-happy path requirements are somewhat in tension with each other, so what is ideal in one environment is not necessarily ideal in another.

And in particular, "easy to understand/modify" is truly subjective. Different people have different ideas of what easy to understand means. Even if we get to a world where AI is writing all our code, "easy to understand/modify for the AI" is still an important question. We've probably all seen prototypes that collapse under their own weight of slop by now.

Iirc Google's solution to this was to make the top of page shopping panel something companies could bid on and then conduct arms length auctions where Google Shopping and its competitors get to bid.

Presumably Google Shopping will be better at matching users to items and be able to big more than most on average, but 3rd parties can still develop an edge in some niches.

Anyway, the EU wasn't satisfied and fined them another 7bn. And obviously competitors like free traffic, not having to pay Google.

I think they have since started removing rich units for things like Flights/Hotels and trying to figure out what product they are allowed to actually provide in the EU.

But in general, they are just going to keep getting sued forever because they have a strong incentive to find the line on how to monetize search and obviously other aggregators do not like this and have the EU on their side.

In general, operating in the EU seems like a mine field where you have to accept that you're going to get shaken down regularly and do the best you can to thread the needle profitably.