HN user

Denzel

1,459 karma

LV

denzel dot morris one at that google service

Posts5
Comments441
View on HN

Writing out my nod of appreciation to counterbalance some of the negativity. I do enjoy the effort and thought you into this essay. I didn’t read the whole thing, just a few sections from the beginning, and I think you do a great job explaining how these concepts evolved from simple beginnings.

You’ve condensed down a lot of concepts I learned through research and trial-and-error over my career as a SWE.

Thank you for sharing Faza. Please continue putting in the effort and sharing these essays!

Yes, DeepSeek is open-weight, but these third-party providers offering similar prices are subsidized with VC money as well. And you can find a range of prices for deepseek-v4-flash going up to and over $1/Mtok.

Even that $1/Mtok provided by Together AI is heavily subsidized by more than $1B in VC money.

This makes it unclear how the true cost curve is progressing. It’s not possible to confidently comment one way or another on the rate that cost is coming down when the entire industry is so heavily subsidized.

Unless your insider is the CFO, I wouldn’t trust these sources to have the access, knowledge, or insight to determine whether they’re running inference at a profit.

Simple test: can they get their hands on a data center contracts and financials?

Search isn’t anywhere near as high profile as the profitability of AI inference, and yet, even aspects of the Search org were walled off from the rest of the company such that other employees couldn’t see what we see.

I say the following not to brag but to offer up some more perspective for you: you’re comp is in the lower-mid range of the market. Comp was there for remote, startup eng positions back in 2017.

As you move up in comp, the market actually gets more difficult, not just because the market is more competitive but because some companies won’t even interview qualified engineers with FAANG on their resume because they don’t believe they can afford them.

So I can understand why you might have an easier time compared to other engs.

Thanks for point-by-point.

Your first two quotes are about targeting in the Iraq War; specifically how the breakdown in careful analysis, precipitated by the new systems, led to the exact mis-targeting they were trying to solve. That’s what the entire article is about.

And your third quote is from an ex-official commenting on the event after the school strike happened.

These quotes contradict your original point, ie they show how careful analysis has been designed out of the system.

We killed young kids, but not on purpose. We targeted a building and intent matters. I refuse to believe anyone in the decision chain would move forward if they believed kids were going to be killed. If you do - how can you? Why would they?

This sounds incredibly naive. For starters, plausible deniability due to diffuse responsibility is a thing.

“Of course we don’t target schools and kill children, this was a system error.” But the message gets sent regardless and meanwhile we have people arguing back-and-forth over grains of sand because they took an action with deliberate plausible deniability.

For a historical analog that involved killing US children “unintentionally”, you can read up on the Ludlow Massacre - https://www.pbs.org/wgbh/americanexperience/features/rockefe...

Of course they didn’t intend to kill the children, they only intended to disperse the strikers by setting their tents on fire. It was simply a mistake.

You’re acting like the U.S. government is a monolithic good faith actor right now. The current administration’s behavior is qualitatively different than past administrations.

Do you also believe this administration will ever officially confirm Renee Good and Alex Pretti were not domestic terrorists?

It’s hard to interpret your points charitably here.

So I read the entire TFA, where do you see “quotes [from] those in the know who believe this should have been eliminated as a target”? I saw no such quotes about the school in TFA. Maybe I missed it.

there was precisely one mis-strike in 1000s of sorties

How did you verify this? Because I’ll remind you, the U.S. administration denied responsibility for some time before owning up to this due to public pressure. Absent public pressure, I guess we would’ve had zero mis-strikes.

so this already is a low error rate

As a father of similarly aged daughters, I can’t express enough how grotesque and disturbing the term “error rate” is here.

We targeted and killed young children. Plain and simple.

However, you have made a very, very strong assumption that these targets were not carefully evaluated.

Let’s take the opposing assumption that this target was carefully evaluated then. Please reason through the implications now?

How much did this cost? Has there ever been an engineering focus on performance for liquid?

It’s certainly cool, but the optimizations are so basic that I’d expect a performance engineer to find these within a day or two with some flame graphs and profiling.

Apologies, I may have misinterpreted the passage below from your repo:

This crate was developed with the assistance of Claude Opus 4.5 initially to answer the shower thought "would the Braille Unicode trick work to visually simulate complex ball physics in a terminal?" Opus 4.5 one-shot the problem, so I decided to further experiment to make it more fun and colorful.

Also, yes, I don’t dispute that human written software takes iteration as well. My point is that the significance of autonomous agentic coding feels exaggerated if I’m holding the LLM’s hand more than I have to hold a senior engineer’s hand.

That doesn’t mean the tech isn’t valuable. The claims just feel over exaggerated.

First, very cool! Thank you for sharing some actual projects with the prompts logged.

I think you and I have different definitions of “one-shotting”. If the model has to be steered, I don’t consider that a one-shot.

And you clearly “broke” the model a few times based on your prompt log where the model was unable to solve the problem given with the spec.

Honestly, your experience in these repos matches my daily experience with these models almost exactly.

I want to see good/interesting work where the model is going off and doing its thing for multiple hours without supervision.

Weird, I broke Opus 4.5 pretty easily by giving some code, a build system, and integration tests that demonstrate the bug.

CC confidently iterated until it discovered the issue. CC confidently communicated exactly what the bug was, a detailed step-by-step deep dive into all the sections of the code that contributed to it. CC confidently suggested a fix that it then implemented. CC declared victory after 10 minutes!

The bug was still there.

I’m willing to admit I might be “holding it wrong”. I’ve had some successes and failures.

It’s all very impressive, but I still have yet to see how people are consistently getting CC to work for hours on end to produce good work. That still feels far fetched to me.

Good points - admittedly, I didn’t put enough effort into building connections through different pipelines back when I was contracting. Upwork and a few personal connections were my sole sources.

It just felt really difficult to do both the engineering work while trying to do customer development at the same time.

The fact that OP has been able to do this for so long, while supporting a family, piqued my interest.

Cool, that potential 5x cost improvement just got delivered this year. A company can continue running the previous generation until EOL, or take a hit by writing off the residual value - either way they’ll have a mixed cost model that puts their token cost somewhere in the middle between previous and current gens.

Also, you’re missing material capex and opex costs from a DC perspective. Certain inputs exhibit diseconomies of scale when your demand outstrips market capacity. You do notice electricity cost is rising and companies are chomping at the bit to build out more power plants, right?

Again, I ran the numbers for simplicity’s sake to show it’s not clear cut that these models are profitable. “I can sort of see how you can get this to work” agrees with exactly what I said: it’s unclear, certainly not a slam dunk.

Especially when you factor in all the other real-world costs.

We’ll find out soon enough.

It's that we're paying more for objectively worse service than we had a decade ago.

I'm not asking for magic, I'm asking where went the reliability we already had, at the prices we're already paying.

My god thank you! My partner and I have been talking about this for the past 2 years in the context of food service and delivery service industry.

Greater than 50% of all our restaurant orders are straight up wrong or missing items, whether it’s from local places, chains, or fast food restaurants.

The unreliability is staggering, especially because we’re paying so much more!

It’s gotten so bad that we’re done with certain services and establishments for good now, or we make sure to QC before leaving the restaurant to ensure everything is in the bag.

Even more ironic, this happened a couple weeks ago at Texas Roadhouse — the same restaurant I worked in decades ago as a teenager, so I remember the process we had to go through for to-go orders.

First, we’d take the order over the phone. We’d repeat the order back to the customer to confirm everything (1st QC). When the food came up in the window, we’d pack the food in bags, crossing off every item on the receipt before stapling it to the bag (2nd QC). When the customer came to pick up their food, we’d have to take every box out of the bag, show the customer the food, and confirm that everything they expected in their order was there (3rd QC).

No customer. Every left. With an incorrect order. Simple.

That process is gone now. We paid more and came home missing my partner’s meal. Wtf.

Uhm, you actually just proved their point if you run the numbers.

For simplicity’s sake we’ll assume DeepSeek 671B on 2 RTX 5090 running at 2 kW full utilization.

In 3 years you’ve paid $30k total: $20k for system + $10k in electric @ $0.20/kWh

The model generates 500M-1B tokens total over 3 years @ 5-10 tokens/sec. Understand that’s total throughput for reasoning and output tokens.

You’re paying $30-$60/Mtok - more than both Opus 4.5 and GPT-5.2, for less performance and less features.

And like the other commenters point out, this doesn’t even factor in the extra DC costs when scaling it up for consumers, nor the costs to train the model.

Of course, you can play around with parameters of the cost model, but this serves to illustrate it’s not so clear cut whether the current AI service providers are profitable or not.

We probably work at the same company, given you used MAANG instead of FAANG.

As one of the WAU (really DAU) you’re talking about, I want to call out a couple things: 1) the LOC metrics are flawed, and anyone using the agents knows this - eg, ask CC to rewrite the 1 commit you wrote into 5 different commits, now you have 5 100% AI-written commits; 2) total speed up across the entire dev lifecycle is far below 10x, most likely below 2x, but I don’t see any evidence of anyone measuring the counterfactuals to prove speed up anyways, so there’s no clear data; 3) look at token spend for power users, you might be surprised by how many SWE-years they’re spending.

Overall it’s unclear whether LLM-assisted coding is ROI-positive.

‘Desirable difficulty’ is the research term. To solve your problem, first understand your users need a mindset change. We need to connect their action to a “satisfying feeling” as you said.

You want your users to be like weight lifters. No lifter comes out the gym saying, “Man that was the best workout, felt so easy,” to the contrary, lifters use progressive overload to induce difficulty because that difficulty connects to the results they want.

For your users, you need some way to measure the outcome, so that you can show them, “hey look, that mild discomfort lead to more progress on what you care about,” and then you need to consistently message that some difficulty is good.

Mindset change takes consistency and time. Won’t happen over night. You’ll know you succeeded when students become aware of “hey, I’m not learning as well if it doesn’t feel difficult”, and then react by increasing the challenge.