HN user

Jcampuzano2

2,013 karma

Software dev

Posts0
Comments478
View on HN
No posts found.

Sure that can be called an attack, but then we must also concede these labs essentially massively attacked everyone else in existence to get the data, and continue attacking as we speak.

In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.

Obviously a bit more complicated than that but it still holds pretty well.

To be fair in every harness I've ever used that has LSP support, they never actually utilize the LSP for more deterministic refactoring tools. And even then when I do enable the LSP in many harnesses oftentimes it doesn't even use it at all.

Maybe they haven't been taught to do so or it's not integrated into the system prompt or the tools but all of them only ever use the LSP to read files/symbols.

Every harness I've used will happily just call the edit tool over and over or do a find and replace via sed or programatically call a python/perl script rather than rename a symbol via other means.

Define "failed".

If what ends up happening is that every listing has misleading AI photos but they have to disclose it, then also what ends up happening is nobody trusts them anymore. Consumers will know by default to not trust the photos.

In my eyes, thats a win since that's a better outcome than them secretly using AI photos.

Of course in my ideal world it would be outlawed altogether, But even if they were still allowed to use AI photos but were forced to disclose it, that's still a good first step.

They're giving everyone their next hit.

What I read on social media about people and these resets gives off literal worst kind of addiction vibes. I've literally seen people talking about "Oh I had an existential crisis without Fable/GPT-5.6"

These people legitimately need help, or alternatively a social life.

Maybe its different on my end because I just use a sub outside of work for fun stuff. At work its not my money so I don't really care. I go to work, maybe use these subs at home every once in a while for a fun personal project and if I hit the limits (I rarely even do) I play video games or hang out with my wife/family.

People are borderline tying their identities to these models it seems, and yet most people aren't even building anything interesting.

Okay I stand corrected then.

Seems strange for a company of your size to have one person push changes that should have easily caught this edge case then. Seems like a change even a small handful of people could have reasonably thought up this side effect of.

It's Thariq from the Claude Code team here. This was my change! I made the AskUserQuestion tool so am generally in charge of maintaining it.

In a sense yes, I think it is actually reasonable to complain that the answer is too human/individualized here because it likely wasn't this individual human who made this decision, but he's making it seem like it is so that we are less likely to blame the company as a whole.

It's counterintuitive but when one singular person owns up to the problems that, at the root, are actually systemic to the decision making of the whole company it plays on the psychology of us as humans.

"I'm in charge of maintaining it" - This is not the same as "I'm in charge of all of the decision-making behind the implementation of how this tool works for users".

I actually agree exactly with your last point that one single person taking blame is counter-intuitive/non-productive here, but it actually seems like what these large companies desire is to have one person be the fall guy to play on people's sympathies.

If this were some small startup it would make sense but this is not that case.

GPT-5.6 13 days ago

Did you not read the second sentence? Obviously I know what sol is given my first language being Spanish. I'm just speaking in a general sense that it can be confusing for others.

I already know plenty who had no clue what the difference between Terra and Luna would be.

GPT-5.6 13 days ago

Previously it was much more obvious which model to reach for depending on your use case because they had the mini and nano naming conventions.

Getting rid of that seems like a step back. Just a personal nit though.

I've seen buzz about this elsewhere as well but to me effort levels seem more like spend limits disguised with another word. I don't think they should even exist.

GPT-5.6 13 days ago

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming.

The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confusing.

That aside all the numbers look amazing, and I'll be happy to probably main this alongside grok-4.5 for a while comparing the two on price and efficiency.

I vastly prefer the direction that OpenAI seems to be going with token efficiency and performance compared to Anthropic who seems to be moving towards a world where you just token-max as much as possible ignoring any and all costs.

Muse Spark 1.1 13 days ago

Competition for cheaper and efficient models is a good thing, regardless of if you don't like SpaceX, Meta, etc. Especially from US based labs

I for one am really glad to get competitive models that will push the major labs to bring prices down. While Chinese open source labs are also great, unfortunately when it comes to US/Western political pressure it won't often have as much of a bearing on labs bringing prices down, especially for enterprises.

Also if these numbers are true, this is truly breaking ground finally for Meta.

Show HN: 18 Words 13 days ago

I'd probably say there instead of a Challenge Mode and a Relax Mode like you said, it could just be a combined mode where there is a timer but after it goes out it simply continues the game on Relax Mode.

Or alternatively every word still has the timer and then at the end if you finish, it tells you how many words you completed under the timer and gives you a score based on that.

And then maybe an option for those who don't want the timer to show at all, since maybe it adds a bit of pressure. You can have just a simple option that removes the timer entirely from view

I mean its the same thing as the data center investments.

Think whatever you want about them, whether they're good or bad when it comes to environment, public health etc.

But one thing cannot be ignored - that they are not built to employ some large swath of people. They can be run with very lean teams, much leaner than the average person thinks for something so large. Any claim that they are employing some measurable amount of people is a sham they try to push onto the public.

If we actually punished corrupt officials and we had some kind of truth serum that forced people to admit yes/no as to whether they are corrupt, I would not be surprised if the majority of officials in the federal government would be culled. Its practically a breeding ground for corruption.

I keep trying to convince directors and executives at my company to look past the cost per token amount but they refuse to do so. Those are the only things that actually give any sort of measurement of the monetary value of a token by these labs, and so its what many go by.

For example there's some benchmarks that show that Opus for any task that requires a higher than `high` level of effort, may have actually been cheaper to use Fable on low even though the cost per token is drastically higher

Similarly with GPT 5.5 vs Opus. They simply look at the dollar amounts the labs assign to each model and run with it.

But part of the issue compounds on the fact that there are many people who simply default to the smartest model/effort and don't actually vary their model per task. So in some sense I don't actually blame them very much.

I'm not gonna lie, I chuckled a bit reading this.

This hasn't been the case for at least a decade now, if not more.

First it was extended out to maybe once every 2 years, then more, and lately at every company I've worked at (primarily large companies) where pay was mentioned the response is "we pay at or above market rates - discuss with your manager."

Claude Sonnet 5 22 days ago

I'm struggling to understand why I'd ever use this instead of just using a lower effort level for opus given on many of the benchmarks listed the cost per task rises above opus at anything higher than medium effort.

Only thing I can think of is for when someone is out of opus credits. Of course there are API billing use cases but I'd probably still just use opus on low.

Not sure why this is being downvoted.

A draft is by definition forced labor and a form of slavery. If you are drafted, you are being forced to do something you did not volunteer for.

If enough people volunteered there would be no need for a draft in the first place.

People can argue whether they think its "correct" or not all day, its still forced labor.

There certainly are arguments against a draft at the individual level - Example being if the people being drafted do not believe in their own governments ability to lead.

A draft is by definition forced labor either way - because if enough people volunteered there would be no need for a draft in the first place.

But usually it doesn't matter either way. Either the people believe in their own government/territory and see it as worth defending, or alternatively they don't believe in it, but their government is authoritarian and will force you at risk of punishment or death into being drafted.

In both cases a draft is still forced labor for those who did not volunteer to participate.

Nowhere in my argument do I contend it may not affect myself. In fact I have basically already accepted its very likely I'll be replaced due to AI in the very near future. Thats just the unfortunate reality of things at the individual level.

So yes, I do agree it sucks for lots of people living in the moment.

I mention in some other comments that yes, AI "visionaries" make the level of replacement seem to be on a scale almost never before seen and so the "benefit" for the majority would have to be absolutely massive (however we define benefit). And currently its hard to see how it could reach that level. I was just noting we cannot "only" see it through the lens of replacement. If for example billionaires (trillionaires now?) did actually spread the benefit and we overhauled the economic systems in much of the world for humanity it _might_ actually be a benefit. Its just hard to see this ever happening given history.

I definitely have not "gotten mine" like the billionaires pushing AI. But other inventions in hindsight have very clearly benefited humanity as a whole even with the unfortunate effects on the people of the time.

I wrote in a separate comment elsewhere that I don't really disagree that the current top brass of AI push it as something that would replace people on a much larger degree than most other technology, but that in my opinion the argument does hold - where if we did see it as a net benefit greater than the loss it would still be worth it, and that is the measurement we should be using. But yes, given the level of replacement the net benefit would have to be absolutely massive.

And yes, "net benefit" is hard to measure for an unrealized/developing product.

So I don't disagree with you. In the current economic system where we need human labor (in the majority of the world) to make a living, its hard to see the current vision of AI by those in charge to lead to anything but mass suffering.

AI will quickly turn the world into an even greater disparity between the "haves" and the "have nots" with its current vision.

I agree was going to comment that yes it does feel slightly different because in most cases technology narrowly targeted specific niches at the time, that when replaced could at least in hindsight be seen to likely benefit the majority at the cost of the existing laborers.

Whereas in this case most of the top brass of AI do push it as something more akin to "we think this will replace practically all human labor". And without the availability of human labor, at least given the current economic system its hard to see how that'd lead to anything but mass suffering.

I do think the argument still holds. If we were able to see it as a net benefit to all, it would still be worth it. Its just that with the level of replacement we're talking about the net benefit would need to be massive (however we define "benefit") The problem is there is plenty of research showing it is still net negative in many cases, especially (in my opinion) when it comes to cognitive ability and early stage development for children/youth.

The closest similarity may be the development of the personal computer or something along those lines.

Why da f*#$ do they have to continue developing a technology which they think will replace droves of people by machines?? There is nothing sexy about it. There is nothing cool about it either.

This argument alone doesn't really work because literally billions of peoples jobs and livelihoods have been replaced throughout history from the explicit development of technology that replaces the need for human labor.

The question should not be whether the technology replaces people by machines, it should be whether it provides a net benefit overall. You could use this argument to say we should have never invented the printing press if you only thought of the people who used to manually transcribe books and documents.

The Coming Loop 29 days ago

I personally have not had good luck with loops due to similar issues as the post author - but if you were to port your flow to "looping" it would be something like:

- An automation that periodically checks for PRD's at a given location that have not yet been implemented.

- If it sees one not implemented, it puts a lock on it (so other agents later don't pick it up while its still working) and implements the PRD in code, assuming it has the figma link and all specs required.

- When its done it makes a PR, waits for if it passes and even in some cases automatically merges into your staging/preview enironments and just pings you with a build/URL. You can then leave feedback or something and it can also also poll for pending feedback. Or you just mark it looks good, the agent then merges the PR, moves the PRD to implemented status, maybe even writes/updates docs and cleans up any temporary work.

- Repeat checking for new PRD's every T unit time. (10 minutes, 1 hour, etc)

This is how people say you should be looping - you never even cared or looked at the code, and also never prompted the agent yourself.

But I find most agents are often pretty bad still at replicating UI vs making something from scratch and most design specs are still not as detailed around how things look at all sizes, in all scenarios etc. Design seems to be one of those things that still requires a human to validate. And then all the things the post author mentions about it not being willing to apply hard constraints, minimize impossible states, validate at edges and prevent horrendous overchecking of things. etc.

Well maybe don't run the call center at all, and actually make things fixable yourself without interacting with a human/LLM.

Example: At least here in the US plenty of companies still require calling in to cancel. Include that by default as a user flow/feature (and we're getting better but many utility, gyms and other companies still require calls) and boom, you've gotten rid of probably 50+% of call volume in many places still requiring this.

But of course they want the best of both worlds as you describe. They want to inconvenience the hell out of you for things that are 100% able to be implemented as just regular user flows so they can get people to just drop it while also saving money on these flows with LLM's.

The problem is a problem of choice I believe.

When we use AI ourselves via tools like chatbots, harnesses etc. we are mostly actively choosing to do so, and have some control. We can always just decide to stop and do the work ourselves if its not working out.

In the call center/situation of companies embedding it in their products, often its not in a way that gives users the choice. They are forcing it onto their users with no other option, or at the very least they are always forced to play along with the LLM until it finally gives up.

Its user hostile since we can't decide to break out of the LLM loop when we want to.

Add on top of that most of these companies are actually forcing the use of the AI related features simply to fulfill someones KPI's/internal metrics.

I actually never mentioned anything about actually using the AI tools integrated into Cursor in my post.

I think I'd generalize my post more to say the more often somebody reaches for the terminal, in my anecdotal experience the more proficient they tend to be.

Never said I was judging, just making an observation. And to answer - yes by book you would be an outlier.

Its just an anecdotal experience.