I agree, this is clearly an indictment against LLMs. If LLMs and agents were capable they'd 100% write it natively but they realize the current limitations.
HN user
avbanks
No it's because LLMs aren't that good yet.
I second this, OP I do truly emphasize with your situation (I graduated right before the 2008 crash and had to join the Army). You should look into creating a startup or make significant contributions on a major open source project. Don't think that just because you're fresh out of school VC's won't be willing to fund you. If you're dead set on being a cog on a wheel feel free to send me your resume and I can give you some feedback.
I've been trying to articulate this exact point. The problem w/ LLM's is that at times they are very capable but always unreliable.
I'm starting to realize that this is most likely what will happen. They'll be available in select major cities, for certain areas, under certain weather conditions.
The drivers who can't handle the edge cases face the justice system. If You or I did that we'd face repercussions.
How exactly will driverless cars work, the edge cases are infinite. They would have to be limited to certain routes, conditions, and cities.
This tutorial was extremely helpful, I was having a hard time grokking the power of zippers.
I don't think people fully realize how good the open source models are and how easy it is to switch.
I've noticed this with TikTok and I'm almost certain YouTube 1P metrics are wildly inaccurate in particular views and non-bot comments.
A lot of big tech companies are being very opportunistic and reducing hiring/laying-off under the guise of A.I. but really it's weak economy.
Its incredible to see that was once possible.
Congrats! A great achievement.
That makes a massive difference, if true.
This is such an overlooked aspect of coding agents, the code review process is significantly harder now because bug/vulnerabilities are being hidden under plausible looking code.
This is why music festivals are so popular.
He's been well ahead of the curve on this, he's been saying for years that scaling/synth data won't work.
This is exactly how I've been seeing it. If you're deeply knowledgable in a particular domain like lets say compiler optimization I'm unsure if LLM's will increase your capabilities (your ceiling), however, if you're working in a new domain LLMs are pretty good at helping you get oriented and thus raising the floor.
This imo is the biggest issue, LLMs can at times be very capable but they always are unreliable.
Capability != Reliability
The effects of Section 174 seem to be understated, it aligns with the layoffs and the size of the layoffs.
This is exactly what I've been trying to point out, while LLM's and coding agents are certainly helpful, they're extremely over-hyped. We don't see a significant bump in open source contributions, optimizations, and innovation in general.
People will always contribute to open source. If AI agents are so good why aren’t people building open source projects around agents? The computing power of many agents would greater than that of a sole agent. As of right now we’re not really seeing anything the sort.
If the writing is on the wall shouldn't we be seeing a massive boost in open source contributions? Shouldn't we be seeing a spike in new kernels, operating systems, network stacks, database, programming languages, frameworks, libraries...?
I still find 3.5 Sonnet the best for my coding tasks (better than o1, o3-mini, and R1). The other models might be trying to game system and fine tune the models for the benchmarks.
LLM based AI tools are the new No/Low Code.
This is actually a very interesting insight, not only do you have to worry about sponsored results but people could game the system by spamming their library/language in a places which will be included in the training set of models. This will also present a significant challenge for security, because I can have a malicious library/package spam it in paths that will be picked up in the training set and have that package be referenced by the LLM.
A lot of people in community are weary of benchmarks for this exact reason.
This is still being litigated I believe.
IMO the first argument is invalid, however, the second one is a completely valid argument.