LLMs aren’t lazy. They don’t cut corners because a simpler solution feels good enough. If they know how to solve something thoroughly, they will.
This is a severe misunderstanding in how LLMs work..
I don't know how this got on my front page....
HN user
LLMs aren’t lazy. They don’t cut corners because a simpler solution feels good enough. If they know how to solve something thoroughly, they will.
This is a severe misunderstanding in how LLMs work..
I don't know how this got on my front page....
I'm more interested in the business impact of this
So you spend billions of dollars training the model, only for it to be used in the US.
Then interesting to see where most of anthropic revenue comes from. If it's the US then they're fine but if it's global then they'll see a drop in revenue?
Then add to this decision, companies are going to significantly reduce their token spend.
So what does all of this mean for their IPO?
Super appreciate you replying to my comment.
I think I understand where you're coming from now. What confused me is that the post is written in a way that it seemed like what Fable was doing was actually better. Maybe I should've looked at post as an exploratory post on Fable instead.
How can a LLM be assigned an emotion as being "proactive"
I can't edit my post, this is wrong. "Proactive" is defined as a behaviour instead of an emotion.
Thanks to everyone pointing it out!
How can a LLM be assigned an emotion as being "proactive". This is highly misleading to anyone that scans just the headlines.
What actually happened is that the user started a prompt, and Claude took $12 worth of tokens to resolve the issue. How it did so was basically looping until it got to the answer
How is this proactive? It's literally being token greedy and maximising revenue for the LLM owner. People really need to be putting on business hats at this stage, because we are being lead to believe that "more tokens = better". It is not, there are efficient ways to solve a problem and there are inefficient ways to do so too.
Each problem solved incurs a cost, and is expected to yield an ROI at some point. This is how we should be viewing things now.
For me it depends on who you listen to
If you're following a bunch of people who are from LLM labs, you're going to be more incentivised to tokenmaxx because it's in the Lab's best interest tonget you to behave that way.
Practically, many companies aren't labs with endless runway. Companies hopefully follow a PnL model. And when you look at things with that lens, many of the times the LLM use case falls apart.
You're seeing a bunch of companies starting to realise that tokenmaxing yields very little ROI.
Even the LLM labs, the guy that spent $1+mil tokens has nothing to show for it in terms of revenue to the company. And you have to keep sinking that much into AI for ... "features".
There are some good use cases for AI. I ended up with a positive ROI on a greenfield project myself, albeit on a small scale.
The way that AI has been making people have totally irrational decisions on executive, pure business and technical standpoints is simply mindblowing. I don't understand how people can't take a step back and see what's actually happening from a macro perspective.
it's a net positive for robinhood, not the trader necessarily
Product market fit and profitability are two different things.
Arguably product market fit was fond last November already. I don't think agents were the turning point that caused this.
Profitability, not yet. For me, it depends on whether companies are seeing a positive ROI from their ai investments. This website is skewed more towards big tech companies, but everyone that's not big tech needs to see positive ROI with using ai tools. Short term might see a profit, but medium to long term we still need to wait a bit.
So the article doesn't mention how the individual should prepare, but rather that government should prepare. Are individuals just that powerless against these perceived outcomes with AI?
My opinion is that AI can be a force for good, but why is everything so aggressively framed as a class war? Why must such a path be taken?
I wonder if this is how things were during the industrial revolution as well
Why would you replace an existing codebase like this instead of forking the repo instead and then making the changes?
Why is there only a fine and not also the seizure/forfeiture of the stolen property and the derivative works/products built on it?
Like the fine means nothing to meta, and they'll still be the beneficiaries of their infringement.
In this current state, you really just need to have enough money to bypass this lawsuit and be on your way.
There's probably a few things to consider
* How much CPU/token usage does openclaw users use in general? Similarly, how much does high volume openclaw users use vs "normal" claude high volume users?
* Are there political elements we can't see that's affecting this? OpenClaw and anthropic doesn't have a good history in general and this is just a continuation of that?
Something I don't understand, there's a lot of complaints yet people are reluctant to stop using the service? Are folks already vendor locked or is it a case of "well, this doesn't seem to affect me?" The consumer behaviour of these complaints is very interesting.
When I went through tough breakups? I lost myself in open source... on GitHub. During college at 4 AM when everyone is passed out? Let me get one commit in. During my honeymoon while my wife is still asleep? Yeah, GitHub. It's where I've historically been happiest and wanted to be.
I've never had such an obsession to a platform or an activity as this. Some might say this is unhealthy, but I admire folks who can reach this level of obsession in their craft. It's just a joy to read about for me
It looks like it's this person's fault?
* you can't blame ai if your production token is on the same machine as the staging/ development environment?
* you can't blame ai if you didn't know that the production api token gave access to all apis.
Like if this is the level of operational thinking going into this app, then I'm sorry no ai agent or platform can prevent this from happening.
Everything else in this "post mortem" is performative at best.
The only real question one could ask railway is why do they have api endpoints that can affect production available? Maybe these should only be performed on the platform itself instead?
Once I wrote the perfect piece of software. It was so perfect that there was literally no bugs for months.
How could this have happened? Well, the code was shipped but no customer was running it in production.
I find HN to be a bad resource to ask for learning resources. I previously asked for help in learning how claude works but no responses.
Maybe one pointer for others is that people are genuinely curious about learning new things, but as experts we choose not to engage these types of posts, why is that?
OP, in your case you need to move from theory into actually building your own agent and make it do things. Start by solving small problems and then make them more complex over time.
The headline seems to be flashy indeed, but ai didn't really solve this imo.
They just seemed to fix their technology choices and got the benefits.
There's existing golang versions of jsonata, so this could have been achieved with those libraries too in theory. There's nothing written about why the existing libraries aren't good enough and why a new one needed to be written. Usually you need to do some due diligence in this area, but no mentions of it in this post
In order to measure the real efficiency, gnata should've been benchmarked against the existing golang libraries. For all we know, the ai implementation is much slower.
The benchmarks in the blog are also weird. The measurement is done within the app, but you're meant to measure the calls within the library itself (e.g calling the js version in its isolated benchmark vs go version in its isolated benchmark). So you don't actually know what the actual performance of the ai written version is?
The only benefit, again, is that they fixed their existing bad technology choice, and based on what is observed, with a lesser bad technology choice. Then it's layered with clickbait marketing titles for others to read.
I'll probably need to expect more of these types of posts in the future.
I think Sora was technically impressive as a concept. The way it was managed as a product wasn't good.
There didn't seem to be any marketing for it. Like I can't even remember an ad for it or any content creator type of person pushing Sora actively.
To get access to Sora I believe you needed to be on a paid plan?
It's really difficult to get user generated content going when it's behind a paywall.
It's also hard to tell if this means that openai is in trouble, or if this is just a badly managed product that deserved to be killed. With the negative sentiment on openai, folks might think the former.
I would add that getting customers, especially paid customers for your app is not easily solved with ai too.
Only a few get lucky with funding, only a few have a profitable business.
It's very interesting to see the opinions in the answers vs the off site opinions.
On the site itself, the opinion seems to be that stackoverflow is not dead and it's okay if there's reduced traffic and questions to the site.
Yet off site, the opinion is that stackoverflow is dead and added to that, the community isn't very friendly anymore.
It's just really interesting to see the perspectives between the two.
In the last 60 days I have written over 600,000 lines of production code — 35% tests — and I am doing 10,000 to 20,000 usable lines of code per day as a part-time part of my day while doing all my duties as CEO of YC.
LOC will never be a good metric of software engineering. Why do we keep accepting this?
I can generate 1 million LOC if I really wanted to.
As long as LOC is the main metric for these setups, they will never be successful.
I think the advice is good but maybe the title could be improved.
But for Engineer A’s work, there’s almost nothing to say. “Implemented feature X.” Three words.
To me, this is the main problem. Engineer A is unable to describe the impact of their work, how the work affected the business. Your manager isn't responsible for promoting your own work, you are.
Engineer B’s work practically writes itself into a promotion packet: “Designed and implemented a scalable event-driven architecture, introduced a reusable abstraction layer adopted by multiple teams, and built a configuration framework enabling future extensibility.” That practically screams Staff+.
Maybe it's just the narrow of the article, but if promotion only looks at complexity and not quality of delivery and impact on the business then this isn't a good engineering team to be in.
There are many cases where simplicity is celebrated and recognized. It's up to the engineer to know what the impact of their work is, if they can't do that then that's on them.
If you are a senior engineer, the bottom line is that I wouldn’t recommend the jump to management right now. I would wait a couple of years to see how things will look like.
BUT, and it’s a big but - if your gut tells you to do it (and not your brain), if it’s truly a path you want to pursue - then go for it!
It feels more like the purpose of this article was to get the sponsored segment out than to actually give useful advice. Like how is this the conclusion?
For my friend specifically, staying on the IC track, becoming a Staff engineer and switching companies would have given him ~20-30% more than the EM promotion he was offered.
Company promotions do not give a higher salary bump than moving companies. The friend could be at a company that pays less for all roles. Additionally, that visualisation does a low-high representation and doesn't take outliers into account. Staff engineer roles tend to have outliers when it comes to salaries. EM roles do not
If anyone wants some advice from an engineering director
* If you only want to become an EM for the money, you probably won't like it. It's the same as an engineer that's only coding for the money. The more you like something, the more you would want to learn it
* The EM title means different things at different companies. Some companies are only/mostly about line management duties. In other companies, you're expected to do project + stakeholder management. In other companies, you're also expected to do operations, budgeting and technical + business strategy. As you can see, it's different to an IC who is building software and there's more of a focus on the things around building software.
* Being hands on is one thing. But what distinguishes one EM from another is engineer empathy. If you're an EM on the team and haven't did a PR (with or without ai), then you have zero empathy for your engineers because you have no idea what it takes to build a feature for your team. Using LLMs improves engineer empathy, but you need to learn it despite it.
* AI/LLMs will change two main things: the ability for an EM to be more hands on and the way EMs design team processes. Just like it changes engineer's ability to code, the EM needs to think holistically on how the development process will change and adapt accordingly. Do you have a path for the team to use AI agents? Do you have ways to reduce meetings and achieve the same level of alignment with LLMs? This is the type of thing EMs will/should be thinking about.
* The career path of an EM is largely dependent on the growth of a company. You will only get "stuck" if your company is not growing. If a company grows, there will be a need to hire engineers, then hire someone that manages those engineers and eventually someone that manages those managers.
* The other thing about EM careers. Advancement also depends on how well you are fitting into the business. For small companies, being more hands on as an EM is better. For larger companies, fitting in well with the company values, culture and leadership principles of the company is better.
I really don't appreciate the author's lack of understanding on how engineering leadership works and the general gatekeeping in this article. Sure AI is changing things, but there's really no need to steer people away and gatekeep roles like this role implies.
OpenClaw has nearly half a million lines of code, 53 config files, and over 70 dependencies. This breaks the basic premise of open source security. Chromium has 35+ million lines, but you trust Google’s review processes. Most open source projects work the other way: they stay small enough that many eyes can actually review them. Nobody has reviewed OpenClaw’s 400,000 lines. It was written in weeks with no proper review process.
Yeah, but the world rewarded this by making it the fastest growing github project. The author gets on the podcasts, gets the high profile jobs from big tech. I'm more encouraged to do things this way than being security minded about all this.
And there's no accountability to this at all. If an agent leaks private data, the user is to blame and not the author. If Google bans your services for using api keys incorrectly, we cast the bad eye towards Google and not the maintainer than enabled and approved it.
There's just so much incentive for for "not reading code" and not developing secure code that is just going to get worse over time. This is the hype and the type of engineering that we all allow either by agreeing or by staying silent.
I agree with the author, but the world works off a different set of principles than what we're used to. I just see the world blindly trusting agents more.
Openai employees already get crazy salaries. What motivates someone to do this?
I would understand a low salaried person doing this, but not someone from a really high paying org
This is pretty fun!
I'm interested in what will happen if you replay the prompts with different LLMs and the same LLM. I wonder how different the games will become?
It has vibe code and dogs in the title
Code has always been expensive. Producing a few hundred lines of clean, tested code takes most software developers a full day or more. Many of our engineering habits, at both the macro and micro level, are built around this core constraint.
...
Writing good code remains significantly more expensive
I think this is a bad argument. Code was expensive because you were trying to write the expensive good code in the first place.
When you drop your standards, then writing generated code is quick, easy and cheap. Unless you're willing to change your standard, getting it back to "good code" is still an equivalent effort.
There are alternative ways to define the argument for agentic coding, this is just a really really bad argument to kick it off.
Feels like this is very hard to measure isn't it?
Are we saying that llm's have zero economic growth, or are we saying that the sum of winners and losers in llm usage are zero or less than zero?
I think there's many examples of llms resulting in winners, or maybe this signal is just very high in the tech space.
But maybe there's not enough reporting on the losers in llms at the moment? (E.g. did llm displace their jobs, they have llm use cases that failed, etc, etc)
So the timeline is basically
* User uses Google oauth to integrate their open claw
* user gets banned from using Google AI services with no warning
* user still gets charged
If you go backwards, getting charged for services you can't access is rough. I feel sorry for those who are deeply integrated into Google services or getting banned on their main accounts. It's not a great situation.
Also, getting banned without warning is rough as well. I wonder if the situation will be different for business accounts as opposed what seems like personal accounts?
The ban itself seems fair though, google is allowed to restrict usage of their services. Even though it's probably not developer friendly, it's within their rights to do so.
I guess there's some level of post mortem to do on the openclaw side too.
* Why did openclaw allow Google anti gravity logins?
* The plugin is literally called "google-antigravity-auth", why didn't that give the signal to the maintainers?
* Why don't the maintainers, for an integration project, do due diligence checks on the terms of service of everything you're integrating with?