This is interesting, would being labelled "applied" be considered an insult to them?
HN user
suttontom
I'm glad it's a few short years instead of a few long years.
I don't think the poster meant applying the death penalty to the corporation.
Agree with you on exercise. But staring at a painting is to meditation what curls are to building your core. People all recommend the same type of meditation for a reason--so that you can close your eyes, focus on something, notice when your mind drifts, and bring the attention back. You didn't mention anything about awareness or noticing when you lose focus, so your advice seems to completely miss the point.
It's crazy that you could actually use the excuse that since it's all vibe-coded, there's no way a human could have written it, so Anthropic bears no responsibility.
Meanwhile humans can pop in and leave little morsels like this and blame it on the model.
These AI images also add to the public mistrust of AI, a growing problem for innovation in a field that is sometimes seen as biased, opaque and extractive.
Oh my, how would anyone ever have gotten that impression?!
Your instinct is correct, and in a lot of cases it's true. However, I've heard from enough doctors by now (a cardiologist, psychiatrist, and epidemiologist/former physician) that they use medical LLMs and find them extremely helpful, mostly as a way to either bring up knowledge they'd forgotten about or as a way to learn something new and then verify it. I'm extremely skeptical about LLMs in general and the connection to Gell-Mann Amnesia is apt, but I wouldn't necessarily write them off completely like that. There are experts using the models that find them genuinely helpful in their field.
How did it help?
It's a blog. He's using hyperbole. It's not a Supreme Court opinion.
I'm asking genuinely, is there a connection between housing, education, and healthcare becoming so much more expensive and them also being the three parts of the economy that have the most government interference (in the US)? If so is it causal?
I don't want to be cynical, but maybe spending hours every day using Claude has made some of us particularly attuned to picking this up. For some reason as soon as I read "The trap was in app/test/index.js," I instantly knew it was Claude. It's too bad, because there will obviously be some false positives, but it makes me immediately disregard the author.
You're wrong in lots of ways.
Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/
The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150
I wouldn't agree with that. The issue with software is that the people you make things for are usually anonymous and you'll never meet them, but if you've ever built software that helped someone and you witnessed it, it feels really good.
This is such a tired, meaningless argument. I've never seen a human in 10 years of professional software engineering at a large company ever so confidently, consistently create and send out seemingly well-reasoned code that's as wrong as what SOTA models using CC or Codex do. If a human did this, they would be fired or perpetually remain a junior who no one wants to work with.
Also, if a human does this, you can replace them and get a human who will not do it. The default for an LLM is to generate plausible-looking text that may or may not be completely incoherent. That is not the default for a human. Again, if you find that your colleague consistently fabricates APIs, you can hire someone who isn't crazy instead, but you cannot do the same with LLMs.
This is commonly known as "LLM-as-a-judge" and anecdotally multiple people I know who write code using OpenRouter or using multiple models say it's surprisingly effective. It's strange that there don't appear to be any major papers on it since ~early 2025, which at this point is basically ancient history.
Ah yes, the magical equivalent of "you are a senior software engineer who writes bug-free code".
IME people would benefit greatly from the process, albeit tedious and time-consuming, of testing out the same prompt sequence/session with the exact same model multiple times. It becomes clear extremely quickly how capable but unreliable and inconsistent a model can be even when given the same context. If you have ever completed a long, complicated task with an agent and then lost the session and tried doing the same thing again from scratch you may have had the experience of seeing the subtle changes that come up in the model's thinking which lead it to accept or reject certain paths and ignore or incorporate prompt instructions like the one you've provided.
Isn't that kind of what they're doing with this rollout? Except they're just hand picking the companies.
What is your problem? Do you think something is an opinion piece just because it has a byline? What about https://www.forrester.com/press-newsroom/forrester-impact-ai...? Is there literally any evidence you'd accept?
You know companies lie and overstate things, right?
Do you know if anyone has trained, say, a pre-2017 model and tried to get it to come up with Attention Is All You Need? If it did, would you say that was only because it's a synthesis of prior art? If so, what isn't?
Are you joking? Is there literally "nothing" you can imagine that Claude can't do?
This is a good example of being bad at writing code.
Not to be cynical but do you think this would matter at all? Are you saying that companies would hold themselves to their missions or even something that's legally binding?
"Google is not a conventional company. We do not intend to become one."
OpenAI being founded as a nonprofit and becoming for profit.
Didn't Anthropic literally say they wouldn't train on your data or keep it for longer than 30 days unless legally required, and then decided to opt people in to having their conversations used for training?
Am I going crazy? Is a PR with 94 commits that adds 1,600 LoC actually considered "very reviewable"? Please someone tell me if I'm crazy?
Models are not innately backwards-compatible. Both OpenAI and Anthropic encourage running evaluations and comparing the performance of your existing agent workflows against new models before just stepping up to the newest one because you may encounter regressions. I myself have seen lengthy/long-horizon multi-agent workflows begin breaking after moving to a newer model because for some reason the prompt containing an instruction to call a tool that worked 99/100 times before suddenly just stops working and needs to be modified.
I think LLMs are extremely useful, mostly for coding. But saying we're extremely close to an AI that can "reliably come up with novel actions for physical robots" feeds into the hype that these tools can do way or are very close to doing more than they're actually capable of, especially when we talk about reliability. That's the kind of rhetoric that has partially created this bubble, because in no world is what you're saying realistic.
The worst thing is when someone cites a video or a demo of an AI doing something and says, "See! It's here!" Remember when the Devin video came out years ago?
You can say "eventually" AI will be able to do xyz, but eventually the sun will blow up, too, so what the fuck are we talking about?
"UniSuper’s production Google Cloud VMware Engine (GCVE) private cloud was automatically deleted one year after it’s creation due to a misconfiguration in how it was created. When it was created, there was a bug in the creation script which passed a null value."
That's pretty amazing. Not due to a cascading failure from someone changing a config deep inside of a system that caused a bunch of unintended effects, just someone who messed up writing a shell script?
They can't even reliably follow instructions from text. I think "it's just around the corner/just wait x months/just wait and see bro" is one of the most telling signs of AI psychosis.
You do know that this was the same thing people said about crypto, right? And that the internet of things where your fridge connects to the Internet is hated by most consumers and had nowhere near the impact that IoT evangelists said it would?
This is such a creepy dystopian thing to say. Don't you realize that? Isn't this "yes, there will be pain, but the future is inevitable and we must go forward into it" attitude straight out of multiple horror and sci-fi stories?
Edit: Nevermind, parent is an LLM/bot.