HN user

suttontom

117 karma
Posts0
Comments71
View on HN
No posts found.

Agree with you on exercise. But staring at a painting is to meditation what curls are to building your core. People all recommend the same type of meditation for a reason--so that you can close your eyes, focus on something, notice when your mind drifts, and bring the attention back. You didn't mention anything about awareness or noticing when you lose focus, so your advice seems to completely miss the point.

It's crazy that you could actually use the excuse that since it's all vibe-coded, there's no way a human could have written it, so Anthropic bears no responsibility.

Meanwhile humans can pop in and leave little morsels like this and blame it on the model.

These AI images also add to the public mistrust of AI, a growing problem for innovation in a field that is sometimes seen as biased, opaque and extractive.

Oh my, how would anyone ever have gotten that impression?!

Your instinct is correct, and in a lot of cases it's true. However, I've heard from enough doctors by now (a cardiologist, psychiatrist, and epidemiologist/former physician) that they use medical LLMs and find them extremely helpful, mostly as a way to either bring up knowledge they'd forgotten about or as a way to learn something new and then verify it. I'm extremely skeptical about LLMs in general and the connection to Gell-Mann Amnesia is apt, but I wouldn't necessarily write them off completely like that. There are experts using the models that find them genuinely helpful in their field.

I'm asking genuinely, is there a connection between housing, education, and healthcare becoming so much more expensive and them also being the three parts of the economy that have the most government interference (in the US)? If so is it causal?

I don't want to be cynical, but maybe spending hours every day using Claude has made some of us particularly attuned to picking this up. For some reason as soon as I read "The trap was in app/test/index.js," I instantly knew it was Claude. It's too bad, because there will obviously be some false positives, but it makes me immediately disregard the author.

You're wrong in lots of ways.

Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...

This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/

The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150

This is such a tired, meaningless argument. I've never seen a human in 10 years of professional software engineering at a large company ever so confidently, consistently create and send out seemingly well-reasoned code that's as wrong as what SOTA models using CC or Codex do. If a human did this, they would be fired or perpetually remain a junior who no one wants to work with.

Also, if a human does this, you can replace them and get a human who will not do it. The default for an LLM is to generate plausible-looking text that may or may not be completely incoherent. That is not the default for a human. Again, if you find that your colleague consistently fabricates APIs, you can hire someone who isn't crazy instead, but you cannot do the same with LLMs.

Ah yes, the magical equivalent of "you are a senior software engineer who writes bug-free code".

IME people would benefit greatly from the process, albeit tedious and time-consuming, of testing out the same prompt sequence/session with the exact same model multiple times. It becomes clear extremely quickly how capable but unreliable and inconsistent a model can be even when given the same context. If you have ever completed a long, complicated task with an agent and then lost the session and tried doing the same thing again from scratch you may have had the experience of seeing the subtle changes that come up in the model's thinking which lead it to accept or reject certain paths and ignore or incorporate prompt instructions like the one you've provided.

Claude Opus 4.8 2 months ago

Do you know if anyone has trained, say, a pre-2017 model and tried to get it to come up with Attention Is All You Need? If it did, would you say that was only because it's a synthesis of prior art? If so, what isn't?

Claude Opus 4.8 2 months ago

Are you joking? Is there literally "nothing" you can imagine that Claude can't do?

Not to be cynical but do you think this would matter at all? Are you saying that companies would hold themselves to their missions or even something that's legally binding?

"Google is not a conventional company. We do not intend to become one."

OpenAI being founded as a nonprofit and becoming for profit.

Didn't Anthropic literally say they wouldn't train on your data or keep it for longer than 30 days unless legally required, and then decided to opt people in to having their conversations used for training?

Models are not innately backwards-compatible. Both OpenAI and Anthropic encourage running evaluations and comparing the performance of your existing agent workflows against new models before just stepping up to the newest one because you may encounter regressions. I myself have seen lengthy/long-horizon multi-agent workflows begin breaking after moving to a newer model because for some reason the prompt containing an instruction to call a tool that worked 99/100 times before suddenly just stops working and needs to be modified.

Gemini Omni 2 months ago

I think LLMs are extremely useful, mostly for coding. But saying we're extremely close to an AI that can "reliably come up with novel actions for physical robots" feeds into the hype that these tools can do way or are very close to doing more than they're actually capable of, especially when we talk about reliability. That's the kind of rhetoric that has partially created this bubble, because in no world is what you're saying realistic.

The worst thing is when someone cites a video or a demo of an AI doing something and says, "See! It's here!" Remember when the Devin video came out years ago?

You can say "eventually" AI will be able to do xyz, but eventually the sun will blow up, too, so what the fuck are we talking about?

"UniSuper’s production Google Cloud VMware Engine (GCVE) private cloud was automatically deleted one year after it’s creation due to a misconfiguration in how it was created. When it was created, there was a bug in the creation script which passed a null value."

That's pretty amazing. Not due to a cascading failure from someone changing a config deep inside of a system that caused a bunch of unintended effects, just someone who messed up writing a shell script?

Gemini Omni 2 months ago

They can't even reliably follow instructions from text. I think "it's just around the corner/just wait x months/just wait and see bro" is one of the most telling signs of AI psychosis.

Google I/O 2 months ago

You do know that this was the same thing people said about crypto, right? And that the internet of things where your fridge connects to the Internet is hated by most consumers and had nowhere near the impact that IoT evangelists said it would?

Google I/O 2 months ago

This is such a creepy dystopian thing to say. Don't you realize that? Isn't this "yes, there will be pain, but the future is inevitable and we must go forward into it" attitude straight out of multiple horror and sci-fi stories?

Edit: Nevermind, parent is an LLM/bot.