HN user

brokencode

3,152 karma
Posts10
Comments740
View on HN

Personally, I often paste screenshots into Claude Code of the application it’s working on. And I’ve even had it work autonomously on something and regularly grab its own screenshots.

Or sometimes I will have tables, charts, or even screenshots of text that I would otherwise have to have another step to OCR or type out.

Multimodal saves me time on a regular basis. Not sure it’s a game changer, but just lets me communicate with the model in all sorts of ways that would be harder otherwise.

Maybe the deaths would have happened, or maybe these places would have become self sufficient or received aid elsewhere. There’s really no way to know.

But we committed to providing this aid. Even if we think it would be better in the long term for others to do it instead, there’s no excuse to cut it off suddenly without making reasonable efforts at a transition.

Imagine you’re the long term caretaker for a family member, then just decide one day you’re tired of it, so stop. You make no attempt to find somebody else to take over, you simply stop bringing them their meds/food or whatever.

Do you think your family should then forgive you after grandma dies? I don’t think so.

I don’t know, does global cancer research shut down when you stop giving the $500? Do kids immediately stop receiving treatment?

USAID literally ran ambulance systems that shut down due to lack of diesel. They delivered lifesaving drugs that stopped.

We made commitments to communities to run these services, then suddenly killed them off. We didn’t try to find other countries to step in. We didn’t try to get the local governments to take over.

We did jack shit to try to preserve lives in this transition process.

What rule about steelmanning? We’re commenting online, not writing peer reviewed research.

And yeah, some people lose the benefit of the doubt. Sorry, but actions have consequences.

Elon doesn’t just get to kill hundreds of thousands of poor people by eliminating USAID and expect everyone to treat him the same way.

He’s made enemies for life, and he deserves it.

I also don’t think one mistake should define a company. But for me it’s just about trust.

Musk has proven time after time that he doesn’t deserve my trust. I will never trust Grok as long as he’s in charge of it.

I agree that the guardrails on the top models have gotten out of hand, though.

Fable for instance won’t answer even basic health questions. As if you are going to take nutrition advice and make a bioweapon with it.

Partly this is due to government interference. Hopefully we get to a better place as competition heats up with open and Chinese models.

Elon himself promoted Grok’s “spicy mode” that allowed generating NSFW content that the other AI vendors wouldn’t touch with a 20 foot pole.

Believe whatever you want. Elon’s beliefs and personality problems have been baked into the core of Grok, so it’s no surprise that it turned out to be a CSAM-generating MechaHitler that steals people’s data.

Anybody surprised when Grok turns out to be trash really should read up on the guy who made it.

How much can you really certify that data is destroyed?

Customer data could live on the computer Elon pretends to play Diablo 4 on for all we know.

Grok 4.5 14 days ago

No you didn’t, and that’s not much of an insult.

Grok 4.5 14 days ago

Yup, the destitute and dying are famous for their highly publicized research.

Grok 4.5 14 days ago

Sure, let’s make jokes about the poor dying of malaria and AIDS. You sound like a Grok power user.

Grok 4.5 14 days ago

As opposed to what? Going down to Africa to count the bodies myself?

Of course I read about it in the media. And there are articles from Harvard, UCLA, and others that say the same thing

Grok 4.5 14 days ago

Did they kill hundreds of thousands of poor people by shutting down USAID too?

Grok 4.5 14 days ago

Grok has a serious credibility problem due to Elon’s decisions and personal insanity.

Will it ever recover? Maybe. But it’s got an uphill battle even compared to the Chinese models, and that’s saying something.

Road to Elm 1.0 16 days ago

Yeah Elm has had a very strange arc, but I think calling it a research language is right.

There was a period where it was heavily evangelized. Many blog posts were written and talks given, and there was a lot of enthusiasm and adoption.

Then the author just kind of disappeared and the project stalled.

Which of course he had a right to do since it’s his project, but I think he should have set expectations better from the beginning.

The heavy evangelism helped spread the ideas, but also set up developers to feel blindsided and abandoned.

Fable 5 is Back 21 days ago

Yup, but apparently our cyborg cats can only be kittens and the cyborg mice are probably going to be like 4 feet tall. At least according to the US government.

Claude Sonnet 5 21 days ago

The graphs do that already. I was expecting them to try to explain how good it was at simple tasks.

Claude Sonnet 5 22 days ago

Kind of crazy how bad this release actually is. I even dug around in the full system card, and every graph showed the same thing.

Low and maybe medium will save money on simpler tasks, but after that it just isn’t worth it compared to Opus.

I wish they would have explained in the blog post why they think anybody would ever want to use this above medium.

Maybe it works well on things that aren’t clear in the benchmarks.

Claude Sonnet 5 22 days ago

That is a bad comparison. Compare Sonnet xhigh against Opus medium, which is both better and cheaper.

You’re expecting me to know your job? Give me a break.

I’m wondering the same thing. You keep talking of some grand poisoning problem but can’t point to any specific public information except an article saying that it’s possible. As if that was ever in doubt.

Guess we’ll just have to agree to disagree.

You’ve seen actual model poisoning? Or have you seen a model return the wrong answer due to what it saw in a search result? Or were they hallucinations perhaps? How do you know it’s due to poisoned training data?

And do you even realize how much data 0.001% of the training data for a frontier models is? They’re trained on 10s of trillions of tokens, meaning you’d need hundreds of millions of tokens of poisoned data.

Some of these problems you mention could become real barriers to models improvements, though there are plenty of countermeasures, such as by focusing on high quality data sources like I mentioned before.

We’ve already probably gotten as much as we’re ever going to get from simply scraping more and more unstructured text from the web as a way to improve model performance.

The type of training being done now is around tool use and solving specific types of problems better, which is the type of training data you simply don’t find lying around on the web.

You are totally misunderstanding my argument then. As I said, garbage in garbage out. Your article is just an example of that. It’s pretty obvious that if you train an LLM on bad data, you will get bad output.

What I’m saying is that the AI labs are handling this not by fixing the “garbage out” part, but by minimizing the “garbage in” part.

The fact that all you could come up with was research (not an actual example of poisoning a real training set) from 2025 kind of proves that this isn’t some kind of widespread, unsolvable problem like you seem to be claiming.

The question is not whether it has happened or will continue to happen. Of course it will always be a problem to some extent.

Your original claim is that this will be enough of a problem to prevent models from improving in expert level knowledge. I completely disagree with this premise.

If the models fail to improve, it will likely be due to limitations in the transformer architecture rather than poisoned training data.

And even then, I doubt that the transformer is the best architecture we will ever come up with.

Clearly it doesn’t learn or think like a human does, since humans don’t need many gigabytes of text samples to learn to talk, so there is some room for improvement.

There are so many better data sources that AI labs can use here that this argument really holds no water at all.

Peer reviewed journals, textbooks, in-house teams of experts, trusted news publications, etc.

The whole idea of scraping large swaths of the internet for training data has always been pretty dubious due to the variable data quality.

I mean, just look at the early Google models that told people to put glue in their pizza due to a joke in the training set. Garbage in, garbage out.

This is one of the first and most obvious problems all of these labs have run into, and countermeasures are only going to improve.