HN user

thorum

2,678 karma
Posts40
Comments428
View on HN
www.youtube.com 2mo ago

The physics slop that YouTube wants me to make [video]

thorum
2pts0
suno.com 7mo ago

Suno AI Partners with Warner Music Group (WMG)

thorum
9pts4
restofworld.org 9mo ago

Malawi's new farmhand: AI that speaks the local language

thorum
2pts0
arxiv.org 12mo ago

Gemini 2.5 Pro Capable of Winning Gold at IMO 2025 with Prompting

thorum
4pts3
www.njkumar.com 12mo ago

Training a Flappy Bird Diffusion World Model to Run in a Web Browser

thorum
3pts0
www.youtube.com 1y ago

I tried to prove I'm not AI [video]

thorum
3pts0
gwern.net 1y ago

Towards Benchmarking LLM Diversity and Creativity

thorum
1pts0
www.youtube.com 1y ago

The sham legacy of Richard Feynman [video]

thorum
8pts1
techcrunch.com 1y ago

Chinese lab DeepSeek releases a 'reasoning' AI model that rivals OpenAI's o1

thorum
8pts1
old.reddit.com 1y ago

I used vision models to help me win at Age Of Empires 2

thorum
4pts0
github.com 1y ago

Claude Artifact Runner

thorum
1pts0
www.symmetrymagazine.org 1y ago

The deconstructed Standard Model equation

thorum
43pts28
github.com 2y ago

Ruler: What's the Real Context Size of Your Long-Context Language Models?

thorum
2pts0
www.rollingstone.com 2y ago

Drake releases diss track with AI Tupac, Snoop Dogg voices

thorum
4pts2
web.archive.org 2y ago

What Is Enlightenment? (1784)

thorum
70pts53
www.nytimes.com 3y ago

OpenAI Worries About What Its Chatbot Will Say About People’s Faces

thorum
7pts1
www.forbes.com 3y ago

Twitch Streamer XQC Moves to Kick in $100 million deal

thorum
4pts0
old.reddit.com 3y ago

Gfycat has been down for two days due to an expired SSL certificate

thorum
96pts36
www.change.org 3y ago

Unplug the Evil AI

thorum
2pts0
news.ycombinator.com 3y ago

Are there any open source or CLI vectorizers as good as Vector Magic?

thorum
4pts0
old.reddit.com 3y ago

Stability AI accuses open source developer of stealing copyrighted code

thorum
2pts0
news.ycombinator.com 3y ago

Ask HN: How to figure out why free users don't convert?

thorum
9pts20
eshoo.house.gov 3y ago

US Rep. Anna Eshoo Urges NSA and OSTP to Address Unsafe AI (Stable Diffusion)

thorum
1pts0
wanderinginn.com 3y ago

On AI and the Future of Writing, by one of the web’s most successful authors

thorum
1pts0
news.ycombinator.com 3y ago

GitHub is removing the Trending Repositories page

thorum
75pts33
help.openai.com 3y ago

OpenAI API pricing update FAQ

thorum
87pts193
www.nytimes.com 4y ago

With Rising Book Bans, Librarians Have Come Under Attack

thorum
20pts4
github.com 4y ago

Google AGI GitHub Repository

thorum
1pts0
nonint.com 4y ago

Cheater Latents

thorum
1pts0
blog.piekniewski.info 4y ago

AI winter is well on its way (2018)

thorum
2pts0

Interesting read! Creating tests is highlighted as something Claude did well, but it strikes me that all the weaker rejected solutions could have been avoided if it were really good at designing intelligent tests for itself. For example, the first solution “was very specific to the reported bug and wouldn’t have fixed the general case” and the third suggestion “prevented the perfectly valid use of as conversion expressions in go commands as well”. I imagine both of these cases could have been noticed and avoided by the agent if it had planned out adequate tests ahead of time.

The “correct”, elegant way for AI to interact with existing software would take decades and billions of dollars to build. Someone would have to do the hard work of building new APIs, solving decades of accessibility issues, etc.

Or you can show an AI screenshots and ask it where to click.

The team with the most star power and hype tends to attract the best young talent.

If the next big breakthrough in AI comes from Anthropic, good chance it comes from some genius you’ve never heard of who decided to work there because of [famous researcher].

Midjourney Medical 1 month ago

I wish them all the best and hope they succeed, but can’t help but suspect they’ve fallen into deep LLM psychosis. Even if you assume they can build this thing and it works as described and then get past all the regulatory hurdles, the scale of infrastructure they’re talking about is enormous.

Unfortunately for the people mad about this, I predict the only thing they will accomplish by pressuring the rsync maintainers, is to discourage everyone else from responsibly disclosing their use of AI. You’re just going to make people disable Claude attribution on their commits to avoid drama.

Somewhat useless article. To summarize, we have anecdotes suggesting they may work but no one has figured out how to prove or disprove it in a study, and the author has some doubts. Meanwhile supplements can be dangerous if you take too much or have a liver condition, or if you buy them from an unreputable source, as with every other substance on earth. Author confirms it tastes good in milk.

You’re probably right in a literal technical sense, but a very large number of people (maybe most?) would choose “no” if properly informed and asked for consent, and lots of people are morally opposed even in principle to downloading a large AI model onto their computer. I’m not one of them, but they’re out there. So in a cultural sense, it is different.

Google Flow Music 3 months ago

The models are primitive right now, but we’re clearly heading toward “AI as sound synthesis, human as artist” - much like how producers currently use a DAW to assemble premade loops and sounds from Splice, but with the producer now able to prompt any sound, filter, or effect they can imagine into existence and then rearrange them into a song.

See for example Suno Studio, which is not very good in my opinion, but shows the direction they’re going.

I have the opposite experience: random HN/Reddit comments saying “this sucks” or “whoa this is a huge improvement” are the only benchmark that means anything. Standard benchmarks are all gamed and don’t capture the complexity of the real world.

Stars have been useless as signals for project quality for a while. They’re mostly bought, at this point. I regularly see obviously vibe-coded nonsense projects on GitHub’s Trending page with 10,000 stars. I don’t believe 10,000 people have even cloned the repo, much less gotten any personal value from it. It’s meaningless.

Ape thinking is a cognitive practice where a human deliberately solves problems with their own mind. Practitioners of ape thinking will typically author thoughts by thinking them with their own brain, using neurons and synapses.

The term was popularized when asking a computer to do it for you became the dominant form of cognition. "Ape thinking" first appeared in online communities as derogatory slang, referring to humans who were unable to outsource all their thinking to a computer. Despite the quick spread of asking a computer to do it for you, institutional inertia, affordability, and limitations in human complacency were barriers to universal adoption of the new technology.

but the number of problems requiring deep creative solutions feels like it is diminishing rapidly.

If anything, we have more intractable problems needing deep creative solutions than ever before. People are dying as I write this. We’ve got mass displacement, poverty, polarization in politics. The education and healthcare systems are broken. Climate change marches on. Not to mention the social consequences of new technologies like AI (including the ones discussed in this post) that frankly no one knows what to do about.

The solution is indeed to work on bigger problems. If you can’t find any, look harder.

I’m honestly surprised LLMs are still screwing up citations. It does not feel like a harder task than building software or generating novel math proofs. In both those cases, of course, there is a verifier, but self-verification with “Does this text support this claim?” seems like it ought to be within the capabilities of a good reasoning model.

But as I understand the situation, even the major Deep Research systems still have this issue.

The article presents AGENTS.md as something distinct from Skills, but it is actually a simplified instance of the same concept. Their AGENTS.md approach tells the AI where to find instructions for performing a task. That’s a Skill.

I expect the benefit is from better Skill design, specifically, minimizing the number of steps and decisions between the AI’s starting state and the correct information. Fewer transitions -> fewer chances for error to compound.

Agree that planning time is the bottleneck, but

3 days

still seems slow! I’m saying what happens in 2028 when your entire project is 5-10 minutes of total agent runtime - time actually spent writing code and implementing your plan? Trying to parallelize 10m of work with a “town” of agents seems like unnecessary complexity.

Am I wrong that this entire approach to agent design patterns is based on the assumption that agents are slow? Which yeah, is very true in January 2026, but we’ve seen that inference gets faster over time. When an agent can complete most tasks in 1 minute, or 1 second, parallel agents seem like the wrong direction. It’s not clear how this would be any better than a single Claude Code session (as “orchestrator”) running subagents (which already exist) one at a time.

I agree that LLMs can be useful companions for thought when used correctly. I don’t agree that LLMs are good at “supplying clean verbal form” of vaguely expressed, half-formed ideas and that this results in clearer thinking.

Most of the time, the LLM’s framing of my idea is more generic and superficial than what I was actually getting at. It looks good, but when you look closer it often misses the point, on some level.

There is a real danger, to the extent you allow yourself to accept the LLM’s version of your idea, that you will lose the originality and uniqueness that made the idea interesting in the first place.

I think the struggle to frame a complex idea and the frustration that you feel when the right framing eludes you, is actually where most of the value is, and the LLM cheat code to skip past this pain is not really a good thing.

Your other comment sounded like you were interested in learning about how AI labs are applying RL to improve programming capability. If so, the DeepSeek R1 paper is a good introduction to the topic (maybe a bit out of date at this point, but very approachable). RL training works fine for low resource languages as long as you have tooling to verify outputs and enough compute to throw at the problem.

Developed by Jordan Hubbard of NVIDIA (and FreeBSD).

My understanding/experience is that LLM performance in a language scales with how well the language is represented in the training data.

From that assumption, we might expect LLMs to actually do better with an existing language for which more training code is available, even if that language is more complex and seems like it should be “harder” to understand.

I remember reading and hearing similar rants from programmers 15 years ago, long before LLMs. The author kept going and figured it out, and probably got some pride and enjoyment from finishing the project in spite of the frustrating moments. That’s what learning to code has always been like.

Hollywood had a ton of issues but it at least had some... class?

It looked that way because they had media training and their public personas were carefully managed, with staged interviews and media appearances. Behind the scenes, it’s a different story.

Influencers are rewarded for seeming authentic. Mr Beast coming across badly in a traditional TV interview just makes his audience think he’s more real.