It's a spectrum. I don't think that they will be able to sell t-shirts using a drawing you've uploaded. But it will probably allow them to defend themselves a bit better if they get sued for selling the data for LLM training.
HN user
josu
@josusanmartin
The console one seems the only relevant one. I'm a casual gamer, but the joycon running out of battery doesn't feel that annoying to me. That 16% has a much smaller impact on the JoyCon than on the main device.
Is it?
Battery capacity: 5172mAh, approximately 1% smaller than current version (5220mAh)
Spain is not Somalia, why not let Indra do it?
The data may be safer with the CCP, at least they won't lose it.
So CPanel's security is just as bad as their UI, who would have thought?
LLMs can also be really good in fields where you are not an expert. You just need to be very aware of your limitations, and start parallel conversation so one agent fact checks the other.
Done! The whole planet is now veggies.
enact similarly stupid laws.
No new law was enacted. The ISPs are enforcing a court order.
I wonder what will happen to the drivers if a large representation of the 1 million+ daily trips are displaced by automation?
If it happens gradually enough, they will just find other jobs. After the transition, society will be producing more with the same labor force, and thus the aggregate utility will increase.
invested in (...) anything else practical.
I don't understand how this is the top comment. LLMs have unlocked a lot of value for me personally, and arguably for the society as a whole. They are also one of the coolest technologies I've tried in years. As a technologist, I'm really glad that money is pouring in and allowing us to find its limits.
How is it cheating? Do you also need to independently discover Bloom filters to be able to use them?
I think you misunderstood, it's not about solving the problem, is about finding the most efficient solution. Give it a shot, and see if you can get to the top 10 on any task.
Thank you. Yeah, I'm doing all those things, which do get you close to the top. The rest of things I'm doing are mostly micro-optimizations such as finding a way to avoid AVX→SSE transition penalty (1-2% improvement).
But I don't want to spoil the fun. The agents are really good at searching the web now, so posting the tricks here is basically breaking the challenge.
For example, chatGPT was able to find Matt's blog post regarding Task 1, and that's what gave me the largest jump: https://blog.mattstuchlik.com/2024/07/12/summing-integers-fa...
Interestingly, it seems that Matt's post is not on the training data of any of the major LLMs.
The above is a software engineering problem. Reimplementing a JSON parser using Opus is not fun nor useful, so that should not be used as a metric.
I've also built a bitorrent implementation from the specs in rust where I'm keeping the binary under 1MB. It supports all active and accepted BEPs: https://www.bittorrent.org/beps/bep_0000.html
Again, I literally don't know how to write a hello world in rust.
I also vibe coded a trading system that is connected to 6 trading venues. This was a fun weekend project but it ended up making +20k of pure arbitrage with just 10k of working capital. I'm not sure this proves my point, because while I don't consider myself a programmer, I did use Python, a language that I'm somewhat familiar with.
So yeah, I get what you are saying, but I don't agree. I used highload as an example, because it is an objective way of showing that a combination of LLM/agents with some guidance (from someone with no prior experience in this type of high performing architecture) was able to beat all human software developers that have taken these challenges.
You are looking at this wrong. Creating a json parser is trivial. The thing is that my one-shot attempt was 10x slower than my final solution.
Creating a parser for this challenge that is 10x more efficient than a simple approach does require deep understanding of what you are doing. It requires optimizing the hot loop (among other things) that 90-95% of software developers wouldn't know how to do. It requires deep understanding of the AVX2 architecture.
Here you can read more about these challenges: https://blog.mattstuchlik.com/2024/07/12/summing-integers-fa...
I know what's like running a business, and building complex systems. That's not the point.
I used highload as an example because it seems like an objective rebuttal to the claim that "but it can't tackle those complex problems by itself."
And regarding this:
"Claude is very useful but it's not yet anywhere near as good as a human software developer. Like an excitable puppy it needs to be kept on a short leash"
Again, a combination of LLM/agents with some guidance (from someone with no prior experience in this type of high performing architecture) was able to beat all human software developers that have taken these challenges.
Lol, the problem is not finding a solution, the problem is solving it in the most efficient way.
If you think you can beat an LLM, the leaderboard is right there.
So my verdict is that it's great for code analysis, and it's fantastic for injecting some book knowledge on complex topics into your programming, but it can't tackle those complex problems by itself.
I don't think you've seen the full potential. I'm currently #1 on 5 different very complex computer engineering problems, and I can't even write a "hello world" in rust or cpp. You no longer need to know how to write code, you just need to understand the task at a high level and nudge the agents in the right direction. The game has changed.
- https://highload.fun/tasks/3/leaderboard
- https://highload.fun/tasks/12/leaderboard
- https://highload.fun/tasks/15/leaderboard
I was able to lead these two competitions using LLM agents, with no prior rust or c++ knowledge. They both have real world applications.
- https://highload.fun/tasks/15/leaderboard
- https://highload.fun/tasks/24/leaderboard
In both cases my score showed other players that there were better solutions and pushed them to improve their scores as well.
This reads like a hit piece. Cars crash, that obvious. Are Robotaxis crashing above or below human-driver rates?
I know what 99% of the people in HN are thinking while reading this post, and I agree, so I would like to hear the contrarian views:
Who is this useful for? What is the strategy behind this?
Read the article, the author is definitely in favor of using git.
Simple is good.
The monetary base expanded (aka "they printed a ton of money") between 2020-2022, so there were 2 options: the monetary base had to contract again, or the prices needed to adjust.
SPF is a Sun Protection Factor, meaning it multiplies the time it takes for your skin to burn. For example, if very light skin normally burns in about 10 minutes, SPF 20 stretches that to ~200 minutes, which is already over 3 hours. Since dermatologists recommend reapplying every 2 hours regardless, going beyond SPF 30–50 (which blocks ~97–98% of UVB) doesn’t add much practical benefit. Even for very fair skin, correct application and reapplication are far more important than chasing SPF 100.
Use Claudia to onboard the first users, gain some traction, and when the inevitable C&D letter comes change the name.
He may have burned the keys, not the coins. The process of burning the coins is by sending them to an address such as: 1111111111111111111114oLvT2
Airplanes. Some airlines offer 30 minutes of free wifi or something.
Ha, thanks for making the most of a bad comment.
Sorry about my comment, it didn't add anything to the conversation.