They sell their products using the credentials they gained.
I’d never heard of socket until they found and reported shai hulud hiding in pytorch lightning. It pays off.
HN user
NLP engineer
They sell their products using the credentials they gained.
I’d never heard of socket until they found and reported shai hulud hiding in pytorch lightning. It pays off.
I think this is such an nefariously unnecessary negative argument.
Most, if not all, of the shai-hulud attacks that hit npm and other ecosystems were preventable with cooldowns. And these were not detected because regular users reported the worms, but because security researchers did. I don’t think I’ve ever seen an attack that was discovered because a user reported it.
Nice! That sounds a lot more like a mission statement than your actual readme.
Not the software, the whole thing.
As an author: show me why you thought this was interesting and why you’re doing it, and why you think it’s relevant. What does it build towards? What does climbing this leaderboard mean to me?
Absent those things, this is just some thing my opus could generate as well.
Sadly 100% generated.
I think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.
For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.
I don’t think this is the right take-away though.
Your fully generated project description makes people lose faith in the actual project. If you went through the trouble of writing the whole code yourself, why generate the comment presenting it to the public wholesale.
Even the title is a Claudeism, it makes me sad
The file drawer effect, except this one maybe should have stayed filed.
I think he is a good example of someone who writes mainly to show he is ready for the next rung of the corporate ladder. That is, his posts are not meant to be useful, but to show higher-ups he is useful to them.
As mentioned by a sibling comment: this is an insensitive take.
It takes a lot of courage to write down one’s struggles for all the world to see. Your analysis denies the OP their self-reflection, and instead reduces it to a thing you happened to find in your own life.
I am interested in why you chose to do this, and publish it with the headline you used. Was it to learn something? Or to get publicity for another project?
Tbh, this sounds like fear mongering to me. Of course the statement “99.9% of servers are not compliant” sounds impressive, but then it turns spec hasn’t even been released yet.
Also some general feedback: the whole thing looks generated, as does the comment I am replying to.
Why do all this work and then let an ai write the blog post.
Ok, thanks!
It’s personal ad, basically. The author is trying to get a job as an evaluator somewhere and is hoping that putting 1000$ on the line will get them enough publicity to land them an interview/get a job somewhere.
Is it even legal to publish excerpts of books like this? Or does this fall under some kind of exemption/fair use clause?
Thanks for introducing me to the article! I’ve experienced this myself but didn’t know it had a name.
A non-autoregressive transformer trained with a classification objective.
The second part of this comment is not what I expected. I also don’t think it is true. I got bit by a CORS error at work recently that passed by Claude, copilot, and another senior engineer.
We’ve been on the receiving end of this complaint with Semble. I think it is a valid complaint, but constructing a benchmark for this kind of thing is just very difficult and expensive because of the (harness) x (model) x (mcp/cli) combination.
With traditional ml/tooling, not showing benchmarks was usually a red flag. But for llm tooling, I’m not so sure.
Ha thanks, it was pink a while ago
Extreme programming in a nutshell. I like doing this to features: build it, then take it down and rebuild but better.
Puppy slush automatically pushed through vents into your codebase
Thanks! This is very similar indeed. Related: I see a lot of “drive-by” PRs by agents, who obviously have no intent of ever maintaining the code they wrote.
I’m not sure I share your view of PRs. I still see submitting PRs as something that puts pressure on maintainers. Even incorrect PRs take time to verify and review.
I also don’t see how this differs between the “gap” and the “fence” part of the metaphor. Whether someone submits a rewrite/removal (fence) or a new feature (gap) for PR review, it’s still going to cost me attention.
It’s an interesting question: I’d say this is more of a vulnerability creator than the actual vulnerability.
Similar to how using very difficult technologies makes you more likely to create code with vulnerabilities: the technologies are not the vulnerability, but it’s easier to cause them.
This paper oversells on the title. Like, what is chronos, which embedding model was used, which reranker, how was the reranking done, why is chronos much better than claude code
Sure, the whole premise is exactly that proof of work reduces the value of scraping, while having negligible impact on users. If the data is so valuable that bot operators are willing to pay 10s of cpu, then other measures are necessary.
Nevertheless even for these high value cases, you can still argue that it disincentivizes the business model, it becomes less efficient.
Because it destroys the economics of scraping. It’s too expensive with proof of work, or at least not as economically viable
I was also at the event and was pretty disappointed. Most of the talks were pretty low on information. I was at the “build” stage, which supposedly was the technical stage, but the talks there didn’t really go into technical specifics.
The papyrus talk was awesome though.
It was directed at the parent who implied that we didn’t think about this.
I agree with your point about the evals and how you can get discontinuities: good search can be worse than bad search when agents can do many searches. We’re working on it