I wouldn't expect their detection of hacked accounts to be 100% correct. Sure, it might be obvious when a human takes a look, but humans can't proactively look at every account's usage.
HN user
srdjanr
I think it depends on how much time you spend learning something and "engaging" with LLM.
I have many times asked it something I was slightly curious about, got the answer after the first or 2nd-3rd prompt, spent 3 minutes in total and forgot it after 15 minutes probably.
But a few times I've spent an hour or more on a topic, asking many questions, thinking between responses, and I actually learned something.
I also have one anecdote - I (think) I understood basics of quantum mechanics intuitively for the first time. Of course it's probably superficial and maybe not 100% correct, but it is better than every time I tried to understand it before.
It definitely wasn't a single prompt, but two hours of back and forth, with a lot of time spent thinking (me, not LLM) in between. There were multiple times where I misunderstood something, so if I just read a book I'd probably get stuck many times.
It's a very lagging metric, and also influenced by sales, market conditions for your customers etc.
Good question, I guess you can run two identical Postgress instances and check if they have a diff
They should call themselves Time Lords
I have a few opinions on this:
1) My (and possibly other people's) last impression of this was when it was merged just based on all tests passing. In many projects, relying just on tests would definitely be reckless (not sure about coverage/quality of Bun tests).
I didn't follow much what where they doing later, maybe it is indeed good enough. For example, Claude Code using it and being fine is reassuring.
2) Making such a big decision that quickly and not (even having the time to) consulting community doesn't really inspire confidence.
3) You can be reckless even if everything ends up being perfect in the end.
Going 80mph in a city, you probably have like 80% chance of not crashing. And if you don't crash, you just had a much faster and more fun trip. Doesn't mean it wasn't reckless.
Of course he has the right to do it, no one disputes that. The users and community also have the right to complain, or stop using and supporting Bun because they don't like his actions
Tutorial was great! And I would like the option to give up.
They could be using this just to throw out the obviously bad CVs, and then manually go over the rest. I'm not sure if they do this in practice, but the tech itself can be useful.
Also if HR was really useless (or actively hurting the company) they wouldn't still have a job (or they'll lose it eventually). No one likes burning money for no reason. So obviously they are doing something useful.
That's a good idea.
The only drawback I see is that you should compare every pair of CVs for best results, and that grows quadraticly with number of CVs. Of course you can settle for fewer comparisons and not perfect results. But then I'm not sure if you can hit a good ratio of quality and token spend.
It makes sense to me intuitively (though I'm not sure if my reasoning is actually correct).
Worse model may not "know" enough to distinguish between a 70 and a 100 candidate, so it's expected that it's output has high variance. But a better model might "know" enough, so it can be more confident and thus more consistent.
Of course, nothing can guarantee the right answer from LLMs
At the end they mentioned they're exploring global trust signals
So this is only one of the reasons, and a relatively small one:
“Our estimate is that about 200 to 400 pedestrians a year would not have died if vehicles had remained approximately the same size over the past quarter-century,” the report continued. “That represents about 10 percent of the recent increase in pedestrian deaths.”
Of course there is. If it increased from 1MB to 4MB, that would definitely be insignificant
How can they even do that without storing plaintext passwords?
Security researcher is a common term, there's also market research which doesn't look like it falls under your definition
The root comment is talking about adding blood, breath, urin, spit... analysis. For body imaging only I agree with you. But if we add all this, I guess we'd be able to rule out many false positives
Are more plants coming? I think I heard it won't be many of them, because it's risky.
If the bubble bursts and RAM demand drops, then they'll have big losses. And that's not an impossible scenario over the few X years that it takes to build a plant
I don't mind it here at all, in fact I didn't even notice it's AI before reading this comment. It's clearly not a one-shot AI slop but a well thought out and edited by a human post.
Not everyone who has something interesting to say is a good writer, and I think it's great if AI can help them tell their stories.
These all look fine to me, and I don't think I'd assume they're AI generated.
There definitely is a distinct AI "design" look, I just don't really see it here.
Nobody is arguing for unsafe models
Then what are people arguing for? I see only two totally distinct options: unsafe models or someone being the arbiter of safety
It seems like they run a classifier model before going to Fable (or falling back to Opus), so it should be fine
My understanding is that it's designed for fixed-size documents? There's a big difference between a layout system for that, and one where size of a document can vary wildly, up to completely opposite aspect ratios
There's a difference between the driver intentionally driving into crowd, and not intentionally but possibly still recklessly (drifting and losing control, falling asleep, etc). In those cases I would probably use "car hits the crowd", at least in my language
We should treat it as attempted attack in the sense of preparing for the next one, but I don't see why we should call it "attack" without any evidence
I frequently tell agent to do something, wait ~10 min (which is just enough that I can't/don't want to start anything else), ask it to change something, wait a few minutes again, and so on. So I'm basically idle while waiting for agent, and it would be great if it was faster.
It's like your compile times were ~10 min. Sure, it's not a huge deal, but it's sooo anoying
Why should the chatbot team necessarily take the blame? For all we know, they could have got approval from the tool team to make it public, and passed additional security review for making it public.
Also, why fire anyone after a single mistake?
Well fun can be the gain