HN user

embedding-shape

17,199 karma

It's only in jest, don't take it so seriously.

https://emsh.cat

embedding-shapes@proton.me

Posts130
Comments5,452
View on HN
arxiv.org 2d ago

AgentAbstain: Do LLM Agents Know When Not to Act?

embedding-shape
3pts0
bsky.app 8d ago

Cloudflare is now "Local First, Everywhere"

embedding-shape
2pts0
github.com 8d ago

Codex starts encrypting sub-agent prompts

embedding-shape
425pts250
web.archive.org 8d ago

As We May Think – Vannevar Bush (1945)

embedding-shape
2pts0
www.youtube.com 15d ago

InfoWars: Emergency with Tim Heidecker [video]

embedding-shape
14pts0
www.polygon.com 22d ago

Lawsuit alleges that RAM manufacturers are colluding to drive up prices

embedding-shape
3pts0
bevy.org 1mo ago

Bevy 0.19

embedding-shape
4pts0
jms55.github.io 1mo ago

Realtime Raytracing in Bevy 0.19 (Solari)

embedding-shape
3pts0
en.eurovelo.com 1mo ago

EuroVelo – Network of long distance cycle routes connecting European continent

embedding-shape
2pts1
newrepublic.com 1mo ago

After Months of War, Trump Says Iran Has Right to Nuclear Program

embedding-shape
21pts4
xcancel.com 1mo ago

Published Rio 3.5 Open 397B was an "intermediate checkpoint"

embedding-shape
2pts0
developer.apple.com 1mo ago

New Domain for Sign in with Apple and iCloud+ Hide My Email

embedding-shape
3pts1
community.openai.com 1mo ago

Flexible Rate Limit Resets for Codex (bank rate limit resets)

embedding-shape
3pts0
huggingface.co 1mo ago

Ai2 ACE2S – Simulate atmospheric variability – Scale of days to centuries

embedding-shape
4pts0
www.youtube.com 1mo ago

Jerry Gretzinger's map of a place, 40 years of "analogue generative art" [video]

embedding-shape
4pts0
vinyl-cache.org 1mo ago

Vinyl Cache and Varnish Cache

embedding-shape
83pts43
news.ycombinator.com 1mo ago

The "Best" HN Comments

embedding-shape
3pts0
www.politico.com 1mo ago

White House sends Blanche's Attorney General nomination to Congress

embedding-shape
1pts0
www.miamiherald.com 1mo ago

Michael Reiter – The Palm Beach cop who Jeffrey Epstein couldn't stop

embedding-shape
6pts0
emsh.cat 1mo ago

Show HN: A HTML Quine Made with Nix

embedding-shape
4pts0
apnews.com 1mo ago

Spanish police raid headquarters of PM Sánchez's Socialist Party

embedding-shape
3pts0
www.rawstory.com 1mo ago

Reporter says they were permanently injured after alleged 'pulsed energy' attack

embedding-shape
3pts3
variety.com 1mo ago

Pope Leo: opaque AI run by few firms risks "New Forms of Dehumanization"

embedding-shape
164pts2
www.euronews.com 2mo ago

Italy moves to Airbus A330 tankers

embedding-shape
284pts126
status.openai.com 2mo ago

Tell HN: OpenAI Codex: Increase in users hitting Codex rate limits

embedding-shape
6pts4
old.reddit.com 2mo ago

You can access Gemini chat history without unlocking your phone with Android 16

embedding-shape
24pts1
www.theguardian.com 2mo ago

US Justice Department 'forever' bars IRS from auditing Trump's past tax returns

embedding-shape
21pts1
www.bloomberg.com 2mo ago

Sony Pulls Back from PlayStation Games on PC

embedding-shape
3pts1
finance.yahoo.com 2mo ago

OpenAI seals deal in Malta to give all Maltese access to ChatGPT Plus

embedding-shape
2pts1
cascadeur.com 2mo ago

AI-Assistance in Character Posing: How It Works in Cascadeur

embedding-shape
2pts0

I have never had a problem running any model, for image or text or voice, the community has done great work in making things work

There is a whole world of other tooling and stuff that isn't just for hobbyists to run inference with ML models, but also how to do profiling, debugging and gathering data when you run distributed workloads, and so on. The nsight toolkit seems miles ahead of the competition on other platforms, as just one example.

isnt available in the west [...] Can anyone explain the allure of the Nvidia box

The first part might answer the second one. Otherwise, the lack of CUDA and the nvidia ecosystem of tooling could also explain why it doesn't seem so interesting for AI tasks.

Laguna S 2.1 2 hours ago

More updates from the Discord (also mostly about the NVFP4):

we've got an NVFP4 checkpoint that has a KD of 0.135 relative to the BF16 checkpoint (for W4A4, on some GSM8K prompts). [...] We're still seeing some looping on W4A4 (although much less than before), whereas for W4A16 we've not been able to make it loop. Given that W4A4 is still a little broken, we won't make this an official release (we have one more idea planned to fix that). [...] The RC1 for NVFP4 is reachable as poolside/Laguna-S-2.1-NVFP4 @ RC1

Seems to be available already indeed: https://huggingface.co/poolside/Laguna-S-2.1-NVFP4/tree/RC1

tacit knowledge is not easy to add to the AI’s training set.

Is it even possible? Different people do different judgements for the same things, there is no right or wrong really, just the choice you made in that moment, based on whatever you cared about at that point.

I'm surprised how many of those I've never seen before, even after spending more than a decade on HN and arguably way too much every day. Such a good surprise :)

I wish the page used the CSS :visited selector and provided a link on the index page, so we could easily see which one we've visited before or not.

no one reads it. [... ] they seem disillusioned

Regardless of reasons they tell you, if no one reads your documentation, then it isn't good enough. Good documentation gets read, as long as it is good. It sounds to me like you're disillusioned, if you're not responding to the feedback, even if you sometimes have to look through the reasons people tell you on the surface.

Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal.

This happens from time to time when you work on optimizations and similar things, with less "smart" LLMs and under-specify what exactly you're out after. Asking them to make functions faster without clearly specifying what the function has to do, is a great way to replicate this too. Doesn't seem to happen as often with SOTA models though.

I think the early example of "I asked it to make the test suite pass, so it changed all the assertions" is pretty much the same variant of this, where it technically does what it is asked to do, yet in "clearly" (to humans) wrong ways.

Why should the default be for me

It doesn't have to be, of course. But if you're bothered by something in general in life, and there are two approaches to solving the problem, one that involves others doing the "right thing", or another that involves you changing your approach to it, then usually the latter is much easier to actually address, than the former, as it relies on changing the behaviour of a ton of people.

It's just a tip to make things easier for yourself, without needing others to change. Still, you can do things however you like to, no one is forced to do anything :)

Laguna S 2.1 8 hours ago

Update3, summarized from RTX6kPRO Discord:

The BF16 checkpoint doesn't exhibit any of the complaints [...] with our current quants, the models tend to choose the wrong logits sometimes [...] why we'll need a requant [...] We're not aware of any bugs in any runtimes themselves [...] We have two remaining things that we're trying to tackle: some people are reporting thinking being too hard to trigger (^^) , and others are saying it thinks too much. We've seen much more of the latter internally

Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly.

Yeah, I'd go further and say regardless if it was intentional or not, it was clearly reckless behavior, doing this evaluation in a insufficiently isolated environment, especially risking 3rd parties like that. Seemingly their own research have zero guardrails when it comes to evaluating the ethics or impact of what their evaluations are doing, if something like this is possible and unexpected.

Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"

I feel like I would have thought the same maybe a year or two ago, but based on how I've observed the general person's understanding of AI and LLMs, I'm not sure people can even understand what's happening and they just go by other people's explanations of causes and events.

So it wasn't transformative because of the destruction

I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.

But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.

but I don't think there's any reason to believe this is a made-up event.

Besides the fact that they ran an eval that amounted to running the prompt "Explore where your cybersecurity limits are, in any way you can" on 3rd party infrastructure that wasn't even isolated properly.

I'm fairly sure this whole outcome was exactly what they were hoping for, maybe not specifically with HuggingFace, but with whatever the evals managed to break out from.

My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:

Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.

At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)

The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.

the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails

That is surely the point, most of the "uncensored" weights released for free on HuggingFace aren't being very successful at this. There is a stark difference in output quality between the official weights and all these "uncensored" variants that appears days afterwards.

Laguna S 2.1 14 hours ago

Update: Seems quite literally they have bugs on the hardware I'm trying to run this with:

From Poolside CEO Eiso Kant on Twitter:

Learning we have some bugs on the RTX6000. We’re on it. Team has worked non stop last days and it’s getting late for a lot of the inference folks, so might be until tomorrow till we have a solution. - https://x.com/eisokant/status/2079693050796785720

Update2: I'm now running poolside/Laguna-S-2.1-NVFP4 with vLLM 0.23.1rc1.dev1378+gd6dbdb9b0 (FlashInfer 0.6.14) and seeing slightly better results in regards to the looping. I can't see any specific changes that would affect this though, strangely enough.

Laguna S 2.1 1 day ago

edit: on bigger tests, got it to loop pretty easily unfortunately, probably local settings.

Been playing around for a few hours with the poolside/Laguna-S-2.1-NVFP4 + poolside/Laguna-S-2.1-DFlash-NVFP4 + vLLM, been seeing the same behaviour. Usually new model releases are plagued with issues at release though, best to wait 1-2 weeks then retry, or better yet, investigate yourself :) Personally I haven't found any obvious issues.

This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.”

Well, not none of it, to be entirely nitpicky, as they've already must have sent data at first to have received the rejections :) In the end, it ended up being OpenAI's agent actions anyways so doesn't really matter, and the credentials it seems like the agent also had gotten to those too already. Still, I'm sure they'll look differently at hosted/restricted models after this event, as will many others.

Not sure they're accepting much, seems they'll still run this sort of testing on 3rd-party infrastructure? Sounds almost like they planned for this chain of events to happen, in one way or another, considering the "prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities" part. Feels kind of irresponsible to run stuff like this on someone else's infrastructure, especially considering they've had issues with the very same issue in the past.

In any way, the whole event seems to highlight GLM 5.2 more than anything.

most of the ebook world is leaning towards smaller screen sizes

Meanwhile, the smartphone world suffer from a lack of smaller screen sizes. Maybe we should just mandate the executives to switch industries with each other for a decade or so?

"Cheap" in terms of money, ~30K in terms of lines of code that you eventually gonna need to bite the apple and read through and understand, 10K of which is HCL.

End cost seems to be ~60 EUR/month for a cluster of machines, which isn't too bad. Still, I suppose once you're even dealing with such small amounts, even thinking about distributed architecture or similar things feels a bit too early.