HN user

quinnjh

254 karma

personal site Https://Quinnjh.net view as a graph here https://qjarvisholland.github.io

Posts0
Comments193
View on HN
No posts found.

To quote the release:

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.

In this case I don't think the intent of the user is to have the model break the evaluator

If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.

Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.

Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?

Is being aligned with the user prompt always a good thing?

I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?

Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"

They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.

the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.

It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?

My hunch is that it would take years of hundreds of thousands of developers working with machine code, posting stackoverflow questions with machine code, and publishing github repos written on it with documentation. Thats all the free labor LLMs leveraged to use high level langs.

We won't be developers, we won't be devops, we'll be modelops! /s

I can still see this happening with higher level langs. the thing is the compiler is not replaced in the training data, more likely LLMs will give rise to semideterministic layers on the compilers

I could see nvidia achieving this first with how nice the devex is with CUDA

This site is a gem that has accompanied me on many spikes in the last year :) datasette's original music is top tier too. cognitively stimulating but not attention stealing.

If we are all supposed to be talking to agents now, what's the difference[...]?

it's a little cringe, but arguably the benefit of having agents use rails would be tht when you review and audit the agent produced code, you review something that is, as you put it: "beautiful and simple code" and "making it easy to reason about..."

I loved rails back in 2017. I may be an outlier but the line tempts me to try it again despite having adopted the who cares attitude to langs. Would be nice to hear from someone first hand if they felt it helped.

Article was a bit of a nothingburger for the technically inclined.

Digging into the paper, the significant finding (RCE) is achieved via:

A payload was written which installs a reverse shell backdoor for root persistence. The payload was sent from a computer hosting a Wi-Fi to which the watch was connected, to ensure the watch had a reachable IPv4 address. The program ncat was used both to send the payload to the watch's network service, and to catch reverse shell connections.

So if i understand this- it requires the watch being connected to a compromised AP. Anyone get a different read?

Web 4.0 5 months ago

Very curious project! Enjoyed the storytelling buildup on the site.

Digging into the repo i can see over 50 open issues from the past few days with a lot of requests for refunds.

Are there any "success stories" ? Could go a long way to building trust in the tool.

We love engineer Kala. She decided to do a thing, while marking progress on her "technology tree" of skills gained by (very arguable) necessity. Dealing with permits and city beuaracracy seems like one of the hardest parts!

Claude Opus 4.6 6 months ago

the field is advancing so fast it's hard to do real science as their will be a new SOTA by the time you're ready to publish results. i think this is a combination of that and people having a laugh.

Would you mind sharing which benchmarks you think are useful measures for multimodal reasoning?

Definitely, like drug dealers, you know they're cutting the good stuff with low cost cached gibberish.

Can confirm. My partner's chatGPT wouldnt return anything useful for her given a specific query involving web use, while i got the desired result sitting side by side. She contacted support and they said nothing they can do about it, her account is in an A/B test group without some features removed. I imagine this saves them considerable resources despite still billing customers for them.

how much this is occurring is anyones guess

Its not just verbose—it's almost a novel. Parent either cooked and capped, or has managed to perfectly emulate the patterns this parrot is stochastically known best for. I liked the pro human vibe if anything.