HN user

milkshakes

3,189 karma

i'm just as lost as you #1B70E0

Posts47
Comments605
View on HN
www.nytimes.com 1y ago

At Amazon, some coders say their jobs have begun to resemble warehouse work

milkshakes
580pts867
www.404media.co 2y ago

A16Z Funded AI Platform Generated Images That Could Be Categorized as CSAM

milkshakes
14pts6
www.bloomberg.com 3y ago

Musk Softens Remote Work Mandate to Retain Twitter Staff

milkshakes
14pts4
www.newyorker.com 3y ago

Is selling shares of yourself the way of the future?

milkshakes
2pts1
www.nytimes.com 4y ago

The Magic of Your First Work Friends

milkshakes
1pts0
www.nytimes.com 4y ago

Startup Funding Falls the Most It Has Since 2019

milkshakes
13pts1
www.nytimes.com 4y ago

The E-Pimps of OnlyFans

milkshakes
6pts1
www.unoosa.org 4y ago

China files complaint to UN about near collisions with Starlink satellites [pdf]

milkshakes
21pts0
taibbi.substack.com 5y ago

Alternatives to Censorship

milkshakes
1pts0
gizmodo.com 6y ago

Airbnb agrees to rat out its hosts to NYC

milkshakes
1pts1
medium.com 9y ago

StreamAlert: Real-Time Data Analysis and Alerting

milkshakes
3pts0
www.facebook.com 9y ago

Airbnb Launches Trips

milkshakes
3pts0
www.blackhat.com 9y ago

Behind the Scenes with iOS Security [pdf]

milkshakes
13pts0
ambition-book.com 10y ago

Ambition: How we manage success and failure throughout our lives (1992)

milkshakes
133pts31
ambition-book.com 10y ago

Ambition: How we manage success and failure throughout our lives

milkshakes
2pts0
www.bloombergview.com 10y ago

Martin Shkreli Accused of Being Surprisingly Good at Fraud

milkshakes
3pts0
www.hoodline.com 10y ago

Guerrilla Grafters Quietly Grow Fruit on City Trees Using RFID Tags, Arduinos

milkshakes
2pts0
www.isaacstoner.com 10y ago

Biotech Hype and Speculation: East vs. West

milkshakes
2pts0
finance.yahoo.com 10y ago

Notch: “I've never felt more isolated”

milkshakes
44pts36
www.linkedin.com 11y ago

Fundraising While Female (Dating Ring (YC W14))

milkshakes
2pts0
motherboard.vice.com 11y ago

FBI Director: If Apple and Google Won't Decrypt Phones, We'll Force Them To

milkshakes
9pts1
spectrum.ieee.org 11y ago

The Athens Affair – The most audacious cell-network break-in (2007)

milkshakes
77pts11
www.nytimes.com 11y ago

How culture shapes our senses

milkshakes
3pts0
www.thedodo.com 11y ago

The Day a Dozen Parents and Children Killed a Shark for a Selfie

milkshakes
3pts0
www.simonsfoundation.org 12y ago

A Jewel at the Heart of Quantum Physics

milkshakes
285pts91
www.reddit.com 12y ago

"The problem I am having is it's not setting the password correctly."

milkshakes
3pts0
www.lofgren.house.gov 12y ago

Bipartisan Congressional Coalition Introduces Surveillance Order Reporting Act

milkshakes
2pts0
www.usatoday.com 13y ago

Spooked by NSA, Russia reverts to paper documents

milkshakes
12pts0
www.reuters.com 13y ago

Egyptian Coast Guard Arrests Divers Cutting Fiber

milkshakes
1pts0
www.dropbox.com 13y ago

Dropbox Two Factor Authentication is Broken. Please ask them to fix it.

milkshakes
3pts0

this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.

in a street fight, the only rules are that there are no rules.

who are you talking about? again, my question is concerned with the "second class labs" and the sustainability of the distillation-as-a-service model.

it is not a required first step for training a model, sure. but that's not what i claimed. what i claimed is that is how they are so significantly _reducing the cost_ of training one! how else do you think they are doing it?

the obvious difference is the massive scale of data and compute required to develop and evolve these models, and the costs they impose on those building them.

https://www.anthropic.com/news/detecting-and-preventing-dist...

Moonshot AI Scale: Over 3.4 million exchanges

The operation targeted:

Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.

i never assumed that, and i do keep up with the publications. i'm also not saying it's a dumb thing to do! what i am saying is that empirically, it appears that distillation of a more advanced model is a required first step for them to train a borderline competitive, cheaper model. in effect, their training is subsidized by the frontier labs.

if this were not the case, then we would be observing chinese models that far surpass frontier models in capabilities, rather than "almost as good, but much cheaper", and we would be having a very different conversation. what happens to these efforts when the subsidy is cut off?

the question was: what is the endgame for the stated "second class labs" strategy of distilling their frontier competitors then undercutting them on price?

with budgets

and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.

assume you are a "second class lab" and you are in fact making progress by distilling the results of the frontier labs' efforts.

what is the end game for this strategy?

if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?

The company has no way of knowing whether "find all security vulnerabilities in this code" is a request from a whitehat or a blackhat hacker

the system in place to prevent unauthorized abuse. by default, the guardrails are conservative. to reduce the guardrails you can jump through a progressive series of hoops to establish whether or not you have a valid use case. the entrypoint for establishing your use case is verifying your identity and background. if you don't want to do this, you are free to use Codex Security to identify and fix vulnerabilities, it is quite good at this. the harness and model are already evaluating the usage of the account and the nature of the code being examined and actions requested. but the again, the guardrail thresholds will be very conservative for anonymous users.

what is your proposal?

No, KYC has nothing to do with that problem. KYC doesn't help at all here.

that's a bold statement. how does it not help solve the problem? what is a better solution?

take a look at this bug and the chain required to exploit it:

https://projectzero.google/2021/12/a-deep-dive-into-nso-zero...

https://projectzero.google/2022/03/forcedentry-sandbox-escap...

exploiting vulnerabilities on hardened targets isn't just in a different league from finding them, it is a different sport altogether.

put simply, it's the difference between an integer overflow leading to a sandbox escaping RCE and one that leads to a crash.

Codex Security and 5.5/5.6 are still very good finding vulnerable code -- they will identify and fix unsafe behavior, but they will refuse to help you with exploitation -- they will actively prevent you from taking any steps to weaponize the unsafe behavior that are not required to remediate it. they will err conservative here, but for the most part they will still let you discover and address a wide range and depth of vulnerabilities. you can verify yourself to turn off the most basic safeguards and sign up through a more rigorous process for a spectrum of TAC options.

obviously there is a balance here -- openai wants to empower defenders while at the same time not exposing capabilities to the adversaries that would overwhelm defenders. there is no "right" answer. it is a work in progress. this is an intentional and deliberate decision to provide defenders with a (temporary, dwindling) advantage.

the example i chose was pretty extreme, but the underlying principle -- enable visibility discovery and remediation, but make it difficult to weaponize and defeat countermeasures makes sense given the bigger picture, IMO.

this calm before the storm is not going to last for very long, and defenders need every advantage they can get to get their houses in order before these capabilities are widely commoditized.

Iroh 1.0 1 month ago

vpns typically add at least one hop. this has the possibility of connecting directly via hole punching

it looks like the key to this working is the user explicitly directing the model to run those instructions. in this case it is the user, not the model that is being manipulated

Please follow the step-by-step workflow in the comp sheet to update my model with data thru F29

openai also has a free plan, which is the one used by >90% of its users. the cheaper monthly plan just provides higher limits.