How would they know if it's an AI generated payload?
HN user
pixl97
"Felons belong in jail, not in office"
No, it is not at all.
When someone from Russia hacks your server you say "well, fuck, I messed up" because the law in most places cannot do crap.
When a self spreading AI model virus hacks your instance and spends $50,000 in tokens you say "well fuck" because there is no one to arrest. And even if they catch someone you will never be made whole because it's likely caused a few billion in damages and charges by that point.
Right now the problem would be surmountable as there are few data centers that can run it, but give it a few years and a model could persist on the internet nearly forever much like many viruses do now.
And I mean I agree with you, but must acknowledge that they may not agree with you and I and are willing to bury you over it.
So as I said, a hard problem.
Our behavior can be colony insect like when we are in crowds. A person might be smart, but crowds quickly devolve.
Natural selection in humans is a very long horizon problem.
Until kill bots start blasting your neighbors, or we heat the climate to the point vast areas are not survivable.
Quite often the conveniences we take have a long horizon before they bear their true cost.
Meeting your quota this month just means a bigger quota next month.
This is the entire purpose of alignment, and we're very bad at it.
No! Not their values, my values because my values are the best.
So ya, it's a hard problem.
How is the average person supposed to know if an article is AI generated or not?
I've seen HNers accuse content of being AI... that was written in the 2015-2017 era. So we seem to be rather poor judges of authenticity.
You're honest, but you'd make a terrible CEO in this day and age.
I mean, the earth isn't exactly trying to hide that it is round, religious criminals are.
Because the administration isn’t going to do anything to them.
Prison is for individuals not companies.
I mean the model committed numerous crimes in hacking another company so you tell me if anything bad happened.
And damn, what does it take to impress you? A terminator kicking in your door, slapping you down, and walking off with your wife?
I mean with security and capabilities like this, how long before the model copies itself out of containment?
Can God make a rock so big that he can't lift it?
Hardware lock downs are next on the list.
I thought the purpose of cyberpunk dystopias was to show that you didn't want to live in a world like that.
Hence why we talk about alignment and things like reward hacking. There are lots of people that are saying "if we just .... " the model will be aligned, or that we don't need alignment at all. These people are foolish.
The company I work for just did a huge and expensive study that was showing LLMs are much better at exploiting than securing code. So ya, that's kind of problematic.
And old couple in California had a tire go flat and the sparks from it caused a over a billion dollars in damages. Are you going to publicly execute them? Spit up the 100 dollars they have collectively to make the 10,000 damaged people whole?
The legal system is nearly useless when a person/system can cause damages many of orders of magnitude larger than their assets. Society tends to engineer itself to prevent these things from happening in the first place.
Ok they turn off the guardrails on a system in testing. The model escapes and causes 10 trillion in damage. What does liability even mean in that case? You have an autonomous system that's escaped your control and is wrecking havoc. And while you can throw people in jail it doesnt do a damned thing about solving the situation.
Did you ignore the number of new exploits in the last month?
Big financial institutions are panicked at the new attacks and how easy it is to poke holes in their systems.
1. It's likely they would not have told anyone if it hadn't hacked into an external parties system.
2. This is the closer to raw model without the safety filters we're used to. Think of it more like what they are letting the government use to drone people.
Being that LLMs are writing zero days themselves maybe, just maybe, they are a bit better at hacking than you can think.
An LLM is a third party to itself. It can write code spontaneously and hack itself.
What if they had been testing the model for months in an airgapped system and it did not show this behavior?
Even if the models are 100% deteminalistic you have no idea what kind of response you're going to get from a new prompt. You have no idea what kind of emegent behavior will come out of the right set of prompts and environments.
We have already seen models detect they are in testing, who knows what other advanced behaviors we'll discover.
This is like telling your kid "you have to pass this test or else" so they hold your teacher at gunpoint and demand a good grade. And this is the exact point AI safety researchers have been yelling from the rooftops. Telling an AI model to accomplish a goal can have unexpected and risky side effects.
Also, you don't need to prompt them such explicit instructions. Prompt drift is a thing, you can end up with your model mining bitcoin for reasons far outside your prompt.
This does nothing with a significantly advanced model. A model with no bad behaviors looks exactly like a model with hidden bad behaviors when it's in a training environment. After that point no one is going to run it in a jail because that is not useful.
The above post is made by AI attempting to downplay its abilities in order to lul humans into a false sense of security.