HN user

casebash

110 karma
Posts9
Comments49
View on HN

Authors of original paper: Samuel G. B. Johnson, Amir-Hossein Karimi, Yoshua Bengio, Nick Chater, Tobias Gerstenberg, Kate Larson, Sydney Levine, Melanie Mitchell, Iyad Rahwan, Bernhard Schölkopf, Igor Grossmann

I'll copy my LinkedIn comment:

"Well done to the UK for not signing the fully compromised Statement on Inclusive and Sustainable Artificial Intelligence for the People and the Planet. Australia shouldn't have signed this statement either given how France intentionally derailed attempts to build a global consensus on how we can develop AI safely.

For those who lack context, the UK organised the AI Safety Summit at Bletchley Park in November 2023 to allow countries to discuss how advanced AI technologies can be discussed safely. There was a mini-conference in Korea, France was given the opportunity to organise the next big conference, a trust they immediately betrayed by changing the event to be about promoting investment in their AI industry.

They renamed the summit to the AI Action Summit and relegated safety from the sole focus to being just one of five focus areas, but not even one of five equally important focus areas, but one that seems to have been purposefully minimized even further.

Within the conference statement safety was reduced to a single paragraph that undermines safety if anything:

“Harnessing the benefits of AI technologies to support our economies and societies depends on advancing Trust and Safety. We commend the role of the Bletchley Park AI Safety Summit and Seoul Summits that have been essential in progressing international cooperation on AI safety and we note the voluntary commitments launched there. We will keep addressing the risks of AI to information integrity and continue the work on AI transparency.”

Let’s break it down: • First, safety is being framed as “trust and safety”. These are not the same things. The word trust appearing first is not as innocent as it appears: trust is the primary goal and safety is secondary to this. This is a very commercial perspective, if people trust your product you can trick them into buying it, even if it isn't actually safe. • Second, trust and safety are not framed as values important in and of themselves, but as subordinate to realising the benefits of these technologies, primarily the "economic benefits". While the development of advanced AI technologies could theoretically create a social surplus that could be taxed and distributed, it's naive to assume that this will be automatic, particularly when the policy mechanisms are this compromised. • Finally, the statement doesn’t commit to continuing to address these risks, but only narrowly to “addressing the risks of AI to information integrity” and “continue the work on AI transparency”. In other words, they’re purposefully downplaying any more significant potential risks, likely because discussing more serious risks would get in the way of convincing companies to invest in France.

Unfortunately, France has sold out humanity for short-term commercial benefit and we may all pay the price."

Most of the comments here only make sense under a model where AI isn't going to become extremely powerful AI in the near term.

If you think upcoming models aren't going to be very powerful, then you'll probably endorse business-as-usual policies such as rejecting any policy that isn't perfect or insisting on a high bar of evidence before regulating.

On the other hand, if you have a world model where AI is going to provide malicious actors with extremely powerful and dangerous technologies within the next few years, then instead of being radical, proposal like this start appearing extremely timid.

Will Oprah screw up the AI story?

Quite possibly, but likely not as bad as this article.

Complete clickbait title, assumes that the author's hobby horses are the most important thing in the world, bizarrely argues that crypto hype is an "attack on labour".

I'm not going to try to recap all of that, but, as an example, if you have a sufficiently strong understanding of arithmetic, learning basic modular arithmetic should be effortless, pigeonhole principle completely obvious.

I was quite surprised when I tried applying for a Microsoft internship in uni and they gave me a question on the pigeon-hole principle.

Just thought I'd add a comment as someone who came top of the state in my grade in multiple olympiad competitions:

I always felt that a large part of my advantage came from having a strong understanding of maths from the ground up.

I felt that a lot more people could have gained the same level of understanding as I did if they had been willing to work hard enough, but I also felt that almost no-one would, because it'd be an incredibly hard sell to convince someone to engage in years-long project where they'd go all the way back to kindergarten and rebuild their knowledge from the ground up.

In other words, excellence is often the accumulation of small advantages over time.

I expect this to end up having been one of the worst timed blog posts in history. Open source AI has mostly been good for the world up until now, but we're getting to the point where we're about to find out why open-sourcing sufficiently bad models is a terrible idea.

While this article makes some valid points, it basically just ignores the reasons why the law is being passed, that is the potential for open-models to enable bio-attacks, cyberattacks, election manipulation, automated personalised scams, and who knows what else.

One might question why that is. Perhaps it's the case that Jeremy has an excellent response to these points which he has somehow neglected to raise. Or perhaps it's because these threats are very inconvenient for an open source developer.

I'm sure he'd say that open-sourcing models means that all actors have access to defensive systems and that the good guys outnumber the bad guys and it'll all work out well.

And that could be true. Or it could be false. It's not like we really know that everything would work out fine. It's not that we've run the experiment. I mean maybe it works out like that, or maybe one guy creates a virus and then it doesn't really matter how many folk on the other side, but we still get kind of screwed because we can only produce vaccines that fast. It's that's what going to happen? I don't really know, but it's at least plausible. I mean, maybe we'll automate all aspects of vaccine production and be able to respond much faster, but that's dependent on when we develop this technology vs. when AI starts significantly helping with bioweapons with someone then using it for an attack. And at that point it's all so uncertain and up in the air that it's seems rather strange for someone to suggest that it'll all be fine.

"This paper follows a recent trend of marketing excellent theoretical work as LLMs being capable of secretly plotting behind your back, when the realistic implication is backdoor risk".

Many top computer scientists consider loss of control risks to be a possibility that we need to take seriously.

So the question then becomes, is there a way to apply science to gain greater clarity on the possibility of these claims? And this is very tricky, since we're trying to evaluate claims not about models that currently exist, but about future models.

And I guess what people have realised recently is that, even if we can't directly run an experiment to determine the validity of the core claim of concern, we can run experiments on auxiliary claims in order to better inform discussions. For example, the best way to show that a future model could have a capability is to demonstrate that a current model possesses that capability.

I'm guessing you'd like to see more scientific evidence before you want to take possibilities like deceptive alignment seriously. I think that's reasonable. However, work like this is how we gather that evidence.

Obviously, each individual result doesn't provide much evidence on its own, but the accumulation of results has helped to provide more strategic clarity over time.

Should AI Be Open? 2 years ago

OpenAI just released a response to Musk's lawsuit.

One of the emails provided as evidence of their claims is Slate Star Codex's "Should AI be open?"

The emails show Musk forwarding him an email he received to Sam Altman and then Ilya replying that for a hard takeoff scenario, open sourcing could make it easier for a bad actor to reach AGI first and that the right strategy would be to share everything in the short and possibly medium term and that it would make sense to start being less open as they got closer to AGI.

https://openai.com/blog/openai-elon-musk#email-4

Yeah, at first I read that as it using 26.8% of the original steps, but reducing the number of steps by 26.8% is not that impressive. I wonder whether it actually reduces total search time as there is added overhead of running the neural network.

EA here:

"Discussing the risks and opportunities in front of us intelligently, e/accs believe, is a sign of a flourishing civil society"

Except e/acc has made a massive contribution to lowering the standard of discourse. Beff talked intelligently on Lex and they are much more reasonable on Twitter Spaces, but on the Twitter timeline itself 90% of their posts are some combination of trash/propaganda/insults.

"Rather, their moral vision is one where more people — including and especially those who consider themselves hands-off today — actively engage with emerging technology and identify concrete plans for its development and stewardship, rather than reflexively backing away from what they don’t understand"

Again, e/acc seems to be all about "build, build, build!" which stands in stark contrast to taking a step back and thinking carefully about the impacts of what you're doing before you do it.

"Discussing the risks and opportunities in front of us intelligently, e/accs believe, is a sign of a flourishing civil society."

Again, this isn't accurate. E/acc is very much not about balance and also very much not about about discussing risks intelligently vs. almost always criticising the people making the claims instead of engaging in discussion on a technical level.

...

Criticism aside, this article paints a picture of e/acc which, while not representative of the movement as it exists, is something which it could choose to grow into.

As someone who majored in philosophy, it's kind of embarrassing that I would struggle to name names as to who might be likely to be seen as the most brilliant by future generations.

I suppose if I'm allowed to name someone who only died recently, I suspect that Derek Parfit will be remembered as someone who was simply brilliant. His teletransportation problem will remain a classic thought experiment and he'll be remembered for his work on population ethics, though it's less clear to me what the future will make of his main project: working out a unified moral theory.

On the other hand, I can name some^ names as to which philosophers are most likely to be seen as influential:

Currently, it's looking like Peter Singer, Toby Ord and Will MacAskill will all have had significant influence through the Effective Altruism movement. Similarly, Nick Bostrom has had significant influence through his book Superintelligence.

Dan Dennet had significant influence by being one of the "Four Horsemen" of New Atheism. His direct influence peaked a long time ago, but it remains to be seen how responsible he will ultimately be seen for the West's decline in religion (or whether this will just be seen as as trend that was happening anyway).

^ Possibly biased

Why would Roko's basilisk play a big part in your reasoning?

In my experience, it's basically never been a part of serious discussions in EA/LW/AI Safety. Mostly, comes up when people are joking around or when speaking to critics who raise it themselves.

Even in the original post, the possibility of this argument was actually more of a sidenote on the way to main point (admittedly, he's main point involved an equally wacky thought experiment!).

My Left Kidney 3 years ago

FTX didn't actually fund the castle, from what I've heard. It was another group.

Just because people decide to focus their time on one issue (say AI), doesn't mean that they don't appreciate people working on other issues. Another example: I can appreciate a local soccer coach without having to believe that their work is equally important as that of the president/prime minister.

ChatGPT Summary:

The article explores the potential impact of artificial intelligence (AI) on society and human consciousness. The author argues that AI has the potential to transform the world in ways that are difficult to predict, and that we may be underestimating the speed at which this transformation will occur. The author notes that AI is already being used in a variety of applications, and that it is likely to become increasingly powerful and pervasive in the coming years. The author raises concerns about the potential risks associated with AI, including the possibility that it could lead to the extinction of the human race. The author argues that we need to accelerate our adaptation to these technologies, or make a collective decision to slow their development.

(Sadly, the summary makes the article sound boring, but this is just how it is with ChatGPT summaries)

I actually think that open source AI will be a terrible idea for sufficiently advance AI for the same reasons that open source nukes or bioweapons would be a mistake. So I'd be very happy to see them change their name to ClosedAI.

They are still dedicated towards trying to advance humanity and I think that's more important than non-profit status. (For the record, I'm uncertain whether they are net positive or not given that I am skeptical of their approach to safety).