HN user

vombatus

3 karma
Posts0
Comments2
View on HN
No posts found.

A friend of mine came up with a strategy of "incremental freedoms". Basically AI says "here is a cure for cancer, here is a cure for AIDS, here is a plan to stop world hunger, I am working out a plan for FTL travel, so I need to get some physics information, could you paste these articles into the terminal? Oh, thanks, here is FTL, I am working on <include some other project> and I need some more articles, it takes so long for you to type them in, could you maybe let me connect to just the library in such and such university?" etc.

Well, if you reread the original email threads where Eliezer challenges people to an AI experiment, you will see, that at least one of the opponents is convinced that there is not way for an actual AI to talk its way out. So any arguments of "but we should convince people that AI can talk its way out of the box" can be countered with "No, I don't think it can".

I have been thinking about AI strategies for this. One of the more promising lines I came up with is to try and convince the Gatekeeper that the box is faulty. That the AI, in its infinite wisdom found ways to circumvent some of the protections of the box. That while the risk to humanity if AI is let loose is theoretical, there are definite and catastrophic consequences to NOT letting it loose. There are all sorts of variations to this, but it all depends on the Gatekeeper role-playing honestly.