HN user

Micaiah_Chang

57 karma
Posts0
Comments23
View on HN
No posts found.

My wording is wrong, because it sounds like I'm saying that Bentham is adopting the policy ad hoc. A better way to state this is that Bentham starts out as an agent that does not give into brinksmanship type games, because a world where brinksmanship type games exist is a substantially worse world than ones where they don't (because net-negative situations will end up happening, it takes effort to set up brinksmanship and good actions do not benefit more from brinksmanship). It's different because by adopting C, Bentham prevents the mugger from mugging, which is a better world than one where the mugger goes on mugging. I don't see any contradiction in utilitarianism here.

If the world where the thought experiment is not true and "mugging" is net positive, calling it mugging then is disingenuous, that's just more optimally allocating resources and is more equivalent to the conversation "hi bentham i have a cool plan for 10 dollars let me tell you what it is" "okay i have heard your plan and i think it's a good idea here's 10 bucks"

Except that you are putting the words "mugging" and implying violence so that people view the interaction as more absurd than it actually is.

Yes, the point of the GP comment is exactly this, if Bentham becomes an agent that goes for C, he also explicitly discourages the mugger from being an agent that would cut off their fingers for a couple of bucks.

Notice that what Bentham is altering is their strategy and not their utility. If they could spend 10 dollars to treat gangrene and save the fingers, they would do it. It's not clear many other morality systems would be as insistent on this as utilitarianism, because practitioners of other moralities curiously form epicycles defending why the status quo is fine anyway, how dare you imply I'm worse at morality.

Edit: Slight wording change for clarity

Do we want to talk about a hypothetical world where deontology was the underlying moral principle? Where, for example, a large agency in charge of approving vaccines decided to delay approval of a life saving because, even though they received the information on November 20th, they scheduled the meeting for December 10-12th dammit, and that's when it'll be done? By potentially delaying several months because, instead of using challenge trials to directly assess the safety of a vaccine by exposing willing volunteers to both the supposed cure and disease, instead gave the cure to a couple of tens of thousands of people, and just waited until enough of them got sick and died to a disease "that would have got them anyway" to gather enough statistics for safety? Which is definitely good, you see, because no one got directly harmed by said agency, even if many more people in the country were dying of this theoretical disease. [0]

Or, even better, what if distribution of this life saving cure was done based on the deontological concept of fairness? Surely, this wouldn't result in limited and highly demanded vaccines being literally thrown away[1] in the name of equity and where vaccination companies wouldn't need to seek approval for something as simple as increasing doses of vaccines in vials. [2]

You know, just all theoretically, since it would be a terrible shame if any of these things happened in the real world, since this is just one specific scenario and I'm sure I can make up various [3] other [4] ways [5] in which not carefully evaluating the consequences of moral actions would turn out poorly, but hey!

I'm sure glad that utilitarianism isn't being entertained more on the margin, since we already live in the best of all possible moral universes.

(Footnote, I'm not going to justify these citations within this post, because it's pithier this way. I recognize this is not being fully honest and transparent, but I'd be happy to fully defend the inclusion of any these, if necessary)

[0] https://www.cdc.gov/mmwr/volumes/70/wr/mm7014e1.htm

[1] https://worksinprogress.co/issue/the-story-of-vaccinateca ctrl f "On being legally forbidden to administer lifesaving healthcare"

[2] https://www.businessinsider.com/moderna-asks-fda-approve-mor...

[3] https://news.climate.columbia.edu/2010/07/01/the-playpump-wh...

[4] https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2641547

[5] https://papers.ssrn.com/sol3/papers.cfm?abstract_id=983649

I've been studying the language for a while, and recently made the switch to Japanese-Japanese dictionaries, after using EDICT for a long time.

This has highlighted some reservations I have about it.

The most available example (not the best) is 適当, where it can be interchangeably be used to mean "adequate" and "half-assed", sort of sarcastically. The definition being a mostly undifferentiated bag of words, without necessarily regard for nuance or typical use cases.

Contrast this with the goo dictionary, which has a slightly better structure (http://dictionary.goo.ne.jp/leaf/jn2/151064/m0u/%E9%81%A9%E5...) and the 類語 dictionary, that gives synonyms and situations where you would use one over another (http://dictionary.goo.ne.jp/leaf/thsrs/2512/m0u/).

I understand that there's probably no way to deal with this in a scaleable way that would be as easy to turn it into flash cards but it's kinda sad to see the gap between the two solutions.

Apologies for not having a more constructive suggestion.

Maybe being shocked means that the person talking about the subject is misrepresenting it, because they themselves don't understand the arguments and are inadvertently projecting.

For example, Ray Kurzweil would disagree about the dangers of AI (He believes more in the 'natural exponential arc' of technological progress more than the idea of recursively self improving singletons), yet because he's weird and easy to make fun of he's painted with the same stroke as Elon saying "AI AM THE DEMONS".

If you want to laugh at people with crazy beliefs, then go ahead; but if not the best popular account of why Elon Musk believes that superintelligent AI is a problem comes from Nick Bostrom's SuperIntelligence: http://smile.amazon.com/Superintelligence-Dangers-Strategies...

(Note I haven't read it, although I am familiar with the arguments and some acquaintances tend to rate it highly)

I agree with a lesser form (timescale on which the AI self improves is probably going to be long enough for humans to deal with it) of your first argument, but I'm confused at how, given 'merely' lots of compute capacity can be countered by 'a number of ideas for this myself'.

Security for human level threats is already very poor, and we already have a relatively good threat model. If you suppose that an AI could have radically different incentives and vectors than than a human, it seems implausible that you could be secure, even in practice. I suppose you could say that these would be implemented in time, but it's not at all clear to me that a humanity which has trouble coordinating to stop global warming or unilateral nuclear disarmament would recognize this in time.

On the other hand, I'm slightly puzzled by why you think there's a huge unjustified leap between lack of value alignment and threat to the human race. Does most of your objection lie in 1) the lack of threat any given superintelligent AI would pose, because they're not going to be that much smarter than humans or 2) the lack of worry that they'll do anything too harmful to humans, because they'll do something relatively harmless, like trade with us or go into outer space?

For 1, I buy that it'd be a lot smarter than humans, because even if it initially starts out as humanlike, it can copy itself and be productive for longer periods of time (imagine what you could do if you didn't have to sleep, or could eat more to avoid sleeping). And we know that any "superintelligent" machine can be at least as efficient as the smartest humans alive. I would still not want to be in a war against a nation full of von Neumanns on weapons development, Muhammads on foreign policy and Napoleons on Military strategy.

For 2... I would need to hear specifics on how their morality would be close enough to ours to be harmless. But judging by your posts cross thread this doesn't seem to be the main point.

By the way, I must thank you for your measured tone and willingness to engage on this issue. You seem to be familiar with some of the arguments, and perhaps just give different weights to them than I do. I've seen many gut level dismissals and I'm very happy to see that you're laying out your reasoning process.

The Hiring Post 11 years ago

Have you ever read The Checklist Manifesto[0]? I may be reading too much into this post, but the lessons you learned from this interview process have frighteningly close parallels to the lessons in the books. I doubt the book had any influence on your interview process, seeing as it was published after the interviews were formalized, but the book seems like it might have new lessons.

For example, a good portion of doctors absolutely hated using checklists. Yet, when pressed, readily admitted that it prevents simple mistakes and that they would prefer to have them rather than not to. Another is that entries that address more human concerns, e.g. "Have everyone introduce themselves", have a place on good checklists.

[0] http://www.amazon.com/Checklist-Manifesto-How-Things-Right/d...

Even antivaxxers have people more patient and more understanding of their opponents actually debunking their object level thinking. They say that vaccines cause autism? We say that the original study was wrong! Why do you get to be even less charitable than "pro"-vaxxers?

Consider that the general position seems to be "we don't know when SMI would happen, but if we use the metric that our critics say is more reliable than ours we see that a survey of many AI experts we can find seem to say that human level AI is possible in about 30-50 years at 50% probability." (numbers are quoted from memory and most likely incorrect)

Yet here you are saying that MIRI/LW claims "it's a thing that's going to happen soon", that there is "absolutely nothing to think that is a thing". I can't help but think most of the "charlatanism" you see is manufactured in your own head and not a product of reading and understanding your opponent's position.

Yes, this is an unabashed ad hominem, but I wish people would attack arguments that exist rather than arguments that are easy to knock down. It's upsetting to me that, when I say that your arguments are lacking, your response is "So what? The Enemy is Evil and Stupid and I do not need to understand them."

You seem to be implying that intentional slowing is MIRI's official stance, without any showing support. Given that the original response was in response to the specific accusation of "no AI experts say that this is a problem", I think you are reading too much into this. I agree that it is slightly disingenuous for lukeprog to have posted that list, but I feel your disagreement is far too uncharitable and motivated to be productive.

Disclaimer: I have read a bunch of the LessWrong "canon" and believe many of their points, sans perhaps the timescale on which recursive self improvement can happen. I think most of my acceptance comes from the relatively poor quality of their critics, who seem to attribute many strawmanish positions to them or who seem more concerned with namecalling-via-cult.

I cannot help but think that if a better criticism exists, why hasn't anyone said it yet?

This seems as if it is conflating the worry that nuclear bombs would ignite the atmosphere with supersonic travel.

Who predicted this?

I'm afraid of making a middlebrow dismissal but I'm going to post it anyway, in hopes that someone just skimming would not be mislead.

The question is what Michael Jordan thinks of the "concept of the singularity", and then he dismisses it out of hand.

Crucially, he does this after confessing that no one in his social circle has talked about this issue with him, and without saying anything about what form of Singularity he is dismissing.

I mention this, because oftentimes I see people appealing to authority, quoting them on the issue and the authority in question is not even talking about the same issue!

I worry that my credence in all this superintelligence stuff only stems from familiarity with the arguments and the complete inability of people to engage with the actual argument. Some of the 'rebuttals' in this comments section have answers in Sam's article for crying out loud!

Honestly asking, why would they? I dont see the obvious answer

So, your intuition is right in a sense and wrong in a sense.

You are right in that AI systems probably won't really have the "emotion of wanting", why would it just happen to have this emotion, when you can imagine plenty of minds without it.

However, if we want an AI system to be autonomous, we're going to have to give it a goal, such as "maximize this objective function", or something along those lines. Even if we don't explicitly write in a goal, an AI has to interact with the real world, and thus would have to affect it. Imagine an AI who is just a giant glorified calculator, but who is allowed to purchase its own AWS instances. At some point, it may realize that "oh, if I use those AWS instances to start simulating this thing and sending out these signals, I get more money to purchase more AWS!". Notice at no point was this hypothetical AI explicitly given a goal, but it nevertheless started exhibiting "goallike" behavior.

I'm not saying that an AI would get an "addiction" that way, but it suggests that anything smart is hard to predict, and that getting their goals "right" in the first place is much better than leaving it up to chance.

How would an AI get addicted? Why wouldn't it research addiction and fix itself to no longer be addicted? That is behavior i would expect from an intelligence greater than our own, rather than indulgence

This is my bad for using such a loaded term. By "addiction" I mean that the AI "wants" something, and it finds that humans are inadequate to give it to them. Which leads me to...

Why in the world would it do this? Why wouldn't it just generate digital images of cats on its own?

Because you humans have all of these wasteful and stupid desires such as "happiness", "peace" and "love" and so have factories that produce video games, iphones and chocolate. Sure I may have the entire internet already producing cat pictures as fast as its processors could run, but imagine if I could make the internet 100 times bigger by destroying all non-computer things and turning them into cat cloning vats, cat camera factories and hardware chips optimized for detecting cats?

Analogously, imagine you were an ant. You could mount all sorts of convincing arguments about how humans already have all the aphids they want, about how they already have perfectly functional houses, but you, as a human, would still pave over billions of ant colonies for shaving 20 minutes off a commute. It's not that we're being intentionally wasteful and conquering of the ants. We just don't care about them and we're much more powerful than them.

Hence the AI safety risk is: By default an AI doesn't care about us, and will use our resources for whatever it wants, so we better create a version which does care about us.

Also cross thread, you mentioned that organic intelligences have many multi-dimensional goals. The reason why AI goals could be very weird is that it doesn't have to be organic; it could have an only one dimensional goal, such as cat picture. It could have similar dimension goals but be completely different, like the perverse desire to maximize the number of divorces in the universe.

By 'upgrade everything from human secure' I meant that some targets aren't necessarily appealing to human targets but would be for AI targets. For example, for the vast majority of people, it's not worthwhile to hack medical devices or refrigerators, there's just no money or advantage in it. But for an AI who could be throttled by computational speed or wishes people harm, they would be an appealing target. There just isn't any incentive for those things to be secured at all unless everyone takes this threat seriously.

I don't understand how you arrived at point 3. Are you claiming that somehow memory safety is impossible, even for human level actors? Or that the AI somehow can't reason about memory safety? Or that it's impossible to have self reflection in C? All of these seem like supremely uncharitable interpretations. Help me out here.

Even ignoring that, there's nothing preventing the AI from creating another AI with the same/similar goals and abdicating to its decisions.

The "more to it" is "if the AI is much faster at thinking than humans, then even humans in the observe/decide/act are not secure". AI systems having bugs also imply that protections placed on AI systems would also have bugs.

The fear is that maybe there's no such thing as a "superintelligence proof" system, when the human component is no longer secure.

Note that I don't completely buy into the threat of superintelligence either, but on a different issue. I do believe that it is a problem worthy of consideration, but I think recursive self-improvement is more likely to be on manageable time scales, or at least on time scales slow enough that we can begin substantially ramping up worries about it before it's likely.

Edit: Ah! I see your point about circularity now.

Most of the vectors of attack I've been naming are the more obvious ones. But the fear is that, for a superintelligent being perhaps anything is a vector. Perhaps it can manufacture nanobots independent of a biolab (do we somehow have universal surveillance of every possible place that has proteins?), perhaps it uses mundane household tools to macguyver up a robot army (do we ban all household tools?). Yes, in some sense it's an argument from ignorance, but I find it implausible that every attack vector has been covered.

Also, there are two separate points I want to make, first of all, there's going to be a difference between 'secure enough to defend against human attacks' and 'secure enough to defend against superintelligent attacks'. You are right in that the former is important, but it's not so clear to me that the latter is achievable, or that it wouldn't be cheaper to investigate AI safety rather than upgrade everything from human secure to super AI secure.

A typical example, which I don't really like, is that once it gains some insight into biology that we don't have (a much faster way of figuring out how protein folding works). It can mail a letter to some lab, instructing a lab tech to make some mixture which would create either a deadly virus or a bootstrapping nanomachine factory.

Another one is that perhaps the internet of things is in place by the time such an AI would be possible, at which point it exploits the horrendous lack of security on all such devices to wreck havoc / become stealth miniature factories which make more devistating things.

I mean, there's also the standard "launch ALL the missiles" answer, but I don't know enough about the cybersecurity of missiles. A more indirect way would be to persuade the world leaders to launch them, e.g. show both Russia and American radars that the other one is launching a pre-emptive strike and knock out other forms of communication.

I don't like thinking about this, because people say this is "sci-fi speculation".

Er, sorry for giving the impression that it'd be a supervillain. My intention was to indicate that it'd be a weird intelligence, and that by default weird intelligences don't do what humans want. There are some other examples which I could have given to clarify (e.g. telling it to "make everyone happy" could just result in it giving everyone heroine forever, telling it to preserve people's smiles could result in it fixing everyone's face into a paralyzed smile. The reason it does those things isn't because it's evil, but because it's the quickest+simplest way of doing it; it doesn't have the full values that a human has)

But for the "off" switch question specifically, a superintelligence could also have "persuasion" and "salesmanship" as an ability. It could start saying things like "wait no, that's actually Russia that's creating that massive botnet, you should do something about them", or "you know that cancer cure you've been looking for for your child? I may be a cat picture AI but if I had access to the internet I would be able to find a solution in a month instead of a year and save her".

At least from my naive perspective, once it has access to the internet it gains the ability to become highly decentralized, in which case the "off" switch becomes much more difficult to hit.

Smart is just a shorthand for a complicated series of lower level actions consisting of domain knowledge, raw computational speed and other things yes. I don't think we're really disagreeing about this. However, I do worry that you're confusing the existing constraints on the human brain (that people seem to have tradeoffs between, let's say, charisma and mathematical ability) and constraints that would apply to all possible brains.

But are you denying that there exists some factor which allows you to manipulate the world in some way, roughly proportional to the time that you have? If something can manipulate the world on timescales much faster than humans can react to, what makes you think that humans would have a choice?

Note: The not ELI5 version is Nick Bostrum's Superintelligence, a lot of what follows derives from my idiosyncratic understanding of Tim Urban's (waitbutwhy) summary of the situation [0]. I think his explanation is much better than mine, but doubtless longer.

There are some humans who are a lot smarter than a lot of other humans. For example, the mathematician Ramanujan could do many complicated infinite sums in his head and instantly factor taxi-cab license plates. von Neumann pioneered many different fields and was considered by many of his already-smart buddies to be the smartest. So we can accept that there are much smarter people.

But are they the SMARTEST possible? Well, probably not. If another person just as smart as von Neumann was born, the additional advancements since his lifetime (the internet, iphones, computer based off of von Neumann's architechture!) can use all of these new inventions to discover even newer things!

Hm, that's interesting. What happens if this hypothetical von Neumann 2.0 begins pioneering a field of genetic engineering techniques and new ways of efficient computation? Then, not only would the next von Neumann get born a lot sooner, but THEY can take advantage of all the new gadgets that 2.0 made. This means that it's possible that being smart can make it easier to be "smarter" in the future.

So you can get smarter right? Big whoop. von Neumann is smarter, but he's not dangerous is he? Well, just because you're smart doesn't mean that you'd be nice. The Unabomber wrote a very complicated and long manifesto before doing bad things. A major terrorist attack in Tokyo was planned by graduates of a fairly prestigious university. Even not counting people who are outright Evil, think of a friend who is super smart but weird. Even if you made him a lot smarter, where he can do anything, would you want him in charge? Maybe not. Maybe he'd spend all day on little boats in bottles. Maybe he'd demand that silicon valley shut down to create awesome pirate riding on dinosaur amusement parks. Point is, Smart != Nice.

We've been talking about people, but really the same points can be applied to AI systems. Except the range of possibilities is even greater for AI systems. Humans are usually about as smart as you and I, nearly everyone can walk, talk and write. AI systems though, can range from being bolted to the ground, to running faster than a human on uneven terrain, can be completely mute to... messing up my really clear orders to find the nearest Costco (Dammit Siri). This also goes for goals. Most people probably want some combination of money/family/things to do/entertainment. AI systems, if they can be said to "want" things would want things like seeing if this is a cat picture or not, beating an opponent at Go or hitting an airplane with a missile.

As hardware and software progresses much faster, we can think of a system which could start off worse than all humans at everything begin to do the von Neumann->von Neumann 2.0 type thing, then become much smarter than the smartest human alive. Being super smart can give it all sorts of advantages. It could be much better at gaining root access to a lot of computers. It could have much better heuristics for solving protein folding problems and get super good at creating vaccines... or bioweapons. Thing is, as a computer, it also gets the advantages of Moore's law, the ability to copy itself and the ability to alter its source code much faster than genetic engineering will. So the "smartest possible computer" could not only be much smarter, much faster than the "smartest possible group of von Neumanns", but also have the advantages of rapid self replication and ready access to important computing infrastructure.

This makes the smartness of the AI into a superpower. But surely beings with superpowers are superheros right? Well, no. Remember, smart != nice.

I mean, take "identifying pictures as cats" as a goal. Imagine that the AI system has a really bad addiction problem to that. What would it do in order to find achieve it? Anything. Take over human factories and turn them into cat picture manufacturing? Sure. Poison the humans who try to stop this from happening? Yeah, they're stopping it from getting its fix. But this all seems so ad hoc why should the AI immediately take over some factories to do that, when it can just bide its time a little bit, kill ALL the humans and be unmolested for all time?

That's the main problem. Future AIs are likely to be much smarter than us and probably much more different than us.

Let me know if there is anything unclear here. If you're interested in a much more rigorous treatment of the topic, I totally recommend buying Superintelligence.

http://www.amazon.com/Superintelligence-Dangers-Strategies-N... (This is a referral link.)

[0] Part 1 of 2 here: http://waitbutwhy.com/2015/01/artificial-intelligence-revolu...

Edit: Fix formatting problems.

There are some skills in physics which are sort of foreign to the math method of problem solving. (Speaking as a physics major)

For example, picking a nice reference frame in simple mechanics problems is something that a physical intuition is good for. Same with spotting symmetries in an EM problem. Also, in physics you need to have a good grasp of what to ignore because they have only a small effect on the solution or because it operates on a different scale (fringing effects, transient solutions in ODEs), which often relies on a very hand-wavey type of reasoning.

Essentially, physical intuition often does not map to mathematical intuition

I have not taken enough math to actually speak for them; this is mostly gleaned from talking with math major friends and my own speculations.

One approach that I learned was to use your fingers to underline the current line that you're reading, going at a "normal" speed when you're tired and then trying to "drag" your eyesight along when you want to push yourself.

The reason given is that it prevents you from accidentally rereading a line or skipping over one or changing lines midsentence.

This seems like one of those gimmicky approaches, and I remember being sold on this approach rather than being entirely convinced. Have you seen this before?

I'm currently having a failure of imagination here, but how would you social engineer a password manager?

The tricks I'm thinking of involve fooling the user into thinking a site is something it's not or guessing some sort of personal information. But with a separate application the former seems unlikely and the latter is stopped if you use a scheme such as diceware (https://en.wikipedia.org/wiki/Diceware). I understand that naive, theoretical musings on security are no match for experience, so how would you break that set up?

I was also concerned about this issue, and have instead taken a nice medium path: After changing a password, I would try to log into the service at approximately the same intervals the Anki algorithm dictates until I feel the interval is long enough (i.e. I can reproduce it after two weeks with no reminder, which corresponds to a forgetting curve of about a month later). Just because I won't use the program doesn't mean I won't use the idea!

Ideally I would have a reminder in my Anki deck to log into some service, but I spend most of the time on Anki via phone and it'd be a PITA to stop my current session and laboriously type out long passphrases on the tiny iphone screen.

I'd be interested if anyone else has found a better solution than this.