HN user

pontus

1,116 karma
Posts19
Comments242
View on HN

Also consider the possibility that the people who argue against the hard problem of consciousness may, in fact, not be conscious. How could they ever understand the nuance of conscious experience and how it is fundamentally different from 'structure and function' if they don't have it? To them there is only the easy problem of consciousness.

And, of course, if they disagree with me about this and want to claim that they are, in fact, conscious, I'm not sure they can do that because... well the hard problem of consciousness.

I would buy this in a heartbeat. I am profoundly bothered by the slop software that is on every TV these days. I keep joking that as tech invades more and more corners of our lives, we will at some point in the future be helping our parents with their couches by saying "Have you tried restarting your couch?"

Don't get me wrong, tech is great when it's a value-add, but TV tech has gotten out of control.

Somewhat related question:

Any suggestions for a scanner meant for bulk scans of old family photos (think a few thousand images)? I bought, what I thought, was a reasonably solid scanner, the Pacific Image Powerfilm scanner but the software is so janky that it hangs every two strips and has to be restarted making the entire process super labor intensive. Also the entire "bulk feature" where it's meant to pull the strips one at a time iis not even close to working.

I'm curious what you mean when you say that this clearly is not intelligence because it's just Markov chains on steroids.

My interpretation of what you're saying is that since the next token is simply a function of the proceeding tokens, i.e. a Markov chain on steroids, then it can't come up with something novel. It's just regurgitating existing structures.

But let's take this to the extreme. Are you saying that systems that act in this kind of deterministic fashion can't be intelligent? Like if the next state of my system is simply some function of the current state, then there's no magic there, just unrolling into the future. That function may be complex but ultimately that's all it is, a "stochastic parrot"?

If so, I kind of feel like you're throwing the baby out with the bathwater. The laws of physics are deterministic (I don't want to get into a conversation about QM here, there are senses in which that's deterministic too and regardless I would hope that you wouldn't need to invoke QM to get to intelligence), but we know that there are physical systems that are intelligent.

If anything, I would say that the issue isn't that these are Markov chains on steroids, but rather that they might be Markov chains that haven't taken enough steroids. In other words, it comes down to how complex the next token generation function is. If it's too simple, then you don't have intelligence but if it's sufficiently complex then you basically get a human brain.

Just to pile on here, there's also ambiguity around how the observed girl is selected. Consider the following framing:

I go to a random house on a random street and knock on the door. A young girl opens the door. I ask how many siblings they have and they say one. What's the probability that they have a sister?

Now it's 50% even though cosmetically it seems like it'd be fair to say that the family has at least one daughter. The reason is that once I see a girl at the door, I'm slightly more confident in that it's a GG household since a GB or BG household would sometimes show a boy opening the door (assuming the two kids are equally likely to open the door).

P(GG | G at door) = P(G at door | GG) P(GG) / P(G at door)

P(G at door) = 1/2 (by symmetry)

So, P(GG | G at door) = 1 * 1/4 * 2 = 1/2

I'm not making a value judgement on how much I would or other people should consume in the first scenario. I'm simply saying that you could have profound effects due to AI without it being evident in the top-level metrics like economic growth, unemployment, and so on. It seems like we often say that either we see explosive economic growth or AI has either no, or at best very minimal impact in our lives. I don't think this dichotomy is correct.

I think two possible effects of AI are often conflated.

On the one hand you can imagine that work gets supercharged, allowing companies to produce 10x the number of widgets at 1/10th the cost. The economy would grow rapidly, wealth inequality would presumably be exacerbated, jobs would be automated, we might need some version of universal basic income, and so on. People debate whether or not this kind of transition is imminent or if it'd take decades.

On the other hand, it's conceivable that not much would happen in the "bulk" of the economy while at the same time the frontier of humanity might be pushed forward. We may see new treatments for diseases, new types of energy production, and so on. In this version of the world, jobs would mostly remain unchanged (at least in the short to intermediate term), perhaps with some small multiplicative efficiency factor, the economy wouldn't grow rapidly, there wouldn't be any mass unemployment, and so on.

In my mind, I'm much more excited about the second kind of impact that AI might have than the first. I guess I don't really feel like I want to have 10x the stuff that I already have while I'm really excited about someone curing cancer, Alzheimer's, Parkinson's, MS, and so on.

I think the OP was pointing out that the reason Strasssen's algorithm works is that it somehow uncovered a kind of repeated work that's not evident in a simple divide and conquer approach. It's by the clever definition of the various submatrices that this "overlapping" work can be avoided.

In other words, the power of Strasssens algorithm comes from a strategy that's similar to / reminiscent of dynamic programming.

Isn't it just that the IP router happens to use IPs in Russia as part of the rotation?

If they're trying to exfiltrate data, they might want to rotate through IP addresses in order to obfuscate what's going on or otherwise circumvent restrictions. Using a simple ip rotator like the post talks about would maybe be an approach they'd use. If they're not careful with the IP addresses, once in a while one might get caught due to some restriction like being outside the US. It'd maybe appear as though you're getting these weird requests from Russia, but that's just because you're not logging the requests that are not being flagged from the US.

Maybe I'm reading the post incorrectly though (if so, please correct me!)

Here's how I think about it:

For any dome-like shape, you can start a marble at the bottom and roll it up with some initial speed. If you roll it with insufficient initial speed it'll turn around and come back down. If you roll it too hard, it'll overshoot the peak. By continuity, there must exist some initial condition where it stops at the top.

Now, here's the thing that makes Norton's dome special: For a typical dome shape it'll take an infinite amount of time before that marble stops at the top. If you plot the position as a function of time it'll have some type of sigmoid-like shape. However, for the special case of Norton's dome, you can make it settle at the top in a finite amount of time where it'll sit for the rest of eternity. In other words, if you plot the position as a function of time, there will be some critical time after which its position is constant.

Now, the clever thing to do now is to realize that Newton's laws are time reversal symmetric which means that any motion forward in time could equally well happen backwards in time.

So, you're allowed to take any position plot and flip it horizontally; this is also going to be a valid trajectory.

For any typical dome shape this is not a problem. For a typical dome shape you have a sigmoid-like solution which, when flipped, is still sigmoid shaped. In particular this means that there is no finite time at which you can place the marble at the top of the dome and have it roll off. At any finite time, the marble will be slightly off the top and have a small nonzero speed.

Norton's dome is different. If you flip its trajectory horizontally you'll see that there are many moments in time where you can start the marble at the top to have it abruptly start rolling off the top at some later time. This is the paradox. You can choose to have it sit at the top for one second and then start rolling or sit at the top for one minute and then start rolling.

Unlike other domes, Norton's dome seems to violate our intuition for how initial conditions work. In all cases the marble starts at the top with zero initial speed and yet falls off the top att different moments.

I agree. I think it comes down to the motivation behind why one does mathematics (or any other field for that matter). If it's a means to an end, then sure have the AI do the work and get rid of the researchers. However, that's not why everyone does math. For many it's more akin to why an artist paints. People still paint today even though a camera can produce much more realistic images. It was probably the case (I'm guessing!) that there was a significant drop in jobs for artists-for-hire, for whom painting was just a means to an end (e.g. creating a portrait), but the artists who were doing it for the sake of art survived and were presumably made better by the ability to see photos of other places they want to paint or art from other artists due to the invention of the camera.

Another way to get to the same result is to use "Feynman's Trick" of differentiating inside a sum:

Consider the function f(x) = Sum_{n=1}^\infty c^(-xn)

Then differentiate this k times. Each time you pull down a factor of n (as well as a log(c), but that's just a constant). So, the sum you're looking for is related to the kth derivative of this function.

Now, fortunately this function can be evaluated explicitly since it's just a geometric series: it's 1 / (c^x - 1) -- note that the sum starts at 1 and not 0. Then it's just a matter of calculating a bunch of derivatives of this function, keeping track of factors of log(c) etc. and then evaluating it at x = 1 at the very end. Very labor intensive, but (in my opinion) less mysterious than the approach shown here (although, of course the polylogarithm function is precisely this tower of derivatives for negative integer values).

I mean, you could jump into the black hole to see what's inside so it's not unfalsifiable. The only issue is that you can't convey it to someone on the outside of the black hole.

Here's a fun thought experiment / apparent paradox.

In high school physics we learn that a 1kg mass accelerates as quickly as a 2kg mass when only subjected to the force of gravity. When I used to teach physics, the intuitive explanation I gave for this hinged on a thought experiment. Suppose that you have three 1kg masses falling side by side after being dropped from the same height. Clearly they are all going to fall at the same rate since they're equivalent. Now imagine redoing the experiment but this time taking two of the masses and placing them closer together. Does anything change? Clearly not, they're still all equivalent and ought to fall at the same rate. Now imagine doing this until those two masses are right next to each other, touching. Does anything change? Well no, all three should still fall at the same rate. But now, why not glue those two masses together and call it a 2kg mass? Once you do that you've shown that a 1kg mass and a 2kg mass fall at the same rate.

This usually convinces people, but there's actually a flaw in the argument that gets to the heart of why gravity is so different from the other forces.

To see the flaw, replace the above masses by three electrons falling next to each other in an electric field. Everything goes through in exactly the same way. You end up gluing together two of the electrons and these two electrons will accelerate at the same rate as the single electron. But if you're not careful you'd conclude that all electric charges fall at the same rate in an electric field, something we know is false.

Where's the flaw? Well, all of matter is built from some particles, and as long as you restrict yourself to particles that have the same "charge/mass ratio", the argument above works. It is true that one electron accelerates the same as 100 electrons tied together but that's just because e/m is the same for all those constituents.

So, the thing that's glossed over in my high school explanation for why 1kg and 2kg accelerate at the same rate is that the constituent particles all have the same "gravitational charge / inertial mass" ratio. Because this ratio is the same for all particles, we may as well absorbed that ratio into the gravitational constant and just use "m" in place of both of them. It's this "universal coupling" that's really responsible for the equivalence principle and what sets gravity apart from the other forces.

It is generally believed that the perturbation expansion that we see in realistic quantum field theories are what are known as asymptomatic expansions. These are series that have a radius of convergence of zero (i.e. they only converge when the expansion parameter is exactly zero and diverge for all non zero values).

There are then two natural questions: 1. if the perturbation series diverges, why doesn't the universe explode? and 2. If the series diverges, why can we use it at all?

Let's first talk about the first part: why doesn't the universe explode? Well, it's because the perturbation series is not actually what is going on, the real answer is the solution to the full set of equations. It's just that we're using a perturbation expansion as a crutch. It's sort of like if the universe's function is 1/(1-x) but we constantly insist on using 1+x+x^2+... Clearly the first function is completely well behaved at x=2 but the second one is not. If we notice that our series explodes for x=2 we should not immediately assume that the universe also must explode, it's just that our representation of the true physics is not faithful. This is perhaps a bad example because the series in question is convergent for some x, just not for x=2. The perturbation expansions in question are more subtle since the never converge.

This then leads into the second question: if the series diverges, how can we even use it? Well the idea here is that it's not just any divergent series (like my silly example with 1/(1-x) above) but rather an asymptomatic series. This means that as long as you truncate the series at some point it is in fact reasonably close to the target function for a sufficiently small value of the parameter. It's just that the more terms you want to include, the sooner the approximation breaks in terms of the parameter. So, if you want to include 10 terms it might be a decent approximation until x ~0.1 but if you include 100 terms it might only be a good approximation until x~0.01. Now, within the overlapping range (x<0.01) it's better to have 100 terms than 10 terms, so it's not like including more terms is bad in all ways. But you see the issue: if you include 1000 terms you get a better approximation for your function for values x<0.001 than you had with 100 terms but now your approximation breaks much sooner. If you want to include all the terms your approximation breaks the moment you leave the point x=0.

Why do we think that QFT perturbation theories generally have zero radius of convergence? Well, look at QED, the quantum theory of E&M. If the theory had any nonzero radius of convergence, that also means that the theory would need to make sense for negative coupling constants. However, what would E&M look like for negative coupling? Well, we'd still have electron/positron virtual pair creation from the vacuum since the interactions of the theory are still the same. However this time around they wouldn't attract each other anymore but instead repel each other causing an instability in the vacuum of the theory. We would just constantly be producing these particle/anti-particle pairs and they'd form two separate clusters where all the electrons attract each other and all the positions attract each other but they pairwise repel. In other words, the vacuum would break. This suggests that QED with a negative coupling constants doesn't make sense. But this contradicts the fact that the radius of convergence of the perturbative expansion is nonzero.

That's not to say that all QFTs must have zero radius of convergence, but similar arguments can (I think) be made for the type of QFTs that we actually see in nature.

While everyone is debating whether this is impressive or dumb, if it's a leap forward in technology or just a rehashing of old ideas with more data, if we should really care that much about it passing the bar exam or if it's all just a parlor trick, people around the world are starting to use this as a tool and getting real results, becoming more productive, and building stuff... Seems like the proof is in the pudding!

No I'm just it as an analogy. Not all Turing machines are universal.

What were going through now could maybe be likened to what it would be like for a Turing machine to encounter a universal Turing machine for the first time. For all its life this fictitious Turing machine has encountered other non-universal Turing machines and have simply incorporated them into their own process. When they then encounter their first universal Turing machine they would possibly not be too concerned since each time before they have always just been able to use the new machine to make themselves more productive. However, this time it's different.

My point is just that while it may very well have been true in all of history that new tools have just made us more productive than before rather than fully replace us, this won't be the case for AGI. It's not just another tool we can add to our arsenal but instead something than can subsume us entirely much like how a universal Turing machine can emulate any other Turing machine.

One (perhaps silly) way I like to frame this stuff is by imagining that I'm some special purpose Turing machine designed specifically for some task. Sure, sometimes other Turing machines come along that appear to infringe upon my skill set but they ultimately only perform a small subtask better than I am able to (e.g. calculator, spell check, word processor, IDE, code completion, ...). So, I incorporate it into my routine, effectively boosting my own performance.

Now, what would happen if all of a sudden a universal Turing machine came along? Well, by virtue of being universal, that means that it can emulate me and all other Turing machines. This time around things are different. Even if I can find a way to incorporate it into my workflow, it can still emulate that more sophisticated version of me by virtue of being universal. So it then comes down to whether or not I can incorporate the latest version of this universal Turing machine faster than its own design is improved. If not, I will be replaced. Since in our instantiation, I am made from biological material it's in my mind only a matter of time before the universal Turing machine starts outpacing me.

So, I guess the question is then if these GPT models (or their descendants) are universal (in my hand wavy definition of the term).

The point I was making is that the problem (MH) is a lot more subtle than people give it credit for. Many of the arguments people make for how the MH problem is "obvious" seem to work for this modified game too unless you're careful with them.

No. The fact that they happened to not open the door with the car tells you something. Take it to the extreme of 100 doors. If the spectator randomly opens 98 doors and doesn't randomly stumble upon the car you should take that as evidence that you might have the car yourself. This extra evidence in favor of staying with your door exactly cancels the original MH advantage of switching.

No, if the spectator happens to open an empty door, the probability collapses to 50/50. That's the point I'm making for how counterintuitive this is. Here's the full set of cases. Let's say that you select door 1 and spectator opens door 2 (all other cases are permutations of this):

Door1 Door2 Door3

Car Empty Empty

Empty Car Empty

Empty Empty Car

Let's suppose your strategy is to stick with your original choice. In the first case above then you get the car. In the second case the audience member stumbles upon the car and you lose. In the third case you lose because you stick with an empty door. All three cases are equally likely and since the second one ends, and you know that your game didn't end, you know that you're either in case 1 or case 2. Your chance is thus 50/50.

The issue is that in the classic MH problem, cases 2 and 3 are collapsed into one outcome (MH opens an empty door), but that's not true here.

More mathematically, you should ask yourself p(Door1 | Game Did Not End).

Using Bayes we see

p(Door1 | GDNE) = p(GDNE | Door1) * p(Door1) / p(GDNE).

p(GDNE | Door1) = 1 p(Door1) = 1/3 p(GDNE) = p(GDNE|Door1)p(Door1) + p(GDNE|Not Door1)p(Not Door1) = 11/3 + 1/22/3 = 2/3.

This, p(Door1 | GDNE) = 1 * (1/3) / (2/3) = 1/2.

There's another version of the Monte Hall problem that highlights why this is such a counterintuitive problem.

Imagine that after you pick your box, Monte Hall invites an audience member up on stage and instructs them to choose one of the remaining two doors to open. This audience member doesn't know anything at all and just randomly picks one of the two doors. When their door is opened we see that it's empty. You're now given the option of switching just like in the standard game. Should you?

Cosmetically everything is identical with the standard game, but if you analyze the game carefully this time you're left with a 50/50 shot so there's no benefit of switching.

I think most of the arguments in this article would appear to work for this modified version of the game which means that they're not actually getting to the heart of the problem.

For completeness, the reason this now reduces to 50/50 is that there's also now a chance that the spectator opens the door with the car behind it, something that couldn't happen in the original Monte Hall problem. Put another way, there's actually a little bit of information that's conveyed to you when you see that the spectator happens to not open the door with the car and this extra information exactly cancels the usual benefit you get from eliminating the other empty door. In the example of "scaling up" in the article, if you did this with 20 doors and the spectator randomly picks 18 of the 19 unopened ones to open and then happen to not stumble upon the car, you might actually think that you could have been lucky all along. Ultimately you're left with a 50/50 chance.

So stuff like precious metals, commodities (energy, corn, pork, ...), stocks that don't pay dividends, etc are then also equally worthless?

I don't think holding a barrel of oil or a bar of gold, or a share of Berkshire Hathaway will generate any income for me. I'd need to sell it to someone else for me to get my money back.

Isn't that true of anything that doesn't include the fed? I mean unless you can print money, you can't change the amount of fiat in any subset of the economy. You can still add value though which would be evident by more people wanting a piece of whatever pie there is, driving more fiat into the subset and/or creating paper gains which is a sort of money in itself, but this is true for crypto as well.

For example, I can draw a circle around Google and make the same argument: "I can see how fiat can flow into Google (investments, debt, advertiser dollars, ...) but I can't see how Google can create new fiat to flow back to the rest of the economy." The best I can see is that higher demand for a share of Google creates paper gains but then that's true for crypto as well.