HN user

felippee

645 karma

Researcher and a developer interested in AI for real world physical applications (read robotics).

http://blog.piekniewski.info

Posts11
Comments115
View on HN

Oh, somebody got triggered here! Yes, there is sarcasm in this post! And if you don't like it, fine. But please, don't give me bullshit about being a jerk. I think you probably have not seen a real jerk in your life yet.

I'm not sure how to interpret these pictures. They don't suggest anything to me. And certainly don't suggest anything about the quality of representations. And BTW how do you measure quality of representations?

Yes it is certainly not fair that the network they spend one page explaining and probably weeks training and researching can be hardwired in 30 lines of python. This is very unfair. But this is the reality, and so the post states.

Also the idea to add coordinate as a feature has been used in the past without giving even much thought.

Toy examples are great. As long as they are not trivial. Some guy, presumably smart, once said that "things should be as simple as possible but not simpler". The toy example they play with is just too simple.

Author of the post here: I think their paper would have been much better if they included the piece of code which I wrote in python to explain that the transformation they are learning is obviously trivial and the fact that it works is not in question. This would leave them a lot more space to focus on something interesting, perhaps explore the GAN's a little further, cause what they did is somewhat rudimentary. But that omission (and lack of context for previous use of such features in the literature) left a vulnerability which I have the full right to exploit in a blog post.

Author of the post here. I totally agree that negative stuff should be published. But without the fanfares. I think they could have changed the tone of that paper and I would not have an issue with it. It is likely that if they did that they'd never go through some idiot reviewer who expects "a positive result" or some similar silliness. This is not a perfect world. The paper as is makes strong claims about the novelty and usefulness of their gimmick. If it turns out your stuff is at least partially hollow and you take on the pompous tone, you have to be ready to take some heat. Science is not about tapping friends on the back (which BTW is what is happening a lot with the so called "deep learning community"). Science is about fighting to get to some truth, even if that takes some heat. People so fragile that they cannot take criticism should just not do it.

The post mocks them primarily for learning the trivial coordinate transform. That is the core of the paper and ridiculing this piece leaves very little left on the table. The ImageNet test is just an appendix, a cherry on the cake, a curiosity one should say.

Author here (of the post, not the paper). I think you don't understand how science works. The whole point of the exercise (which indeed may have been forgotten these days) is to attack ideas/papers. The first line of attack should be your friends to make sure you don't put anything out there that is silly. The second line of attack are the reviewers, who may or may not be idiots themselves, but in the perfect world should serve the same purpose. The third line attack are independent readers, people like me. I found it to be trivial, took my liberty to attack it. It is not personal and should not be taken so. These guys may in the future publish the most amazing piece of research ever. But this one is not it. They should realize this and my blog post serves this purpose. If somebody gets offended and takes it personally, so be it. I think people should have a bit thicker skin, especially in science. I took quite a bit of bullshit myself (and I'm sure I will have to take more) and never complained. So relax, read the paper, read the post, learn something from both and go on.

No, there is something special to machine learning and AI in general. E.g. SIGGRAPH papera are different. No one there claims to solve more than they actually do solve. DL is soaked with hype and self congratulatory BS. The best way to spot it is to check the citations. Typically they solve an already solved problem, skipping entirely any pre deep learning literature on it (or if they do cite it, only to dump BS on it) and then just cite a few of their own more or less relevant papers. I'm aware I'm overgeneralizing here and not every paper is like that, but I've seen enough to detect a trend.

It is as if defending or advertising "deep learning" was the purpose of the paper. It is not. The purpose of a paper is to show a solution to a problem. Much of DL literature (again not all) is a "solution in a desperate search of a problem" rather than the opposite.

I think many of these papers (including this one) would make a great blog post, but just isn't quite enough in terms of scientific content for a full blown paper. A curiosity, nice gimmick, but nothing more. Not really a solution to a problem, not really any idea of non trivial universality.

Right, I totally agree. What is more, I think this same result could actually be sold without the pompous deep learning bullshit and be received quite differently. If they did not claim to invent the wheel, but rather modestly noted their observation (which in a limited way is actually quite cool - that is from a known - at least on this forum - deep learning skeptic like me), it would make a much better impression.

Same is true actually for many DL papers. They'd be actually cool, if they weren't oversold.

Author here: seriously I'm here at the front page for the second day in the a row!?

The sheer viral popularity of this post, which really was just a bunch of relatively loose thoughts indicates that there is something in the air regarding AI winter. Maybe people are really sick of all that hype pumping...

Just a note: I'm a bit overwhelmed so I can't address all the criticism. One thing I would like to state however, is that I'm actually a fan of connectionism. I think we are doing it naively though and instead of focusing on the right problem we inflate a hype bubble. There are applications where DL really shines and there is no question about that. But in case of autonomy and robotics we have not even defined the problems well enough, not to mention solving anything. But unfortunately, those are the areas where most best/expectations sit, therefore I'm worried about the winter.

It's not like humanity really needs another chess playing program 20 years after IBM solved that problem (but now utilizing 1000x more compute power). I just find all these game playing contraptions really uninteresting. There are plenty real world problems to be solved of much higher practicality. Moravec's paradox in full glow.

Also why does every single result has to be breathtaking?

If you build the hype like say Andrew Ng it better be. Also if you consume more money per month than all the CS departments of a mid sized country, it better be.

I'm skeptical, and side with Rodney Brooks on this one. First, reinforcement learning is incredibly inefficient. And sure, humans and animals have forms of reinforcement learning, but my hunch it that it works on an already incredibly semantically relevant representation and utilize the forward model. That model is generated by unsupervised learning (which is way more data efficient). Actually I side with Yann Lecun on this one, see some of his recent talks. But Yann is not a robotics guy, so I don't think he fully appreciates the role of a forward model.

Now using models for RL is the obvious choice, since trying to teach a robot a basic behavior with RL is just absurdly impractical. But the problem here, is that when somebody build that model (a 3d simulations) they put in a bunch of stuff they think is relevant to represent the reality. And that is the same trap as labeling a dataset. We only put in the stuff which is symbolically relevant to us, omitting a bunch of low level things we never even perceive.

This is a longer subject, and a HN is not enough to cover it, but there is also something about the complexity. Reality is not just more complicated than simulation, it is complex with all the consequences of that. Every attempt to put a human filtered input between AI and the world will inherently loose that complexity and ultimately the AI will not be able to immunize itself to it.

This is not an easy subject and if you read my entire blog you may get the gist of it, but I have not yet succeeded in verbalizing it concisely to my satisfaction.

I sure agree there are many interesting things going on, there is no question about that. Also most of them are toy problems focused in some restricted domains, while a huge bag of equally interesting real world problems is sitting untouched. And let me tell you, all those VC's that put in probably way north of $10B are not looking forward to more NIPS papers or yet another style transfer algorithm.

Sure, but have you heard about Moravec's paradox? And if so, don't you find it curious that over the 30 years of Moore's law exponential progress in computing almost nothing improved on that side of things, and we kept playing fancier games?

Sure it is not technically a loan. But it carries the same sentiment change when it blows up. People get extremely cautious, to the point of skipping some really good ideas. And not just those VC's that made the bets but everyone else too. Fear spreads just as effectively as hype.

Yes. Sensationalist.

Yes, perhaps. But I'm entitled to my opinion just as you are entitled to yours. And time will tell who was right.

So your explicit reason for omitting Waymo, as I understand it, is that it didn't support your argument?

You see, when you make any argument, you always omit the infinite number of things that don't support it and focus on the few things that do. The fact that something does not support my argument, does not mean it contradicts it.

You might also note that this is not a scientific paper, but an opinion. Yes, nothing more than an opinion. May I be wrong? Sure. And yet this opinion appears to shared by quite a few people, and makes a bunch of other people feel insecure. Perhaps there is something to it? We will see.

But in the worst case it will make some people think a bit and make an argument either for or against it. I may learn today a good argument against it, that will make me think about it more and perhaps I will change my opinion, or I'll be able to defend it.

So far you have not provided such an argument, but I wholeheartedly encourage you to do so.

Hi, it appears that "sensationalist garbage" triggered quite a bit of a discussion. This is typically indicative that the topic is "sensitive". Perhaps because many people feel the winter coming as well. Maybe, maybe not, time will tell.

And FYI, Tesla is in the business of making self driving car. If you read the article, you might learn that Tesla is actually the first company to sell that option to customers. You can go to their website right now and check that out.

Uber, like it or not is one of the big players of this game. I agree they may have somewhat toxic culture, but I guarantee you there are plenty of really smart people there who know exactly the state of the art. And their failure is therefore indicative of that state of the art.

I also omitted Cruise automation and a bunch of other companies, perhaps because they have more responsible backup drivers that so far avoided fatal crashes. But I analyze the California DMV disengagement reports in another post if you care to look. And by no means any of these cars is safe for deployment yet.

OK, the definition of scalable is crucial here and it causes lots of trouble (this is also response to several other posts so forgive me if I don't address your points exactly).

Let me try once again: an algorithm is scalable if it can process bigger instances by adding more compute power.

E.g. I take a small perceptron and train it on pentium 100, and then take a perceptron with 10x parameters on Core I7 and get better output by some monotonic function of increase in instance size (it is typically a sub linear function but it is OK as long as it is not logarithmic).

DL does not have that property. It requires modifying the algorithm, modifying the task at hand and so on. And it is not that it requires some tiny tweaking. It requires quite a bit of tweaking. I mean if you need a scientific paper to make a bigger instance of your algorithm this algorithm is not scalable.

What many people here are talking about is whether an instance of the algorithm can be created (by a great human effort) in a very specific domain to saturate a given large compute resource. And yes, in that sense deep learning can show some success in very limited domains. Domains where there happens to be a boatload of data, particularly labeled data.

But you see there is a subtle difference here, similar in some sense to difference between Amdahl's law and Gustafson's law (though not literal).

The way many people (including investors) understand deep learning is that: you build a model A, show it a bunch of pictures and it understands something out of them. Then you buy 10x more GPU's, build model B that is 10x bigger, show it those same pictures and it understands 10x more from them. Look I, and many people here understand this is totally naive. But believe me, I talked to many people with big $ that have exactly that level of understanding.

Hey, a small advice for the future: never build your belief entirely on a youtube video of a demo. In fact, never build your belief based on a demo, period.

This is notorious with current technology: you can demonstrate anything. A few years ago Tesla demonstrated a driverless car. And what? Nothing. Absolutely nothing.

I'm willing to believe stuff I can test myself at home. If it works there, it likely actually works (though possibly needs more testing). But demo booths and youtube - never.

but it’s moving down the right path

Time will tell. I think DL is amazing, but is no the right path towards solving problems such as autonomy. I think if you enter this field today, you should definitely take a look at other methods than DL. I actually spent a few years reading neuroscience. It was painful, and I certainly can't tell I learned how the brain works, but I'm pretty certain it has nothing to do with DL.

I certainly encourage everybody to consult the source material! Man, this is a blog, opinion by default not perfect.

But when I hear the keyword "major advances" I'm highly suspicious. I had seen already so many such "major advances" that never went beyond a circle of self citing clique.

I'll happily read your next post where you will include all of those. In fact amount of VC money spent in that field would only support my claim. And the number of papers is irrelevant. There were thousands of papers about Hopfield network in the 90's and where are all of them now? You see, all the things you point out is the surface. What really matters is that self driving cars crash and kill people, and no one has any idea how to fix it.

The difference between the current AI renaissance and the past pre-winter AI ecosystems is the level of economic gain realized by the technology

I would argue this is well discounted by level of investment made against the future. I don't think the winter depends on the amount that somebody makes today on AI, rather on how much people are expecting to make in the future. If these don't match, there will be a winter. My take is that there is a huge bet against the future. And if DL ends up bringing just as much profit as it does today, interest will die very, very quickly.

Thanks for that, that is essentially my point. Agree it is not very rigorous, but it gets the idea across. By scalable we'd typically think "you throw more gpu's at it and it works better by some measure". Deep learning does that only in extremely specific domains, e.g. games and self play as in alpha go. For majority of other applications it is architecture bound or data bound. You can't throw more layers, more basic DL primitives and expect better results. You need more data, and more phd students to tweak the architecture. That is not scalable.