HN user

vundervul

32 karma
Posts0
Comments5
View on HN
No posts found.

Bayesian optimization > random search > grid search

Grid search is nothing like a uniform prior since you would never get a grid search-like set of test points in a sample from a uniform prior.

I didn't really want to write a list of criticism for what is presumably a smart and earnest gentleman and the similarly smart and earnest woman who summarized the tips from his talk, but here goes:

The H2O architecture looks like a great way to get a marginal benefit from lots of computers and is not something that actually solves the parallelization problem well at all.

Using reconstruction error of an autoencoder for anomaly detection is wrong and dangerous so it is a bad example to use in a talk.

Adadelta isn't necessary and great results can be obtained with much simpler techniques. It is a perfectly good thing to use, but it isn't a great tip in my mind. This isn't something I would put on a list of tips.

In general, the list of tips doesn't just doesn't seem very helpful.

Who is Arno Candel and why should we pay attention to his tips on training neural networks? Anyone who suggests grid search for metaparameter tuning is out of touch with the consensus among experts in deep learning. A lot of people are coming out of the woodwork and presenting themselves as experts in this exciting area because it has had so much success recently, but most of them seem to be beginners. Having lots of beginners learning is fine and healthy, but a lot of these people act as if they are experts.

Ray Kurzweil is a laughingstock in the machine learning community and he hasn't been a legitimate researcher for many years. Putting him in the same sentence as Andrew Ng is a bit insulting.