2-parts. features a ton of interviews with high level executives including the facebook "head of growth" and alex stamos who was the security chief. I thought stamos came out of the interview seeming pretty forthright and direct (he no longer works at facebook). the other executives did not come off so well...
HN user
skeptic_69
I think the green-washing comment is definitely on-point. co-opt the opposition and there is no opposition.
everything old is new again
I agree.
Are you referring to master switch or attention merchants? been meaning to read both.
I don't know if THIS article was directed by Facebook related PR but the times reporting says there is other PR basically identical to this. So I think we can decisively label your conjecture as "credible"
:/ yeah.
not just jaron lanier! there are multiple thinkers who are fellow travelrs. Cal newport, Tristan harris, Nicholas Carr...
LOL. indeed.
how would hear about this story unless you are personally friends with alex stamos or zuckerberg or sandenberg? I am all for skepticism but blanket rejection of responsible journalism seems like an over-reaction. the reporting of the new york times on facebook has been continuously borne out by events.
yep.
I rendered my disapproval of Facebook by deleting it. That is the only action they care about.
too little too late but better than nothing.
1.mmmmmmmmm ok I am willing to accept you meant the quadratic loss instead of 0-1 error. that seems reasonable.
2. this is paper is centered in a research thrust that IS focused on generalization. see my below comment.
I don't know who most people are but this paper COULD be important in understanding why stochastic gradient works well in practice.
Personally I doubt it very much.
3. massively overfitting to the training dataset BUT generalizing well is a real phenomenon and yes it is very weird. happens in deep nets and i believe adaboost. i.e. continuing to train after you have zero 0-1 loss. I agree this is a weird way to communicate this idea but that is what the community uses.
have you heard of something called ERM? uniform convergence?
The typical way of showing generalization in ML is to show that if we have some low or zero error solution on the test data-set, for a large enough dataset, with high probability, the error on our training data set is close to the error on the real and unknown distribution. The first step which is basically "find a low error hypothesis on the training data" is called the ERM principle.
In practice we observe stochastic gradient descent works pretty well in solving the ERM problem and the solutions generalize well (perform well when deployed).
This is very weird since neural networks are really weird objects with very non-linear and non-convex behavior and gradient descent shouldn't play well with weird bumps and curves and valleys.
People want to show mathematically that stochastic gradient descent does well on neural networks.
This paper claims gradient descent is effective at minimizing quadratic loss on the training data.
If we could improve the results to show that on the true distribution we also have low loss-that might be compelling that gradient descent converges to the minimum error solution.
None of this explicitly stated since this is a well understood part of basic literature in learning theory.
Showing an algorithm can do erm on the hypothesis class is the first and (easier ) part of showing generalization.
If you want a good reference that explains this in a more coherent way I recommend looking at the first 4 chapters of understanding machine learning theory by Shai-Shalev Schwartz.
If you still think the comments I was responding to are not totally incoherent-take note of the fact that the very first sentence in the paper is "One of the mysteries in deep learning is random initialized first order methods like gradient descent achieve zero training loss"
1. people overfit the baby datasets to zero training loss (MNIST) all the time. maybe you meant a "hard" dataset.
2. You clearly have no idea what you are talking about. This paper is trying to argue a bit about why neural networks generalize well by showing with math that a nn with some of their conditions converges to the zero training loss. It isn't remotely meant to be practical. IT IS A THEORETICAL PAPER.
And comparing it to nearest neighbors of 1 is so so so so so silly it isn't even wrong.
edit. #1 is actually an entire research direction in the theory of machine learning fyi.
It is possible to get neural networks that massively overfit but still generalize (which Is weird).
https://arxiv.org/pdf/1611.03530.pdf
That paper was really famous. It showed you can get zero training loss on data when you replace the labels with random noise.
edit 2: I am sorry to be harsh. It is just hard to read such arrant nonsense.
Well, maybe because it's all done by a priest-caste cartel, shielded from reality in ivory towers, and then presented to the public in a similar way religion once was forced on peasants (taxes included). Academia is the modern priesthood: usurping the right to all knowledge, in bed with the state, corrupted to the bones, fighting heretics.
I don't mean to be rude....I actually fuck that. I mean to be rude. You are an irrefutable argument against public-review.
do you put your work in the arxiv?
There is a serious and well documented basis in neutral american and european media-going back decades.
It reminds me of how so many smart people make decisions based on an anecdotal basis, including American media like NYT. I implore people to not do that again.
side note: your background is wild. you have been around the block a time or two..I think you have some good stories to tell.
I don't understand why you would publish this in nature instead of FOCs/STOC
there is a basis to at least some of these claims.
Seals do a lot of direct action. Someone more qualified should explain how much direct action they do.
I think a lot.
i think become a seal/green beret is pretty challenging !
only tangentially related but I have always thought that Executive Outcomes is such an ominous and great name..
obviously learning things effectively makes you a better researcher. I am talking about the value to the academic culture of the department.
ehhhhhhh I am not so sure the main limiting factor in going on to grad school is access to quality undergraduate education. a diligent student at the top public school in every state is probably qualified to go do research in graduate school.
I'd imagine a more limiting factor is the a) willingness to work really hard for 6+ years for uncertain rewards for the joy of research with very low wages. especially when you can go into industry and make 100k+ b) student loans-see low wages as an academic.
As an academic in ML at least-I think there are more than enough academics in the field or trying to get here....look at NIPs submissions!
I think if you want to be a leader in AI research you want to attract and nurture the best graduate students and researchers. I don't really know how much doing good research has to do with good online education. I am sure online education is useful but these issues don't really have much in common.
CMU does have a department of machine learning fyi. which is probably what the other guy was referring to since to most people ML = AI and AI = ML