Training adjusts all whole net's weights. And you are right, we get the joint changes and decompression machinery to work together.
HN user
irodov_rg
2 karma
Posts0
Comments4
No posts found.
New Method for Compressing Neural Networks Better Preserves Accuracy 8 years ago
This is shown in one of the plots, the less you compress the less you loose. While there is some analysis in the paper on how the computations reduce etc, the results are mostly emperical.
Its mostly the lookup table which takes up the most space. This work is about breaking it into 2 layers and continuing to train to gain accuracy. The output model becomes 90% smaller compared to the original model.
BERT is more computationally expensive. It might end up giving better results on the task mentioned in the paper but we don't know. At the time of writing this all of the contextual word embedding techniques were fairly new and were not tried.