HN user

ankeshanand

250 karma

AI Researcher https://twitter.com/ankesh_anand

Posts6
Comments26
View on HN
Llama 2 3 years ago

Has anyone in this subthread actually read the papers and compared the benchmarks? LLama2 is behind PALM-2 on all major benchmarks, I mean they spell this out in the paper explicitly.

Sorry if it wasn't clear, I do mention the linear classification protocol several times in the post. If you want to evaluate performance on a classification task, you have to show it labels during evaluation, otherwise it's an impossible task. Note that the encoder is freezed during evaluation, and only a linear classifier is trained on top. Now, even when evaluated on a limited set of labels (as low as 1%), contrastive pretraining outperforms purely supervised training by a large margin (check out Figure 1 in the Data-Efficient CPC paper: https://arxiv.org/abs/1905.09272.

I did not get the second part unfortunately, could you elaborate more and clarify if you are talking about a specific paper?

I didn't mean to convey that we should abandon generative self-supervised methods, but I can see how comparing them gives that impression.

Agree that using them in conjunction would make sense, since generative methods could capture some features better and vice versa.

One thing to note is that the camera viewpoint (it's position, roll, pitch, and yaw) is fed along with the images during training. Requiring access to this ground truth makes this method very constraining to use in practice.

"Our method still has many limitations when compared to more traditional computer vision techniques, and has currently only been trained to work on synthetic scenes. However, as new sources of data become available and advances are made in our hardware capabilities, we expect to be able to investigate the application of the GQN framework to higher resolution images of real scenes."

Bag of Words is not actually a great approach to understand text because it ignores the semantics of the word. For example, 'hotel' and 'motel' which are similar words have completely different vector representations in the BoW model.

A popular alternative is to use a distributed word embedding such as word2vec[1], where similar words are grouped together in the vectorspace.

Edit: If there are few observations, like in this case, we don't need to train the word2vec model on the dataset itself. We can use pre-trained word embeddings such as the one publicly released by Google which was trained on the Google News dataset.

[1]https://word2vec.googlecode.com/

Location: Kolkata, India

Remote: Yes

Willing to relocate: Yes

Technologies: Python, Djagno, C/C++, Qt, ROS, Javascript, d3.js, Backbone.js, PHP, HTML5, CSS3

Resume: http://goo.gl/fdKHtR

Email: ankeshanand1994 at gmail dot com

GitHub: https://github.com/ankeshanand

Areas of Interest: Software Engineering, Web Development, Privacy, Social Computing, Data Visualizations, Data Science

* Google Summer of Code fellow for BRL-CAD in 2014.

* Mathematics and Computer Science undergrad at the Indian Institute of Technology, Kharagpur.

I am actively looking for Software Engineering Internships starting from May 2015.

Location: West Bengal, India. Actively looking to relocate, Bay Area preferably.

Remote: No

Languages: Python, PHP, Javascript, HTML5, CSS3, C++, C, Matlab, MySQL, Assembly

Framweorks / Libraries: Django, Flask, ROS, Qt, OpenCV, d3.js, Leaflet, Bootstrap

- Google Summer of Code student for BRL-CAD in 2014.

- Maths and CS undergrad at IIT Kharagpur.

- Spent a summer researching at the Max Planck Institute of Software Systems, Germany

Resume: https://www.dropbox.com/s/ankl9g7x3njlye7/AnkeshCV.pdf?dl=0

Github: https://github.com/ankeshanand

Personal Website: http://ankeshanand.com/

I am looking for Software Engineering Internships starting from May 2015.

Email: ankeshanand1994 [at] gmail dot com