If you're an individual developer and not an enterprise, just go straight to Google AIStudio or GeminiAPI instead: https://aistudio.google.com/app/apikey. It's dead simple getting an API key and calling with a rest client.
HN user
ankeshanand
AI Researcher https://twitter.com/ankesh_anand
We've done extensive comparisons against GPT-4V for video inputs in our technical report: https://storage.googleapis.com/deepmind-media/gemini/gemini_....
Most notably, at 1FPS the GPT-4V API errors out around 3-4 mins, while 1.5 Pro supports upto an hour of video inputs.
Has anyone in this subthread actually read the papers and compared the benchmarks? LLama2 is behind PALM-2 on all major benchmarks, I mean they spell this out in the paper explicitly.
You can also rent a cloud TPU-v4 pod (https://cloud.google.com/tpu) which 4096 TPUv-4 chips with fast interconnect, amounting to around 1.1 exaflops of compute. It won't be cheap though (excess of 20M$/year I believe).
It's important in the context that RL does not have performance ceilings.
Looks like any Github pages served with CloudFlare are getting blocked, I am trying out a fix.
Yep, Karpathy has mentioned this multiple times in their AI talks.
If you carefully curate who you follow, Twitter can be more like a bunch of subreddits, with the added signal of knowing who's posting. So it ends up a being great way to keep up with small communities.
Sorry if it wasn't clear, I do mention the linear classification protocol several times in the post. If you want to evaluate performance on a classification task, you have to show it labels during evaluation, otherwise it's an impossible task. Note that the encoder is freezed during evaluation, and only a linear classifier is trained on top. Now, even when evaluated on a limited set of labels (as low as 1%), contrastive pretraining outperforms purely supervised training by a large margin (check out Figure 1 in the Data-Efficient CPC paper: https://arxiv.org/abs/1905.09272.
I did not get the second part unfortunately, could you elaborate more and clarify if you are talking about a specific paper?
I didn't mean to convey that we should abandon generative self-supervised methods, but I can see how comparing them gives that impression.
Agree that using them in conjunction would make sense, since generative methods could capture some features better and vice versa.
Great piece, you might want to update the article with the mention of PyTorch Mobile that released today: https://pytorch.org/mobile/home/
There's active research in Model-Based RL right now that tries to tackle 1) and 2) together.
Bear has no Android app, and no LaTeX support either, so this works better for me.
Neat idea, but isn't the $5k amount for 3 months kinda low?
One thing to note is that the camera viewpoint (it's position, roll, pitch, and yaw) is fed along with the images during training. Requiring access to this ground truth makes this method very constraining to use in practice.
"Our method still has many limitations when compared to more traditional computer vision techniques, and has currently only been trained to work on synthetic scenes. However, as new sources of data become available and advances are made in our hardware capabilities, we expect to be able to investigate the application of the GQN framework to higher resolution images of real scenes."
The new version does use MCTS, you should read the paper again. :)
Twitter, Reddit, Facebook, HN, Twitch
Probably referring to the Master-Slave terminology
We could always use pre-trained word embeddings, the few observations won't matter then.
Bag of Words is not actually a great approach to understand text because it ignores the semantics of the word. For example, 'hotel' and 'motel' which are similar words have completely different vector representations in the BoW model.
A popular alternative is to use a distributed word embedding such as word2vec[1], where similar words are grouped together in the vectorspace.
Edit: If there are few observations, like in this case, we don't need to train the word2vec model on the dataset itself. We can use pre-trained word embeddings such as the one publicly released by Google which was trained on the Google News dataset.
Location: Kharagpur, India
Remote: No
Willing to relocate: Yes (San Francisco / Bay area preferably)
Technologies: Python, Javascript, Django, Flask, MySQL, MongoDB, Redis, HTML, CSS, ReactJS (Full stack); scikit-learn, Pandas, Numpy, NetworkX (Data science)
Resume/CV: http://ankeshanand.com/CV.pdf
Email: ankeshanand@iitkgp.ac.in
Location: Kolkata, India
Remote: Yes
Willing to relocate: Yes
Technologies: Python, Djagno, C/C++, Qt, ROS, Javascript, d3.js, Backbone.js, PHP, HTML5, CSS3
Resume: http://goo.gl/fdKHtR
Email: ankeshanand1994 at gmail dot com
GitHub: https://github.com/ankeshanand
Areas of Interest: Software Engineering, Web Development, Privacy, Social Computing, Data Visualizations, Data Science
* Google Summer of Code fellow for BRL-CAD in 2014.
* Mathematics and Computer Science undergrad at the Indian Institute of Technology, Kharagpur.
I am actively looking for Software Engineering Internships starting from May 2015.
Location: West Bengal, India. Actively looking to relocate, Bay Area preferably.
Remote: No
Languages: Python, PHP, Javascript, HTML5, CSS3, C++, C, Matlab, MySQL, Assembly
Framweorks / Libraries: Django, Flask, ROS, Qt, OpenCV, d3.js, Leaflet, Bootstrap
- Google Summer of Code student for BRL-CAD in 2014.
- Maths and CS undergrad at IIT Kharagpur.
- Spent a summer researching at the Max Planck Institute of Software Systems, Germany
Resume: https://www.dropbox.com/s/ankl9g7x3njlye7/AnkeshCV.pdf?dl=0
Github: https://github.com/ankeshanand
Personal Website: http://ankeshanand.com/
I am looking for Software Engineering Internships starting from May 2015.
Email: ankeshanand1994 [at] gmail dot com