Do you know how to publicly comment? I couldn't find a way on the press release or their website.
HN user
bcaine
Working on autonomous driving.
Specialties: Computer Vision, Machine (Deep) Learning
It's actually not linear, its a power law. That means we need exponentially more compute, data, and model parameters to see linear improvements in performance.
Not to pour too much cold water on this, but the claim of 100% accuracy has a huge caveat. In the paper (Page 4) they state:
Interaction. The original question may not be a prompt that synthesizes a program whose execution results in the correct answer. In addition, the answer may require multiple steps with clear plots or other modalities. We therefore may interactively prompt Codex until reaching the correct answer or visualizations, making the minimum necessary changes from the original question
Which to me basically sounds like they had a human in the loop (that knows how to solve these math problems) that kept changing the question until it gave the correct answer. They do measure the distance (using a sentence embedding model) of the original question to the one that yielded the correct answer, but that feels a bit contrived to me.
Nevertheless, its still really cool that the correct answer is indeed inside the model.
It looks like this is just Pilocarpine, which has been used for decades to treat glaucoma, available at pharmacies everywhere, and is commonly used (perhaps off label) to shrink pupils. I wonder what if anything they changed compared to the generic version?
I've been using Pilocarpine off label to shrink my pupils at night after ICL surgery (an alternative to lasik) to solve debilitating halos caused by my pupils growing larger than the implanted lens.
In my experience, it does increase close range vision (at some minor expense to long range vision). That said, it also gives a mild headache, and blurs your vision substantially for the first 5-15 minutes after use. I don't really see the appeal of using it daily unless you really have to.
While I sort of agree that machine learning will end up as an experimental science, it's way, way too early to say whether the theory relating deep learning to kernel methods (e.g. Neural Tangent Kernels) will be useful or not.
As an example, just last week a (huge) paper [1] was put on arXiv that used these theoretical methods to analyze a bunch of common architecture building blocks (skip connections, normalization, etc), and then applied their theoretical findings to figure out how to train Resnet like models in similar training time without these seemingly "required" building blocks.
Deep Learning is still in its infancy in many ways, and this type of research takes time, slowly building on successive results.
Another absolutely fantastic resource is this Jupyter Notebook based textbook on Kalman Filters and related topics: https://github.com/rlabbe/Kalman-and-Bayesian-Filters-in-Pyt...
An addition to your correction: there are many other ways to solve the SLAM problem beyond (kalman/information/particle) filters. Optimization based approaches are very popular (search terms: Graph SLAM, Factor Graphs, Pose Graphs).
Was there any statistically significant changes in compensation at each level based on more unique skills or more granular classifications of engineers?
In particular I'm curious about Machine Learning/Data Science, Robotics, Distributed Systems etc, but I would imagine Web vs Mobile vs DevOps vs Data may look different too.
My first ever internship in software engineering (in 2012) was in the MUMPS world! I worked on a project for a contracting company building an automated testing framework (in python) to interface with and allow the VA to refactor their massive VistA EHR that was entirely written in MUMPS.
The fact that the system works at all is total magic to me, with hundreds of subsystems and millions of lines of code, all with the same shared global variable pool. I remember having to spend a few days digging through hundreds of pages of kernel documentation (it has its own kernel!) to simply find out how to write to a file..
What you can remember seems largely correct. I think commands and syntax were case insensitive, and variables were case sensitive too? All kinds of insanity like that.
In this specific case, it was probably further from the strike zone than the pitcher wanted. That said, pitchers throw balls often to see if hitters will swing at them, and some hitters are more apt to swing at bad pitches (and therefore pitchers exploit this).
Pitches at the bottom of the strikezone, or pitches that are low and drop below the strikezone (that would be called balls) are generally hard to hit, so a lot of pitchers throw sinking pitches there hoping to either get a very borderline strike or a swing and miss.
No, the position most definitely had absolutely nothing to do with longest path or combinatorial optimization.
Anyway, my larger point is that what I've been seeing interviewing is that these tests are becoming much more common at US startups without companies removing/reducing the rest of their technical evaluation process, nor really structuring the problems to be a good signal.
In an ideal world where companies do take home tests right, I think its a great solution. But what I've been seeing more often than not doesn't support that, making it hard to support.
I'm really curious what you've been seeing at Starfighter. Are partnering companies still going on to do a full technical interview? Or does Starfighter largely replace their normal technical evaluation?
Ignoring the fun of the challenges themselves (which probably isn't entirely fair), the latter makes it very compelling for a candidate. The former does not.
I think that's fair. I've had both the former and the latter, but unfortunately most of my experiences fall into latter case, where it's simply been hoop jumping. Most of my friends (all about to graduate, so a good number of examples) are experiencing the same.
For example one company gave a problem with five parts, with the final part being solve longest path on a bipartite weighted graph (which is quite a hard and time consuming problem). After that, the next step was a phone technical screen, then an on-site with 4-5 more interviews, most being white-boarding. It was basically hazing instead of an evaluation criteria.
An alternative is my last job, which had a take home test that took about 6 hours, but that was the whole technical part of the process. Being on the other side reviewing them, the problem absolutely gave enough information.
I totally get there's a right way to do it, but like most interviewing trends, companies seem to just be adding this as a step instead of revamping their process.
I agree that this is "better" than the alternative, but it can be absolutely exhausting for candidates actively searching for a job. I feel like it's recently become much, much more common (from my small-ish sample of me and some friends).
My issue with this approach is fourfold:
1. Most companies have no idea how to structure a problem that is both informative to them and also not abusive to the candidates time.
2. Companies generally do this right after the recruiter phone screen, which most likely doesn't give the candidate enough information to decide if the next steps are worth their time.
3. Most companies still do a whole suite of normal tech screens after you work on a take home problem.
4. If you're actively looking, getting a bunch of these over a short period of time is likely. I know during my full time search, more than 50% of companies had a take home test right after the recruiter screen. Most of these were 4-8 hours of work each, due within the week.
A lot of startups structure it more like hazing or a barrier to entry than an evaluation criteria. I have some fun (read: horrifying) anecdotes from my recent search that illustrate the problems above, but I don't think any of my points are surprising.
A nice alternative would have been to simply have one or two projects completed that are straightforward to evaluate and walk companies through them, letting them ask me questions.
Two courses I've found are very good for Neural Networks & Deep Learning are Karpathy's CS231n from Stanford[0] and Nando de Freitas's Deep Learning class from Oxford[1][2].
Been through parts of both and they seem really good. That said, I'd bet that both of these classes would be a lot more meaningful after either Hastie's class (listed here) or Andrew Ng's course.
[1] https://www.youtube.com/playlist?list=PLE6Wd9FR--EfW8dtjAuPo...
Is the Boston team and office part of The Echo Nest's in Somerville, or somewhere new?
If you design the system correctly, flickering should not be visible to the human eye. Obviously that depends on factors like clock rate of your transmission and your coding technique. As long as your signal is reasonably DC balanced and your modulation rate switches faster than the human eye can detect (say > 60Hz), you should be in the clear. The IEEE 802.15.7 working group [0] proposes clock rates of 200 kHz to 120 MHz, which all should be fine.
As for a separate IR emitter/photosensor, in theory it could increase range. Range ends up being decided by a bunch of tradeoffs and channel conditions that effect your Bit Error Rate (BER), which increases the further you move away from the source. The major factors that effect BER is the noise in the room (the ambient irradiance levels of other light sources), your transmission power, photodiode sensitivity to your wavelength, filtering quality, area of photodiode, and forward error correction quality. So moving to a wavelength with less noise certainly would help tremendously, assuming everything else is equal.
[0] http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=616358...
Fun read. Topic modeling can be fascinating to work with.
Curious how they measured performance of their model, and whether they found a "best" number of topics for LDA where their model stopped getting much benefit by having more topics.
I'd imagine increased number of topics would have some interesting side effects where it would create too narrow of recommendations.
How is this different from Covester (http://covestor.com/), which has been around (and seemingly struggling) for awhile?
Nice read. I did something sort of similar with the same dataset about a year ago. I compared LDA (Latent Dirichlet Allocation) to TF-IDF as tools to find similar beers based on their review text. Lots of intuitive and funny topics discovered.
I suggest you play with LDA, it seemed to work really well at generating topics. There is also a lot of fascinating, very readable research using it. Check out SNAPs work on the same dataset [1] and some of the Yelp Dataset challenge winners [2]. If you end up interested in doing so, Gensim [3] was pleasant enough to work with.
[1] http://snap.stanford.edu/data/web-BeerAdvocate.html
[2] http://www.yelp.com/dataset_challenge
[3] https://radimrehurek.com/gensim/wiki.html#latent-dirichlet-a...
The problem is that it's near impossible for a student to pay attention during longer classes. I think you're trying to measure accomplishing things by the amount of material you work through during a given class, instead of by the amount of information the student actually absorbs during that time.
That seemed to be one of the main points in the article, just getting through material was a terrible way to look at classroom learning.
I also experimented with it a few years ago trying to build an automated GUI testing framework, and it turns out that its way too fragile and non-portable to be usable for that use case.
I remember running into issues as soon as anything changed regarding resolution, scaling, graphic settings, color scheme etc.
Pretty fun to make toy programs in to automate stuff with though.
Here is a video from a year ago showing the virtual shrink in action (it may be an older version):
I'd be curious how the study measured 40 pages of reading each week. My guess is that most, if not all of my classes (Computer Engineering student) have > 40 pages of reading associated with the topics taught each week, but for many of them the reading is never assigned, nor is it ever performed.
In my experience, reading only really is necessary when you are really confused on a topic from lecture, miss class, or have a terrible professor.
I'd assume its similar across a lot of STEM disciplines (with respect to not actually reading textbooks much).
It depends what kind of students you want to get. A lot of start ups can make that same argument and pay for the top quality talent (with more than a few thousand dollars).
What I would look at is teaming up with universities and professors and see if you can get students to work in teams on a non-profit project for school credit (possibly as capstone projects?). It allows you to side-step the whole salary and competing with internships thing, and gives students a chance to get real world experience during the school year. Of course that creates the new problem of finding a progressive enough university to sponsor that kind of program, but it's an interesting avenue to explore.
I would argue that a good portion of snapchats use happens in extended conversations (for example, flirting). Adding the ability to quickly drop into video to share something will feel pretty organic and add to the experience.
It may not add to their original value prop, but it is a value add to the way people are actually using the product.
This sounds like a great program, I just wish it was offered year-round. Even though I think Northeastern University and Waterloo are the only schools with a completely integrated, well defined Co-op program, it seems like its a growing trend.
I'd assume having year round interns and a continuous recruitment process would be less disruptive to the team's work velocity and give you a bit bigger reach for students too.
Plus, I'm a bit jealous of some of the summer-only internships at a lot of interesting companies. Can't complain about graduating with 18+ months of interesting work experience pretty much guaranteed though.
Awesome,thanks for the reply. I'll keep an eye on it going into my last Co-op next Winter/Spring.
Any chance this will be an ongoing program not limited just to summers?
I ask this because there are a growing number of schools (including mine) that have full time Intern/Co-op programs in during the Fall and Spring semesters that I know would have interested and talented students.