For comparison, google offers competitive fresh-out-of-top-tier NLP Ph.D.s ~150k cash + 50k /yr bonus.
HN user
speechduh
I saw a great talk about this @ SIAM data mining by Dave Madigan, for http://omop.org/
Basically, they took a TON of demographic research in the health sciences, explored all possible hyperparameter tunings, and found that they could get p-values of 0.05 in either direction for most of the papers depending on choice of data sources and many other types of hyperparameters.
Far field speech recognition is _really_ hard. You can't just throw any random set of mic / speaker into a tube and expect it to work. The echo is almost exclusively mic / speaker, in a very specific configuration. There's no way the product they're shipping will be able to do far field speech recognition (much less while playing music).
I tend to think computer literacy is the wrong approach. I think the right approach looks like "interfaces that don't require literacy" (one of the reasons why I think human-quality NLU for command & control systems are going to be amazing). Most people don't need most of the functionality of a computer.
So, for anyone not aware: Schmidhuber is _obsessed_ with this. He wrote an enormous literature review of deep learning [0] basically because he felt that people weren't crediting ideas enough. This isn't a one-off essay, for him, he's been banging this drum for quite a while.
Not saying he's wrong, just FYI.
When you have that much information, the chances for spurious conclusions are also endless.
IVONA is text to speech only.
state of the art is still very much using WFSTs and DNNHMMs. IBM and Google are still beating baidu.
Even if the speech-to-text doesn't, the natural language understanding does.
That particular stuff is actually pretty typical. I have a textbook that shows similar results on Shakespeare using N-grams from years ago.
I do think it's fair to call it out, though. Articles in second or third tier journals deserve higher levels of scrutiny. I'd rather see this replicated a couple of times before acting on it.
As you yourself said, the claim is "qualitatively" different, not "quantitatively" different. If you think about what qualitatively means, I really have no idea what you're trying to say. "Just" being "more extreme" is sufficient for something to be qualitatively different. Different symptoms start manifesting, especially as compensatory systems start to fail. As far as I can tell your point is vacuous; please clarify, otherwise.
If you're instead trying to say that they're biologically / mechanically similar phenomena, well, that's a different discussion we could have.
Have you ever experienced severe depression (i.e., the type that prevents you from getting out of bed for months or causes you to be hospitalized)? Because it's absolutely fucking awful. I'd love some clarification of where you're going with this hypothetical, because right now it sounds like you're denying the experience of a whole lot of people in a whole lot of pain, without having much of a point.
+1.
Brother was (fresh out of college) employee 20 at a company that IPO'd > $1 billion, and he got about $150k out of his options from 3 years of work. You have to get lucky to find a company IPOing enough to make you a millionaire.
Uhm, yes, duh. Isn't this common knowledge? How do people think speech recognition systems are trained? What's more upsetting is that this person is so uninformed about speech systems that they think this is weird, AND they have access. The only people who should have access are the people who are actually doing science.
Ahhh. Reading clarified it: this is someone who got hired to do transcription.
"I'm given an audio file (sound bite) and the corresponding text based translation (how the phone translated the speech). My job is to listen to the file, compare it to the text and provide feed back on how correctly the sound bite was interpreted by the phone. If the text and speech are a perfect match, I just move on. However, if the phone either translated something incorrectly due to a heavy accent or loud background noise, I note that in my evaluation. "