HN user

picheny

28 karma
Posts0
Comments21
View on HN
No posts found.

It is understandable; different people may see different performance depending on what speech service is used. In addition, please keep in mind this is just a beta service; we have only been out there less than a week and know we still have some tuning to do. We also recognize the need for customization and personalization abilities; keep watching our website for continuing developments.

We really appreciate the comments and will try to fix the problems. To be honest, sometimes developers can't see even the most obvious flaws in documentation. If you can highlight even one incomprehensible point it would help a lot to accelerate the revision process.

One thing a machine learning system can do that any one human cannot do is ingest lots of data. For example, for some tasks in which I have tried to compare human vs machine speech recognition performance the machine actually does better because the machine may - for example - know a singer's name that an individual human may not recognize.

Yeah it was. At the risk of waxing about old history, when we demoed the first large vocabulary speech recognition system back in 1984 it ran on a bunch of IBM mainframes. Within two years it was running on a PC with some special purpose cards. Today much more powerful recognizers run locally on smartphones. We have always found we can shrink something down once we solve the basic problem, and it is important not to let computational limitations prevent you from seeing the best solution.

As a speech technologist, I am amazed and proud about how far long the technology has progressed, especially over the last few years. Even my wife now uses speech input on mobile devices (and may finally think I may be doing something productive...). With that said, speech input is still a surprisingly finicky technology and different people will see different beahviors across systems from different providers.

We know we have strong core speech technology based on various comparisons we have done in the context of competitive evaluations done in conjunction with various government funded speech programs. However, our service is still very new. We could have waited for months to tune it, but our primary goal here is to solicit feedback from the community for how to make our services easier to use, especially in the context of our other platform services. We don't want to wait till the design is so mature that it is impossible to change - so any and all feedback is very welcome!

We have worked on audio analytics in the past for things such as outdoor sound detection and vehicle identification. We are currently focusing on speech-based analytics such as language ID and affect recognition. The statistical methodologies we are using for speech are easily extended to such domains. We hope that by puuting out these initial speech services we will get feedback from the community about related problems and welcome your suggestions.