HN user

adrianbg

208 karma

https://twitter.com/AdrianBPatter

[ my public key: https://keybase.io/adrianb; my proof: https://keybase.io/adrianb/sigs/kITb4pm-MtYKLO4tk86T1ch5Sy37O-SEIrULxl7ekL4 ]

Posts10
Comments126
View on HN

I ranted about this in the unofficial Alexa Slack team and the people there gave some plausible (but unsatisfying) reasons why it works like that. I'm not a huge fan of Amazon though, so I'm happy to blame them. Essentially it boils down to different regions being treated as different "languages" as well as having features rolled out to them at different rates.

Google's speech recognition is supposed to be much better, though not sure if they'd allow "one intent to rule them all" like you want.

For free-form speech recognition in Alexa, the best option I've seen mentioned on the public Alexa Slack team is using the "SearchQuery" slot. So you'd still have to make a weird catch-all intent that would eat up some of the words (and you wouldn't be able to see them). At the same time, you shouldn't assume that Alexa will give you very good results with such loose constraints. Even in my simple skill it's very bad about confusing certain pairs of words.

Sometimes I wish it didn't. They don't give you the original audio, any kind of confidence score, or even alternative hypotheses. It's really a pretty rigid platform. A lot of things that seem like they should be reasonable are impossible. Eg., I'd prefer to just say a list of post titles and let people interrupt Alexa when they hear something they like. That is impossible right now without pretty serious hacks.

Yeah.. turns out that's actually a friend of a friend. I'm curious to see what people prefer for obvious reasons :). The advantages of my approach are personalization, interactivity, and scalability beyond just HN. Theirs is probably going to be a nicer experience than I can do for non-interactive use.

Alexa does most of the hard stuff: speech recognition / intent detection, and speech synthesis.

My back end is a simple Python service on GCP that handles HTTP requests from Alexa. The same service also downloads the HN front page from the FireBase mirror and gets summaries from this API:

https://rapidapi.com/textanalysis/api/Text%20Summarization

It's not perfect though, so I may switch to a more expensive summarization API, supplement it with manual summaries, and/or train my own summarization model.

Yes they are already penalized. I personally am kinda happy about this blacklisting.

The startup doesn't need to file the H1B themselves. The employee can transfer it from another company.

Yes, exactly. I do wonder whether a similarly good end-to-end system could be trained by constraining the alignments as I've seen done in some papers.

That's fair. I think both approaches are useful in different situations. Mozilla seems to be focussing on Siri-like use-cases as opposed to dictation. Even for dictation, for many people having to train the system themselves is more work than they're willing to do. I'm sure what you want will exist eventually :)

Yeah. I don't get the impression that the Kaldi core team has been trying very hard recently to get SOTA on eval2k/switchboard. This number uses one acoustic model with a trigram LM decode + fourgram rescoring -- there isn't even a neural net language model in there. If I remember correctly, Microsoft's first "human parity" result used something like three acoustic models and at least four types of language models. This Kaldi model is competitive with the best single acoustic model Microsoft used.