HN user

ftyers

499 karma
Posts8
Comments140
View on HN
Common Voice 3 years ago

If you want to maximize the utility of a dataset like this, you really would want to let each speaker at least assign a lot of tags/labels to their profile; even if you don't want to deal with the hornet nest of trying to figure out all the distinctions, even unstructured labels would be a start, and ideally allowing people to tag individual recordings as well, because there are a lot more variations than just "language" and "accent" here.

This is exactly what the freeform accent (actually "variant") field is. You can add as many tags as you like. https://foundation.mozilla.org/en/blog/how-we-are-making-com...

Common Voice 3 years ago

Target segment was a was of including specific subdatasets. For example the digits dataset which was just the digits 0-9 and yes/no.

Yeah, that's what I mean, e.g. it's mostly existing published stuff. The new stuff is some partial summarisation in Eastern Huasteca Nahuatl, and some spoken audio (by EHN speakers), although, it's unclear what the audio gains. Without training, it's not really intelligible to most speakers of modern varieties.

The sad thing is that there isn't really anything new here. It's the Anderson and Dibble translation, and some random extra stuff. For 15 years work it's quite a limited contribution. In addition, it's not freely licensed. I'm working on a free/open-source licensed edition with linguistic annotation. If anyone is interested, ask for the link, it's on GitHub.

Yeah, this is insane e.g. I've seen cents/fl.oz, dollars/litre, cents/ml, dollars/unit. For the same product. And yes, totally malicious. Fresh Thyme does this. It's ugly.