HN user

iansimon

112 karma
Posts12
Comments24
View on HN

Hey Kyle, we didn't try anything more advanced than next-step sampling. You probably have a better sense than I do how much improvement such techniques are likely to yield. My unfounded suspicion is that we're close to the limit of generation quality from this dataset, and so I'm most interested in trying to gather 10-100x more skilled performances, one way or another.

There's also no consensus on whether the high- or low-temperature samples sound better. I've heard both opinions from several people.

Sageev did the final rendering, not sure what he used but I'm pretty sure it was nothing too fancy.

By far my favorite use of Hangouts is the crossword puzzle app. It's much better than having 5 people try to collaboratively solve a crossword in person.

I have no idea why I haven't seen more games like this in Hangouts...

Other people being willing to pay for something doesn't necessarily make it positive-sum or economically productive. A lot of startups (certainly not all) seem to be trying to win tournaments: http://en.wikipedia.org/wiki/Tournament_theory

In that sense, SF is is similar to Hollywood. And the result is that society devotes too many resources to the next big social network or the next big summer blockbuster, and not enough to, say, making comfortable chairs.

> If there is a 1% chance the world will end unless we do x, we shouldn’t do a cost-benefit analysis. Instead, assuming x is feasible, we should simply do it.

What is the chance the world will end even if we do x? What does feasible mean? Is killing 10% of the population feasible? Is this potential end of the world happening in a year, or in a thousand years?

I'd go on, but this is starting to look a bit like a cost-benefit analysis.

Here's how to avoid that: when you realize you're watching for that reason, go on Wikipedia and read the entire plot synopsis.

One nontrivial part is transforming the spectrogram into some representation that is robust to the things that can affect the query audio, like background noise. Another nontrivial part is figuring out, given this representation, how to quickly match the query with a song or database of songs.

I'm fairly sure that SoundHound does not match hummed or sung queries with album recordings, but with other hummed melodies submitted by users and labeled with the song ID.

I would imagine the algorithm is different, since the Shazam algorithm is designed to find exact matches corrupted by noise, EQ, etc., and two hummed versions of the same melody may vary in key, tempo, and timbre, and have small rhythmic differences.

Yes, keeping the model up to date is definitely an important issue. Don't forget that this also involves capturing new aerial imagery periodically.

Another approach that may have more potential than automatically extracting 3D structure from aerial imagery is using aerial depth maps from something like LIDAR. There's a research group at USC doing this, and you can see a video describing the process here:

http://graphics.usc.edu/~qianyizh/miscs/cvpr09_streaming_ful...