HN user

steffon

57 karma
Posts8
Comments16
View on HN

"I still don't see what I can do with how we group those preferences. "

The way in which this site will allow users to group their preferences seems like a slight organizational difference when compared with other recommendation sites that use collaborative filters, but it has huge implications. This post is meant to give people a taste of what I'm starting to try and find others who are interested. I'd love to talk about any and all specifics and their implications especially if you are programmer. If you want to chat my AIM is rocksld3.

That was my thought too so I have a demo under development. Unfortunately, it won't be ready by Oct.11, and of course, it's all about the right team anyway which is what I'm really interested in. The main reason for posting this idea is to find like-mided entrepreneurs interested in the idea so we can build a team and, as you said, "demonstrate the idea by showing...build an app".

This page explains many of netflix's limitations well: http://harry.hchen1.com/2006/10/03/391. But more importantly, look at these limitations in light of how the discovery engine is organizing its preference data and how it's collecting preference data.

The critical difference with the discovery engine is the idea of a group of content that users fill themselves with content based on criteria they see as relevant. Yes, a users aggregate preference composition is important, but what is more important is their set of preferences regarding a specific collection of content. This way, a user can be really into classical music, horror movies, and modern furniture, and get relevant recommendations for each interest, connecting with people who are most in the know regarding each interest.

Here are two large differences between IRC:

1.) One of the main problems I am trying to solve is in your second sentence: "Once you find a chan you like..." It's time consuming to find new content you like, especially in a setting as nebulous (for most people) as IRC. With the discovery engine, all a user needs to do is rate recommendations based on what they are interested in, and they are immediately connected with groups of like-minded people and what they know. Fast and easy.

2.) In addition to connecting to groups of like-minded people, the user gets recommendations from those who are most influential and most in the know. All content is not equally desirable, and users using the discovery engine will only get recommended that content which has empirically shown is wanted by people that think like them (relevant).

"groups of content" are collections of content that users create themselves (the social bookmarking aspect of the site). Users that want to build a reputation for being in the know and influential are incentivized to make these collections relevant. The groups also become lists users can form to bookmark content they find around the web and want to save it one spot.

Based on what you have put into your collection of content, and collections others have made, the algorithm aggregates those similar people into the same network. People in the same network get content recommended to each other from their like-minded peers that they have not discovered on their own.

With regards to the feasability of such an algorithm, I've talked about it with many mathematicians and machine learning programmers and the wheel does not have to be reinvented for this application. The tools already exist, and just have to be customized and tweaked for this application.

I've got a few ideas, such as front-loading the site with content and preference data by using differnt API's, such as Flickr's API, Youtube's, del.icio.us, Last.fm's. This way when you showup, lots of your preferences and rating "work" comes along with you. Additionally, the preference dataset is jumpstarted.

What you would tell your little sister about the site is "find more things you like, and if you're good at finding new content, be recognized for it."

I agree: I rarely find that suggestions based off of anonymous, aggregrate consumer "social" data helps "pick" worthwhile suggestions for me. If a category of shopping is completely new to me, then what everyone else knows can make me more aware of the commonplace options (which is helpful). But if I know anything about the category of items I am looking at, I really don't care what the anonymous masses already know because I know it already too or it's a very "obvious" connection.

The examples you use as counter-arguments to my approach are not really counter-arguments, but deal with a different issue than what I was talking addressing in my last response. I was talking about finding music based on personal preferences. Your counter-examples of "making out with my girl" or dancing "at a club" are group situations where group preferences are most important.

You are right when you say that compilations like "Ibiza Club" or "Ellington for Lovers" make a lot of money: so does selling Muzak (the background music in supermarkets and most commercial spaces with music). Background music is the ultimate "genre" that is compatible with group preferences: no one is offended... but at the same time, nobody really cares.

I agree with you that musical preferences have more to them than "people like you also like"... this would disregard the reality that people sometimes do categorize music by situation/mood, categories like listening to music to dance, exercise, study, make-out, host cocktail parties etc.

But as I suggested before, domain specific categories of music do not have to be mutually exclusive to personal preferences (currently they are). Let's say your goal is to have romantic music. Why not use a "people like you also like" function bounded by the category of romantic music? This way you could get romantic music, i.e. music everyone thinks is romantic, and romantic music that you like as an individual. Even better would be to use a "people like you and your girlfriend also like" function within the specific category of romantic music :)

"Also, with all the choice available, it's easy to overestimate people's desire to even HAVE choice." This is pretty fatalistic don't you think? I think the explosion of choice online frustrates people because they KNOW something is out there that they will really love, but they can't FIND it. This screams opportunity for a website to act as a choice agent to direct people to the music, video, merchandise etc. that they want but can't find themselves among the infinite choices. Infinite choice results in infinite search costs without a decision agent.

I don't think that it is entirely accurate to equate news delivery with music delivery, nor do I think the lessons from journalists are entirely relevant when it comes to building a social platform for music playlists. Cultural objects (e.g. music,art,fashion,film,literature etc.) are substantively different than news articles.

The news is largely valued for its accuracy, timeliness, and topical relevance to a reader. Music and other cultural objects are valued for their enjoyment, which is contingent on individual personal preferences. I'm skeptical of a platform that would use domain-specific talent to create playlists for what comes down to individual personal preference. Domain specific talent makes much more sense regarding news where the relevance can be easily determined because it is largely topical, whereas the relevance for music is based on personal preference, which can't be determined through domain specific talent.

I'm also not convinced that more meta-tagging is the answer either. If I tag a song with a mood, genre, style, etc., it will help a user find that song based on those new categories. But in the end, finding new songs based on categories isn't that helpful because what matters is finding songs based on individual personal preferences. For example, I might know I want an energetic, intense, electro-rock ballad, but defining that domain doesn't guarantee I will like the results. People rarely like all the songs on an album, and as far as domains go, albums are very tightly defined domains. Likewise, using Uber-Dj's as domains don't seem to be the best answer. After all, their playlists are based off of their preferences, and, regardless of how much they "know" about music, unless their preferences are like mine, their list will not help much.

I completely agree with this article. And a startup that could efficiently tap into people's hunger for fame and recognition on a broad basis could become a destination for user generated content that far outdoes current web 2.0 offerings.

"We observed that users cite a variety of reasons for posting content online--chief among them, a hunger for fame, the urge to have fun, and a desire to share experiences with friends"

What's interesting about the notion of reputation and fame is that, for a user to feel like they belong or are "having fun" or a held in esteem by a group of people, that group needs to be a compatible reference group meaning, they need to have things in common (this idea was touched upon in the article "The Problem with Social News" http://news.ycombinator.com/item?id=50015 and the discussion of homophily, meaning loosely birds of a feather flock together). For examples, entrepreneurs are much more likely to be concerned about the opinions of other entrepreneurs than other people. I take more seriously the music suggestions of people that have the same tastes as me (obvious, right?).

So if everyone on a site could be plugged into a compatible reference group, the fame and reputation motivation could be leveraged across much larger populations of users, not just the ones that want to be, for example, on the top viewed videos on youtube. Each user would be compared against those users with which they have common interests, so the idea of having a reputation becomes more salient, so people spend more time contributing, and the content becomes better for everyone at an individual level. Lots of sites are able to harness the reputation motive on subject specific areas (like Y Hacker news), but no site has tried to harness the motivation for reputation across all interest groups. Definitely email me if this interests you because it fascinates me and I have plans to build such a site.

But would the Digg of music be enough? iJigg, which is something like that, launcher earlier this year. http://www.techcrunch.com/2007/01/18/jigg-that-music/. Although it doesn't seem to have moderaters with real music knowledge. But even if it did, that wouldn't be enough to make results relevant for everyone because the criteria for relevance is an individuals concept of "good". Sure moderators might know what is "good" for certain audiences, but if results are organized in a digg-esque fashion, it becomes a top-ten list, meaning that no matter how good your moderators are, they can't generate recommendations that will be "good" for individuals (and the strategy to appeal to the largerst number of individuals is to appeal to the lowest common denominator, which one could argue radio already does well.)

This makes sense to me for publishers and writers that aren't in the top 25% of books being sold. Marketing for many consumables can be thought of as a two-step diffusion process: Diffusion step 1. Utilize the mass-media to send your product message toward a target audience. Diffusion step 2. Hope that the early adopters and opinion leaders within those audiences tell their friends and exercise their networks of influence so that "word-of-mouth" perpetuates the message first sent by the mass-media.

For many cultural product industries, like publishing, they release many more products then they can afford to buy mass-media advertising campaigns for. For those products without mass-media coverage, it makes a lot of sense to give away free copies of the book. The innovators and early adopters that crave and promote new products have the material they need to jumpstart the second step of the diffusion process (word-of-mouth), and a publisher and writer could see higher sales numbers than they otherwise would have: they still benefit from second stage diffusion without the mass-media costs.

And if there is a concern that the "message would get out" that the book is free online, the publisher could only give away the first 12,000 free online. It seems like putting a cap on free copies could maximize the potential of "word-of-mouth" diffusion and limit the risk of lost sales.

What iTunes does not reduce are the search costs incurred by a user. If a user isn't planning on purchasing what is on the top ten list, a user has to know what he wants beforehand at iTunes, just like shopping at a Virgin Megastore. So before a purchase is made, a user still has to pay search costs to find what he wants before he makes a purchase.

So as you said, you would go to Pandora or last.fm instead of illegal downloads to find new music and when the new "good" song is discovered, you go to iTunes to purchase it. In this situation, Pandora and last.fm decrease your search costs by presenting what you found to be a "good" song in less time than it would have taken you to find a "good" song on your own. So as long as Pandora and last.fm can reduce your search costs enough so that the value of the "good" song outweighs the cost, then Pandora and last.fm replace the record companies as better agents and you derive more value from your music.

But are Last.fm and Pandora good enough agents for everyone? Or do they really only work as good agents for some people (i.e. better solutions exist)? For example, record companies can and do play a strong hand in determining the results on Last.fm. As long as there are mass-media outlets like the radio and MTV, record companies can influence the "tastes" of mass numbers of people, which influence what music they play and share on file networks, which influence dramatically the scrobbling statistics on Last.fm and the recommendations Last.fm produces with its collaborative filter, skewing results dramatically and consistently toward major label artists: getting results of mainstream artists on Last.fm does not reduce search costs as these artists were easy to find regardless.

It seems to me that the search costs for discovering music on Pandora are also very high, as I have to listen to Pandora like a radio station as opposed to just getting results, and the database is set-up by what musicologists think. Sure the sonic characteristics and structure of a song influence the song, but it doesn't determine whether I will think it is "good". As an extreme example, take two piano sonatas: one is written by Chopin, the other by Tom Dick last week in Ohio. A musicologist at Pandora would classify them the same as they have the same sonic characteristics, arrangement, instrumentation, structure etc. but that does not help me reduce my search costs because I still have to listen to Tom Dick's sonata on their web radio before I get to Chopin.

I think audiences that consume a lot of music could significantly reduce their search costs on a site like oink, because everyone there consumes a lot of music. Because everyone in the community is always on the hunt for the new "good" music (and are technologically savvy), their top ten lists are predictably very relevant to their community, probably much higher than any of them would ever find on Pandora or Last.fm.