Very good example of why developers are dumb for writing mobile apps. It makes zero business sense with the exception of the very few who sell hundreds of thousands.
Apple screws the developers and we all bend over. What a joke!
HN user
Very good example of why developers are dumb for writing mobile apps. It makes zero business sense with the exception of the very few who sell hundreds of thousands.
Apple screws the developers and we all bend over. What a joke!
Not really a lot of relevant experience. Would like to see some Hadoop or heck, even some C#. Let's talk again in a year or two.
Successful high tech CEO here.
I don't know where all of this bad advice came from, but you should never, ever have an equal cofounder.
same here...
I'm in the course as well and don't really understand what the fuss about Octave is. The help system is abysmal and third-party library support seems far behind other popular platforms. It's also quite unstable on Windows.
I've been playing around with R for the past two weeks and have been more or less happy (with the exception of the memory and speed limitations in the GNU implementation).
Nevertheless, this will be a really exciting field for the next decade or so. Amazing possibilities right now!
this
I see such bad financial advice coming from those who obviously have never made money (shouldn't be surprising, I guess). And then when someone comes along who has made a lot of money, most people ignore what they have to say.
$630? Why is this on the front page? Why does anyone even care? This is not important.
And assuming they could (despite your arguments against) determine the method of monetization on a site
I can assure you that that they cannot do this accurately. You would be amazed at the scummy business models that openly advertise on Google and are not caught. It would be incredibly difficult to do so as some of them are downright ingenious (one example: free software that updates your drivers.)
It sounds to me that you are positing some type of magical technology that doesn't exist predicated on Google's seeming omniscience. Of course, I am eager to stand corrected...?
No...
Exactly, right? I mean, I read that and thought "I'd kill OP to be 20 again".
20... no job, no ambition, no problems.
I want you to put this in perspective.
I read your previous post and you were working a shit job. I am a CEO of a tech company and I can assure you I know what you are talking about. I've never heard of such crazy hours (on a regular basis).
You're 20, so that necessarily means your job opportunities are limited. But so what? I was 24 before I got my first "real" job. And today I'm doing pretty well by any standards.
You have the next 20 years to do manifest whatever kind of crazy dream you have in your head. So go do it.
Edit: Thought maybe I would also suggest to add a few words about your location, what type of coding you like to do (languages, algorithms, domain space, w/e), and how you would rate your own level of emotional maturity. Maybe there are a few people out there who might be willing to give you a shot if it's the right fit... but be honest! Nothing could be worse than misrepresenting yourself and ending up similarly unhappy all for a few bucks. Way better to live honestly, trudge through the short-term pain, and find a long term solution that makes you really happy.
I know you were looking for an answer from Matt but I thought I would offer up my opinion here (as you might have noticed, I love talking about this stuff).
Our models currently suggest that the presence of contextual advertising is a significant predictive factor of webspam.
We use 10-fold bagging and classification trees, so it's not all that easy to generalize. But I pulled one model out at random for fun.
The top predictive factor in this particular model is the probability outcome of the bigrams (word pairs) extracted the visible text on the page. Here are a few significant bigrams:
relocationcompanies productsproviding productspure qualitybook recruitmentwebsite ticketsour thesetraffic representingclients todayplay tourshigh registryrepair rentproperties weddingportal printingcanvas prhuman privacyprotection providingefficient waytrade printingstationery priceseverything website*daily
Next, this model looks for tokens extracted from the URL and particular meta tags from the page. Similar to above, but I believe unigrams only. A few examples follow. Please keep in mind that none of these phrases are used individually... they are each weighted and combined with all other known factors on the page:
offer review book more Management into Web Library blog Joomla forums
The model then looks at the outdegree of the page (number of unique domains pointed to).
From there, it breaks down into TLD (.biz, .ru, .gov, etc)
The file gets pretty hard to decipher at this point (it's a huge XML file) but contextual advertising is used as a predictive variable throughout.
Just from eyeballing it, it appears to be more or less as significant as the precision and recall rate of high value commercial terms, average word length (western languages only), and visible text length.
Based on what I'm looking at right now, my answer would be that sponsored posts are going to be far more harmful to the user experience than advertising.
Can't answer the rest of your question which I assume relates to the number of ad blocks or amount of space taken up by ads... we don't measure it.
Edit: Just realized that Google will probably delist this page within 24 hours. Should've used a gif for those bigrams. Oh well ;-)
This really made me think. But I took exception to your comment:
I agree but don't you think that from an algorithmic point of view Google would be better looking at what the user wants and what monetization models they prefer[...]
No. In fact, I am a rather loudmouthed opponent to Google's somewhat clumsy attempts to measure this ala "Quality Score".
In addition to webspam detection and machine learning, I have spent way too much time in marketing (I have a master's degree in marketing, in fact.)
A neat thing I learned along the way was the value of market research.
There are so many nuances in every line of business. Segments, preferences, pricing, even down to minutia (now well studied) such as fonts, gutter widths, copy styles, and so on.
You can learn a lot by combining large amounts of data and well chosen machine learning algos. But even with a few thousand businesses in most categories in a particular country (far less outside of the US), that doesn't give an outsider enough data to truly distinguish what can be a winning formula from a spammy one. This knowledge is hard won through carefully executed experiments and research.
A few years ago I was researching the topic of landing page formulas by category. One example that stuck out most in my mind was mortgages. There were a few tried and true "formulas" that significantly outperformed the rest. Two stuck out:
1) Man, woman, and sometimes child standing on a green lawn in front of home. Arrow pointing down from top left of landing page to mid/lower right positioned form. Form limited to three fields.
Edit: http://imgur.com/90VmB
2) Picture of home/s docked to bottom of lead gen page. No people. Light/white background. Arrow pointing down from top left of landing page to centrally located form.
Edit: http://imgur.com/JkLlH
These sites were incredibly successful. More than a few of them had to contend with quality score issues over the years. Can an algorithm capture nuances such as the ones I mentioned? In theory... they could. But today, they don't. All of Google's QS algorithms to date have been failed attempts and have caused an incredible amount of harm and distrust.
You finished that sentence with:
to see versus the averages in terms of monetization models on spam sites.
I'm not at all sure what this means. Could you explain? Is it even possible to directly model the monetization model of a site without having direct access to their metrics?
Matt,
You are of course correct. The fault is mine for miscommunicating... I find myself becoming less self-editorial these days when I write on the web and tend to think everyone is on the same page as I am.
I was actually referring to an informal study I did earlier this year. I measured sites which were receiving an average of 50,000 or more visitors from Google US search (organic) per month over a six month period. Then I compared those with a similar set from a subsequent six month period to see which had significantly dropped off in traffic and rankings. The purpose of this was to estimate the number of significant sites which were penalized over that period of time. The final estimate came to about 700 sites/year which were penalized. There are lots of uncontrolled variables here of course... but I was looking for an "order of magnitude" answer simply for curiosity's sake.
The 1 million spam pages created per day were of course excluded from consideration as they never received much traffic from Google in the first place.
So just to clarify my earlier response, I am advocating for a policy that would apply to websites exceeding a certain threshold of organic traffic for a significant period of time.
We do mean well, and for the most part when people say we're arrogant it's because we didn't hire them, or they're unhappy with our policies, or something along those lines. They're inferring arrogance because it makes them feel better.
I LOL'd when I read this. How ironically arrogant.
This is a fair comment but I believe the issues stem more from the process Google follows.
It is obvious that Google cannot communicate exact reasons why a site was penalized as that would help spammers. However, there is nothing that prevents them from adding a step to warn the offending website and give them a heads up before the ban/penalty takes place, along with an explanation of the policy that is/was being violated.
Most of these heads up would go ignored, some would not and yes, it would incur a support cost. However, the number of websites which are significantly penalized isn't onerous... I believe fewer than 1,000 each year?
When a company has become the defacto gateway to the internet, I believe they have a responsibility to webmasters. Google has lost a lot of goodwill over the years because of these seemingly arbitrary penalties... Instituting such a practice would be a worthwhile investment.
That's the ultimate check on Google: if we start to act too abusive or "evil" we know that people can desert us. So it's in our enlightened self-interest to try to act in our users' long-term interests.
I hear this line thrown about quite a bit. And while it's true with regular users, it's certainly not true for webmasters or advertisers. Google controls around 67% of all US search share. If an advertiser doesn't play by your rules, they forfeit a significant amount of natural search traffic.
All good for most of your policies but there are some real gray areas. I had a site years ago that got hit with a javascript hack on an obscure page. StopBadware found it within a day or two and suddenly we were blacklisted... virtually all organic traffic disappeared overnight. It took weeks to get the warning lifted and that was only as a result of a significant viral PR campaign (such as this).
Maybe this has been addressed, but there are other areas. Affiliate sites are also penalized quite heavily by Google. It's one thing to take a stand due to the supposed quality of many of these sites (which frankly has little correlation with the presence of affiliate links... most sites suck). It's quite another when Google has a large affiliate advertising practice in house and a significant investment in an affiliate link tracking/cloaking company.
I mention this with all due respect and I hope you take it as constructive criticism. I think the organic side of the house does a great job overall. The paid side is another story IMHO. Part of this is organizational stupidity... I struggle with this every day and I have a much smaller organization.
Sadly, quite true.
"True" - because my current understanding (which Matt_Cutts can elucidate on if he chooses to) is that Google has looked into - but does not currently incorporate - the presence of advertising as a spam signal.
"Sadly" - because my independent research has shown that advertising - most notably the presence of Google AdSense - is a reliable predictive variable of a page being spam.
All things being equal, a page with AdSense blocks on it is far more likely to be spam. Yet as of a few months ago, that does not appear to weigh very heavily into the equation.
I think %50 of the problem is the arbitrary picking of sites to block (and it's not working, btw[1]) and %50 of it is that google seems uninterested in explaining or advising people when it happens to them.
This. Assume you work on the webspam team and you have a 92% spam detection rate but a 99.9% accuracy rate on what you do detect.
There are around 40,000,000 active domains in a given month listed on Google. That means 40,000 sites on average are being penalized without reason.
Hi Matt,
We don't know each other but I think we know of each other. I'm rather immersed in webspam detection and found this incredibly interesting.
You imply that the challenge is finding a solution that scales. Yet it sounds to me from your response that this site was flagged via manual review. Did I misunderstand?
If I heard you correctly, then is manual review a significant equation in the webspam detection methodology? You guys are boiling the ocean so I find that rather hard to swallow.
The more likely conclusion I can draw is that he had a significant number of (auto-generated) pages on his site flagged as spam and that in turn raised some eyebrows.
BTW, you and your team are doing some amazing work. I wish the paid side was up to the standards you set.
Google:Social::Microsoft:Search
True, I thought I mentioned the color thing, but I guess not. You definitely could not read a TRS-80 clearly on a TV. You could on an Apple, but who would? Most people - remember the buyers were predominantly hobbyists and schools - bought the monitors.
I was working on a TRS-80 Model I in the late 70s, I want to say 1977 but maybe it was '78. I remember it vividly, even down to the massive 4kb memory expansion (which weighed around 10lbs and threw off massive amounts of heat).
Later on we got a Model III. It didn't have nearly as much character as the I. I didn't like the monolithic looks of it much but it was admittedly a much cleaner machine with the built-in disk drives (the Model I eventually supported 5 1/4" floppies but they were humongous standalone units.)
What fun!
I have to disagree. I used the Apple IIe and the TRS-80 extensively during this period. I coded games and other software in BASIC and assembly on both machines. I poured over the schematics of both machines for hours at a time. You could literally say that I knew those machines inside and out.
What you refer to as "user experience" was not at all uncommon. The TRS-80, the Sinclair, and many other computers shipped in plastic boxes. I don't recall the Apple natively supporting a hookup to a TV (the TRS-80 did not), but this was certainly not a positive at the time... TVs were far more difficult to read and work on than monitors.
Both the TRS-80 and the Apple shipped with BASIC and connected to a cassette recorder.
There were dozens of other computers, but there were just a few that were commonly used.
Interesting service and a nice UI. I just gave it a shot. Interested to see how well it works!