Outpan here,
We use a combination of polite massive-scale web crawling and user contribution. We had an in-house web crawler for years until recently when we released it as a stand-alone service.
HN user
A key-attribute-value store for everything.
https://outpan.mixnode.com
Outpan here,
We use a combination of polite massive-scale web crawling and user contribution. We had an in-house web crawler for years until recently when we released it as a stand-alone service.
Semantic web is too good of an idea not too iterate on constantly :)
haha. We should use this as the project mantra.
Would you be open to having a little chat via email? hi@outpan.com
This is great! thanks. I will look into adding the dbpedia.org data.
"a journey of a thousand miles begins with a single step"
As for your question about real world applications: Outpan was only a product database up until last week. It is used in over a hundred apps, some with more than a million users.
I expect this to work on a larger key space as well. It is interesting to see how the expansion works out in terms of usage patterns.
I think graph overlap is what actually determines what data is "accurate". There are currently bots writing to the database by people who are not connected and I have yet to see the overlap (since the key space is too large for the number of bots right now). I'm excited to see how that plays out.
I will send you an email :)
I was driving so couldn't stay on the phone. The number of keys for different concepts is just very large (if not infinite for the sake of avoiding philosophical debates). It will take a some time to have enough data to cover even popular keys. Fortunately I have a lot of time :D
Up until last week, outpan was strictly used for gathering data on product barcodes. We are in the process of adding data in other categories.
Twitter url seems like a strong key since it is referenced in many other contexts.
As for curation, it is intended to provide examples of what key, attr and values are regardless of their content. This is an experimental feature and might be removed/tweaked...
Common Crawl is great! however, some use cases require larger crawls with a higher frequency.
Thanks a lot! This sounds reasonable. Did you guys look into professional services for this?
Would you be able to share what your stack was? and the resources it took? Thanks a lot.
I'm not sure how he manages to crawl with this speed using such low amount of resources.
We did a benchmark on Nutch and couldn't really pass the 10-14 M(B)ps on a $1200/month machine. Even though we hired a professional to optimize the setup. The same is roughly true about Heritrix.
Just wondering if there is something missing in his setup, such as domain/ip rate limiting.
That post is what triggered my Ask post.
The problem is the huge contrast with https://www.quora.com/How-much-would-it-cost-to-crawl-1-bill...
Even taking into account the drop in prices on AWS. Also, if you take a quick look at companies that provide such services the prices are orders of magnitude higher than deusu's costs.
Awesome job!
For the life of me I can't figure out how you manage to crawl over a billion web pages (even in 2-3 months), index the data and run the server with €300 per month. Especially the crawler part...
Up to 2000 calls/minute. We will hopefully increase this soon.
That's true; the mandatory sing-up is a temporary measure, I'm sure there are a lot of ways to fight spammers including the method you mentioned. Thank you for your suggestion.
wow, thanks a lot for your thorough advice :) tbh I really hate putting road blocks in front of people who would like to contribute to the database, however we were the victim of a huge spam attack with completely random ip rotations. I will eventually remove the mandatory sign up once we have a strong moderation system in place, this shouldn't be too far from now. Thanks again, I really appreciate your comment :)
user contributions, purchased databases, paid moderation and web crawlers.