If there were canonical sports stats data sources to include, it wouldn't be difficult. Email us at info@kemvi.com if there's something specific you have in mind.
HN user
kemvi
You can't get a complete list the way it's currently set up. But SomeEntity.property(), without an argument, will give you a sample of them.
Not yet, but drop us a line at info@kemvi.com.
Coming soon, along with R and Javascript!
Thanks, it's fixed -- that shouldn't have slipped through testing.
Currently, it's great for socioeconomic data on countries, as well as for high-res socioeconomic data for the US (county resolution). There's data from the World Bank, Dept of Health, Dept of Labor, Dept of Justice, and several other public sources, with all entities reconciled.
We're building a pipeline to make it so that ETL involves as little human time as possible, but it's in its early stages now.
It's highly interconnected. For example, you can make hops like USCounty--USState--Senator--Vote on economic stimulus bill.
The data backend is indeed NoSQL, but we didn't choose a graph db because there are no really good graph solutions that are easily parallelized.
The idea is that you'd use this via the python shell (and soon, from your R code or Javascript code), and not through the website. Looks like we need to make that clearer. The website stack is Postgres+Django+Python.
I had heard of these guys, but hadn't seen their site in months. Sounds like they're doing something pretty similar, which is exciting, because it means they've tested the business model.
The freemium model is exactly what I was thinking. And, yes: Rapidminer, Weka, etc, are all on the desktop. Thanks for the motivation.
Great, thanks! I knew of timetric, mentioned below, but not about these. I suppose there's a strong selection bias.
The business model probably should be B2B: I imagine the users would be companies that need to do business intelligence, journalists, universities, and research organizations.