What is state of the art currently?
HN user
agencies
You say that if someone has the chops to be a real mathlete they won't need Polya's _How To Solve It_
I'll say I went to college 25 years ago with people who had competed internationally in high school and who placed competitively on the Putnam, and they LOVED Polya's book.
I think whether you enjoy seeing strategies laid out well--—whether or not you've been able to figure some of it out yourself---depsnds more on your personality than on how good you are at solving creative math problems.
Is the code or expanded explanation available?
Here's a write-up from a relatively small/personal perspective
Yeah several threads on HN that have lists of tools and pros/cons.
Have you tried https://historio.us/ ??
Depends on required features like cross browser support, cross device support, handling pdfs, ocr images, etc. Some of the mentioned features already exist in the browser. Not sure if the browser vendors are incentivized to develop and maintain such features.
How much would you be willing to pay for such a service?
Interesting intro/overview in "What every software engineer should know about search" https://scribe.rip/p/what-every-software-engineer-should-kno...
There is definitely demand and folks are willing to pay:
If it were easy/cheap enough to host, can a model like shared game servers or web/email hosting work? People pay $20 a month without thinking for web hosting. What does it take to make "search hosting" a thing, where cheap search hosting companies can crop up both at the low end with bare bones offerings and others climb up the value chain with offerings like squarespace...
Many domains have expired or content is no longer available.
If a HN story is a link to Wikipedia, the HN api serves the content of the Wikipedia page??
To clarify I'm not asking about HN itself but articles linked from HN.
As you said the HN api is great and there are at least 2 existing published crawls of it that help a lot.
Nice! Have you considered publishing your crawl data?
HN continues to pay lip service to wanting to make this better but does NOTHING to move the ball forward. Simple things like providing barrier less access to HN content and harder things like crawling all content linked from HN would be a great resource to bootstrap new search engines.
Can you describe how context would work differently than additional search terms?
What were your failed searches?
What's the source/breakdown of the 350 million pages? Thanks!
Please comment with your recent failed searches in this thread.
Any new search engine is going to need a niche with lots of users.
I haven't seen a comprehensive list of actual failed queries that new search engines could focus on solving.
How big is a medium sized corpus? 10m documents?
Does anyone know of code for simplified triangle geometry rendering of dem data? The color data from satellite images blows up storage cost but simplified geometry might be small enough to both be useful and pretty to look at.
Who has concrete steps to make this better? Seems like one or two people are making their own engines, but moving the needle is going to take a lot more than that...
Re: crawling being too hard
Have you contributed your crawl data to common crawl?
I've heard (but can't find the reference) that it doesn't have good enough coverage (independent of freshness).
Does anyone know of code for simplified triangle geometry rendering of dem data? Usually these renderings use the color data from satellite images which blows up storage cost. But simplified geometry might be small enough to both be useful and pretty to look at.
I continue to want a communal set of 50-100 million urls and data that are "good" (i.e. for any value of good and more complete than common crawl) that are accessible enough to be easy to work with that can be used to experiment with different web search tech. There are enough separate problems to tackle in web search that breaking it down would maybe move the needle. We have lots of kaggle competitions about ranking, but using closed data. What other types of kaggle projects would help web search?
What can we do to foster a sustainable bazaar of projects to make it easier to build web search engines?
Does anyone have a simple way to make rendered images from osm data? Usually it involves postgres, but I wish there was a renderer that could use the vector tiles directly...
I'm familiar. Sounds like lots of folks are (as in the thread), but still have friction being able to leverage it.
Sounds to me like there are still are gaps that new services could fill to make it easier to use.
Yeah osm definitely has some of that data. Email is in my bio if you want to chat more.