Only Big Tech (Microsoft,Google,Facebook) can crawl the web at scale because they own the major content companies and they severly throttle the competition's crawlers, and sometimes outright block them. I'm not saying it's impossible to get around, but it is certainly very difficult, and you could be thrown in prison for violating the CFAA.
HN user
gbmatt
We are just robots in a human simulator, reliving our creation.
Q: how might an AI algorithm be modified in order to return citations with its response?
A: There are several ways in which an AI algorithm could be modified to return citations with its responses. Here are a few possibilities:
One approach would be to add a citation generation module to the AI algorithm. This module could be trained on a dataset of citations and could be used to generate appropriate citations for the information provided by the AI algorithm.
Another option would be to modify the AI algorithm to include a database of references, with each reference being associated with a specific piece of information. When the AI algorithm provides a response, it could also return the appropriate reference(s) from the database.
A third possibility would be to modify the AI algorithm to search external sources of information (such as online databases or libraries) for relevant citations whenever it provides a response. The AI could then return the relevant citations along with its response.
It is worth noting that adding a citation generation module or incorporating external sources of information into an AI algorithm can be a complex task, as it requires the development of additional algorithms and data structures to support these features. It may also require the acquisition of additional data or resources to support the citation generation process.I just posted this same comment on the ddg story, but I'm going to post it here as well.
Google forced my search engine (gigablast) basically out of business. I had ixquick.com as a big client at one time; I was providing them with search results from my custom web search engine. Then their CEO called me one day and told me he was cancelling, even though he'd been a client for over 10 years. He said it was because of some change Google had made to their agreement. Ixquick needed Google's results and ads for their startpage.com website, and, even though my results were shown on their ixquick.com and later ixquick.eu site, apparently Google wasn't good with that.
Yeah, Google forced my search engine basically out of business. I had ixquick.com as a big client at one time; I was providing them with search results from my custom web search engine. Then their CEO called me one day and told me he was cancelling, even though he'd been a client for over 10 years. He said it was because of some change Google had made to their agreement. Ixquick needed Google's results and ads for their startpage.com website, and, even though my results were shown on their ixquick.com and later ixquick.eu sites, apparently Google wasn't good with that.
everyone needs equal access to public data. right now only big tech can download the many web pages (without thottling or being ip banned) on linkedin (microsoft), youtube (google), facebook, github (microsoft) and billions of more pages. this also leads to a gap on AI training sets to give big tech even more entrenchment. for instance, only microsoft can build that ai coding application they did because other companies can't access all of github without being throttled or ip banned (last time i checked - but i could be wrong now)[microsoft owns github]. regardless, we need some sort of bot 'bill of rights' to ensure equal access going forward. perhaps the answer is legislation or perhaps it is some massive p2p proxy net. i think it is legislation because the p2p proxy net is too hard to implement, and it would have to solve turing tests.
but perhaps web 3.0 (dweb) can just bypass all this nonsense and make its own versions of these popular services with baked-in accessibility for all.
the complexity of the search algorithm has also increased substantially since 2005 And, in 2005, a billion page index was pretty big. Now it's closer to 100 billion.
thanks ben, you are too kind.
the javascript is run by your browser, so you can fully audit it.
hey thanks for the recognition, people. :) finally, all my problems are solved. this comment is here for hacker news karma points.
I'd argue that a level playing field and more competition in the search space is a good thing.
100% custom.
that's tripped out. where did you hear about that?
I'll admit I had not been working on the quality of single term queries as much as I should have lately. However, especially for such simple queries, having a database of link text (inbound hyperlinks and the associated hypertest) is very, very important. And you don't get the necessary corpus of link text if you have a small index. So in this particular case the index size is, indeed, quite likely a factor.
And thank you for the elaborate breakdown. It is quite useful and very informative, and was nice of you to present.
And I'm not saying that index size is the only obstacle here. I just feel it's the biggest single issue holding Gigablast's quality back. Certainly, there are other quality issues in the algorithm and you might have touched on some there.
Cloudflare is not the only gatekeeper, too. Keep that in mind. There's many others and, as an upstart search engine operator, it's quite overwhelming to have to deal with them all. Some of them have contempt for you when you approach them. I've had one gatekeeper actually list my bot as a bad actor in an example in some of their documentation. So, don't get me wrong, this is about gatekeepers in general, not just only Cloudflare and Cloudfront.
It's not quite that easy. Have you ever tried it? See my post below. Basically, yes, I've done it, but i had to go through a lot and was lucky enough to even get them to listen to me. I just happened to know the right person to get me through. So, super lucky there. Furthermore, they have an AI that takes you off the whitelist if it sees your bot 'misbehave', whatever that is. So if you have a certain kind of bug in your spider, or your bot 'misbehaves', whatever that means is anyone's guess, then you're going to get kicked off the list. So then what? You have to try to get on the whitelist again? They have Bing and Google on some special short lists so those guys don't have to sweat all these hurdles. Lastly, their UI and documentation is heavily centered around Google and Bing, so upstart search engines aren't getting the same treatment.
there's some stuff here : https://github.com/gigablast/open-source-search-engine
brave 'falls back' to bing. which in my experience is most of the time. in fact, out of all the queries i did a while back, they all seemed to come directly from bing. is there a way to disable the reliance on bing and get pure 'brave only' results? and can you be more specific as to what this fraction is? do you blend at all?
yes, large proxy networks are potential solutions. but they cost money, and you are playing a cat and mouse game with turing tests, and some sites require a login. furthermore, people have tried to use these to spider linkedin (sometimes creating fake accounts to login) only to be sued by microsoft who swings the CFAA at them. so you start off with an intellectual desire to make a nice search engine and end up getting sidetracked into this pit of muck and having microsoft try to put you in jail. and, no, i'm not the one microsoft was suing.
it's both storage and computational. they go hand in hand.
both ddg and brave are bing (microsoft) in disguise.
it's continually spidering. just not at a high rate. actually, back in the day i had real time updates while google was doing the 'google dance'. that caused quite a stir in the web dev community because people could see their pages in the index being updated in real time whereas google took up to 30 days to do it.
I've had extensively dealing with Cloudflare. They have a complex whitelisting system that is difficult to get on, and they also have an 'AI' system that determines if you should be kicked off that whitelist for whatever reason.
Furthermore, they give Google preferred treatment in their UIs and backend algos because it is the incumbent and nobody cares about other smaller search engines. So there's a lot of detail to how they work in this domain.
It's 100% Cloudflare's fault, and it's up to them to give everyone a fair shot. They just don't care. Also, you are overlooking the fact that Google is a major investor (and so is Bing and Baidu). So really this exacerbates the issue. Should Google be allowed (either directly or indirectly) to block competing crawlers from dowloading web pages?
it should be. there should be some sort of 'bots rights' to level the playing field. perhaps this is something the FTC can look into. but, as it is right now big tech continues to keep their iron grip on the web and i don't see that changing any time soon. big tech has all the money and controls access to all the data and supply chains to prevent anyone else from being a competitive threat.
look at linkedin (owned by microsoft unspiderable by all but google/bing). github (now microsoft using this to fuel its AI coding buddy, but if you try to spider this at capacity your IP is banned) facebook (unspiderable) .. the list goes on and on ..
and as you can see, data is required to train advanced AI systems, too. So big tech has the advantage there as well. especially when they can swoop in and corrupt once non-profit companies like openai, and make them [partially] for-profit.
and to rant on (yes, this is what i do :)) it very difficult to buy a computer now. have you tried to buy a raspberry pi or even a jetson nano lately? Who is getting preferred access to the chip factories? Does anyone know? Is big tech getting dibs on all the microchips now too?
and i work with rasengan on private.sh so yes there's some issue there. one of the back end servers is returning a max capacity error of sorts... we are checking into it.
Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive.
I believe my algorithms are decent, but the biggest problem for Gigablast is now the index size. You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough because I don't have the cash for the hardware. btw, I've been working on this engine for over 20 years and have coded probably 1-2M lines of code on it.
A project I'm involved with https://private.sh/ has search privacy that is more verifiable than ddg. It uses client-side cryptography in the javascript, and routes your request through an anonymizing proxy, similar to how TOR works. The encrypted query is only readable by the search engine, in this case, gigablast, whereas the independent proxy scrubs your IP address.
I'm pretty desperate at this point. My current expenses are not very high though and I can keep it running probably as long as I am alive, but I am going crazy trying to figure out how to monetize. Does anyone have any ideas? I will be extremely grateful if you could share. It currently has close to 5B pages indexed. Any advice?
gigablast coder here. i noticed if you search for 'pyopengl glclipplane' it gets it. ('python' does not occur anywhere on the page!) i think this is a synonym issue. 'opengl python' should be synonymous with 'pyopengl'. i think google gets this because they have billions of queries from which to derive synonyms. i'll be adding more synonyms to gigablast over time and hope to close this synonym gap.
that being said, i think 90% of the bad search results come from the index being too small (straight up lacking the best web page), or not having good enough synonyms.
if the doj, etc. breaks google's stranglehold on search ads up into 2 or more independent search ad companies, then other non-big-tech search engines might have a chance at bidding for some of the 'default' provider lists for cell phones, browsers, etc. but as it is right now, google does not allow their search ads to be displayed on other search engines' search results. sucks!