Kagi employee here. The search business is sustainable, we don't need a push in AI things to make it work. We're looking at these features/ideas because we think they complement search well, not because we need them from a cashflow perspective :)
HN user
TisButMe
[ my public key: https://keybase.io/tisbutme; my proof: https://keybase.io/tisbutme/sigs/74z5cJ6u55lz0fHQ9jGzJNyxRDYR2CtRjGiRsUdj86Y ]
Heya, I work at Kagi, and I did the math for this. Our unit economics are sound, and we don't plan on subsiding usage with ties-attached money. We have abuse prevention mechanisms, and we regularly review heavy use use-cases so we can either optimize for them or offer alternative workflows.
Heya, I work at Kagi. This sounds like a great idea, thanks for the input! We'll put it on the requested features list (which our users contribute to).
If you think we should word the privacy policy differently, please do submit some feedback on kagifeedback.org with the specifics - we have changed it in the past through exactly this process. It has been written by engineers, mostly for engineers at the beginning. I'm sure we can improve the wording to make it more binding, we're not trying to squirrel away from it. If you have enough legal knowledge to harden it, we'd welcome the contribution :)
https://kagi.com/privacy - it's not written in legalese.
I work at Kagi. We don't have KPIs that track search quality in that way (we don't have any frontend tracking at all, so we can't know if you click any of the results), and we haven't touched much the sources of data we're using over the last few months. We've also had conflicting reports about this problem, so I'm wondering if it's not variable quality over time and place of our upstreams that's a problem.
That said, we're aware that it's currently something we have to improve, "!g best cafe" is not where we want to be. We're working on it, but if you have specific examples/suggestions, please do submit them to kagifeedback.org so we can track them.
I work at Kagi. I think this is partly a function of who our early user base was, and the tons of feedback they gave which helped improve our search. Because we don't keep the data, we can't do a lot of the ML-powered search improvements that eg. Google could over all search domains.
The good news is as the community expand, so will the feedback, and we'll be able to improve overall. The even better (and quite surprising) news is that we could be better than eg. Google on any topic at all, which tells me that the ML magic is not actually magic, and we will be able to outperform them in other domains as the community grows, the feedback increases, and our code becomes better.
I work at Kagi. I'm as blown away by the recent coverage as you are. We're a 15ish people company, which is staying away from VC money and only raised a relatively small amount of money directly from our users. We don't have the time nor the money to spend on astroturfing. We do have an incredibly supportive community, and I'm sure they help in spreading the word. I think it also help that we have a good product :)
In fact, we even axed our referral bonus program a couple months back to ensure that no third party had anything to gain in promoting Kagi. Recently, that included saying "no" to an independent journalist who explicitly asked for referral bonus for their readers. All you see is organic.
I work at Kagi. This is a very reasonable fear to have. We do actually not store the data, but of course you'd need to take me at my word for this.
That said, if we did lie or change that, we'd be in immediate breach of our privacy policy (https://kagi.com/privacy), and as a result be a very easy target for a lawsuit. Given that we're intentionally not VC backed, between the horrible press this would be and the actual costs of fighting such a lawsuit, I expect not much would be left of Kagi afterwards. We are liable to users, in a pretty existential way.
Heya, I work at Kagi. This is correct, we do not personalize searches other than by respecting the user's customizations (eg. domain preferences, lenses, etc...) which are all entirely user-controlled.
Hey daveoc64. I work at Kagi. Really cool stuff is coming for Ultimate, and being a early Ultimate user you'll get it before the newer users. If that doesn't work for you (fair enough) and you switch to pro, we'll prorate all your credits. If that also doesn't work, contact support@kagi.com and we'll help.
We do - we have lenses which basically search subsets of the web (and you can make your own). Academic is a default lens :)
I do, we don't log searches.
We do read the comments, but as far as I'm concerned this is a feature, not a bug :)
Possibly yes - I think that's my point with predicting peak oil wrong for 50 years. Still, right now it seems every time OpenAI/someone else adds a new content filter, someone figures out a prompt escape that works.
I'm not too worried about GPTs trained on GPTs, maybe that's an LLM analogy to AlphaGo playing itself a lot to learn how to play go. I'm more worried about people specifically trying to get into the training corpus with biased/wrong/misleading/security-risk content.
I now did in the parent comment :P
(author here) How do you know what's a prompt injection vs actual content? If you train another LLM to tell you what's a prompt injection, how do you know it has 100% coverage of all possible injections? OpenAI has been battling people trying to bypass their prompt re-write filter, and as far as I can see, not really winning, just constantly adding stuff to their blocklist until the next thing gets discovered.
Agreed, and I mentioned that solution in the article, but I'm not so convinced this is true. It reads a bit like the "if you're a great programmer, the lack of memory safety of C isn't a problem!" argument. In theory sure, but in practice it seems CVEs keep on popping up.
(Author here) that's what I thought originally, but then it means that LLMs never get to learn from new content - current ones stop in 2021, they don't know that Russia invades Ukraine, or that Arc is a cool browser or the API of any libraries released after their end date (which has been an issue for me for code generation using fast moving libraries). I don't think it's good enough to stop acquiring new content.
Maybe, but your interface is smooth and slick! It was a weekend project for me, so didn't have time to refine it too much, and it does what I use it to do, but it definitely looks like I could learn a lot in design from you :D
Looks like we had a similar idea :D http://bandmap.me/
Where are you in France?
I've got unlimited call + sms + 4G data (limited bandwidth after 20Gb) for 20€/month, and 300Mb/s fibre connection for 35€/month. I have not seen similar rates in the UK/US.
But those people will always be a net negative for society, basic income or not. Empowering those who would make a positive difference if they had enough time/money sounds like a good idea to me.
I think we'll have to agree to disagree on most of those points then.
I do not think there are trivial/uninteresting questions. You have to prioritise, but you can't just sweep stuff under the rug and call it a day. I'm not even using the "it might be huge!" argument, just that science is about curiosity. Most math won't end up as useful as cryptography, but it doesn't matter.
I do think that it is part of your job, as a scientist, to document what you do, and what you observe. If a software engineer on my team didn't document his code/methodology correctly, he'd be reprimanded, for good reason. Yeah, it takes time, but it's part of the job. This way, we avoid having 4 people independently rediscovering how to set up the build tools.
So you just accept for no reason that tap water is bad somehow, and discard the result you've just gotten?
I do understand that you have a limited amount of time, and can't just go after everything, but when something happens in science, it needs to be documented. Yeah, maybe someone else should investigate, but someone should. Maybe that particular phenomena that lead to the water influencing your result will give you knowledge about cell metabolism. Who knows? If it has that much of an effect on cell growth that you need to deal with it, it's already more active than a lot of compounds we try out, anyway...
To go back to the computer analogy, it feels like my program is bugged, and to debug it I'm changing variables names (which as far as I know shouldn't matter), and then the code magically works again. Sure, some days, I'll go "Ok, compiler magic, got it", but most days I'd be pretty intrigued, and I'd look into it, because yeah, I might just have found a GCC bug.
I agree, no one cares, but I did. I don't know what I don't know yet, and I don't want to presume anything. The tap water thing might actually lead us to solid models which would explain why tap water breaks the experiment. That's why I really think we should start a movement of publishing everything, and trying to deal with simpler models/systems we do understand before going up to models with so many unknowns that the results are basically a dice roll.
But wet lab scientists should be even more careful! They have way less control over the system they're trying to study than you do, so stats are the only security net we have to even attempt to do anything with the data we produce.
I also agree on the innocent until proven guilty part, but by now I've seen and talked to hundreds of people with the best intentions, who do not realise how important careful examination of the data is, so I'm growing a bit disillusioned.
I've been lucky enough to work in outstanding labs, with people published in Nature, and other journals of that quality. I've worked in 2 countries, and for 4 different labs. I've also talked with people from all over the world, who have worked everywhere, from Harvard to the Pasteur Institute to Cambridge University. The stories are all the same. I hoped I would find some place where people were trying to do things the right way, but what I found is that currently, you don't need to to be published in top journals, so why bother?
It's really refreshing to hear you talk about trying to troubleshoot why an experiment didn't work the way you expected, I hear mostly of people retrying blindly until it "succeeds". What did you do with what you learned with the water causing the failure? Did you publish this, so that someone (or you!) could try to figure out why water was a problem, or at least so that no one would have the same issue? This is the other point: when people do bother about finding about why things fail, I've never seen any of them try to follow up on that, and figure out not only what made it fail, but why it made it fail. "Yeah, the annealing temp was not the right one". Ok, but why?
Of course playing with systems we don't understand is the point, but we have to be very careful about them. We should be varying 1 parameter at a time. This is mostly impossible in biology, but right now we're not even trying to do anything about it.
I hoped that it wasn't as bad in computational biology, or ecology, or any other biology field where systems and models are actually defined. It saddens me to read that your experience was as bad as mine...
I absolutely agree that sometimes, you need to redo an experiment for good reasons.
In most cases I've seen, people do not know why they redo the experiment, though. They know it hasn't produced the data they expected, so they redo it. Maybe it was because a reagent was bad, or a co-worker left the incubator open overnight, or maybe it was because the model is stupid. Who knows?
That's my point, actually. Biologists are playing with systems they do not understand, changing parameters somewhat randomly without any control over them, and they then try to interpret whatever comes out, but ONLY it fits what they wanted. If it doesn't, then "Oh, the PCR machine is at it again!", and they throw the results away.