It is a fine thing when a man who thoroughly understands a subject is unwilling to open his mouth, and only speaks when he is questioned.
Yoshida Kenko, Essays in Idleness
HN user
It is a fine thing when a man who thoroughly understands a subject is unwilling to open his mouth, and only speaks when he is questioned.
Yoshida Kenko, Essays in Idleness
been a lot of these RAG abstractions posted recently. As someone working on this problem, it's unclear to me whether the calculation and ingestion of embeddings from source data should be abstracted into the same software package as their search and retrieval. I guess it probably depends on the complexity of the problem. This does seem interesting in that it does make intuitive sense to have a built-in db extension if the source data itself is coming from the same place as the embeddings are going. But so far I have preferred a separation of concerns in this respect, as it seems that in some cases the models will be used to compute embeddings outside the db context (for example, the user search query needs to get vectorized. why not have the frontend and the backend query the same embedding service?) Anyone else have thoughts on this?
There's an issue in the pgvector repo about someone having several ~10-20million row tables and getting acceptable performance with the right hardware and some performance tuning: https://github.com/pgvector/pgvector/issues/455
I'm in the early stages of evaluating pgvector myself. but having used pinecone I currently am liking pgvector better because of it being open source. The indexing algorithm is clear, one can understand and modify the parameters. Furthermore the database is postgresql, not a proprietary document store. When the other data in the problem is stored relationally, it is very convenient to have the vectors stored like this as well. And postgresql has good observability and metrics. I think when it comes to flexibility for specialized applications, pgvector seems like the clear winner. But I can definitely see pinecone's appeal if vector search is not a core component of the problem/business, as it is very easy to use and scales very easily
After seeing raw source text performance, I agree that representational learning of higher-level semantic "context clusters" as you say seems like an interesting direction.
For those who don't know: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
I agree with you for the NLP domain, but I wonder if there will also be a bitter lesson learned about the perceived generality of language for universal applications.
i don't disagree with the premise that Google should be responsible and explicitly acknowledge that the average computer-interested person trying out bigquery has no clue how sharp of a knife it is and they actually do need to be protected from themselves. I was in this boat only a few months ago. One thing I will say though is that I think the documentation is actually quite comprehensive, and personally after taking the time to RTFM and actually understand things like columnar storage, partitioned and clustered tables, etc., I was able to optimize costs quite a bit for our use case and am quite pleased with the product overall. Just takes time to learn, it's a (necessarily imo) intricate machine.
to store and retrieve
retrieve != query
this project is extremely simplistic in regards to its vector search tech. pgvector is an open source implementation of an _index_ (multiple algos actually), this uses Cloudflare's completely proprietary index with a single call.
I am not sure what you mean specifically by 'overlapping'. But high-dimensional vector space is really "big" in the sense that everything is way closer together compared to low dimensions (this is the curse of dimensionality for euclidean norm), and this is already something one has to think about regardless of the similarity of the source documents. From reading wikipedia it seems like it's been argued that the curse is the worst with independent and uniformly distributed features.
i think the confusion comes from the mixup between the words "database", "store", and "index". Vector "store" is trivial, even for hundreds of millions of vectors you are still in the realm of what is possible on a single disk. Vector "index" to enable efficient aNN is not trivial for large numbers of high-dimensional vectors, and this is usually the proposed value add of someone providing a vector "database", which combines the two. I think this is also how the words are understood more generally. This project is a wrapper over Cloudflare's infrastructure, which does provide a vector index, though it is not clear how well their index performs in real-world use cases.
Thank you for the response, it is helpful. I am not used to Python and know that I am not using the language well. So I think it is worth focusing on this.
not to make huge (also architectural) mistakes
I'm noticing it's the first time I'm even having to make significant architectural decisions, which is difficult because I don't have much experience to draw from, so even the smallest decisions often require a lot of research.
Young engineer here in charge of a project and feeling quite out of their depth, I agree with this. Currently no mentorship and it will take a couple months for a senior hire. Do you have any advice? What does a senior engineer love/hate to see when they come onto a project started by engineers earlier in their career? How can I be most helpful?
I agree, vector search on music seems pretty obvious at this point.
the doctrine is that only things that are 'novel' and 'non-obvious' (to a 'person having ordinary skill in the art') can be patented, and any accessible published material (among other things) is considered 'prior art' for determining this. So unless your 'one small change' fundamentally alters the behavior of the 'black box' in a novel way, one would probably say that it is not novel and pretty obvious for a person having ordinary skill in the art to emulate a slightly modified version the hardware in software. of course, vector search on music also seems pretty obvious to me. I don't actually know how this plays into infringement specifically rather than trying to patent something though.
this combo technique sounds useful, is it something like reviewing memory palaces with spaced repetition, instead of reviewing standalone information?
i don't like _auto_complete, i find that the nondeterminstic text changes on the screen are pretty distracting. however, I do really like quick lookup of relevant names at my own will from the editor. this is possible with even e.g. ed and ctags.
i'd wager nearly every development environment is integrated to some extent, piping output in the unix shell is integrating.
I don't understand your simile. The limitations of a bicycle are infrastructure: bicycles can go a lot of places very efficiently, but rugged mountain terrain without trails (precisely the place condors thrive) would be difficult to traverse. (this is totally off topic but now i am curious about the energy efficiency of condor flight vs bicycle travel in ideal conditions)
i find SRS very useful. but i personally think it is not well suited to reference information. i'm skeptical of putting all my appointments in SRS for example (maybe birthdays would be worth it). or say I had collected a bunch of papers related to a topic I was very interested in. I don't necessarily think I'd want to memorize the list of them via SRS, but having the titles written somewhere for reference would be great. I think the point of this paper is that if that 'somewhere' is an SRS card, it is almost completely devoid of context (other than 'studying flashcards on computer' context), but if that somewhere is a notebook that contains lists of papers related to all the topics i'm interested in, it's much easier to find. (though computers are good at searching fast)
Me neither. Memory palaces also show that this effect can be emulated (not sure if that's the right word) in the mind. I wonder if there is a way for computers to encourage developing a 'mind map'
hi, as a meditator and epileptic my curiosity is piqued. When you say 'this link', did you mean OP or is there a link you meant to add to your comment? if the latter I would be keen to look at it.
maybe neurofeedback? https://web.archive.org/web/20110710194522/www.epilepsyhealt...
wikipedia claims there is not much recent research when it references that link.
it is being shown that meditation can affect brain waves, I wonder if it is in a way that is beneficial for preventing seizures.
i have the same troubles in social situations with too many inputs. i am quite frankly asocial because of it, though i don't necessarily want to be. do you mind if i ask what your reaction feels like? my reaction feels a lot like dissociation, but it also could be absence seizure. it is hard to introspect in those moments.
not parent but as someone with temporal lobe epilepsy, in addition to what parent describes, difficulty recalling words is definitely something i experience. Probably the most annoying thing that i have had to learn to deal with is people trying to fill in or guess what word i am trying to say while i pause. The hippocampus is located in the temporal lobe and is heavily involved in the processing and storage of short-term memory as well as retrieval of long-term memory, so it makes sense that if the hippocampus is a bit fried from seizures it would manifest as memory-related cognitive issues.
i recently had my first generalized seizure and once i knew and talked with doctors who confirmed it as primarily TLE i realized i had been having focal seizures for years.
i think i may have a similar feeling in long meetings/conversations too. it's like i am forced to stop paying attention for several seconds. i can generally even notice that i am not paying attention, but I am unable to focus even if i try. i think i've noticed it tends to be worse if there are lots of people talking. it feels a lot like dissociation. perhaps it is an 'aura', some sort of absence seizure...
Primarily medication. Fasting has a long (ancient) history as a treatment for epilepsy, the modern revival and investigation of the ketogenic diet specifically relates to certain types drug-resistant childhood epilepsy. The Wikipedia article on the ketogenic diet is quite informative.