GPU retrieval can rival HNSW/IVFPQ on both latency and recall. LinkedIn's LiNR (Feed OON recsys) and SJS and Meta's MoL already deploy exhaustive k-NN at production scale on A100/H100 nodes. See my full post for the technical details
nzhiltsov
Posts28
Comments11