HN user

lookforr

1 karma
Posts0
Comments3
View on HN
No posts found.

this doesn't work in the case where there are two (and only two) similar documents get ingested into the system as new singleton clusters at the same time; this case is very rare so it is not a big issue to you, i guess.

nice article. a few questions here

1. scalability: does your system ingest multiple documents in parallel? if so, how often do you observe over-segmentation, if any?

2. thresholds: how did you set the thresholds at various parts of the systems?