Ask HN: Algorithms for text fingerprinting?
https://news.ycombinator.com/item?id=9716837I remember reading an article a year or so ago about (the NSA) identifying users based on how they write: vocabulary, spelling mistakes, grammar, dialect, and so on.
This is interesting to me because it is extremely difficult to change the vocabulary I use in writing and speaking. Being able to estimate the amount of similarity between two pieces of text would be useful.
The closest I can think of right now would be the proprietary algorithms used to check for plagiarism (for schools and universities, for instance).
Are there any publicly available algorithms for this? Where can I go to learn more? (Academic journals?) Am I just DDGing the wrong search terms?