Deduplicating Massive Datasets with Locality Sensitive Hashing
This article is part of a media partnership with PyData Berlin, a group helping support open-source data science libraries and tools. To learn more about this topic, please consider attending our fourth annual PyData Berlin conference on June 30-July 2, 2017. Matti Lyra and other experts will be giving talks on Natural Language Processing, Machine Learning, AI Ethics and many related data fields. You can also find a more detailed blog post on Matti Lyra's personal blog at https://mattilyra.github.io Many online platforms that deal with natural language documents face a big problem: thousands of duplicate documents.
Sep-21-2017, 18:31:20 GMT
- Technology: