Self-supervised Document Clustering Based on BERT with Data Augment
–arXiv.org Artificial Intelligence
Our contributions (NLP) with less supervision or without supervision are as follows to illustrate our explorations in has been a continuously attractive problem. The unsupervised how to improve clustering accuracy: approaches can be roughly classified as - We find that to additionally introduce generative and discriminative (Chen et al., 2020). A UDA (Xie et al., 2019) in CL can further improve generative approach may use an encoding network clustering accuracy; to learn latent representations for input texts, and feeds the latent representations into another generating - Our experimental results suggest that PCL network. Given an objective of generation can achieve almost the same performance similarity, the most proper latent representations with supervised learning using partial items in can be learned. The generative approaches always dataset; introduce novel neural architectures, for example - We also find that using multi-language back Chiu et al. (Chiu et al., 2020) compared a series translation to generate positive sample in SCL of graph neural network in document clustering can further eliminate the contrast processing accuraccy of 20newsgroup (Lang, 1995). On the in PCL with some performance deteriorations; contrary, the models that are commonly seen in supervised learning are preferred to be reused for a discriminative method.
arXiv.org Artificial Intelligence
Nov-26-2020