Robust Document Representations using Latent Topics and Metadata

Raman, Natraj, Nourbakhsh, Armineh, Shah, Sameena, Veloso, Manuela

arXiv.org Artificial Intelligence 

As an example, datasets are very often accompanied by metadata tags that include some signal about the nature and content Task specific fine-tuning of a pre-trained neural language model of each document in the corpus. Corporate communications, financial using a custom softmax output layer is the de facto approach of late reports, regulatory disclosures, policy guidelines, Wikipedia when dealing with document classification problems. This technique entries, social media messages, and many other forms of textual is not adequate when labeled examples are not available at records often bear tags indicating their source, type, purpose, or an training time and when the metadata artifacts in a document must enterprise categorization standard. This metadata is often stripped be exploited. We address these challenges by generating document before models are applied to the text, in order to avoid convoluted representations that capture both text and metadata artifacts in a and bespoke architectures. In some cases, the metadata is simply task agnostic manner. Instead of traditional auto-regressive or autoencoding concatenated to the document [27] without specific controls on based training, our novel self-supervised approach learns how representations are generated from raw text versus metadata a soft-partition of the input space when generating text embeddings.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found