Goto

Collaborating Authors

 etm




BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study

arXiv.org Artificial Intelligence

As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modeling Hindi short texts, an area that has been under-explored in existing research. Using contextual embeddings, BERTopic can capture semantic relationships in data, making it potentially more effective than traditional models, especially for short and diverse texts. We evaluate BERTopic using 6 different document embedding models and compare its performance against 8 established topic modeling techniques, such as Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), Latent Semantic Indexing (LSI), Additive Regularization of Topic Models (ARTM), Probabilistic Latent Semantic Analysis (PLSA), Embedded Topic Model (ETM), Combined Topic Model (CTM), and Top2Vec. The models are assessed using coherence scores across a range of topic counts. Our results reveal that BERTopic consistently outperforms other models in capturing coherent topics from short Hindi texts.


Embedded Topic Models Enhanced by Wikification

arXiv.org Artificial Intelligence

Topic modeling analyzes a collection of documents to learn meaningful patterns of words. However, previous topic models consider only the spelling of words and do not take into consideration the homography of words. In this study, we incorporate the Wikipedia knowledge into a neural topic model to make it aware of named entities. We evaluate our method on two datasets, 1) news articles of \textit{New York Times} and 2) the AIDA-CoNLL dataset. Our experiments show that our method improves the performance of neural topic models in generalizability. Moreover, we analyze frequent terms in each topic and the temporal dependencies between topics to demonstrate that our entity-aware topic models can capture the time-series development of topics well.


A modified model for topic detection from a corpus and a new metric evaluating the understandability of topics

arXiv.org Artificial Intelligence

This paper presents a modified neural model for topic detection from a corpus and proposes a new metric to evaluate the detected topics. The new model builds upon the embedded topic model incorporating some modifications such as document clustering. Numerical experiments suggest that the new model performs favourably regardless of the document's length. The new metric, which can be computed more efficiently than widely-used metrics such as topic coherence, provides variable information regarding the understandability of the detected topics.


A Joint Learning Approach for Semi-supervised Neural Topic Modeling

arXiv.org Machine Learning

Topic models are some of the most popular ways to represent textual data in an interpret-able manner. Recently, advances in deep generative models, specifically auto-encoding variational Bayes (AEVB), have led to the introduction of unsupervised neural topic models, which leverage deep generative models as opposed to traditional statistics-based topic models. We extend upon these neural topic models by introducing the Label-Indexed Neural Topic Model (LI-NTM), which is, to the extent of our knowledge, the first effective upstream semi-supervised neural topic model. We find that LI-NTM outperforms existing neural topic models in document reconstruction benchmarks, with the most notable results in low labeled data regimes and for data-sets with informative labels; furthermore, our jointly learned classifier outperforms baseline classifiers in ablation studies.


Exclusive Topic Modeling

arXiv.org Machine Learning

Two wellknown challenges in topic modeling are: 1)the predominance of the frequently appearing words in the estimated topics; 2) topics are overlapped with common words, making the structure and interpretation difficult. We propose an Exclusive Topic Model (ETM) to tackle these two issues. ETM can identify field-specific keywords and deliver well-structured topics with exclusive words. More specifically, a weighted Lasso penalty is imposed to reduce the predominance of the frequently appearing yet less relevant words automatically and a pairwise Kullback-Leibler divergence penalty is used to implement topics separation. Topic modeling makes use of the word co-occurrence information to estimate topics. Due to the human language habit and structure, certain words appear more frequently than others, e.g. the Zipf's law. Consequently, the frequently appearing words co-occur with more words and thus are predominant in the estimated topics. The phenomenon makes topic interpretation difficult, as general and frequently appearing words take the place of the true exclusive topic words. The semantic coherence of the estimated topics also deteriorates.


Topic Modeling in Embedding Spaces

arXiv.org Machine Learning

Topic modeling analyzes documents to learn meaningful patterns of words. However, existing topic models fail to learn interpretable topics when working with large and heavy-tailed vocabularies. To this end, we develop the Embedded Topic Model (ETM), a generative model of documents that marries traditional topic models with word embeddings. In particular, it models each word with a categorical distribution whose natural parameter is the inner product between a word embedding and an embedding of its assigned topic. To fit the ETM, we develop an efficient amortized variational inference algorithm. The ETM discovers interpretable topics even with large vocabularies that include rare words and stop words. It outperforms existing document models, such as latent Dirichlet allocation (LDA), in terms of both topic quality and predictive performance.


Project of the Year ZDNet

AITopics Original Links

The business challenge was, therefore, to optimize the utilization of MTR's limited resources--people, tools, workspace and time (four non-traffic hours every day)--and yet be able to comply with the statutory and safety regulations. In 2005, MTR embarked on a project called the Engineering Works & Traffic Information Management System (ETMS) which uses artificial intelligence (AI) for planning, scheduling and managing engineering works. The business challenge was, therefore, to optimize the utilization of MTR's limited resources--people, tools, workspace and time (four non-traffic hours every day)--and yet be able to comply with the statutory and safety regulations. In 2005, MTR embarked on a project called the Engineering Works & Traffic Information Management System (ETMS) which uses artificial intelligence (AI) for planning, scheduling and managing engineering works. The ETMS helps MTR to efficiently plan and execute preventive and corrective engineering works during the limited time available in the non-traffic hours.