media analysis
UPLME: Uncertainty-Aware Probabilistic Language Modelling for Robust Empathy Regression
Hasan, Md Rakibul, Hossain, Md Zakir, Krishna, Aneesh, Rahman, Shafin, Gedeon, Tom
Abstract--Noisy self-reported empathy scores challenge supervised learning for empathy regression. While many algorithms have been proposed for learning with noisy labels in textual classification problems, the regression counterpart is relatively under-explored. We propose UPLME, an uncertainty-aware probabilistic language modelling framework to capture label noise in empathy regression tasks. One of the novelties in UPLME is a probabilistic language model that predicts both empathy scores and heteroscedastic uncertainty, and is trained using Bayesian concepts with variational model ensembling. We further introduce two novel loss components: one penalises degenerate Uncertainty Quantification (UQ), and another enforces similarity between the input pairs on which empathy is being predicted. UPLME achieves state-of-the-art performance (Pearson Correlation Coefficient: 0.558 0.580 and 0.629 0.634) in terms of the performance reported in the literature on two public benchmarks with label noise. Through synthetic label noise injection, we demonstrate that UPLME is effective in distinguishing between noisy and clean samples based on the predicted uncertainty. UPLME further outperform (Calibration error: 0.571 0.376) a recent variational model ensembling-based UQ method designed for regression problems.
Improving Narrative Classification and Explanation via Fine Tuned Language Models
Tyagi, Rishit, Bouri, Rahul, Gupta, Mohit
Understanding covert narratives and implicit messaging is essential for analyzing bias and sentiment. Traditional NLP methods struggle with detecting subtle phrasing and hidden agendas. This study tackles two key challenges: (1) multi-label classification of narratives and sub-narratives in news articles, and (2) generating concise, evidence-based explanations for dominant narratives. We fine-tune a BERT model with a recall-oriented approach for comprehensive narrative detection, refining predictions using a GPT-4o pipeline for consistency. For narrative explanation, we propose a ReACT (Reasoning + Acting) framework with semantic retrieval-based few-shot prompting, ensuring grounded and relevant justifications. To enhance factual accuracy and reduce hallucinations, we incorporate a structured taxonomy table as an auxiliary knowledge base. Our results show that integrating auxiliary knowledge in prompts improves classification accuracy and justification reliability, with applications in media analysis, education, and intelligence gathering.
What Real Deep Learning Applied to Social Media Tells Us About the Crypto Market
Social and news media plays a relevant role in the dissemination of information related to crypto-assets. In a nascent financial market without established disclosure mechanisms, a lot of the relevant events about crypto-assets are distributed first in news and social media channel and, not surprisingly, the market remains incredibly susceptible to those channels. The result is an ecosystem in which social and news media becomes a first-class source of intelligence about the behavior of crypto-assets. Unfortunately, most of the techniques used to analyze social and news media fees for crypto-assets remain incredibly simplistic producing ineffective and often misleading results. In the last few months, our team at IntoTheBlock started different research efforts focused on producing a more sophisticated analysis of social and news media for crypto-assets.