Goto

Collaborating Authors

 Media


What your HUGS reveal about you, according to science

Daily Mail - Science & tech

Bill Clinton's birthday letter to Jeffrey Epstein praising'childlike curiosity' is revealed along with creepy drawing of billionaire pedophile Revealed: Squalid campsite where fugitive New Zealand father's kids were found hiding after four years on the run'She was so f***ed up': Carolyn Bessette's friends tell MAUREEN CALLAHAN of her secret Daddy issue, JFK Jr's murder brag that drove her mad... and why everything we know about her is a lie Infamous'Uranus in retrograde' will turn the lives of three zodiac signs upside-down Greta Thunberg's Gaza flotilla was NOT hit by a drone and their claims'have no basis in truth', Tunisian authorities say - after the group said they were attacked Turn back the clock with the K-beauty retinol cream Amazon shoppers say leaves their skin'silky smooth' - and it's now $10 Idyllic Midwest town is torn apart as mystery tech giant plots $1.6B takeover They were locked in a dungeon inside a house of horrors. But incredible footage shows five kids' daring acts while their parents were out... and it left neighbors speechless This is the extraordinary story of the Kansas City Chiefs' secret weapon Charming town with just 5,000 people named Georgia's prettiest place to live Sharia patrols' spotted in Texas demanding stores stop selling alcohol and pork Glamorous TikToker charged with using medic's identity to carry out cosmetic procedures without a license New photo from Epstein'birthday book' shows joke about Trump'buying girl' after his'lewd birthday message' was revealed Billionaire turns his back on Trump as he blasts President's'risky' financial move that could cost Americans their savings Woman's butt dial voicemail exposes plot to help man dump a dead body '90s TV star cuts a youthful figure at 81 on rare sighting... can you guess who? Whether it's an affectionate cuddle or an awkward squeeze, everyone has their own style of hug. But the way you embrace could reveal parts of your personality, according to a new study. Experts used advanced AI video analysis technology to investigate hugs carried out between friends and romantic partners.


Google's AI Mode to offer Japanese language support

The Japan Times

Google's AI Mode to offer Japanese language support Google has said that its AI Mode will soon be available in Japanese, Korean, Hindi, Indonesian and Brazilian Portuguese globally. Google said Monday its AI-powered search engine AI Mode, which launched in May in only English, is set to be available in Japanese and four other languages as it looks to broaden its global reach. Aside from Japanese, the company said it is set to be available in Korean, Hindi, Indonesian and Brazilian Portuguese globally. At the time of writing, it was still unavailable in Japanese. "Building a truly global Search goes far beyond translation -- it requires a nuanced understanding of local information," Hema Budaraju, vice president of Google Search's product management, wrote in a blog post announcing the news. "With ... our custom version of Gemini 2.5 in Search, we've made huge strides in language understanding, so our most advanced AI search capabilities are locally relevant and useful in each new language we support."


Calibrated Recommendations with Contextual Bandits

arXiv.org Machine Learning

Spotify's Home page features a variety of content types, including music, podcasts, and audiobooks. However, historical data is heavily skewed toward music, making it challenging to deliver a balanced and personalized content mix. Moreover, users' preference towards different content types may vary depending on the time of day, the day of week, or even the device they use. We propose a calibration method that leverages contextual bandits to dynamically learn each user's optimal content type distribution based on their context and preferences. Unlike traditional calibration methods that rely on historical averages, our approach boosts engagement by adapting to how users interests in different content types varies across contexts. Both offline and online results demonstrate improved precision and user engagement with the Spotify Home page, in particular with under-represented content types such as podcasts.


EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

arXiv.org Artificial Intelligence

Large Language Models (LLMs), trained on extensive datasets using advanced deep learning architectures, have demonstrated remarkable performance across a wide range of language tasks, becoming a cornerstone of modern AI technologies. However, ensuring their trustworthiness remains a critical challenge, as reliability is essential not only for accurate performance but also for upholding ethical, cultural, and social values. Careful alignment of training data and culturally grounded evaluation criteria are vital for developing responsible AI systems. In this study, we introduce the EPT (Evaluation of Persian Trustworthiness) metric, a culturally informed benchmark specifically designed to assess the trustworthiness of LLMs across six key aspects: truthfulness, safety, fairness, robustness, privacy, and ethical alignment. We curated a labeled dataset and evaluated the performance of several leading models - including ChatGPT, Claude, DeepSeek, Gemini, Grok, LLaMA, Mistral, and Qwen - using both automated LLM-based and human assessments. Our results reveal significant deficiencies in the safety dimension, underscoring the urgent need for focused attention on this critical aspect of model behavior. Furthermore, our findings offer valuable insights into the alignment of these models with Persian ethical-cultural values and highlight critical gaps and opportunities for advancing trustworthy and culturally responsible AI. The dataset is publicly available at: https://github.com/Rezamirbagheri110/EPT-Benchmark.


Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning

arXiv.org Artificial Intelligence

The rapid growth of visual content consumption across platforms necessitates automated video classification for age-suitability standards like the MPAA rating system (G, PG, PG-13, R). Traditional methods struggle with large labeled data requirements, poor generalization, and inefficient feature learning. To address these challenges, we employ contrastive learning for improved discrimination and adaptability, exploring three frameworks: Instance Discrimination, Contextual Contrastive Learning, and Multi-View Contrastive Learning. Our hybrid architecture integrates an LRCN (CNN+LSTM) backbone with a Bahdanau attention mechanism, achieving state-of-the-art performance in the Contextual Contrastive Learning framework, with 88% accuracy and an F1 score of 0.8815. By combining CNNs for spatial features, LSTMs for temporal modeling, and attention mechanisms for dynamic frame prioritization, the model excels in fine-grained borderline distinctions, such as differentiating PG-13 and R-rated content. We evaluate the model's performance across various contrastive loss functions, including NT-Xent, NT-logistic, and Margin Triplet, demonstrating the robustness of our proposed architecture. To ensure practical application, the model is deployed as a web application for real-time MPAA rating classification, offering an efficient solution for automated content compliance across streaming platforms.


AnalysisGNN: Unified Music Analysis with Graph Neural Networks

arXiv.org Artificial Intelligence

Recent years have seen a boom in computational approaches to music analysis, yet each one is typically tailored to a specific analytical domain. In this work, we introduce AnalysisGNN, a novel graph neural network framework that leverages a data-shuffling strategy with a custom weighted multi-task loss and logit fusion between task-specific classifiers to integrate heterogeneously annotated symbolic datasets for comprehensive score analysis. We further integrate a Non-Chord-Tone prediction module, which identifies and excludes passing and non-functional notes from all tasks, thereby improving the consistency of label signals. Experimental evaluations demonstrate that AnalysisGNN achieves performance comparable to traditional static-dataset approaches, while showing increased resilience to domain shifts and annotation inconsistencies across multiple heterogeneous corpora.


BEAM: Brainwave Empathy Assessment Model for Early Childhood

arXiv.org Artificial Intelligence

Empathy in young children is crucial for their social and emotional development, yet predicting it remains challenging. Traditional methods often only rely on self-reports or observer-based labeling, which are susceptible to bias and fail to objectively capture the process of empathy formation. EEG offers an objective alternative; however, current approaches primarily extract static patterns, neglecting temporal dynamics. To overcome these limitations, we propose a novel deep learning framework, the Brainwave Empathy Assessment Model (BEAM), to predict empathy levels in children aged 4-6 years. BEAM leverages multi-view EEG signals to capture both cognitive and emotional dimensions of empathy. The framework comprises three key components: 1) a LaBraM-based encoder for effective spatio-temporal feature extraction, 2) a feature fusion module to integrate complementary information from multi-view signals, and 3) a contrastive learning module to enhance class separation. Validated on the CBCP dataset, BEAM outperforms state-of-the-art methods across multiple metrics, demonstrating its potential for objective empathy assessment and providing a preliminary insight into early interventions in children's prosocial development.


Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation

arXiv.org Artificial Intelligence

Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. This paper explores a new paradigm: simulating virtual survey respondents using Large Language Models (LLMs). We introduce two novel simulation settings, namely Partial Attribute Simulation (PAS) and Full Attribute Simulation (FAS), to systematically evaluate the ability of LLMs to generate accurate and demographically coherent responses. In PAS, the model predicts missing attributes based on partial respondent profiles, whereas FAS involves generating complete synthetic datasets under both zero-context and context-enhanced conditions. We curate a comprehensive benchmark suite, LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), that spans 11 real-world public datasets across four sociological domains. Our evaluation of multiple mainstream LLMs (GPT-3.5/4 Turbo, LLaMA 3.0/3.1-8B) reveals consistent trends in prediction performance, highlights failure modes, and demonstrates how context and prompt design impact simulation fidelity. This work establishes a rigorous foundation for LLM-driven survey simulations, offering scalable and cost-effective tools for sociological research and policy evaluation. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark


DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

arXiv.org Artificial Intelligence

Abstract--With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short on precisely controlling fine-grained acoustic characteristics of specific sounds. As a result, users that need specific sound content may find it challenging to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTT A). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the customized systems. The experiments show that the proposed model, DreamAudio, generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTT A cases as the benchmark for customized generation tasks. Udio generation, as a crucial technology for enabling artificial intelligence generated content (AIGC) [1], has gained significant interest from the research community.


Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics

arXiv.org Artificial Intelligence

The rise of synthetic media has blurred the boundary between reality and fabrication under the evolving power of artificial intelligence, fueling an infodemic that erodes public trust in cyberspace. For digital imagery, a multitude of editing applications further complicates the forensic analysis, including semantic edits that alter content, photometric adjustments that recalibrate colour characteristics, and geometric projections that reshape viewpoints. Collectively, these transformations manipulate and control perceptual interpretation of digital imagery. This susceptibility calls for forensic enquiry into reconstructing the chain of events, thereby revealing deeper evidential insight into the presence or absence of criminal intent. This study seeks to address an inverse problem of tracing the underlying generation chain that gives rise to the observed synthetic media. A tell-tale watermarking system is developed for explanatory reasoning over the nature and extent of transformations across the lifecycle of synthetic media. Tell-tale watermarks are tailored to different classes of transformations, responding in a manner that is neither strictly robust nor fragile but instead interpretable. These watermarks function as reference clues that evolve under the same transformation dynamics as the carrier media, leaving interpretable traces when subjected to transformations. Explanatory reasoning is then performed to infer the most plausible account across the combinatorial parameter space of composite transformations. Experimental evaluations demonstrate the validity of tell-tale watermarking with respect to fidelity, synchronicity and traceability.