Goto

Collaborating Authors

 Media


Towards Film-Making Production Dialogue, Narration, Monologue Adaptive Moving Dubbing Benchmarks

arXiv.org Artificial Intelligence

Movie dubbing has advanced significantly, yet assessing the real-world effectiveness of these models remains challenging. A comprehensive evaluation benchmark is crucial for two key reasons: 1) Existing metrics fail to fully capture the complexities of dialogue, narration, monologue, and actor adaptability in movie dubbing. 2) A practical evaluation system should offer valuable insights to improve movie dubbing quality and advancement in film production. To this end, we introduce Talking Adaptive Dubbing Benchmarks (TA-Dubbing), designed to improve film production by adapting to dialogue, narration, monologue, and actors in movie dubbing. TA-Dubbing offers several key advantages: 1) Comprehensive Dimensions: TA-Dubbing covers a variety of dimensions of movie dubbing, incorporating metric evaluations for both movie understanding and speech generation. 2) Versatile Benchmarking: TA-Dubbing is designed to evaluate state-of-the-art movie dubbing models and advanced multi-modal large language models. 3) Full Open-Sourcing: We fully open-source TA-Dubbing at https://github.com/woka- 0a/DeepDubber- V1 including all video suits, evaluation methods, annotations. We also continuously integrate new movie dubbing models into the TA-Dubbing leaderboard at https://github.com/woka- 0a/DeepDubber-V1 to drive forward the field of movie dubbing.


Mapping the Italian Telegram Ecosystem: Communities, Toxicity, and Hate Speech

arXiv.org Artificial Intelligence

Telegram has become a major space for political discourse and alternative media. However, its lack of moderation allows misinformation, extremism, and toxicity to spread. While prior research focused on these particular phenomena or topics, these have mostly been examined separately, and a broader understanding of the Telegram ecosystem is still missing. In this work, we fill this gap by conducting a large-scale analysis of the Italian Telegram sphere, leveraging a dataset of 186 million messages from 13,151 chats collected in 2023. Using network analysis, Large Language Models, and toxicity detection tools, we examine how different thematic communities form, align ideologically, and engage in harmful discourse within the Italian cultural context. Results show strong thematic and ideological homophily. We also identify mixed ideological communities where far-left and far-right rhetoric coexist on particular geopolitical issues. Beyond political analysis, we find that toxicity, rather than being isolated in a few extreme chats, appears widely normalized within highly toxic communities. Moreover, we find that Italian discourse primarily targets Black people, Jews, and gay individuals independently of the topic. Finally, we uncover common trend of intra-national hostility, where Italians often attack other Italians, reflecting regional and intra-regional cultural conflicts that can be traced back to old historical divisions. This study provides the first large-scale mapping of the Italian Telegram ecosystem, offering insights into ideological interactions, toxicity, and identity-targets of hate and contributing to research on online toxicity across different cultural and linguistic contexts on Telegram.


GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

arXiv.org Artificial Intelligence

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality. Baseline systems are provided for popular speech recognition toolkits, namely Athena, ESPnet, Kaldi and Pika.


I visited Apple's secret testing labs - here's what REALLY happens behind-the-scenes at the Cork campus

Daily Mail - Science & tech

Apple is best known for its futuristic, spaceship-like headquarters in Cupertino, California. But what many people don't know is that the tech giant also has a huge campus in Ireland. Apple's Cork campus opened its doors in 1980 with a single manufacturing facility and just 60 employees. Fast-forward to today, the site is home to more than 6,000 employees, and serves as Apple's European headquarters. The tech giant is usually extremely private about what happens behind closed doors.


EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models

arXiv.org Artificial Intelligence

As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given the growing variety of explainability methods and diverse stakeholder requirements, frameworks that help stakeholders select appropriate explanations tailored to their specific use cases are increasingly important. To address this need, we introduce EvalxNLP, a Python framework for benchmarking state-of-the-art feature attribution methods for transformer-based NLP models. EvalxNLP integrates eight widely recognized explainability techniques from the Explainable AI (XAI) literature, enabling users to generate and evaluate explanations based on key properties such as faithfulness, plausibility, and complexity. Our framework also provides interactive, LLM-based textual explanations, facilitating user understanding of the generated explanations and evaluation outcomes. Human evaluation results indicate high user satisfaction with EvalxNLP, suggesting it is a promising framework for benchmarking explanation methods across diverse user groups. By offering a user-friendly and extensible platform, EvalxNLP aims at democratizing explainability tools and supporting the systematic comparison and advancement of XAI techniques in NLP.


Clustering Internet Memes Through Template Matching and Multi-Dimensional Similarity

arXiv.org Artificial Intelligence

Meme clustering is critical for toxicity detection, virality modeling, and typing, but it has received little attention in previous research. Clustering similar Internet memes is challenging due to their multimodality, cultural context, and adaptability. Existing approaches rely on databases, overlook semantics, and struggle to handle diverse dimensions of similarity. This paper introduces a novel method that uses template-based matching with multi-dimensional similarity features, thus eliminating the need for predefined databases and supporting adaptive matching. Memes are clustered using local and global features across similarity categories such as form, visual content, text, and identity. Our combined approach outperforms existing clustering methods, producing more consistent and coherent clusters, while similarity-based feature sets enable adaptability and align with human intuition. We make all supporting code publicly available to support subsequent research.


Towards Explainable Temporal User Profiling with LLMs

arXiv.org Artificial Intelligence

Accurately modeling user preferences is vital not only for improving recommendation performance but also for enhancing transparency in recommender systems. Conventional user profiling methods, such as averaging item embeddings, often overlook the evolving, nuanced nature of user interests, particularly the interplay between short-term and long-term preferences. In this work, we leverage large language models (LLMs) to generate natural language summaries of users' interaction histories, distinguishing recent behaviors from more persistent tendencies. Our framework not only models temporal user preferences but also produces natural language profiles that can be used to explain recommendations in an interpretable manner. These textual profiles are encoded via a pre-trained model, and an attention mechanism dynamically fuses the short-term and long-term embeddings into a comprehensive user representation. Beyond boosting recommendation accuracy over multiple baselines, our approach naturally supports explainability: the interpretable text summaries and attention weights can be exposed to end users, offering insights into why specific items are suggested. Experiments on real-world datasets underscore both the performance gains and the promise of generating clearer, more transparent justifications for content-based recommendations.


White House celebrates 'Star Wars Day' with AI image of muscular Trump wielding a lightsaber

FOX News

Charles McBee stops by Fox News Saturday Night With Jimmy Failla to give his take on actor John Boyega calling out the "Star Wars" franchise for its overwhelming whiteness. The White House slammed the "radical left" in a social media post Sunday, showing an AI-generated image of President Donald Trump wielding a lightsaber in celebration of May the Fourth, or "Star Wars Day." May 4 has long been regarded as a day to celebrate the iconic movie franchise as fans post on social media "May the Fourth be with you," an offshoot of the memorable Star Wars quote "May the force be with you." On Sunday, the White House took an opportunity to celebrate the popular day with a post on X, while also taking digs at the Trump administration's biggest critics. "Happy May the 4th to all, including the Radical Left Lunatics who are fighting so hard to bring Sith Lords, Murderers, Drug Lords, Dangerous Prisoners, & well known MS-13 Gang Members, back into our Galaxy. You're not the Rebellion--you're the Empire," the White House wrote.


When Star Wars becomes REALITY: Scientists reveal how you really could be frozen in 'carbonite' like Han Solo

Daily Mail - Science & tech

In George Lucas's classic 1980 film'The Empire Strikes Back', hero Han Solo (Harrison Ford) is frozen in carbonite by the evil Darth Vader. The fictional metal hardened around the heroic space smuggler as it cooled – sealing him in a state of'perfect hibernation'. Carbonite is of course a fictional material, consigned to the realms of the Star Wars galaxy far, far away. But according to one scientist, this scene is not completely the stuff of science-fiction. Dr Alex Baker, a chemist at the University of Warwick, thinks humans could potentially be frozen like Solo with a real-life equivalent.


Nuclear EMP attack moves to big screen as author reflects on 'invisible lifeline'

FOX News

Author William R. Forstchen's bestselling novel "One Second After" – which imagines the devastating effects of an EMP (electromagnetic pulse) strike on the United States – is being adapted into a feature film. The screenplay will be written by renowned sci-fi writer J. Michael Straczynski, with Forstchen himself serving as an executive producer. Fox News Digital spoke with Forstchen about the real-world inspiration behind his work and why he warns that an EMP attack is a looming threat, not just science fiction. "I wanted to write an accurate, a very accurate story of what would happen in a small town in North Carolina if the power went off, and it never came back on," he said. Electromagnetic pulse expert William R. Forstchen speaks at the rally against North Korea on San Francisco's Golden Gate Bridge and Yerba Buena Gardens to support the new Homefront video game on March 2, 2011, in San Francisco, Calif.