Goto

Collaborating Authors

 Media


Unsupervised outlier detection to improve bird audio dataset labels

arXiv.org Artificial Intelligence

The Xeno -Canto bird audio repository is an invaluable resource for those interested in vocalizations and other sounds made by birds around the world. This is particularly the case for machine learning researchers attempting to improve on the bird species r ecognition accuracy of classification models. However, the task of extracting labeled datasets from th e recordings found in this crowd -sourced repository faces several challenges. One challenge of particular significance to machine learning practitioners i s that one bird species label is applied to each audio recording, but frequently other sounds are also captured including other bird species, other animal sounds, anthropogenic and other ambient sounds . These non -target bird species sounds can result in dataset labeling discrepanc ies referred to as label noise . In this work we present a cleaning process consisting of audio preprocessing followed by dimensionality reduction and unsupervised outlier detection (UOD) to reduce the label noise in a dataset derived from Xeno -Canto recordings . We investigate three neural network dimensionality reduction techniques: two flavors of convolutional autoencoder s and variational deep embedding (VaDE (Jiang, 2017)) . While both methods show some degree of effectiveness at detecting outliers for most bird species datasets, we f ound significant variation in the performance of the methods from one species to the next. We believe that the results of this investigation demonstrate that the application of our cleaning process can meaningfully reduce the label noise of bird species datasets derived from Xeno-Canto audio repository but results vary across species.


Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning

arXiv.org Artificial Intelligence

Speaker diarization, a core problem in speech processing, entails partitioning a given audio stream according to the speakers. Even though progress has been made in the development of the models for high - resource languages, there is still a set of specific difficulties in going through a similar process for low - resource languages such as Kurdish: there are very few annotated datasets available; the language has dialects; speakers use code - switching a lot. These challenges are met in this study by training the Wav2V ec 2.0 SSL model on a Ku rdish dataset prepared for this purpose. Thanks to transfer learning, it was possible to transfer multiling ual representations learnt in other languages to the phonetic and acoustic features of Kurdish speech. The general Diarization Error Rate (DER) was reduced by 7.2%, and the cluster purity increased by 13% when compared to the baseline algorithm. They show that making improvements in any state - of - the - art model can help in enhancing the performance of under - resourced languages. Implications of this work include transcription services for Kurdish - language media programs, as well as speaker segmentation in multilingual call centers, teleconferencing, and videoconferencing systems. Therefore, this work demonstrates that self - supervised and transfer techniques can improve speaker diarization for Kurdish and other low - resource languages with diverse features. The approach provides a ba se for building effective diarization systems in other understudied languages, which remai ns essential for speech technology's equity.


Multi-view autoencoders for Fake News Detection

arXiv.org Artificial Intelligence

Given the volume and speed at which fake news spreads across social media, automatic fake news detection has become a highly important task. However, this task presents several challenges, including extracting textual features that contain relevant information about fake news. Research about fake news detection shows that no single feature extraction technique consistently outperforms the others across all scenarios. Nevertheless, different feature extraction techniques can provide complementary information about the textual data and enable a more comprehensive representation of the content. This paper proposes using multi-view autoencoders to generate a joint feature representation for fake news detection by integrating several feature extraction techniques commonly used in the literature. Experiments on fake news datasets show a significant improvement in classification performance compared to individual views (feature representations). We also observed that selecting a subset of the views instead of composing a latent space with all the views can be advantageous in terms of accuracy and computational effort. For further details, including source codes, figures, and datasets, please refer to the project's repository: https://github.com/ingrydpereira/multiview-fake-news.


Application and Optimization of Large Models Based on Prompt Tuning for Fact-Check-Worthiness Estimation

arXiv.org Artificial Intelligence

Application and Optimization of Large Models Based on Prompt Tuning for Fact-Check-Worthiness Estimation Yinglong Y u 1, Hao Shen 2 Zhengyi Lyu 3 and Qi He 4 Communication University of China, Beijing, China 1 yuyingling@cuc.edu.cn 2 shenhao@cuc.edu.cn 3 lyuzhengyi@cuc.edu.cn 4 heqi654321@126.com Abstract --In response to the growing problem of misinformation in the context of globalization and informatization, this paper proposes a classification method for fact-check-worthiness estimation based on prompt tuning. We construct a model for fact-check-worthiness estimation at the methodological level using prompt tuning. By applying designed prompt templates to large language models, we establish in-context learning and leverage prompt tuning technology to improve the accuracy of determining whether claims have fact-check-worthiness, particularly when dealing with limited or unlabeled data. Through extensive experiments on public datasets, we demonstrate that the proposed method surpasses or matches multiple baseline methods in the classification task of fact-check-worthiness estimation assessment, including classical pre-trained models such as BERT, as well as recent popular large models like GPT - 3.5 and GPT -4. Experiments show that the prompt tuning-based method proposed in this study exhibits certain advantages in evaluation metrics such as F1 score and accuracy, thereby effectively validating its effectiveness and advancement in the task of fact-check-worthiness estimation. I NTRODUCTION In today's interconnected world characterized by globalization and informatization, the complexity of multilingual environments and the challenges posed by misinformation have become increasingly severe. With the deepening of international exchanges and the expanding influence of social media, rumors and false information spread rapidly across cyberspace, exacerbating the uncertainty in the global public discourse.


AI Ethics and Social Norms: Exploring ChatGPT's Capabilities From What to How

arXiv.org Artificial Intelligence

Using LLMs in healthcare, Computer-Supported Cooperative Work, and Social Computing requires the examination of ethical and social norms to ensure safe incorporation into human life. We conducted a mixed-method study, including an online survey with 111 participants and an interview study with 38 experts, to investigate the AI ethics and social norms in ChatGPT as everyday life tools. This study aims to evaluate whether ChatGPT in an empirical context operates following ethics and social norms, which is critical for understanding actions in industrial and academic research and achieving machine ethics. The findings of this study provide initial insights into six important aspects of AI ethics, including bias, trustworthiness, security, toxicology, social norms, and ethical data. Significant obstacles related to transparency and bias in unsupervised data collection methods are identified as ChatGPT's ethical concerns.


RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

arXiv.org Artificial Intelligence

Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on standard LLMs, which means we know little about how RAG use cases change a model's safety profile. We conduct a detailed comparative analysis of RAG and non-RAG frameworks with eleven LLMs. We find that RAG can make models less safe and change their safety profile. We explore the causes of this change and find that even combinations of safe models with safe documents can cause unsafe generations. In addition, we evaluate some existing red teaming methods for RAG settings and show that they are less effective than when used for non-RAG settings. Our work highlights the need for safety research and red-teaming methods specifically tailored for RAG LLMs.


VEU-Bench: Towards Comprehensive Understanding of Video Editing

arXiv.org Artificial Intelligence

Widely shared videos on the internet are often edited. Recently, although Video Large Language Models (Vid-LLMs) have made great progress in general video understanding tasks, their capabilities in video editing understanding (VEU) tasks remain unexplored. To address this gap, in this paper, we introduce VEU-Bench (Video Editing Understanding Benchmark), a comprehensive benchmark that categorizes video editing components across various dimensions, from intra-frame features like shot size to inter-shot attributes such as cut types and transitions. Unlike previous video editing understanding benchmarks that focus mainly on editing element classification, VEU-Bench encompasses 19 fine-grained tasks across three stages: recognition, reasoning, and judging. To enhance the annotation of VEU automatically, we built an annotation pipeline integrated with an ontology-based knowledge base. Through extensive experiments with 11 state-of-the-art Vid-LLMs, our findings reveal that current Vid-LLMs face significant challenges in VEU tasks, with some performing worse than random choice. To alleviate this issue, we develop Oscars, a VEU expert model fine-tuned on the curated VEU-Bench dataset. It outperforms existing open-source Vid-LLMs on VEU-Bench by over 28.3% in accuracy and achieves performance comparable to commercial models like GPT-4o. We also demonstrate that incorporating VEU data significantly enhances the performance of Vid-LLMs on general video understanding benchmarks, with an average improvement of 8.3% across nine reasoning tasks.


'Godfather of AI' reveals the startling odds that artificial intelligence will take over humanity

Daily Mail - Science & tech

Scientist and physicist Geoffrey Hinton believes there could be a one in five chance that humanity will eventually be taken over by artificial intelligence. Hinton, a Nobel laureate in physics who's been dubbed the'godfather of AI', made the startling prediction in an April 1 interview with CBS News that was aired on Saturday morning. 'I'm in the unfortunate position of happening to agree with Elon Musk on this, which is that there's a 10 to 20 percent chance that these things will take over, but that's just a wild guess,' Hinton said. Besides his cost-cutting responsibilities in the federal government, Musk is the chief executive of xAI, the company that made the AI chatbot Grok. Musk has said AI will become smarter than the entire human race by 2029.


How to watch LlamaCon 2025, Meta's first generative AI developer conference

Engadget

After a couple years of having its open-source Llama AI model be just a part of its Connect conferences, Meta is breaking things out and hosting an entirely generative AI-focused developer conference called LlamaCon on April 29. The event is entirely virtual, and you'll be able to watch along live on the Meta for Developers Facebook page. LlamaCon kicks off at 1PM ET / 10AM PT with a keynote address from Meta's Chief Product Officer Chris Cox, Vice President of AI Manohar Paluri and research scientist Angela Fan. The keynote is supposed to cover developments in the company's open-source AI community, "the latest on the Llama collection of models and tools" and offer a glimpse at yet-to-be released AI features. The keynote address will be followed by a conversation at 1:45PM ET / 10:45PM ET between Meta CEO Mark Zuckerberg and Databricks CEO Ali Ghodsi on "building AI-powered applications," followed by a chat at 7PM ET / 4PM PT about "the latest trends in AI" between Zuckerberg and Microsoft CEO Satya Nadella. It doesn't seem like either conversation will be used to break news, but Microsoft and Meta have collaborated before, so anything is possible.


US government defunds research on misinformation

New Scientist

The US National Science Foundation (NSF) has terminated government research grants for studying misinformation and disinformation. The defunding comes at a time when propaganda and scams fuelled by the latest artificial intelligence technologies are flooding social media networks, and tech companies are abandoning content moderation efforts and eliminating fact-checking teams. The grant cancellations began on 18 April when the NSF published a statement saying it would not support research on misinformation or disinformation "that could be used to infringe on the constitutionally protected speech rights…