Media
Deep Transfer Learning for Automatic Speech Recognition: Towards Better Generalization
Kheddar, Hamza, Himeur, Yassine, Al-Maadeed, Somaya, Amira, Abbes, Bensaali, Faycal
Automatic speech recognition (ASR) has recently become an important challenge when using deep learning (DL). It requires large-scale training datasets and high computational and storage resources. Moreover, DL techniques and machine learning (ML) approaches in general, hypothesize that training and testing data come from the same domain, with the same input feature space and data distribution characteristics. This assumption, however, is not applicable in some real-world artificial intelligence (AI) applications. Moreover, there are situations where gathering real data is challenging, expensive, or rarely occurring, which can not meet the data requirements of DL models. deep transfer learning (DTL) has been introduced to overcome these issues, which helps develop high-performing models using real datasets that are small or slightly different but related to the training data. This paper presents a comprehensive survey of DTL-based ASR frameworks to shed light on the latest developments and helps academics and professionals understand current challenges. Specifically, after presenting the DTL background, a well-designed taxonomy is adopted to inform the state-of-the-art. A critical analysis is then conducted to identify the limitations and advantages of each framework. Moving on, a comparative study is introduced to highlight the current challenges before deriving opportunities for future research.
The covert intelligence group covering up UFOs: New documentary lifts lid on 'Collins Elite' - secret Pentagon group that believe craft buzzing around in our skies are 'demonic'
A film premiering next month will shine a light on a secret U.S. group which UFO researchers claim is helping to cover up the discovery of alien spacecraft. The film, God versus Aliens, has interviews with two experts about the'Collins Elite', a supposed secretive group within the U.S. military which has helped to cover up alien abductions and crashed spacecraft since the 1950s. British director of the film Mark Christopher Lee said: 'A lot of people know about Majestic 12, a supposed committee of military leaders and politicians interested in UFOs: it's out there in pop culture, along with Area 51. 'But two of my interviewees believe that this is a smokescreen, and the real organization is the Collins Elite, based in the Wright Patterson Air Base [in Ohio]. 'These people are said to work in a private organization on behalf of the government, because there are no freedom of information requests (FOIA) to private companies.' The Wright Patterson Air Base was home to the Project Blue Book investigation into UFO reports which began in 1947 - and there have been previous rumors of a secret'UFO room' at the base.
Doom busters: why some things aren't (quite) as bad as we think
"AI for Good is about building AI in the right way and using it for social good. We've learned there are good business reasons for building this technology safely โ if you want people to adopt it and use it, they need to be able to trust it. We've done a lot of work around helping domestic abuse victims in South Africa, with a chatbot called rAInbow. It was designed to help people understand their legal rights. It can be quite overwhelming to take that first step to getting help and trusted information if you don't know where to begin. I think it's important to acknowledge the risks this technology brings, but there are also tremendous positive opportunities. I spend a lot of my time building AI that helps improve the justice system and helps people understand their legal rights. With this technology, we can produce legal drafts in minutes that used to take days. Courts can function better and faster, so people can get their hearing dates and we can make the system more efficient. What excites me is that the new generation of technologists don't have to have my background. I went to geek school after geek school, but the newest programming language is human language. This means we can bring in people from many different backgrounds to build it. If we do this right, we will be opening up the profile of people who work in technology and AI." Fanning the flames of a "culture war" might drive ratings, generate clicks and provide politicians with election fodder, but the idea there are ever-deepening divides in British social attitudes is misleading.
Anatomy of an AI-powered malicious social botnet
Yang, Kai-Cheng, Menczer, Filippo
Concerns have been raised that they could be utilized to produce fake content with a deceptive intention, although evidence thus far remains anecdotal. This paper presents a case study about a Twitter botnet that appears to employ ChatGPT to generate human-like content. Through heuristics, we identify 1,140 accounts and validate them via manual annotation. These accounts form a dense cluster of fake personas that exhibit similar behaviors, including posting machine-generated content and stolen images, and engage with each other through replies and retweets. ChatGPT-generated content promotes suspicious websites and spreads harmful comments. While the accounts in the AI botnet can be detected through their coordination patterns, current state-of-the-art LLM content classifiers fail to discriminate between them and human accounts in the wild. These findings highlight the threats posed by AI-enabled social bots.
Proposing a conceptual framework: social media listening for public health behavior
Tsao, Shu-Feng, Chen, Helen, Meyer, Samantha, Butt, Zahid A.
Existing communications and behavioral theories have been adopted to address health misinformation. Although various theories and models have been used to investigate the COVID-19 pandemic, there is no framework specially designed for social listening or misinformation studies using social media data and natural language processing techniques. This study aimed to propose a novel yet theory-based conceptual framework for misinformation research. We collected theories and models used in COVID-19 related studies published in peer-reviewed journals. The theories and models ranged from health behaviors, communications, to misinformation. They are analyzed and critiqued for their components, followed by proposing a conceptual framework with a demonstration. We reviewed Health Belief Model, Theory of Planned Behavior/Reasoned Action, Communication for Behavioral Impact, Transtheoretical Model, Uses and Gratifications Theory, Social Judgment Theory, Risk Information Seeking and Processing Model, Behavioral and Social Drivers, and Hype Loop. Accordingly, we proposed the Social Media Listening for Public Health Behavior Conceptual Framework by not only integrating important attributes of existing theories, but also adding new attributes. The proposed conceptual framework was demonstrated in the Freedom Convoy social media listening. The proposed conceptual framework can be used to better understand public discourse on social media, and it can be integrated with other data analyses to gather a more comprehensive picture. The framework will continue to be revised and adopted as health misinformation evolves.
Moisesdb: A dataset for source separation beyond 4-stems
Pereira, Igor, Araรบjo, Felipe, Korzeniowski, Filip, Vogl, Richard
In this paper, we introduce the MoisesDB dataset for musical source separation. It consists of 240 tracks from 45 artists, covering twelve musical genres. For each song, we provide its individual audio sources, organized in a two-level hierarchical taxonomy of stems. This will facilitate building and evaluating fine-grained source separation systems that go beyond the limitation of using four stems (drums, bass, other, and vocals) due to lack of data. To facilitate the adoption of this dataset, we publish an easy-to-use Python library to download, process and use MoisesDB. Alongside a thorough documentation and analysis of the dataset contents, this work provides baseline results for open-source separation models for varying separation granularities (four, five, and six stems), and discuss their results.
Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation
Zhang, Jiaming, Yang, Kailun, Shi, Hao, Reiร, Simon, Peng, Kunyu, Ma, Chaoxiang, Fu, Haodong, Torr, Philip H. S., Wang, Kaiwei, Stiefelhagen, Rainer
In this paper, we address panoramic semantic segmentation which is under-explored due to two critical challenges: (1) image distortions and object deformations on panoramas; (2) lack of semantic annotations in the 360-degree imagery. To tackle these problems, first, we propose the upgraded Transformer for Panoramic Semantic Segmentation, i.e., Trans4PASS+, equipped with Deformable Patch Embedding (DPE) and Deformable MLP (DMLPv2) modules for handling object deformations and image distortions whenever (before or after adaptation) and wherever (shallow or deep levels). Second, we enhance the Mutual Prototypical Adaptation (MPA) strategy via pseudo-label rectification for unsupervised domain adaptive panoramic segmentation. Third, aside from Pinhole-to-Panoramic (Pin2Pan) adaptation, we create a new dataset (SynPASS) with 9,080 panoramic images, facilitating Synthetic-to-Real (Syn2Real) adaptation scheme in 360-degree imagery. Extensive experiments are conducted, which cover indoor and outdoor scenarios, and each of them is investigated with Pin2Pan and Syn2Real regimens. Trans4PASS+ achieves state-of-the-art performances on four domain adaptive panoramic semantic segmentation benchmarks. Code is available at https://github.com/jamycheung/Trans4PASS.
AI news recap: While Hollywood strikes, is ChatGPT getting worse?
Artificial intelligence can now create images, novels and source code from scratch. Except it isn't really from scratch, because a vast amount of human-generated examples are needed to train these AI models โ something that has angered artists, programmers and writers and led to a series of lawsuits. Hollywood actors are the latest group of creatives to turn against AI. They fear that film studios could take control of their likeness and have them "star" in films without ever being on set, perhaps taking on roles they would rather avoid and uttering lines or acting out scenes they would find distasteful. Worse still, they might not get paid for it.
Testing the Depth of ChatGPT's Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5's Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking
Over the eight months since its release, ChatGPT and its underlying model, GPT3.5, have garnered massive attention, due to their potent mix of capability and accessibility. While a niche-industry of papers have emerged examining the scope of capabilities these models possess, the information fed to and extracted from these networks has been either natural language text or stylized, code-like language. Drawing inspiration from the prowess we expect a truly human-level intelligent agent to have across multiple signal modalities, in this work we examine GPT3.5's aptitude for visual tasks, where the inputs feature content provided as ASCII-art without overt distillation into a lingual summary. We conduct experiments analyzing the model's performance on image recognition tasks after various transforms typical in visual settings, trials investigating knowledge of image parts, and tasks covering image generation.
Fast but multi-partisan: Bursts of communication increase opinion diversity in the temporal Deffuant model
Zarei, Fatemeh, Gandica, Yerali, Rocha, Luis Enrique Correa
Human interactions create social networks forming the backbone of societies. Individuals adjust their opinions by exchanging information through social interactions. Two recurrent questions are whether social structures promote opinion polarisation or consensus in societies and whether polarisation can be avoided, particularly on social media. In this paper, we hypothesise that not only network structure but also the timings of social interactions regulate the emergence of opinion clusters. We devise a temporal version of the Deffuant opinion model where pairwise interactions follow temporal patterns and show that burstiness alone is sufficient to refrain from consensus and polarisation by promoting the reinforcement of local opinions. Individuals self-organise into a multi-partisan society due to network clustering, but the diversity of opinion clusters further increases with burstiness, particularly when individuals have low tolerance and prefer to adjust to similar peers. The emergent opinion landscape is well-balanced regarding clusters' size, with a small fraction of individuals converging to extreme opinions. We thus argue that polarisation is more likely to emerge in social media than offline social networks because of the relatively low social clustering observed online. Counter-intuitively, strengthening online social networks by increasing social redundancy may be a venue to reduce polarisation and promote opinion diversity.