Goto

Collaborating Authors

 Media


Automatic Speech Recognition with BERT and CTC Transformers: A Review

arXiv.org Artificial Intelligence

This review paper provides a comprehensive analysis of recent advances in automatic speech recognition (ASR) with bidirectional encoder representations from transformers BERT and connectionist temporal classification (CTC) transformers. The paper first introduces the fundamental concepts of ASR and discusses the challenges associated with it. It then explains the architecture of BERT and CTC transformers and their potential applications in ASR. The paper reviews several studies that have used these models for speech recognition tasks and discusses the results obtained. Additionally, the paper highlights the limitations of these models and outlines potential areas for further research. All in all, this review provides valuable insights for researchers and practitioners who are interested in ASR with BERT and CTC transformers.


On Goodhart's law, with an application to value alignment

arXiv.org Machine Learning

``When a measure becomes a target, it ceases to be a good measure'', this adage is known as {\it Goodhart's law}. In this paper, we investigate formally this law and prove that it critically depends on the tail distribution of the discrepancy between the true goal and the measure that is optimized. Discrepancies with long-tail distributions favor a Goodhart's law, that is, the optimization of the measure can have a counter-productive effect on the goal. We provide a formal setting to assess Goodhart's law by studying the asymptotic behavior of the correlation between the goal and the measure, as the measure is optimized. Moreover, we introduce a distinction between a {\it weak} Goodhart's law, when over-optimizing the metric is useless for the true goal, and a {\it strong} Goodhart's law, when over-optimizing the metric is harmful for the true goal. A distinction which we prove to depend on the tail distribution. We stress the implications of this result to large-scale decision making and policies that are (and have to be) based on metrics, and propose numerous research directions to better assess the safety of such policies in general, and to the particularly concerning case where these policies are automated with algorithms.


Ghost in the Shell's rad PS1 soundtrack is finally coming to the West

Engadget

The soundtrack to the spider-bot-crawling 1997 Ghost in the Shell game adaptation is coming to the West for the first time. Titled Ghost in the Shell: Megatech Body (as an ode to the Fuchikoma mech you pilot in the game), the soundtrack was produced by Takkyu Ishino. The PS1 game adaptation had late-90s gamers piloting a spider-like mech (first appearing in the 1991 manga), blasting enemies to smithereens with twin machine guns and guided missiles. Masamune Shirow, the original manga's author, wrote and illustrated its story and art design. But as 90s shooters often figured out, firing guns nonstop for hours on end is much better with a badass techno soundtrack pumping in the background like an energy drink for your ears. In addition to Ishino, it includes "warehouse-shaking bangers" from Mijk Van Dijk, The Advent, Joey Beltram and Brother from Another Planet (among others).


Engadget Podcast: Hunting data center vampires with Paris Marx

Engadget

What's that feature called on pixel phones? I forget what Android in general about Android specifics. But yes, there there was like a magic erase option there, too Yeah, I was going to say magic eraser, but that is a that's a clean thing it's something like that too, but It works really well like in terms of highlighting a specific object and removing it there are instances where it's too big and it can't like extrapolate like what should be a background so it looks really messy but sometimes like it just like smooths out a bright ugly object in the background was just like general unfocused stuff and that actually may be better.


Elon Musk showcases army of 30,000 'Optimus' robots designed to help with household chores including 'babysitting your kids' ... drawing comparisons to dystopian future depicted in I, Robot

Daily Mail - Science & tech

Elon Musk has showcased his army of 30,000 Tesla Optimus robots that are designed to help with household chores, prompting people to draw comparisons to the dystopian future depicted in I, Robot. In shocking and impressive footage, the humanoid robots were seen stiffly walking in single file across a stage while viewers stood jaw-dropped on the sidelines. Musk said attendees could walk up to the Optimus robots who would do things like serve drinks. 'At scale, you should be able to buy an Optimus robot for 20,000 to 30,000,' he said. 'It can walk your dog, mow your lawn, get the groceries, just be your friend.'


Efficient Neural Music Generation

Neural Information Processing Systems

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obtain the fine-grained acoustic tokens, making it computationally expensive and prohibitive for a real-time generation. Efficient music generation with a quality on par with MusicLM remains a significant challenge.In this paper, we present MeLoDy (M for music; L for LM; D for diffusion), an LM-guided diffusion model that generates music audios of state-of-the-art quality meanwhile reducing 95.7\% to 99.6\% forward passes in MusicLM, respectively, for sampling 10s to 30s music. MeLoDy inherits the highest-level LM from MusicLM for semantic modeling, and applies a novel dual-path diffusion (DPD) model and an audio VAE-GAN to efficiently decode the conditioning semantic tokens into waveform. DPD is proposed to simultaneously model the coarse and fine acoustics by incorporating the semantic information into segments of latents effectively via cross-attention at each denoising step.


When Counterpoint Meets Chinese Folk Melodies

Neural Information Processing Systems

Counterpoint is an important concept in Western music theory. In the past century, there have been significant interests in incorporating counterpoint into Chinese folk music composition. In this paper, we propose a reinforcement learning-based system, named FolkDuet, towards the online countermelody generation for Chinese folk melodies. With no existing data of Chinese folk duets, FolkDuet employs two reward models based on out-of-domain data, i.e. An interaction reward model is trained on the duets formed from outer parts of Bach chorales to model counterpoint interaction, while a style reward model is trained on monophonic melodies of Chinese folk songs to model melodic patterns.


StratLearner: Learning a Strategy for Misinformation Prevention in Social Networks

Neural Information Processing Systems

Given a combinatorial optimization problem taking an input, can we learn a strategy to solve it from the examples of input-solution pairs without knowing its objective function? In this paper, we consider such a setting and study the misinformation prevention problem. Given the examples of attacker-protector pairs, our goal is to learn a strategy to compute protectors against future attackers, without the need of knowing the underlying diffusion model. To this end, we design a structured prediction framework, where the main idea is to parameterize the scoring function using random features constructed through distance functions on randomly sampled subgraphs, which leads to a kernelized scoring function with weights learnable via the large margin method. Evidenced by experiments, our method can produce near-optimal protectors without using any information of the diffusion model, and it outperforms other possible graph-based and learning-based methods by an evident margin.


CNN {2}: Viewpoint Generalization via a Binocular Vision

Neural Information Processing Systems

The Convolutional Neural Networks (CNNs) have laid the foundation for many techniques in various applications. Despite achieving remarkable performance in some tasks, the 3D viewpoint generalizability of CNNs is still far behind humans visual capabilities. Although recent efforts, such as the Capsule Networks, have been made to address this issue, these new models are either hard to train and/or incompatible with existing CNN-based techniques specialized for different applications. Observing that humans use binocular vision to understand the world, we study in this paper whether the 3D viewpoint generalizability of CNNs can be achieved via a binocular vision. We propose CNN {2}, a CNN that takes two images as input, which resembles the process of an object being viewed from the left eye and the right eye.


Hespi: A pipeline for automatically detecting information from hebarium specimen sheets

arXiv.org Artificial Intelligence

Specimen associated biodiversity data are sought after for biological, environmental, climate, and conservation sciences. A rate shift is required for the extraction of data from specimen images to eliminate the bottleneck that the reliance on human-mediated transcription of these data represents. We applied advanced computer vision techniques to develop the `Hespi' (HErbarium Specimen sheet PIpeline), which extracts a pre-catalogue subset of collection data on the institutional labels on herbarium specimens from their digital images. The pipeline integrates two object detection models; the first detects bounding boxes around text-based labels and the second detects bounding boxes around text-based data fields on the primary institutional label. The pipeline classifies text-based institutional labels as printed, typed, handwritten, or a combination and applies Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for data extraction. The recognized text is then corrected against authoritative databases of taxon names. The extracted text is also corrected with the aide of a multimodal Large Language Model (LLM). Hespi accurately detects and extracts text for test datasets including specimen sheet images from international herbaria. The components of the pipeline are modular and users can train their own models with their own data and use them in place of the models provided.