Media
'Godfather of AI' shortens odds that new technology will wipe out human race over the next 30 years
The British-Canadian computer scientist dubbed the'Godfather of AI' has shortened the odds of artificial intelligence (AI) wiping out humans over the next 30 years, warning the technology could one day'take control'. Professor Geoffrey Hinton said we need to be'very careful' and'very thoughtful' about the development of AI which he says is'potentially very dangerous'. He had previously said there was a 10 per cent chance of the technology causing the extinction of the human race - but now predicts that figure to be '10 per cent to 20 per cent', because of the rapid pace at which AI is developing. Speaking on BBC Radio 4's Today programme, Professor Hinton said: 'You see, we've never had to deal with things more intelligent than ourselves before.' He continued: 'And how many examples do you know of a more intelligent thing being controlled by a less intelligent thing?
No Preference Left Behind: Group Distributional Preference Optimization
Yao, Binwei, Cai, Zefan, Chuang, Yun-Shiuan, Yang, Shanglin, Jiang, Ming, Yang, Diyi, Hu, Junjie
Preferences within a group of people are not uniform but follow a distribution. While existing alignment methods like Direct Preference Optimization (DPO) attempt to steer models to reflect human preferences, they struggle to capture the distributional pluralistic preferences within a group. These methods often skew toward dominant preferences, overlooking the diversity of opinions, especially when conflicting preferences arise. To address this issue, we propose Group Distribution Preference Optimization (GDPO), a novel framework that aligns language models with the distribution of preferences within a group by incorporating the concept of beliefs that shape individual preferences. GDPO calibrates a language model using statistical estimation of the group's belief distribution and aligns the model with belief-conditioned preferences, offering a more inclusive alignment framework than traditional methods. In experiments using both synthetic controllable opinion generation and real-world movie review datasets, we show that DPO fails to align with the targeted belief distributions, while GDPO consistently reduces this alignment gap during training. Moreover, our evaluation metrics demonstrate that GDPO outperforms existing approaches in aligning with group distributional preferences, marking a significant advance in pluralistic alignment.
Tell What You Hear From What You See -- Video to Audio Generation Through Text
Liu, Xiulong, Su, Kun, Shlizerman, Eli
The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the generated audio. While Video-to-Audio generation is a well-established generative task, existing methods lack such controllability. In this work, we propose VATT, a multi-modal generative framework that takes a video and an optional text prompt as input, and generates audio and optional textual description of the audio. Such a framework has two advantages: i) Video-to-Audio generation process can be refined and controlled via text which complements the context of visual information, and ii) The model can suggest what audio to generate for the video by generating audio captions. VATT consists of two key modules: VATT Converter, a LLM that is fine-tuned for instructions and includes a projection layer that maps video features to the LLM vector space; and VATT Audio, a transformer that generates audio tokens from visual frames and from optional text prompt using iterative parallel decoding. The audio tokens are converted to a waveform by pretrained neural codec. Experiments show that when VATT is compared to existing video-to-audio generation methods in objective metrics, it achieves competitive performance when the audio caption is not provided. When the audio caption is provided as a prompt, VATT achieves even more refined performance (lowest KLD score of 1.41). Furthermore, subjective studies show that VATT Audio has been chosen as preferred generated audio than audio generated by existing methods. VATT enables controllable video-to-audio generation through text as well as suggesting text prompts for videos through audio captions, unlocking novel applications such as text-guided video-to-audio generation and video-to-audio captioning.
OpenAI whistleblower's mother wants suicide death investigation reopened
If you or someone you know is having thoughts of suicide, please contact the Suicide & Crisis Lifeline at 988 or 1-800-273-TALK (8255). Balaji's death on November 26 was ruled a suicide, and Fox News Digital previously reported that the San Francisco Police Department found no evidence of foul play. But the 26-year-old's mother is urging police to reopen their investigation, saying it "doesn't look like a normal situation." Bereaved mother Poornima Ramarao told Business Insider that a private autopsy commissioned by Balaji's family and completed in early December produced concerning results. Now, they are working with an attorney to urge the department to conduct a "proper investigation."
'Godfather of AI' shortens odds of the technology wiping out humanity over next 30 years
The British-Canadian computer scientist often touted as a "godfather" of artificial intelligence has shortened the odds of AI wiping out humanity over the next three decades, warning the pace of change in the technology is "much faster" than expected. Prof Geoffrey Hinton, who this year was awarded the Nobel prize in physics for his work in AI, said there was a "10% to 20%" chance that AI would lead to human extinction within the next three decades. Previously Hinton had said there was a 10% chance of the technology triggering a catastrophic outcome for humanity. Asked on BBC Radio 4's Today programme if he had changed his analysis of a potential AI apocalypse and the one in 10 chance of it happening, he said: "Not really, 10% to 20%." Hinton's estimate prompted Today's guest editor, the former chancellor Sajid Javid, to say "you're going up", to which Hinton replied: "If anything. You see, we've never had to deal with things more intelligent than ourselves before."
Tom Hanks' New Movie Totally Bombed. I Loved It.
A great thing about catching a cold in December, as a critic, is that it's a perfect time to play NyQuil-induced catch-up with all the screeners I'd yet to watch. Cynthia Erivo is as good as everyone says in Wicked. Hundreds of Beavers is funny and incredibly well calculated, astute in its ability to shape-shift just enough to never get tedious. The Wild Robot is emotionally satisfying--but it made me lament a world in which even a robot has to have her programming overridden by the American social imperative to be a "mother." The Remarkable Life of Ibelin is a worthy reminder of what the old internet, the internet of my own upbringing, used to feel like: communal, social, mysterious.
Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
Yang, Chih-Kai, Fu, Yu-Kuan, Li, Chen-An, Lin, Yi-Cheng, Lin, Yu-Xiang, Chen, Wei-Chih, Chung, Ho Lam, Kuan, Chun-Yi, Huang, Wei-Ping, Lu, Ke-Han, Lin, Tzu-Quan, Wang, Hsiu-Hsuan, Hu, En-Pei, Hsu, Chan-Jan, Tseng, Liang-Hsuan, Chiu, I-Hsiang, Sanga, Ulin, Chen, Xuanjun, Hsu, Po-chun, Yang, Shu-wen, Lee, Hung-yi
This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end model incorporates a decoder-only transformer architecture and aims to achieve seamless interaction while preserving the conversational flow, including full-duplex capabilities allowing simultaneous speaking and listening. The paper also details the training process, including data preparation with synthesized dialogues and adjustments for real-time interaction. We also developed a platform to evaluate conversational fluency and response coherence in multi-turn dialogues. We hope the release of the report can contribute to the future development of spoken LLMs in Taiwanese Mandarin.
ETTA: Elucidating the Design Space of Text-to-Audio Models
Lee, Sang-gil, Kong, Zhifeng, Goel, Arushi, Kim, Sungwon, Valle, Rafael, Catanzaro, Bryan
Recent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the effects of data, model architecture, training objective functions, and sampling strategies on target benchmarks are not well understood. With the purpose of providing a holistic understanding of the design space of TTA models, we set up a large-scale empirical experiment focused on diffusion and flow matching models. Our contributions include: 1) AF-Synthetic, a large dataset of high quality synthetic captions obtained from an audio understanding model; 2) a systematic comparison of different architectural, training, and inference design choices for TTA models; 3) an analysis of sampling methods and their Pareto curves with respect to generation quality and inference speed. We leverage the knowledge obtained from this extensive analysis to propose our best model dubbed Elucidated Text-To-Audio (ETTA). When evaluated on AudioCaps and MusicCaps, ETTA provides improvements over the baselines trained on publicly available data, while being competitive with models trained on proprietary data. Finally, we show ETTA's improved ability to generate creative audio following complex and imaginative captions -- a task that is more challenging than current benchmarks.
Multi-view Fake News Detection Model Based on Dynamic Hypergraph
With the rapid development of online social networks and the inadequacies in content moderation mechanisms, the detection of fake news has emerged as a pressing concern for the public. Various methods have been proposed for fake news detection, including text-based approaches as well as a series of graph-based approaches. However, the deceptive nature of fake news renders text-based approaches less effective. Propagation tree-based methods focus on the propagation process of individual news, capturing pairwise relationships but lacking the capability to capture high-order complex relationships. Large heterogeneous graph-based approaches necessitate the incorporation of substantial additional information beyond news text and user data, while hypergraph-based approaches rely on predefined hypergraph structures. To tackle these issues, we propose a novel dynamic hypergraph-based multi-view fake news detection model (DHy-MFND) that learns news embeddings across three distinct views: text-level, propagation tree-level, and hypergraph-level. By employing hypergraph structures to model complex high-order relationships among multiple news pieces and introducing dynamic hypergraph structure learning, we optimize predefined hypergraph structures while learning news embeddings. Additionally, we introduce contrastive learning to capture authenticity-relevant embeddings across different views. Extensive experiments on two benchmark datasets demonstrate the effectiveness of our proposed DHy-MFND compared with a broad range of competing baselines.
The Sex Scenes in This Season's Hottest Movie Are Just โฆ Oh My God
In Sex Reviews, writers offer a sober critical assessment of the sex scenes in new films and television series. This installment contains spoilers for Babygirl. Nestled amid a nice little set of Christmas releases is Babygirl, an erotic thriller set during the holidays, written and directed by Halina Reijn of Bodies, Bodies, Bodies fame. The film stars Nicole Kidman as Romy, a work-addicted CEO of a robotics company, who lives with her play-directing, gray-goatee-sporting husband Jacob (Antonio Banderas) and two teenage daughters in a gorgeous Manhattan apartment. Romy's creeping dissatisfaction with the rounds of Botox, therapy, and meetings that make up her life comes to a head when she meets Samuel, an intern at her company, played by hyper-handsome English actor Harris Dickinson. In fits and starts, the Gen X Romy and Gen Z Samuel discover that they have a very particular type of chemistry: She wants to be told what to do, and he's willing to tell her.