Media
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
An, Ruichuan, Yang, Sihan, Lu, Ming, Zeng, Kai, Luo, Yulin, Chen, Ying, Cao, Jiajun, Liang, Hao, She, Qi, Zhang, Shanghang, Zhang, Wentao
Current vision-language models (VLMs) show exceptional abilities across diverse tasks including visual question answering. To enhance user experience in practical applications, recent studies investigate VLM personalization to understand user-provided concepts. However, existing studies mainly focus on single-concept personalization, neglecting the existence and interplay of multiple concepts, which limits the real-world applicability of personalized VLMs. In this paper, we propose the first multi-concept personalization method named MC-LLaVA along with a high-quality multi-concept personalization dataset. Specifically, MC-LLaVA uses a joint training strategy incorporating multiple concepts in a single training step, allowing VLMs to perform accurately in multi-concept personalization. To reduce the cost of joint training, MC-LLaVA leverages visual token information for concept token initialization, yielding improved concept representation and accelerating joint training. To advance multi-concept personalization research, we further contribute a high-quality dataset. We carefully collect images from various movies that contain multiple characters and manually generate the multi-concept question-answer samples. Our dataset features diverse movie types and question-answer types. We conduct comprehensive qualitative and quantitative experiments to demonstrate that MC-LLaVA can achieve impressive multi-concept personalized responses, paving the way for VLMs to become better user-specific assistants. The code and dataset will be publicly available at https://github.com/arctanxarc/MC-LLaVA.
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
Hรคrle, Ruben, Friedrich, Felix, Brack, Manuel, Deiseroth, Bjรถrn, Schramowski, Patrick, Kersting, Kristian
Large Language Models (LLMs) have demonstrated remarkable capabilities in generating human-like text, but their output may not be aligned with the user or even produce harmful content. This paper presents a novel approach to detect and steer concepts such as toxicity before generation. We introduce the Sparse Conditioned Autoencoder (SCAR), a single trained module that extends the otherwise untouched LLM. SCAR ensures full steerability, towards and away from concepts (e.g., toxic content), without compromising the quality of the model's text generation on standard evaluation benchmarks. We demonstrate the effective application of our approach through a variety of concepts, including toxicity, safety, and writing style alignment. As such, this work establishes a robust framework for controlling LLM generations, ensuring their ethical and safe deployment in real-world applications.
Benchmarking Foundation Models on Exceptional Cases: Dataset Creation and Validation
Kang, Suho, Park, Jungyang, Ha, Joonseo, Kim, SoMin, Kim, JinHyeong, Park, Subeen, Song, Kyungwoo
Foundation models (FMs) have achieved significant success across various tasks, leading to research on benchmarks for reasoning abilities. However, there is a lack of studies on FMs performance in exceptional scenarios, which we define as out-of-distribution (OOD) reasoning tasks. This paper is the first to address these cases, developing a novel dataset for evaluation of FMs across multiple modalities, including graphic novels, calligraphy, news articles, and lyrics. It includes tasks for instance classification, character recognition, token prediction, and text generation. The paper also proposes prompt engineering techniques like Chain-of-Thought (CoT) and CoT+Few-Shot to enhance performance. Validation of FMs using various methods revealed improvements. The code repository is accessible at: https://github.com/MLAI-Yonsei/ExceptionalBenchmark
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Deitke, Matt, Clark, Christopher, Lee, Sangho, Tripathi, Rohun, Yang, Yue, Park, Jae Sung, Salehi, Mohammadreza, Muennighoff, Niklas, Lo, Kyle, Soldaini, Luca, Lu, Jiasen, Anderson, Taira, Bransom, Erin, Ehsani, Kiana, Ngo, Huong, Chen, YenSung, Patel, Ajay, Yatskar, Mark, Callison-Burch, Chris, Head, Andrew, Hendrix, Rose, Bastani, Favyen, VanderBilt, Eli, Lambert, Nathan, Chou, Yvonne, Chheda, Arnavi, Sparks, Jenna, Skjonsberg, Sam, Schmitz, Michael, Sarnat, Aaron, Bischoff, Byron, Walsh, Pete, Newell, Chris, Wolters, Piper, Gupta, Tanmay, Zeng, Kuo-Hao, Borchardt, Jon, Groeneveld, Dirk, Nam, Crystal, Lebrecht, Sophie, Wittlif, Caitlin, Schoenick, Carissa, Michel, Oscar, Krishna, Ranjay, Weihs, Luca, Smith, Noah A., Hajishirzi, Hannaneh, Girshick, Ross, Farhadi, Ali, Kembhavi, Aniruddha
Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well-tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog.
OpenAI may launch Sora, its text-to-video model, very soon
OpenAI will start announcing new features and demos tomorrow for 12 days through livestreams. Sources familiar with the matter told The Verge that these new products will allegedly include OpenAI's long-awaited text-to-video tool, Sora, and a new reasoning model. The announcement for "12 Days of OpenAI", as the company puts it, was made public on X yesterday. The first livestream will broadcast tomorrow, but the announcements themselves remain unconfirmed That said, in addition to the sources that spoke more recently with The Verge, the Wall Street Journal previously reported Sora was likely to come out before the end of 2024. Sora was revealed early this year, and shared with a small group of testers. But 20 or so of those artists leaked the model to the public in protest of "unpaid labor," The Washington Post reported.
Fox News AI Newsletter: AI catches cancer that mammogram misses
MAMMO MISHAP: A U.K. woman is thanking artificial intelligence for saving her life. The technology picked up cancer cells in the patient's screening that were undetectable by the human eye, according to SWNS. READY AND WILLING: Sam Altman, CEO of OpenAI, the creator of ChatGPT, on Sunday said he is looking forward to working with the incoming Trump administration, adding that he thinks President-elect Trump will succeed at helping to make America a world-leading force in artificial intelligence infrastructure. SEEING IS REPEATING: In a groundbreaking development, researchers at Johns Hopkins University and Stanford University have successfully trained a robotic surgical system to perform complex tasks with the skill of human doctors. "Like all technology, there's the potential for incredible innovation and a real threat and obviously needs to be highly regulated," she told Fox News Digital.
Spotify users SLAM Spotify Wrapped for being 'boring' this year - as one vents 'this stinks of AI'
After feverish anticipation from fans, Spotify Wrapped is finally here, giving you a look at your most-listened-to music of 2024. Spotify's annual Wrapped feature reveals the songs and artists you've played the most over the year โ regardless of whether they're cool or cringey. The viral marketing campaign presents each user's listening habits โ including favourite songs and artists โ as a slick slideshow lasting a few minutes. However, users have slammed Spotify Wrapped for being'boring' and'ugly' this year, while another angry commentator has complained that it'stinks of AI'. On X (Twitter), one user posted: 'spotify making us wait all that time and wrapped has the most boring visuals and slideshow in years.'
Spotify Wrapped Now Includes an AI-Generated Podcast Analyzing Your Listening Habits
Spotify Wrapped's animated yearly recap of your listening habits--at once beloved and reviled--is back again. But in 2024, the flashy visuals will be accompanied by a brand-new audio add-on crafted with artificial intelligence. Starting today, Spotify users will now get the chance to listen to their annual report as a personalized, AI-powered podcast, in which two synthetic hosts discuss the user's most-played tracks and favorite artists with enthusiasm. The new podcast recap is powered by Google's NotebookLM. If you've used Google's AI tool to generate an audio podcast about a topic you're researching, then the two AI-generated voices in the Wrapped podcast will sound familiar.
From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
Mou, Xinyi, Ding, Xuanwen, He, Qi, Wang, Liang, Liang, Jingcong, Zhang, Xinnong, Sun, Libo, Lin, Jiayu, Zhou, Jie, Huang, Xuanjing, Wei, Zhongyu
Traditional sociological research often relies on human participation, which, though effective, is expensive, challenging to scale, and with ethical concerns. Recent advancements in large language models (LLMs) highlight their potential to simulate human behavior, enabling the replication of individual responses and facilitating studies on many interdisciplinary studies. In this paper, we conduct a comprehensive survey of this field, illustrating the recent progress in simulation driven by LLM-empowered agents. We categorize the simulations into three types: (1) Individual Simulation, which mimics specific individuals or demographic groups; (2) Scenario Simulation, where multiple agents collaborate to achieve goals within specific contexts; and (3) Society Simulation, which models interactions within agent societies to reflect the complexity and variety of real-world dynamics. These simulations follow a progression, ranging from detailed individual modeling to large-scale societal phenomena. We provide a detailed discussion of each simulation type, including the architecture or key components of the simulation, the classification of objectives or scenarios and the evaluation method. Afterward, we summarize commonly used datasets and benchmarks. Finally, we discuss the trends across these three types of simulation. A repository for the related sources is at {\url{https://github.com/FudanDISC/SocialAgent}}.
Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries
Ehtesham, Abul, Kumar, Saket, Singh, Aditi, Khoei, Tala Talaei
Generative AI is reshaping the media landscape, enabling unprecedented capabilities in video creation, personalization, and scalability. This paper presents a comprehensive SWOT analysis of Metas Movie Gen, a cutting-edge generative AI foundation model designed to produce 1080p HD videos with synchronized audio from simple text prompts. We explore its strengths, including high-resolution video generation, precise editing, and seamless audio integration, which make it a transformative tool across industries such as filmmaking, advertising, and education. However, the analysis also addresses limitations, such as constraints on video length and potential biases in generated content, which pose challenges for broader adoption. In addition, we examine the evolving regulatory and ethical considerations surrounding generative AI, focusing on issues like content authenticity, cultural representation, and responsible use. Through comparative insights with leading models like DALL-E and Google Imagen, this paper highlights Movie Gens unique features, such as video personalization and multimodal synthesis, while identifying opportunities for innovation and areas requiring further research. Our findings provide actionable insights for stakeholders, emphasizing both the opportunities and challenges of deploying generative AI in media production. This work aims to guide future advancements in generative AI, ensuring scalability, quality, and ethical integrity in this rapidly evolving field.