Generative AI
Microsoft Designer and its AI art will soon land on your PC
Microsoft Designer's AI art and editing capabilities are becoming more formally integrated into Windows and Microsoft's services today, as they move into Photos, Word, and PowerPoint. Designer is the next stage of Microsoft's evolution of AI art, which began in 2022 with Bing Image Creator, migrated to the more advanced Dall-E 2 model, then became part of Microsoft Designer, the wonderful AI-powered design tool that debuted in 2022. Designer's layout elements compete directly with Canva, but Microsoft isn't confining the Designer elements to just the app. The most simple integration is within Word and PowerPoint. You'll need a Copilot Pro subscription, but if you have one, you'll be able to use AI to generate a background for a PowerPoint slide or an integrated graphic inside of a Word document.
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
Yoon, Jinsung, Sinha, Raj, Arik, Sercan O, Pfister, Tomas
Embeddings from Large Language Models (LLMs) have emerged as critical components in various applications, particularly for information retrieval. While high-dimensional embeddings generally demonstrate superior performance as they contain more salient information, their practical application is frequently hindered by elevated computational latency and the associated higher cost. To address these challenges, we propose Matryoshka-Adaptor, a novel tuning framework designed for the customization of LLM embeddings. Matryoshka-Adaptor facilitates substantial dimensionality reduction while maintaining comparable performance levels, thereby achieving a significant enhancement in computational efficiency and cost-effectiveness. Our framework directly modifies the embeddings from pre-trained LLMs which is designed to be seamlessly integrated with any LLM architecture, encompassing those accessible exclusively through black-box APIs. Also, it exhibits efficacy in both unsupervised and supervised learning settings. A rigorous evaluation conducted across a diverse corpus of English, multilingual, and multimodal datasets consistently reveals substantial gains with Matryoshka-Adaptor. Notably, with Google and OpenAI Embedding APIs, Matryoshka-Adaptor achieves a reduction in dimensionality ranging from two- to twelve-fold without compromising performance across multiple BEIR datasets.
From Principles to Practices: Lessons Learned from Applying Partnership on AI's (PAI) Synthetic Media Framework to 11 Use Cases
Leibowicz, Claire R., Cardona, Christian H.
2023 was the year the world woke up to generative AI, and 2024 is the year policymakers are responding more firmly. Importantly, this policy momentum is taking place alongside real world creation and distribution of synthetic media. Social media platforms, news organizations, dating apps, image generation companies, and more are already navigating a world of AI-generated visuals and sounds, already changing hearts and minds, as policymakers try to catch up. How, then, can AI governance capture the complexity of the synthetic media landscape? How can it attend to synthetic media's myriad uses, ranging from storytelling to privacy preservation, to deception, fraud, and defamation, taking into account the many stakeholders involved in its development, creation, and distribution? And what might it mean to govern synthetic media in a manner that upholds the truth while bolstering freedom of expression? What follows is the first known collection of diverse examples of the implementation of synthetic media governance that responds to these questions, specifically through Partnership on AI's (PAI) Responsible Practices for Synthetic Media - a voluntary, normative Framework for creating, distributing, and building technology for synthetic media responsibly, launched in February 2023. In this paper, we present a case bank of real world examples that help operationalize the Framework - highlighting areas synthetic media governance can be applied, augmented, expanded, and refined for use, in practice. Read together, the cases emphasize distinct elements of AI policymaking and seven emergent best practices supporting transparency, safety, expression, and digital dignity online: consent, disclosure, and differentiation between harmful and creative use cases.
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
Majumder, Navonil, Hung, Chia-Yu, Ghosal, Deepanway, Hsu, Wei-Ning, Mihalcea, Rada, Poria, Soujanya
Generative multimodal content is increasingly prevalent in much of the content creation arena, as it has the potential to allow artists and media personnel to create pre-production mockups by quickly bringing their ideas to life. The generation of audio from text prompts is an important aspect of such processes in the music and film industry. Many of the recent diffusion-based text-to-audio models focus on training increasingly sophisticated diffusion models on a large set of datasets of prompt-audio pairs. These models do not explicitly focus on the presence of concepts or events and their temporal ordering in the output audio with respect to the input prompt. Our hypothesis is focusing on how these aspects of audio generation could improve audio generation performance in the presence of limited data. As such, in this work, using an existing text-to-audio model Tango, we synthetically create a preference dataset where each prompt has a winner audio output and some loser audio outputs for the diffusion model to learn from. The loser outputs, in theory, have some concepts from the prompt missing or in an incorrect order. We fine-tune the publicly available Tango text-to-audio model using diffusion-DPO (direct preference optimization) loss on our preference dataset and show that it leads to improved audio output over Tango and AudioLDM2, in terms of both automatic- and manual-evaluation metrics.
In the Age of A.I., How Much Is Silicon Valley Prepared to Give Back?
For the last couple of years, the tech community has tested no-strings-attached payments of 500 or 1,000 a month to those in dire need. Some of these experiments have happened in the heart of Silicon Valley, where a one-bedroom apartment rents for 3,000 a month and a modest house is often an unaffordable luxury. Silicon Valley's backing of these efforts has propelled the idea of a guaranteed income -- also known as cash transfers, unconditional cash and, in its most utopian form, universal basic income -- into the mainstream. But a bipartisan political consensus around the movement is fracturing even though the data seems to show that the programs are effective. In recent months, the Texas attorney general went to court to prevent public funds from being used in a basic income program in Houston.
Hong Kong Testing ChatGPT-Style Tool After OpenAI Took Steps to Block Access
Hong Kong's government is testing the city's own ChatGPT -style tool for its employees, with plans to eventually make it available to the public, its innovation minister said after OpenAI took extra steps to block access from the city and other unsupported regions. Secretary for Innovation, Technology and Industry Sun Dong said on a Saturday radio show that his bureau was trying out the artificial intelligence program, whose Chinese name translates to "document assistance application for civil servants," to further improve its capabilities. He plans to have it available for the rest of the government this year. The program was developed by a generative AI research and development center led by the Hong Kong University of Science and Technology in collaboration with several other universities. Sun said the model would provide functions like graphics and video design in the future.
Disney investigating massive leak of internal messages
The leak was first reported in the gaming press and then picked up by the Wall Street Journal, which said some of the leaked material related to advertising campaigns and interview candidates, with some dating back as far as 2019. There has been growing concern amongst performers, artists and other creatives that the rapid spread of generative AI will undermine their livelihoods and damage the creative environment. Generative AI is trained on vast bodies of existing material - including texts, images, music and video. It is then able to produce new work of a standard that can be hard to distinguish from human-generated material. Nullbulge describes itself as "a hacktivist group protecting artists' rights and ensuring fair compensation for their work".
A hacking group reportedly leaked confidential data from thousands of Disney Slack channels.
A hacking group leaked over a terabyte of confidential data from more than 10,000 Slack channels belonging to Disney, the Wall Street Journal reported on Monday. The leaked information includes discussions about ad campaigns, computer code, details about unreleased projects and discussion about interview candidates among other things. "Disney is investigating this matter," a company spokesperson told the Journal. Nullbulge calls itself a hacktivist group advocating for the rights of artists. A spokesperson for the group told the Journal that it targeted Disney due to concerns about the company's handling of artist contracts and its approach to generative AI.
AIGC for Industrial Time Series: From Deep Generative Models to Large Generative Models
Ren, Lei, Wang, Haiteng, Tang, Yang, Yang, Chunhua
With the remarkable success of generative models like ChatGPT, Artificial Intelligence Generated Content (AIGC) is undergoing explosive development. Not limited to text and images, generative models can generate industrial time series data, addressing challenges such as the difficulty of data collection and data annotation. Due to their outstanding generation ability, they have been widely used in Internet of Things, metaverse, and cyber-physical-social systems to enhance the efficiency of industrial production. In this paper, we present a comprehensive overview of generative models for industrial time series from deep generative models (DGMs) to large generative models (LGMs). First, a DGM-based AIGC framework is proposed for industrial time series generation. Within this framework, we survey advanced industrial DGMs and present a multi-perspective categorization. Furthermore, we systematically analyze the critical technologies required to construct industrial LGMs from four aspects: large-scale industrial dataset, LGMs architecture for complex industrial characteristics, self-supervised training for industrial time series, and fine-tuning of industrial downstream tasks. Finally, we conclude the challenges and future directions to enable the development of generative models in industry.
Generating Multi-Modal and Multi-Attribute Single-Cell Counts with CFGen
Palma, Alessandro, Richter, Till, Zhang, Hanyi, Lubetzki, Manuel, Tong, Alexander, Dittadi, Andrea, Theis, Fabian
Generative modeling of single-cell RNA-seq data has shown invaluable potential in community-driven tasks such as trajectory inference, batch effect removal and gene expression generation. However, most recent deep models generating synthetic single cells from noise operate on pre-processed continuous gene expression approximations, ignoring the inherently discrete and over-dispersed nature of single-cell data, which limits downstream applications and hinders the incorporation of robust noise models. Moreover, crucial aspects of deep-learning-based synthetic single-cell generation remain underexplored, such as controllable multi-modal and multi-label generation and its role in the performance enhancement of downstream tasks. This work presents Cell Flow for Generation (CFGen), a flow-based conditional generative model for multi-modal single-cell counts, which explicitly accounts for the discrete nature of the data. Our results suggest improved recovery of crucial biological data characteristics while accounting for novel generative tasks such as conditioning on multiple attributes and boosting rare cell type classification via data augmentation. By showcasing CFGen on a diverse set of biological datasets and settings, we provide evidence of its value to the fields of computational biology and deep generative models.