Goto

Collaborating Authors

 Generative AI


The Download: rise of the multimodal robots, and the SEC's new climate rules

MIT Technology Review

The news: In the summer of 2021, OpenAI quietly shuttered its mulrobotics team, announcing that progress was being stifled by a lack of data necessary to train robots in how to move and reason using artificial intelligence. Now three of OpenAI's early research scientists say the startup they spun off in 2017, called Covariant, has solved that problem. They've unveiled a system that combines the reasoning skills of large language models with the physical dexterity of an advanced robot. How it works: The new model, called RFM-1, was trained on years of data collected from Covariant's small fleet of item-picking robots, as well as words and videos from the internet. Users can prompt the model using five different types of input: text, images, video, robot instructions, and measurements.


Employees at Top AI Labs Fear Safety Is an Afterthought, Report Says

TIME - Tech

Workers at some of the world's leading AI companies harbor significant concerns about the safety of their work and the incentives driving their leadership, a report published on Monday claimed. The report, commissioned by the State Department and written by employees of the company Gladstone AI, makes several recommendations for how the U.S. should respond to what it argues are significant national security risks posed by advanced AI. Read More: Exclusive: U.S. Must Move'Decisively' To Avert'Extinction-Level' Threat from AI, Government-Commissioned Report Says The report's authors spoke with more than 200 experts for the report, including employees at OpenAI, Google DeepMind, Meta and Anthropic--leading AI labs that are all working towards "artificial general intelligence," a hypothetical technology that could perform most tasks at or above the level of a human. The authors shared excerpts of concerns that employees from some of these labs shared with them privately, without naming the individuals or the specific company that they work for. OpenAI, Google, Meta and Anthropic did not immediately respond to requests for comment. "We have served, through this project, as a de-facto clearing house for the concerns of frontier researchers who are not convinced that the default trajectory of their organizations would avoid catastrophic outcomes," Jeremie Harris, the CEO of Gladstone and one of the authors of the report, tells TIME. One individual at an unspecified AI lab shared worries with the report's authors that the lab has what the report characterized as a "lax approach to safety" stemming from a desire to not slow down the lab's work to build more powerful systems.


An OpenAI spinoff has built an AI model that helps robots learn tasks like humans

MIT Technology Review

The new model, called RFM-1, was trained on years of data collected from Covariant's small fleet of item-picking robots that customers like Crate & Barrel and Bonprix use in warehouses around the world, as well as words and videos from the internet. In the coming months, the model will be released to Covariant customers. The company hopes the system will become more capable and efficient as it's deployed in the real world. In a demonstration I attended last week, Covariant cofounders Peter Chen and Pieter Abbeel showed me how users can prompt the model using five different types of input: text, images, video, robot instructions, and measurements. For example, show it an image of a bin filled with sports equipment, and tell it to pick up the pack of tennis balls. The robot can then grab the item, generate an image of what the bin will look like after the tennis balls are gone, or create a video showing a bird's-eye view of how the robot will look doing the task.


AI talent war heats up in Europe

The Japan Times

An influx of artificial intelligence (AI) startups is heating up the battle for technical talent in Europe, leaving companies like Google DeepMind to choose between paying big or losing out on the region's best minds. The runaway success of OpenAI's ChatGPT has energized investors, who have been pouring money into promising AI startups, eager to uncover the next overnight success. Riding the investment wave, a crop of foreign AI firms -- including Canada's Cohere and U.S.-based Anthropic and OpenAI -- opened offices in Europe last year, adding to pressure on tech companies already trying to attract and retain talent in the region.


Proliferating 'news' sites spew AI-generated fake stories

The Japan Times

A sensational story about the Israeli prime minister's "psychiatrist" has exploded online, but it was AI-generated, originating on one of hundreds of websites researchers warn are churning out tech-enabled fiction masquerading as news. Propaganda-spewing websites have typically relied on armies of writers, but generative artificial intelligence tools now offer a significantly cheaper and faster way to fabricate content that is often hard to decipher from authentic information. Hundreds of AI-powered sites mimicking news outlets have cropped up in recent months, fueling an explosion of false narratives -- about everything from war to politicians -- that researchers say is stoking alarm in a year of high-stake elections around the world.


Grid Monitoring and Protection with Continuous Point-on-Wave Measurements and Generative AI

arXiv.org Machine Learning

Purpose This article presents a case for a next-generation grid monitoring and control system, leveraging recent advances in generative artificial intelligence (AI), machine learning, and statistical inference. Advancing beyond earlier generations of wide-area monitoring systems built upon supervisory control and data acquisition (SCADA) and synchrophasor technologies, we argue for a monitoring and control framework based on the streaming of continuous point-on-wave (CPOW) measurements with AI-powered data compression and fault detection. Methods and Results: The architecture of the proposed design originates from the Wiener-Kallianpur innovation representation of a random process that transforms causally a stationary random process into an innovation sequence with independent and identically distributed random variables. This work presents a generative AI approach that (i) learns an innovation autoencoder that extracts innovation sequence from CPOW time series, (ii) compresses the CPOW streaming data with innovation autoencoder and subband coding, and (iii) detects unknown faults and novel trends via nonparametric sequential hypothesis testing. Conclusion: This work argues that conventional monitoring using SCADA and phasor measurement unit (PMU) technologies is ill-suited for a future grid with deep penetration of inverter-based renewable generations and distributed energy resources. A monitoring system based on CPOW data streaming and AI data analytics should be the basic building blocks for situational awareness of a highly dynamic future grid.


Narrating Causal Graphs with Large Language Models

arXiv.org Artificial Intelligence

The use of generative AI to create text descriptions from graphs has mostly focused on knowledge graphs, which connect concepts using facts. In this work we explore the capability of large pretrained language models to generate text from causal graphs, where salient concepts are represented as nodes and causality is represented via directed, typed edges. The causal reasoning encoded in these graphs can support applications as diverse as healthcare or marketing. Using two publicly available causal graph datasets, we empirically investigate the performance of four GPT-3 models under various settings. Our results indicate that while causal text descriptions improve with training data, compared to fact-based graphs, they are harder to generate under zero-shot settings. Results further suggest that users of generative AI can deploy future applications faster since similar performances are obtained when training a model with only a few examples as compared to fine-tuning via a large curated dataset.


AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production

arXiv.org Artificial Intelligence

The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on Story-to-Video Production. AesopAgent is a practical application of agent technology for multimodal content generation. The system integrates multiple generative capabilities within a unified framework, so that individual users can leverage these modules easily. This innovative system would convert user story proposals into scripts, images, and audio, and then integrate these multimodal contents into videos. Additionally, the animating units (e.g., Gen-2 and Sora) could make the videos more infectious. The AesopAgent system could orchestrate task workflow for video generation, ensuring that the generated video is both rich in content and coherent. This system mainly contains two layers, i.e., the Horizontal Layer and the Utility Layer. In the Horizontal Layer, we introduce a novel RAG-based evolutionary system that optimizes the whole video generation workflow and the steps within the workflow. It continuously evolves and iteratively optimizes workflow by accumulating expert experience and professional knowledge, including optimizing the LLM prompts and utilities usage. The Utility Layer provides multiple utilities, leading to consistent image generation that is visually coherent in terms of composition, characters, and style. Meanwhile, it provides audio and special effects, integrating them into expressive and logically arranged videos. Overall, our AesopAgent achieves state-of-the-art performance compared with many previous works in visual storytelling. Our AesopAgent is designed for convenient service for individual users, which is available on the following page: https://aesopai.github.io/.


SMART: Automatically Scaling Down Language Models with Accuracy Guarantees for Reduced Processing Fees

arXiv.org Artificial Intelligence

The advancement of Large Language Models (LLMs) has significantly boosted performance in natural language processing (NLP) tasks. However, the deployment of high-performance LLMs incurs substantial costs, primarily due to the increased number of parameters aimed at enhancing model performance. This has made the use of state-of-the-art LLMs more expensive for end-users. AI service providers, such as OpenAI and Anthropic, often offer multiple versions of LLMs with varying prices and performance. However, end-users still face challenges in choosing the appropriate LLM for their tasks that balance result quality with cost. We introduce SMART, Scaling Models Adaptively for Reduced Token Fees, a novel LLM framework designed to minimize the inference costs of NLP tasks while ensuring sufficient result quality. It enables users to specify an accuracy constraint in terms of the equivalence of outputs to those of the most powerful LLM. SMART then generates results that deviate from the outputs of this LLM only with a probability below a user-defined threshold. SMART employs a profiling phase that evaluates the performance of multiple LLMs to identify those that meet the user-defined accuracy level. SMART optimizes the tradeoff between profiling overheads and the anticipated cost savings resulting from profiling. Moreover, our approach significantly reduces inference costs by strategically leveraging a mix of LLMs. Our experiments on three real-world datasets show that, based on OpenAI models, SMART achieves significant cost savings, up to 25.6x in comparison to GPT-4.


Active Generation for Image Classification

arXiv.org Artificial Intelligence

Recently, the growing capabilities of deep generative models have underscored their potential in enhancing image classification accuracy. However, existing methods often demand the generation of a disproportionately large number of images compared to the original dataset, while having only marginal improvements in accuracy. This computationally expensive and time-consuming process hampers the practicality of such approaches. In this paper, we propose to address the efficiency of image generation by focusing on the specific needs and characteristics of the model. With a central tenet of active learning, our method, named ActGen, takes a training-aware approach to image generation. It aims to create images akin to the challenging or misclassified samples encountered by the current model and incorporates these generated images into the training set to augment model performance. ActGen introduces an attentive image guidance technique, using real images as guides during the denoising process of a diffusion model. The model's attention on class prompt is leveraged to ensure the preservation of similar foreground object while diversifying the background. Furthermore, we introduce a gradient-based generation guidance method, which employs two losses to generate more challenging samples and prevent the generated images from being too similar to previously generated ones. Experimental results on the CIFAR and ImageNet datasets demonstrate that our method achieves better performance with a significantly reduced number of generated images.