Generative AI
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
Li, Baolin, Jiang, Yankai, Gadepally, Vijay, Tiwari, Devesh
The rapid advancement of Generative Artificial Intelligence (GenAI) across diverse sectors raises significant environmental concerns, notably the carbon emissions from their cloud and high performance computing (HPC) infrastructure. This paper presents Sprout, an innovative framework designed to address these concerns by reducing the carbon footprint of generative Large Language Model (LLM) inference services. Sprout leverages the innovative concept of "generation directives" to guide the autoregressive generation process, thereby enhancing carbon efficiency. Our proposed method meticulously balances the need for ecological sustainability with the demand for high-quality generation outcomes. Employing a directive optimizer for the strategic assignment of generation directives to user prompts and an original offline quality evaluator, Sprout demonstrates a significant reduction in carbon emissions by over 40% in real-world evaluations using the Llama2 LLM and global electricity grid data. This research marks a critical step toward aligning AI technology with sustainable practices, highlighting the potential for mitigating environmental impacts in the rapidly expanding domain of generative artificial intelligence.
Automated data processing and feature engineering for deep learning and big data applications: a survey
Mumuni, Alhassan, Mumuni, Fuseini
Modern approach to artificial intelligence (AI) aims to design algorithms that learn directly from data. This approach has achieved impressive results and has contributed significantly to the progress of AI, particularly in the sphere of supervised deep learning. It has also simplified the design of machine learning systems as the learning process is highly automated. However, not all data processing tasks in conventional deep learning pipelines have been automated. In most cases data has to be manually collected, preprocessed and further extended through data augmentation before they can be effective for training. Recently, special techniques for automating these tasks have emerged. The automation of data processing tasks is driven by the need to utilize large volumes of complex, heterogeneous data for machine learning and big data applications. Today, end-to-end automated data processing systems based on automated machine learning (AutoML) techniques are capable of taking raw data and transforming them into useful features for Big Data tasks by automating all intermediate processing stages. In this work, we present a thorough review of approaches for automating data processing tasks in deep learning pipelines, including automated data preprocessing--e.g., data cleaning, labeling, missing data imputation, and categorical data encoding--as well as data augmentation (including synthetic data generation using generative AI methods) and feature engineering--specifically, automated feature extraction, feature construction and feature selection. In addition to automating specific data processing tasks, we discuss the use of AutoML methods and tools to simultaneously optimize all stages of the machine learning pipeline.
Diffusion Model for Data-Driven Black-Box Optimization
Li, Zihao, Yuan, Hui, Huang, Kaixuan, Ni, Chengzhuo, Ye, Yinyu, Chen, Minshuo, Wang, Mengdi
Generative AI has redefined artificial intelligence, enabling the creation of innovative content and customized solutions that drive business practices into a new era of efficiency and creativity. In this paper, we focus on diffusion models, a powerful generative AI technology, and investigate their potential for black-box optimization over complex structured variables. Consider the practical scenario where one wants to optimize some structured design in a high-dimensional space, based on massive unlabeled data (representing design variables) and a small labeled dataset. We study two practical types of labels: 1) noisy measurements of a real-valued reward function and 2) human preference based on pairwise comparisons. The goal is to generate new designs that are near-optimal and preserve the designed latent structures. Our proposed method reformulates the design optimization problem into a conditional sampling problem, which allows us to leverage the power of diffusion models for modeling complex distributions. In particular, we propose a reward-directed conditional diffusion model, to be trained on the mixed data, for sampling a near-optimal solution conditioned on high predicted rewards. Theoretically, we establish sub-optimality error bounds for the generated designs. The sub-optimality gap nearly matches the optimal guarantee in off-policy bandits, demonstrating the efficiency of reward-directed diffusion models for black-box optimization. Moreover, when the data admits a low-dimensional latent subspace structure, our model efficiently generates high-fidelity designs that closely respect the latent structure. We provide empirical experiments validating our model in decision-making and content-creation tasks.
Elon Musk Just Added a Wrinkle to the AI Race
Yesterday afternoon, Elon Musk fired the latest shot in his feud with OpenAI: His new AI venture, xAI, now allows anyone to download and use the computer code for its flagship software. No fees, no restrictions, just Grok, a large language model that Musk has positioned against OpenAI's GPT-4, the model powering the most advanced version of ChatGPT. Sharing Grok's code is a thinly veiled provocation. Musk was one of OpenAI's original backers. He left in 2018 and recently sued for breach of contract, arguing that the start-up and its CEO, Sam Altman, have betrayed the organization's founding principles in pursuit of profit, transforming a utopian vision of technology that "benefits all of humanity" into yet another opaque corporation.
The Next Tech Backlash Will Be About Hygiene
For centuries it was biology that made humans sick. Today, it is often stress. So argues Dr Gabor Maté about the unrecognized toll that "normal" modern life has on your mental and physical health. Dr. Maté's research, which struck a chord in 2023, invites reflection on the roll out of generative AI into daily life in 2024. As half of British teens report feeling addicted to social media, and as the U.S. surgeon general offers a rare caution against its health risks, the infusion of generative AI into social media appears to threaten our basic hygiene, meaning "the conditions or practices conducive to maintaining health and preventing disease."
Of course Apple wants to bring Google's Gemini AI to iPhones
Apple is reportedly in talks with Google to integrate its Gemini AI in iPhones, Bloomberg reports, a move that should help both companies compete with OpenAI and its (heavily invested) partner Microsoft. While it might seem like an admission that Apple is lagging behind on AI, the partnership fits if you think of generative AI models as an evolution of web searching, something Google already provides to all of Apple's devices. According to the report, Gemini could be the cloud-based generative AI engine for Siri and other iPhone apps, while Apple's models could be woven into the upcoming iOS 18 for on-device AI tasks. Bloomberg notes that Apple has also had discussions with OpenAI about using its own models, and it could still end up partnering with another AI outfit, like Anthropic. Apple could conceivably even work with multiple partners until its own generative models are up to snuff.
Using AI to spot edible mushrooms could kill you
Despite the risks, budding foragers seem to increasingly turn to apps for help identifying mushroom species. According to Google Trends, three of the five top searches related to "mushroom identification" mention apps or software. A search for "mushroom" on OpenAI's GPT Store -- where users find specialized chatbots -- immediately surfaces suggestions such as Mushroom Guide, which claims to identify mushrooms from pictures and tell whether they're edible. On the Apple or Google apps stores you'll find dozens of apps claiming to identify mushrooms, some with "AI" in the names or descriptions.
A Comparative Investigation of Compositional Syntax and Semantics in DALL-E 2
Murphy, Elliot, de Villiers, Jill, Morales, Sofia Lucero
In this study we compared how well DALL-E 2 visually represented the meaning of linguistic prompts also given to young children in comprehension tests. Sentences representing fundamental components of grammatical knowledge were selected from assessment tests used with several hundred English-speaking children aged 2-7 years for whom we had collected original item-level data. DALL-E 2 was given these prompts five times to generate 20 cartoons per item, for 9 adult judges to score. Results revealed no conditions in which DALL-E 2-generated images that matched the semantic accuracy of children, even at the youngest age (2 years). DALL-E 2 failed to assign the appropriate roles in reversible forms; it failed on negation despite an easier contrastive prompt than the children received; it often assigned the adjective to the wrong noun; it ignored implicit agents in passives. This work points to a clear absence of compositional sentence representations for DALL-E 2.
Collage Prompting: Budget-Friendly Visual Recognition with GPT-4V
Xu, Siyu, Wang, Yunke, Liu, Daochang, Xu, Chang
Recent advancements in generative AI have suggested that by taking visual prompt, GPT-4V can demonstrate significant proficiency in image recognition task. Despite its impressive capabilities, the financial cost associated with GPT-4V's inference presents a substantial barrier for its wide use. To address this challenge, our work introduces Collage Prompting, a budget-friendly prompting approach that concatenates multiple images into a single visual input. With collage prompt, GPT-4V is able to perform image recognition on several images simultaneously. Based on the observation that the accuracy of GPT-4V's image recognition varies significantly with the order of images within the collage prompt, our method further learns to optimize the arrangement of images for maximum recognition accuracy. A graph predictor is trained to indicate the accuracy of each collage prompt, then we propose an optimization method to navigate the search space of possible image arrangements. Experiment results across various datasets demonstrate the cost-efficiency score of collage prompt is much larger than standard prompt. Additionally, collage prompt with learned arrangement achieves clearly better accuracy than collage prompt with random arrangement in GPT-4V's visual recognition.
Psittacines of Innovation? Assessing the True Novelty of AI Creations
We examine whether Artificial Intelligence (AI) systems generate truly novel ideas rather than merely regurgitating patterns learned during training. Utilizing a novel experimental design, we task an AI with generating project titles for hypothetical crowdfunding campaigns. We compare within AI-generated project titles, measuring repetition and complexity. We compare between the AI-generated titles and actual observed field data using an extension of maximum mean discrepancy--a metric derived from the application of kernel mean embeddings of statistical distributions to high-dimensional machine learning (large language) embedding vectors--yielding a structured analysis of AI output novelty. Results suggest that (1) the AI generates unique content even under increasing task complexity, and at the limits of its computational capabilities, (2) the generated content has face validity, being consistent with both inputs to other generative AI and in qualitative comparison to field data, and (3) exhibits divergence from field data, mitigating concerns relating to intellectual property rights. We discuss implications for copyright and trademark law.