Goto

Collaborating Authors

 Generative AI


OpenAI and Microsoft are funding 10 million in grants for AI-powered journalism

Engadget

OpenAI and Microsoft are funding projects to bring more AI tools into the newsroom. The duo will give grants of up to 10 million to Chicago Public Media, the Minnesota Star Tribune, Newsday (in Long Island, NY), The Philadelphia Inquirer and The Seattle Times. Each of the publications will hire a two-year AI fellow to develop projects for implementing the technology and improving business sustainability. Three more outlets are expected to receive fellowship grants in a second round. OpenAI and Microsoft are each contributing 2.5 million in direct funding as well as 2.5 million in software and enterprise credits.


Thom Yorke and Julianne Moore join thousands of creatives in AI warning

The Guardian

Abba's Bjรถrn Ulvaeus, the actor Julianne Moore, the Radiohead singer Thom Yorke are among 10,500 signatories of a statement from the creative industries warning artificial intelligence companies that unlicensed use of their work is a "major, unjust threat" to artists' livelihoods. "The unlicensed use of creative works for training generative AI is a major, unjust threat to the livelihoods of the people behind those works, and must not be permitted," reads the statement. Thousands of creative professionals from the worlds of literature, music, film, theatre and television have given their backing to the statement, with authors including Kazuo Ishiguro, Ann Patchett, and Kate Mosse, musicians including the Cure's Robert Smith as well as the composer Max Richter and actors including Kevin Bacon, Rosario Dawson and F Murray Abraham. The organiser of the letter, the British composer and former AI executive Ed Newton-Rex, said people who make a living from creative work are "very worried" about the situation. "There are three key resources that generative AI companies need to build AI models: people, compute, and data. They spend vast sums on the first two โ€“ sometimes a million dollars per engineer, and up to a billion dollars per model. But they expect to take the third โ€“ training data โ€“ for free," he said.


Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups

arXiv.org Artificial Intelligence

AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly concerning when evaluating the safety of generative AI, where perceptions and associated harms can vary significantly across socio-cultural contexts. While recent research has studied the impact of demographic differences on annotating text, there is limited understanding of how these subjective variations affect multimodal safety in generative AI. To address this, we conduct a large-scale study employing highly-parallel safety ratings of about 1000 text-to-image (T2I) generations from a demographically diverse rater pool of 630 raters balanced across 30 intersectional groups across age, gender, and ethnicity. Our study shows that (1) there are significant differences across demographic groups (including intersectional groups) on how severe they assess the harm to be, and that these differences vary across different types of safety violations, (2) the diverse rater pool captures annotation patterns that are substantially different from expert raters trained on specific set of safety policies, and (3) the differences we observe in T2I safety are distinct from previously documented group level differences in text-based safety tasks. To further understand these varying perspectives, we conduct a qualitative analysis of the open-ended explanations provided by raters. This analysis reveals core differences into the reasons why different groups perceive harms in T2I generations. Our findings underscore the critical need for incorporating diverse perspectives into safety evaluation of generative AI ensuring these systems are truly inclusive and reflect the values of all users.


A Comparative Study on Reasoning Patterns of OpenAI's o1 Model

arXiv.org Artificial Intelligence

Enabling Large Language Models (LLMs) to handle a wider range of complex tasks (e.g., coding, math) has drawn great attention from many researchers. As LLMs continue to evolve, merely increasing the number of model parameters yields diminishing performance improvements and heavy computational costs. Recently, OpenAI's o1 model has shown that inference strategies (i.e., Test-time Compute methods) can also significantly enhance the reasoning capabilities of LLMs. However, the mechanisms behind these methods are still unexplored. In our work, to investigate the reasoning patterns of o1, we compare o1 with existing Test-time Compute methods (BoN, Step-wise BoN, Agent Workflow, and Self-Refine) by using OpenAI's GPT-4o as a backbone on general reasoning benchmarks in three domains (i.e., math, coding, commonsense reasoning). Specifically, first, our experiments show that the o1 model has achieved the best performance on most datasets. Second, as for the methods of searching diverse responses (e.g., BoN), we find the reward models' capability and the search space both limit the upper boundary of these methods. Third, as for the methods that break the problem into many sub-problems, the Agent Workflow has achieved better performance than Step-wise BoN due to the domain-specific system prompt for planning better reasoning processes. Fourth, it is worth mentioning that we have summarized six reasoning patterns of o1, and provided a detailed analysis on several reasoning benchmarks.


An Eye for an AI: Evaluating GPT-4o's Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions

arXiv.org Artificial Intelligence

CG (Computer Graphics) is a popular field of CS (Computer Science), but many students find this topic difficult due to it requiring a large number of skills, such as mathematics, programming, geometric reasoning, and creativity. Over the past few years, researchers have investigated ways to harness the power of GenAI (Generative Artificial Intelligence) to improve teaching. In CS, much of the research has focused on introductory computing. A recent study evaluating the performance of an LLM (Large Language Model), GPT-4 (text-only), on CG questions, indicated poor performance and reliance on detailed descriptions of image content, which often required considerable insight from the user to return reasonable results. So far, no studies have investigated the abilities of LMMs (Large Multimodal Models), or multimodal LLMs, to solve CG questions and how these abilities can be used to improve teaching. In this study, we construct two datasets of CG questions requiring varying degrees of visual perception skills and geometric reasoning skills, and evaluate the current state-of-the-art LMM, GPT-4o, on the two datasets. We find that although GPT-4o exhibits great potential in solving questions with visual information independently, major limitations still exist to the accuracy and quality of the generated results. We propose several novel approaches for CG educators to incorporate GenAI into CG teaching despite these limitations. We hope that our guidelines further encourage learning and engagement in CG classrooms.


Hybrid Generative AI for De Novo Design of Co-Crystals with Enhanced Tabletability

arXiv.org Artificial Intelligence

Co-crystallization is an accessible way to control physicochemical characteristics of organic crystals, which finds many biomedical applications. In this work, we present Generative Method for Co-crystal Design (GEMCODE), a novel pipeline for automated co-crystal screening based on the hybridization of deep generative models and evolutionary optimization for broader exploration of the target chemical space. GEMCODE enables fast de novo co-crystal design with target tabletability profiles, which is crucial for the development of pharmaceuticals. With a series of experimental studies highlighting validation and discovery cases, we show that GEMCODE is effective even under realistic computational constraints. Furthermore, we explore the potential of language models in generating co-crystals. Finally, we present numerous previously unknown co-crystals predicted by GEMCODE and discuss its potential in accelerating drug development.


IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing

arXiv.org Artificial Intelligence

Unlike professional Business-to-Consumer (B2C) e-commerce platforms (e.g., Amazon), Consumer-to-Consumer (C2C) platforms (e.g., Facebook marketplace) are mainly targeting individual sellers who usually lack sufficient experience in e-commerce. Individual sellers often struggle to compose proper descriptions for selling products. With the recent advancement of Multimodal Large Language Models (MLLMs), we attempt to integrate such state-of-the-art generative AI technologies into the product listing process. To this end, we develop IPL, an Intelligent Product Listing tool tailored to generate descriptions using various product attributes such as category, brand, color, condition, etc. IPL enables users to compose product descriptions by merely uploading photos of the selling product. More importantly, it can imitate the content style of our C2C platform Xianyu. This is achieved by employing domain-specific instruction tuning on MLLMs and adopting the multi-modal Retrieval-Augmented Generation (RAG) process. A comprehensive empirical evaluation demonstrates that the underlying model of IPL significantly outperforms the base model in domain-specific tasks while producing less hallucination. IPL has been successfully deployed in our production system, where 72% of users have their published product listings based on the generated content, and those product listings are shown to have a quality score 5.6% higher than those without AI assistance.


ACPBench: Reasoning about Action, Change, and Planning

arXiv.org Artificial Intelligence

There is an increasing body of work using Large Language Models (LLMs) as agents for orchestrating workflows and making decisions in domains that require planning and multi-step reasoning. As a result, it is imperative to evaluate LLMs on core skills required for planning. In this work, we present ACPBench, a benchmark for evaluating the reasoning tasks in the field of planning. The benchmark consists of 7 reasoning tasks over 13 planning domains. The collection is constructed from planning domains described in a formal language. This allows us to synthesize problems with provably correct solutions across many tasks and domains. Further, it allows us the luxury of scale without additional human effort, i.e., many additional problems can be created automatically. Our extensive evaluation of 22 LLMs and OpenAI o1 reasoning models highlights the significant gap in the reasoning capability of the LLMs. Our findings with OpenAI o1, a multi-turn reasoning model, reveal significant gains in performance on multiple-choice questions, yet surprisingly, no notable progress is made on boolean questions. The ACPBench collection is available at https://ibm.github.io/ACPBench.


Rupert Murdoch's Dow Jones and New York Post sue AI firm for 'illegal copying'

The Guardian

"This suit is brought by news publishers who seek redress for Perplexity's brazen scheme to compete for readers while simultaneously freeriding on the valuable content the publishers produce," according to the lawsuit filed in the southern district of New York by the Wall Street Journal parent Dow Jones and the New York Post. Perplexity did not immediately respond to emails from Reuters seeking comment. The AI company is among the leading startups attempting to uproot the search engine market dominated by Alphabet's Google. It assembles information from webpages it deems to be authoritative, then provides a summary directly within Perplexity's own tool. Perplexity uses a variety of large language models (LLMs) to generate its summaries, from OpenAI to Meta's open-source model Llama.


TikTok owner sacks intern for allegedly sabotaging AI project

The Guardian

The owner of TikTok has sacked an intern for allegedly sabotaging an internal artificial intelligence project. ByteDance said it had dismissed the person in August after they "maliciously interfered" with the training of artificial intelligence (AI) models used in a research project. Thanks to the video-sharing app TikTok and its Chinese counterpart, Douyin, which rank among the world's most popular mobile apps, ByteDance has risen to become one of the world's most important social media companies. Like other big players in the tech sector, ByteDance has raced to embrace generative AI. Its Doubao chatbot earlier this year took over from the competitor Baidu's Ernie in the race to produce a Chinese rival to OpenAI's ChatGPT.