Goto

Collaborating Authors

 Generative AI


How Sam Altman sidestepped Elon Musk to win over Donald Trump

The Japan Times

At President Donald Trump's inauguration, Sam Altman, CEO of OpenAI, was relegated to the overflow room while other tech billionaires like Elon Musk and Mark Zuckerberg took prime spots on the dais under the Capitol rotunda. But days earlier, before flying to Washington, Altman was on the phone with Trump, preparing an announcement that would outflank Musk and put Altman's company at the center of the new administration's agenda for artificial intelligence. On the 25-minute call, Altman appealed to Trump's love of a big story and of a big deal. He told the president-elect that the tech industry would achieve artificial general intelligence -- the hypothetical moment when technology matches human intelligence -- during the Trump administration, according to three people familiar with the call. And to get there before Chinese competitors, OpenAI, Oracle and SoftBank had completed a 100 billion deal to build data centers across the country.


PyPotteryInk: One-Step Diffusion Model for Sketch to Publication-ready Archaeological Drawings

arXiv.org Artificial Intelligence

Archaeological ceramics are a valuable source of information for reconstructing the customs, exchanges and social relationships of ancient populations, as well as for dating archaeological contexts (Sinopoli 1991; Peroni 1994; Steiner and Allason-Jones 2005; Vidale 2007; Orton and Hughes 2013; Hunt 2016). However, in order to turn a ceramic fragment into a rich source of scientific information, a long process of study and elaboration is required: once recovered in an excavation, the ceramic fragment is washed, catalogued, drawn and made ready for publication through the preparation of tables and figures that allow its correct interpretation and comparison with other archaeological contexts. Archaeological drawing is a fundamental and well-established tool in archaeological practice, and new technologies and methods are emerging to automate, standardise and speed up this process as much as possible. An example of this is the LAD (Laser Aided Profiler - Demjรกn, Pavรบk, and Roosevelt 2023), a tool that allows ceramic fragments to be'drawn' quickly and accurately using a laser beam. Over time, however, many drawings were made by hand using traditional tools such as pencils and then had to be'inked' and made ready for publication. Traditionally, this post-process was done by hand with Indian ink, and nowadays digital drawing programmes are used. This process is however extremely time-consuming and can often discourage the publication of new contexts due to the difficulties in terms of time and resources needed for inking. Generative AI can help to achieve this task, using complex image translation operation. Today, AI is permeating business, creativity and everyday life (Elliott 2019; Le et al. 2020; Varghese, Raj, and Venkatesh 2022; Azatbekova


Fact-or-Fair: A Checklist for Behavioral Testing of AI Models on Fairness-Related Queries

arXiv.org Artificial Intelligence

The generation of incorrect images, such as depictions of people of color in Nazi-era uniforms by Gemini, frustrated users and harmed Google's reputation, motivating us to investigate the relationship between accurately reflecting factuality and promoting diversity and equity. In this study, we focus on 19 real-world statistics collected from authoritative sources. Using these statistics, we develop a checklist comprising objective and subjective queries to analyze behavior of large language models (LLMs) and text-to-image (T2I) models. Objective queries assess the models' ability to provide accurate world knowledge. In contrast, the design of subjective queries follows a key principle: statistical or experiential priors should not be overgeneralized to individuals, ensuring that models uphold diversity. These subjective queries are based on three common human cognitive errors that often result in social biases. We propose metrics to assess factuality and fairness, and formally prove the inherent trade-off between these two aspects. Results show that GPT-4o and DALL-E 3 perform notably well among six LLMs and four T2I models. Our code is publicly available at https://github.com/uclanlp/Fact-or-Fair.


Generating 3D Binding Molecules Using Shape-Conditioned Diffusion Models with Guidance

arXiv.org Artificial Intelligence

Drug development is a critical but notoriously resource- and time-consuming process. In this manuscript, we develop a novel generative artificial intelligence (genAI) method DiffSMol to facilitate drug development. DiffSmol generates 3D binding molecules based on the shapes of known ligands. DiffSMol encapsulates geometric details of ligand shapes within pre-trained, expressive shape embeddings and then generates new binding molecules through a diffusion model. DiffSMol further modifies the generated 3D structures iteratively via shape guidance to better resemble the ligand shapes. It also tailors the generated molecules toward optimal binding affinities under the guidance of protein pockets. Here, we show that DiffSMol outperforms the state-of-the-art methods on benchmark datasets. When generating binding molecules resembling ligand shapes, DiffSMol with shape guidance achieves a success rate 61.4%, substantially outperforming the best baseline (11.2%), meanwhile producing molecules with novel molecular graph structures. DiffSMol with pocket guidance also outperforms the best baseline in binding affinities by 13.2%, and even by 17.7% when combined with shape guidance. Case studies for two critical drug targets demonstrate very favorable physicochemical and pharmacokinetic properties of the generated molecules, thus, the potential of DiffSMol in developing promising drug candidates.


Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

arXiv.org Artificial Intelligence

While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not align with the given visual input, which restricts their practical applicability in real-world scenarios. In this work, inspired by the observation that the text-to-image generation process is the inverse of image-conditioned response generation in LVLMs, we explore the potential of leveraging text-to-image generative models to assist in mitigating hallucinations in LVLMs. We discover that generative models can offer valuable self-feedback for mitigating hallucinations at both the response and token levels. Building on this insight, we introduce self-correcting Decoding with Generative Feedback (DeGF), a novel training-free algorithm that incorporates feedback from text-to-image generative models into the decoding process to effectively mitigate hallucinations in LVLMs. Specifically, DeGF generates an image from the initial response produced by LVLMs, which acts as an auxiliary visual reference and provides self-feedback to verify and correct the initial response through complementary or contrastive decoding. Extensive experimental results validate the effectiveness of our approach in mitigating diverse types of hallucinations, consistently surpassing state-of-the-art methods across six benchmarks. Code is available at https://github.com/zhangce01/DeGF.


Big Tech whistleblower's parents sue, sounding alarm over son's unexpected death

FOX News

If you or someone you know is having thoughts of suicide, please contact the Suicide & Crisis Lifeline at 988 or 1-800-273-TALK (8255). The parents of a young California tech whistleblower whose 2024 death was ruled a suicide are now suing the City and County of San Francisco, alleging they violated public records laws by refusing to fulfill their requests for information about their son's death. Suchir Balaji, 26, was an employee at OpenAI, the artificial intelligence company behind ChatGPT, at the time of his Nov. 26, 2024, death. A San Francisco County medical examiner concluded the next day he died from a self-inflicted gunshot wound inside his apartment. "In the two-plus months since their son's passing, Petitioners and their counsel have been stymied at every turn as they have sought more information about the cause of and circumstances surrounding Suchir's tragic death. This petition, they hope, is the beginning of the end of that obstruction," the lawsuit states.


20 million OpenAI users hacked? Here's how to stay safe, just in case

PCWorld

Have you ever tried ChatGPT? You may want to take a quick moment to freshen up your account's security. A Russian hacker is claiming to have login data for over 20 million OpenAI users--and the information includes email addresses and passwords. On Friday, samples of OpenAI logins emerged on the dark web, along with an offer to sell the full trove of data. Currently, OpenAI says it has not yet found evidence of compromised systems (as per The Independent).


Neural Genetic Search in Discrete Spaces

arXiv.org Artificial Intelligence

Effective search methods are crucial for improving the performance of deep generative models at test time. In this paper, we introduce a novel test-time search method, Neural Genetic Search (NGS), which incorporates the evolutionary mechanism of genetic algorithms into the generation procedure of deep models. The core idea behind NGS is its crossover, which is defined as parent-conditioned generation using trained generative models. This approach offers a versatile and easy-to-implement search algorithm for deep generative models. We demonstrate the effectiveness and flexibility of NGS through experiments across three distinct domains: routing problems, adversarial prompt generation for language models, and molecular design.


2025: The Year of the AI App

WIRED

What a great idea I had for the first Plaintext of 2025. After following the frantic competition between OpenAI, Google, Meta, and Anthropic to churn out brainier and deeper "frontier" foundation models, I settled on a thesis about what's ahead: In the new year, those mighty trailblazers will consume billions of dollars, countless gigawatts, and all the silicon Nvidia can muster in their pursuit of AGI. We'll be bombarded by press releases boasting advanced reasoning, more tokens, and maybe even assurances that their models won't make up crazy facts. But people are tired of hearing about how AI is transformational and seeing few transformations to their day-to-day existence. Getting an AI summary of Google search results or having Facebook ask if you want to pose a follow-up question on a post doesn't make you a traveler to the neo-human future.


Review for NeurIPS paper: Accelerating Reinforcement Learning through GPU Atari Emulation

Neural Information Processing Systems

Weaknesses: My main concern is that results seem to be contradictory to what the authors claimed as the benefit of leveraging GPU accelerations. Specifically, in the "impact statement" the authors described CuLE can "provide access to an accelerated training environment to researchers with limited computational capabilities," but the results show the acceleration won't take into effect unless you use more computation---Figure 2, CuLE runs slower than OpenAI when using a fewer number of environments. If someone can only afford to run 100 environments, would this mean CuLE is not useful here? The limitation of the memory has been noted in the paper which is good. I was confused when looking at Table 3. First, why is there no 120 envs experiment for CuLE?