Goto

Collaborating Authors

 Generative AI


Multi-objective generative AI for designing novel brain-targeting small molecules

arXiv.org Artificial Intelligence

The strict selectivity of the blood-brain barrier (BBB) represents one of the most formidable challenges to successful central nervous system (CNS) drug delivery. Computational methods to generate BBB permeable drugs in silico may be valuable tools in the CNS drug design pipeline. However, in real-world applications, BBB penetration alone is insufficient; rather, after transiting the BBB, molecules must bind to a specific target or receptor in the brain and must also be safe and non-toxic. To discover small molecules that concurrently satisfy these constraints, we use multi-objective generative AI to synthesize drug-like BBB-permeable small molecules. Specifically, we computationally synthesize molecules with predicted binding affinity against dopamine receptor D2, the primary target for many clinically effective antipsychotic drugs. After training several graph neural network-based property predictors, we adapt SyntheMol (Swanson et al., 2024), a recently developed Monte Carlo Tree Search-based algorithm for antibiotic design, to perform a multi-objective guided traversal over an easily synthesizable molecular space. We design a library of 26,581 novel and diverse small molecules containing hits with high predicted BBB permeability and favorable predicted safety and toxicity profiles, and that could readily be synthesized for experimental validation in the wet lab. We also validate top scoring molecules with molecular docking simulation against the D2 receptor and demonstrate predicted binding affinity on par with risperidone, a clinically prescribed D2-targeting antipsychotic. In the future, the SyntheMol-based computational approach described here may enable the discovery of novel neurotherapeutics for currently intractable disorders of the CNS.


E3: Ensemble of Expert Embedders for Adapting Synthetic Image Detectors to New Generators Using Limited Data

arXiv.org Artificial Intelligence

As generative AI progresses rapidly, new synthetic image generators continue to emerge at a swift pace. Traditional detection methods face two main challenges in adapting to these generators: the forensic traces of synthetic images from new techniques can vastly differ from those learned during training, and access to data for these new generators is often limited. To address these issues, we introduce the Ensemble of Expert Embedders (E3), a novel continual learning framework for updating synthetic image detectors. E3 enables the accurate detection of images from newly emerged generators using minimal training data. Our approach does this by first employing transfer learning to develop a suite of expert embedders, each specializing in the forensic traces of a specific generator. Then, all embeddings are jointly analyzed by an Expert Knowledge Fusion Network to produce accurate and reliable detection decisions. Our experiments demonstrate that E3 outperforms existing continual learning methods, including those developed specifically for synthetic image detection.


The Evolution of Learning: Assessing the Transformative Impact of Generative AI on Higher Education

arXiv.org Artificial Intelligence

Generative Artificial Intelligence (GAI) models such as ChatGPT have experienced a surge in popularity, attracting 100 million active users in 2 months and generating an estimated 10 million daily queries. Despite this remarkable adoption, there remains a limited understanding to which extent this innovative technology influences higher education. This research paper investigates the impact of GAI on university students and Higher Education Institutions (HEIs). The study adopts a mixed-methods approach, combining a comprehensive survey with scenario analysis to explore potential benefits, drawbacks, and transformative changes the new technology brings. Using an online survey with 130 participants we assessed students' perspectives and attitudes concerning present ChatGPT usage in academics. Results show that students use the current technology for tasks like assignment writing and exam preparation and believe it to be a effective help in achieving academic goals. The scenario analysis afterwards projected potential future scenarios, providing valuable insights into the possibilities and challenges associated with incorporating GAI into higher education. The main motivation is to gain a tangible and precise understanding of the potential consequences for HEIs and to provide guidance responding to the evolving learning environment. The findings indicate that irresponsible and excessive use of the technology could result in significant challenges. Hence, HEIs must develop stringent policies, reevaluate learning objectives, upskill their lecturers, adjust the curriculum and reconsider examination approaches.


Shaping Realities: Enhancing 3D Generative AI with Fabrication Constraints

arXiv.org Artificial Intelligence

Generative AI tools are becoming more prevalent in 3D modeling, enabling users to manipulate or create new models with text or images as inputs. This makes it easier for users to rapidly customize and iterate on their 3D designs and explore new creative ideas. These methods focus on the aesthetic quality of the 3D models, refining them to look similar to the prompts provided by the user. However, when creating 3D models intended for fabrication, designers need to trade-off the aesthetic qualities of a 3D model with their intended physical properties. To be functional post-fabrication, 3D models have to satisfy structural constraints informed by physical principles. Currently, such requirements are not enforced by generative AI tools. This leads to the development of aesthetically appealing, but potentially non-functional 3D geometry, that would be hard to fabricate and use in the real world. This workshop paper highlights the limitations of generative AI tools in translating digital creations into the physical world and proposes new augmentations to generative AI tools for creating physically viable 3D models. We advocate for the development of tools that manipulate or generate 3D models by considering not only the aesthetic appearance but also using physical properties as constraints. This exploration seeks to bridge the gap between digital creativity and real-world applicability, extending the creative potential of generative AI into the tangible domain.


AI video is heading to Adobe Premiere Pro

PCWorld

Adobe is ushering in the next generation of AI art with an upcoming version of Premiere Pro. It's been about two years since Midjourney ushered in AI art, consisting of art generated entirely from scratch as well as "inpainting" and "outpainting." Outpainting attracted attention because AI art was being used to essentially extend the boundaries of photographs and paintings, creating a plausible addition to what wasn't there. Now Adobe is doing the same with Premiere Pro. On Monday, Adobe showed off a video version of what it calls Generative Fill, the same technique that it uses for Adobe Photoshop and its Adobe Firefly generative AI art.


Adobe previews AI object addition and removal for Premiere Pro

Engadget

Last year Adobe launched Firefly, its latest generative AI model building on its previous SenseiAI, and now the company is showing how it'll be used its video editing app, Premiere Pro. In an early sneak, it demonstrated a few key features arriving later this year, including Object Addition & Removal, Generative Extend and Text to Video. The new features will likely be popular, as video cleanup is one a common (and painful) task. The first feature, Generative Extend, addresses a problem editors face on nearly every edit: clips that are too short. "Seamlessly add frames to make clips longer, so it's easier to perfectly time edits and add smooth transitions," Adobe states.


AI creates Japan ruling party's new poster slogan

The Japan Times

The ruling Liberal Democratic Party on Monday unveiled its first poster featuring a catchphrase created using generative artificial intelligence. The slogan, written in red on a white background, pledges to the public a real feeling of economic revitalization, amid Prime Minister and LDP President Fumio Kishida's drive to raise wages to fuel economic growth. Generative AI tools, including ChatGPT, studied Kishida's remarks and party policy documents over the past three years to draw up drafts, according to people familiar with the matter. The AI-crafted slogan was chosen after LDP executives screened more than 500 candidate phrases, including ones proposed by copywriters. "This doesn't mean at all that an election will be called soon," Takuya Hirai, chair of the LDP's Public Relations Headquarters, told a news conference, referring to speculation that Kishida will call a snap general election as early as June.


TEL'M: Test and Evaluation of Language Models

arXiv.org Artificial Intelligence

It is assumed that readers are already familiar with Language Models of various flavors such as: Transformer-based Language Models (currently the most promising and studied LMs) [78]; Multimodal Foundation Models such as Blip-2 [48] and CLIP [61]; Auto-regressive Language Models [15, 51]; Recurrent Neural Network Language Models [75]; State space language models [40]; Hybrid Models [24] as well as the current and proposed use cases and the various technologies underlying them [1, 65, 70]. There is growing interest in LM performance and benchmarks [13, 16, 18,46, 47, 64,72, 74, 80] with recent acknowledgement that this is a hard problem [53]. Many suggestions are proposed in the commercial literature [17] and a large number of benchmark-based methods have surfaced (Big Bench [67], GLUE Benchmark, SuperGLUE Benchmark, OpenAI Moderation API, MMLU, EleutherAI LM Eval, OpenAI Evals Adversarial NLI, LIT, ParlAI, CoQA, LAMBADA, HellaSwag, LogiQA, MultiNLI, SQUAD to name a few). A review of existing approaches demonstrates that they are not quantitative or rigorous enough to past muster with respect to accepted testing requirements [3, 55]. In particular, existing use of benchmarks do not investigate the extent to which a benchmark can predict or quantify certain properties on future prompts (that is, statistical soundness of any conclusions) and do not identify factors affecting performance dependence as would be possible with more rigorous experimental design and test execution. LMs can be black box, gray box or white box according to the visibility into the architecture and training data used to create an LM (see Table 1). Remote Black Box LMs typically throttle the number of prompts so sustained access for testing could be difficult unless priority access to an API is given. For example, ChatGPT limits users to a small number of free prompts but allows unlimited prompts on its subscription option. Additionally, reproducability may not be guaranteed because of randomness in the response generation and/or continuous adaptation of the LM platform.


HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

arXiv.org Artificial Intelligence

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. The project page is https://thefllood.github.io/HQEdit_web.


Evaluating Text-to-Image Synthesis: Survey and Taxonomy of Image Quality Metrics

arXiv.org Artificial Intelligence

Recent advances in text-to-image synthesis enabled through a combination of language and vision foundation models have led to a proliferation of the tools available and an increased attention to the field. When conducting text-to-image synthesis, a central goal is to ensure that the content between text and image is aligned. As such, there exist numerous evaluation metrics that aim to mimic human judgement. However, it is often unclear which metric to use for evaluating text-to-image synthesis systems as their evaluation is highly nuanced. In this work, we provide a comprehensive overview of existing text-to-image evaluation metrics. Based on our findings, we propose a new taxonomy for categorizing these metrics. Our taxonomy is grounded in the assumption that there are two main quality criteria, namely compositionality and generality, which ideally map to human preferences. Ultimately, we derive guidelines for practitioners conducting text-to-image evaluation, discuss open challenges of evaluation mechanisms, and surface limitations of current metrics.