Goto

Collaborating Authors

 Generative AI


Free to play: UN Trade and Development's experience with developing its own open-source Retrieval Augmented Generation Large Language Model application

arXiv.org Artificial Intelligence

Generative artificial intelligence (AI), and in particular Large Language Models (LLMs), have exploded in popularity and attention since the release to the public of ChatGPT's Generative Pre-trained Transformer (GPT)-3.5 model in November of 2022. Due to the power of these general purpose models and their ability to communicate in natural language, they can be useful in a range of domains, including the work of official statistics and international organizations. However, with such a novel and seemingly complex technology, it can feel as if generative AI is something that happens to an organization, something that can be talked about but not understood, that can be commented on but not contributed to. Additionally, the costs of adoption and operation of proprietary solutions can be both uncertain and high, a barrier for often cost-constrained international organizations. In the face of these challenges, United Nations Trade and Development (UNCTAD), through its Global Crisis Response Group (GCRG), has explored and developed its own open-source Retrieval Augmented Generation (RAG) LLM application. RAG makes LLMs aware of and more useful for the organization's domain and work. Developing in-house solutions comes with pros and cons, with pros including cost, flexibility, and fostering institutional knowledge. Cons include time and skill investments and gaps and application polish and power. The three libraries developed to produce the app, nlp_pipeline for document processing and statistical analysis, local_rag_llm for running a local RAG LLM, and streamlit_rag for the user interface, are publicly available on PyPI and GitHub with Dockerfiles. A fourth library, local_llm_finetune, is also available for fine-tuning existing LLMs which can then be used in the application.


Retrieval-Augmented Generation for Generative Artificial Intelligence in Medicine

arXiv.org Artificial Intelligence

Generative artificial intelligence (AI) has brought revolutionary innovations in various fields, including medicine. However, it also exhibits limitations. In response, retrieval-augmented generation (RAG) provides a potential solution, enabling models to generate more accurate contents by leveraging the retrieval of external knowledge. With the rapid advancement of generative AI, RAG can pave the way for connecting this transformative technology with medical applications and is expected to bring innovations in equity, reliability, and personalization to health care.


Evaluating Text-to-Visual Generation with Image-to-Text Generation

arXiv.org Artificial Intelligence

Despite significant progress in generative AI, comprehensive evaluation remains challenging because of the lack of effective metrics and standardized benchmarks. For instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce reliable scores for complex prompts involving compositions of objects, attributes, and relations. One reason is that text encoders of CLIP can notoriously act as a "bag of words", conflating prompts such as "the horse is eating the grass" with "the grass is eating the horse". To address this, we introduce the VQAScore, which uses a visual-question-answering (VQA) model to produce an alignment score by computing the probability of a "Yes" answer to a simple "Does this figure show '{text}'?" question. Though simpler than prior art, VQAScore computed with off-the-shelf models produces state-of-the-art results across many (8) image-text alignment benchmarks. We also compute VQAScore with an in-house model that follows best practices in the literature. For example, we use a bidirectional image-question encoder that allows image embeddings to depend on the question being asked (and vice versa). Our in-house model, CLIP-FlanT5, outperforms even the strongest baselines that make use of the proprietary GPT-4V. Interestingly, although we train with only images, VQAScore can also align text with video and 3D models. VQAScore allows researchers to benchmark text-to-visual generation using complex texts that capture the compositional structure of real-world prompts. We introduce GenAI-Bench, a more challenging benchmark with 1,600 compositional text prompts that require parsing scenes, objects, attributes, relationships, and high-order reasoning like comparison and logic. GenAI-Bench also offers over 15,000 human ratings for leading image and video generation models such as Stable Diffusion, DALL-E 3, and Gen2.


Generative Artificial Intelligence-Guided User Studies: An Application for Air Taxi Services

arXiv.org Artificial Intelligence

User studies are crucial for meeting user needs. In user studies, real experimental scenarios and participants are constructed and recruited. However, emerging and unfamiliar studies face limitations, including safety concerns and iterative efficiency. To address these challenges, this study utilizes a large language model (LLM) to create generative AI virtual scenarios for user experience. By recruiting real users to evaluate this experience, we can collect feedback that enables rapid iteration in the early design phase. The air taxi is particularly representative of these challenges and has been chosen as the case study for this research. The key contribution was designing a virtual ATJ using OpenAI's GPT-4 model and AI image and video generators. Based on the LLM-generated scripts, key visuals were created for the air taxi, and the ATJ was evaluated by 72 participants. Furthermore, the LLM demonstrated the ability to identify and suggest environments that significantly improve participants' attitudes toward air taxis. Education level and gender significantly influenced participants' attitudes and their satisfaction with the ATJ. Our study confirms the capability of generative AI to support user studies, providing a feasible approach and valuable insights for designing air taxi user experiences in the early design phase.


Extracting Training Data from Unconditional Diffusion Models

arXiv.org Artificial Intelligence

As diffusion probabilistic models (DPMs) are being employed as mainstream models for generative artificial intelligence (AI), the study of their memorization of the raw training data has attracted growing attention. Existing works in this direction aim to establish an understanding of whether or to what extent DPMs learn by memorization. Such an understanding is crucial for identifying potential risks of data leakage and copyright infringement in diffusion models and, more importantly, for more controllable generation and trustworthy application of Artificial Intelligence Generated Content (AIGC). While previous works have made important observations of when DPMs are prone to memorization, these findings are mostly empirical, and the developed data extraction methods only work for conditional diffusion models. In this work, we aim to establish a theoretical understanding of memorization in DPMs with 1) a memorization metric for theoretical analysis, 2) an analysis of conditional memorization with informative and random labels, and 3) two better evaluation metrics for measuring memorization. Based on the theoretical analysis, we further propose a novel data extraction method called \textbf{Surrogate condItional Data Extraction (SIDE)} that leverages a classifier trained on generated data as a surrogate condition to extract training data directly from unconditional diffusion models. Our empirical results demonstrate that SIDE can extract training data from diffusion models where previous methods fail, and it is on average over 50\% more effective across different scales of the CelebA dataset.


OpenAI-Backed Nonprofits Have Gone Back on Their Transparency Pledges

WIRED

A Sam Altmanโ€“funded nonprofit studying the effects of giving monthly checks of up to 1,000 to lower-income households in the US espouses transparency in its operations. "We aim to share data, findings, and insights widely," OpenResearch says on its website, which describes its work as a "public good." But like at least two other Altman-linked organizations--OpenAI and UBI Charitable--OpenResearch has decided to withhold information about its finances and governance. In several years of filings to US tax authorities since their founding, each of the organizations has answered a question about their voluntary disclosure of financial statements, governing documents, and conflict-of-interest policies by stating that the public can review them upon request. It remains unclear whether anyone took them up on the offer in those years.


Why Microsoft, OpenAI and Nvidia are facing anti-monopoly probes

Al Jazeera

The United States Department of Justice and the Federal Trade Commission (FTC) have reportedly reached a deal on how they will pursue an antitrust investigation into tech giants Microsoft, Nvidia, and Open AI. The companies are all major players in generative AI: OpenAI is the nonprofit startup behind ChatGPT, the blockbuster AI-powered chatbot. Microsoft, the world's largest company by market capitalisation, has invested more than 13bn in OpenAI and holds a 49 percent stake in the company's for-profit subsidiary. Chipmaker Nvidia is a global leader in graphic processing units (GPU), a key piece of hardware needed in AI. The company recently hit a 3 trillion valuation, surpassing Apple to become the world's second-largest company.


Microsoft's Japan chief sees country accelerating its use of AI

The Japan Times

Japan has been one of the fastest countries to embrace the use of new artificial intelligence tools and has the potential to accelerate its economy and tech sector by going further, according to Microsoft Japan President Miki Tsusaka. The country's digitalization push got a boost during the pandemic as businesses adapted to new work-from-home arrangements, and Tsusaka believes Japan has made up lost ground after previously being a laggard. "The Japanese have caught up. And I think it will continue to accelerate at this point because the technology enables things that we haven't been able to do," Tsusaka said in an interview. "We don't have enough people, our population is aging, and yet generative AI has the power to accelerate growth."


Estimating the Increase in Emissions caused by AI-augmented Search

arXiv.org Artificial Intelligence

Abstract--AI-generated answers to conventional search queries dramatically increase the energy consumption. This is a based on an updated estimate of energy consumption for conventional search and recent work on the energy demand of queries to the BLOOM model, a 176B parameter model, and OpenAI's GPT -3, which is of similar complexity. The new trend in search engines, to provide an AI-generated answer to the search query, has a considerable impact on the energy consumption and therefore CO2 emissions per query. To illustrate the impact of AI augmented search queries more clearly, I compare the energy consumption and emission of a query to Google's BLOOM model with that of a conventional Google search-style query. If all search queries are replac ed by AI-augmented queries, what does that mean for energy consumption and emissions?


Development of an Adaptive Multi-Domain Artificial Intelligence System Built using Machine Learning and Expert Systems Technologies

arXiv.org Artificial Intelligence

Producing an artificial general intelligence (AGI) has been an elusive goal in artificial intelligence (AI) research for some time. An AGI would have the capability, like a human, to be exposed to a new problem domain, learn about it and then use reasoning processes to make decisions. While AI techniques have been used across a wide variety of problem domains, an AGI would require an AI that could reason beyond its programming and training. This paper presents a small step towards producing an AGI. It describes a mechanism for an AI to learn about and develop reasoning pathways to make decisions in an a priori unknown domain. It combines a classical AI technique, the expert system, with a its modern adaptation - the gradient descent trained expert system (GDTES) - and utilizes generative artificial intelligence (GAI) to create a network and training data set for this system. These can be created from available sources or may draw upon knowledge incorporated in a GAI's own pre-trained model. The learning process in GDTES is used to optimize the AI's decision-making. While this approach does not meet the standards that many have defined for an AGI, it provides a somewhat similar capability, albeit one which requires a learning process before use.