Goto

Collaborating Authors

 Generative AI


Uncertain Boundaries: Multidisciplinary Approaches to Copyright Issues in Generative AI

arXiv.org Artificial Intelligence

In the rapidly evolving landscape of generative artificial intelligence (AI), the increasingly pertinent issue of copyright infringement arises as AI advances to generate content from scraped copyrighted data, prompting questions about ownership and protection that impact professionals across various careers. With this in mind, this survey provides an extensive examination of copyright infringement as it pertains to generative AI, aiming to stay abreast of the latest developments and open problems. Specifically, it will first outline methods of detecting copyright infringement in mediums such as text, image, and video. Next, it will delve an exploration of existing techniques aimed at safeguarding copyrighted works from generative models. Furthermore, this survey will discuss resources and tools for users to evaluate copyright violations. Finally, insights into ongoing regulations and proposals for AI will be explored and compared. Through combining these disciplines, the implications of AI-driven content and copyright are thoroughly illustrated and brought into question.


OpenAI Can Re-Create Human Voices--but Won't Release the Tech Yet

WIRED

Voice synthesis has come a long way since 1978's Speak & Spell toy, which once wowed people with its state-of-the-art ability to read words aloud using an electronic voice. Now, using deep-learning AI models, software can create not only realistic-sounding voices but can also convincingly imitate existing voices using small samples of audio. Along those lines, OpenAI this week announced Voice Engine, a text-to-speech AI model for creating synthetic voices based on a 15-second segment of recorded audio. It has provided audio samples of the Voice Engine in action on its website. This story originally appeared on Ars Technica, a trusted source for technology news, tech policy analysis, reviews, and more.


Microsoft Copilot has reportedly been blocked on all Congress-owned devices

Engadget

The publication said it obtained a memo from House Chief Administrative Officer Catherine Szpindor, telling Congress personnel that the AI chatbot is now officially prohibited. Apparently, the Office of Cybersecurity has deemed Copilot to be a risk "due to the threat of leaking House data to non-House approved cloud services." While there's nothing stopping them from using Copilot on their own phones and laptops, it will now be blocked on all Windows devices owned by the Congress. Almost a year ago, the Congress also set a strict limit on the use of ChatGPT, which is powered by OpenAI's large language models, just like Copilot. It banned staffers from using the chatbot's free version on House computers, but it allowed them to continue using the paid (ChatGPT Plus) version for research and evaluation due to its tighter privacy controls.


OpenAI previews new audio tool that can read text and mimic voices

The Japan Times

OpenAI is sharing early results from a test for a feature that can read words aloud in a convincing human voice -- highlighting a new frontier for artificial intelligence and raising the specter of deepfake risks. The company is sharing early demos and use cases from a small-scale preview of the text-to-speech model, called Voice Engine, which it has shared with about 10 developers so far, a spokesperson said. OpenAI decided against a wider rollout of the feature, which it briefed reporters on earlier this month. A spokesperson for OpenAI said the company decided to scale back the release after receiving feedback from stakeholders such as policymakers, industry experts, educators and creatives. The company had initially planned to release the tool to as many as 100 developers through an application process, according to the earlier press briefing.


MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions

arXiv.org Artificial Intelligence

The increasing significance of large models and their multi-modal variants in societal information processing has ignited debates on social safety and ethics. However, there exists a paucity of comprehensive analysis for: (i) the interactions between human and artificial intelligence systems, and (ii) understanding and addressing the associated limitations. To bridge this gap, we propose Model Autophagy Analysis (MONAL) for large models' self-consumption explanation. MONAL employs two distinct autophagous loops (referred to as ``self-consumption loops'') to elucidate the suppression of human-generated information in the exchange between human and AI systems. Through comprehensive experiments on diverse datasets, we evaluate the capacities of generated models as both creators and disseminators of information. Our key findings reveal (i) A progressive prevalence of model-generated synthetic information over time within training datasets compared to human-generated information; (ii) The discernible tendency of large models, when acting as information transmitters across multiple iterations, to selectively modify or prioritize specific contents; and (iii) The potential for a reduction in the diversity of socially or human-generated information, leading to bottlenecks in the performance enhancement of large models and confining them to local optima.


Generative AI for Architectural Design: A Literature Review

arXiv.org Artificial Intelligence

Generative Artificial Intelligence (AI) has pioneered new methodological paradigms in architectural design, significantly expanding the innovative potential and efficiency of the design process. This paper explores the extensive applications of generative AI technologies in architectural design, a trend that has benefited from the rapid development of deep generative models. This article provides a comprehensive review of the basic principles of generative AI and large-scale models and highlights the applications in the generation of 2D images, videos, and 3D models. In addition, by reviewing the latest literature from 2020, this paper scrutinizes the impact of generative AI technologies at different stages of architectural design, from generating initial architectural 3D forms to producing final architectural imagery. The marked trend of research growth indicates an increasing inclination within the architectural design community towards embracing generative AI, thereby catalyzing a shared enthusiasm for research. These research cases and methodologies have not only proven to enhance efficiency and innovation significantly but have also posed challenges to the conventional boundaries of architectural creativity. Finally, we point out new directions for design innovation and articulate fresh trajectories for applying generative AI in the architectural domain. This article provides the first comprehensive literature review about generative AI for architectural design, and we believe this work can facilitate more research work on this significant topic in architecture.


OpenAI says it can clone a voice from just 15 seconds of audio

Engadget

OpenAI just announced that it recently conducted a small-scale preview of a new tool called Voice Engine. This is a voice cloning technology that can mimic any speaker by analyzing a 15-second audio sample. The company says it generates "natural-sounding speech" with "emotive and realistic voices." The technology is based on the company's pre-existing text-to-speech API and it has been in the works since 2022. OpenAI has already been using a version of the toolset to power the preset voices available in the current text-to-speech API and the Read Aloud feature. There are a bunch of samples on the company's official blog and they sound eerily close to the real thing.


A conversation with OpenAI's first artist in residence

MIT Technology Review

Officially, the appointment started in January and lasts three months. But Reben's relationship with the San Francisco–based AI firm seems casual: "It's a little fuzzy, because I'm the first, and we're figuring stuff out. I'm probably going to keep working with them." In fact, Reben has been working with OpenAI for years already. Five years ago, he was invited to try out an early version of GPT-3 before it was released to the public. "I got to play around with that quite a bit and made a few artworks," he says.


Could OpenAI's Sora text-to-video generator kill off jobs in Hollywood?

Al Jazeera

Artificial intelligence startup OpenAI has been teasing its new AI video generator, Sora, on social media in recent weeks. Last week, it revealed that it had also given actors and directors in Hollywood a first look at the technology – and a chance to try it out – before Sora is launched publicly. OpenAI published a blog post on March 24 titled Sora's First Impressions, showcasing the work that several creative studios and directors had produced using the video generator. Some media experts speculate that Sora will be extremely disruptive to the film creative industry. Al Jazeera spoke to one executive who works in Hollywood, who asked us not to reveal his identity due to the sensitive nature of the subject.


China turns to AI in propaganda mocking the 'American Dream'

Al Jazeera

They say it's for all, but is it really?" So begins a 65-second, AI-generated animated video that touches on hot-button issues in the United States ranging from drug addiction and imprisonment rates to growing wealth inequality. As storm clouds gather over an urban landscape resembling New York City, the words "AMERICAN DREAM" hang in a darkening sky as the video ends. The message is clear: Despite its promises of a better life for all, the United States is in terminal decline. The video, titled American Dream or American Mirage, is one of a number of segments aired by Chinese state broadcaster CGTN – and shared far and wide on social media – as part of its A Fractured America animated series. Other videos in the series contain similar titles that invoke images of a dystopian society, such as American workers in tumult: A result of unbalanced politics and economy, and Unmasking the real threat: America's military-industrial complex. CGTN and the Chinese embassy in Washington, DC did not respond to requests for comment. The Fractured America series is just one example of how artificial intelligence (AI), with its ability to generate high-quality multimedia with minimal effort in seconds, is beginning to shape Beijing's propaganda efforts to undermine the United States' standing in the world. Henry Ajder, a UK-based expert in generative AI, said while the CGTN series does not attempt to pass itself off as genuine video, it is a clear example of how AI has made it far easier and cheaper to churn out content. "The reason that they've done it in this way is, you could hire an animator, and a voiceover artist to do this, but it would probably end up being more time-consuming.