Generative AI
Generative AI Training and Copyright Law
Dornis, Tim W., Stober, Sebastian
Training generative AI models requires extensive amounts of data. A common practice is to collect such data through web scraping. Yet, much of what has been and is collected is copyright protected. Its use may be copyright infringement. In the USA, AI developers rely on "fair use" and in Europe, the prevailing view is that the exception for "Text and Data Mining" (TDM) applies. In a recent interdisciplinary tandem-study, we have argued in detail that this is actually not the case because generative AI training fundamentally differs from TDM. In this article, we share our main findings and the implications for both public and corporate research on generative models. We further discuss how the phenomenon of training data memorization leads to copyright issues independently from the "fair use" and TDM exceptions. Finally, we outline how the ISMIR could contribute to the ongoing discussion about fair practices with respect to generative AI that satisfy all stakeholders.
A Critical Assessment of Modern Generative Models' Ability to Replicate Artistic Styles
Asperti, Andrea, George, Franky, Marras, Tiberio, Stricescu, Razvan Ciprian, Zanotti, Fabio
In recent years, advancements in generative artificial intelligence have led to the development of sophisticated tools capable of mimicking diverse artistic styles, opening new possibilities for digital creativity and artistic expression. This paper presents a critical assessment of the style replication capabilities of contemporary generative models, evaluating their strengths and limitations across multiple dimensions. We examine how effectively these models reproduce traditional artistic styles while maintaining structural integrity and compositional balance in the generated images. The analysis is based on a new large dataset of AI-generated works imitating artistic styles of the past, holding potential for a wide range of applications: the "AI-pastiche" dataset. The study is supported by extensive user surveys, collecting diverse opinions on the dataset and investigation both technical and aesthetic challenges, including the ability to generate outputs that are realistic and visually convincing, the versatility of models in handling a wide range of artistic styles, and the extent to which they adhere to the content and stylistic specifications outlined in prompts. This paper aims to provide a comprehensive overview of the current state of generative tools in style replication, offering insights into their technical and artistic limitations, potential advancements in model design and training methodologies, and emerging opportunities for enhancing digital artistry, human-AI collaboration, and the broader creative landscape.
Evaluating Multimodal Generative AI with Korean Educational Standards
This paper presents the Korean National Educational Test Benchmark (KoNET), a new benchmark designed to evaluate Multimodal Generative AI Systems using Korean national educational tests. KoNET comprises four exams: the Korean Elementary General Educational Development Test (KoEGED), Middle (KoMGED), High (KoHGED), and College Scholastic Ability Test (KoCSAT). These exams are renowned for their rigorous standards and diverse questions, facilitating a comprehensive analysis of AI performance across different educational levels. By focusing on Korean, KoNET provides insights into model performance in less-explored languages. We assess a range of models - open-source, open-access, and closed APIs - by examining difficulties, subject diversity, and human error rates. The code and dataset builder will be made fully open-sourced at https://github.com/naver-ai/KoNET.
ChatGPT reaches 400M weekly active users
ChatGPT has surpassed 400 million weekly active users. "We feel very fortunate to serve 5 percent of the world every week," OpenAI COO Brad Lightcap said on X about the new audience stat. This figure is twice the weekly active user count reported by the company in August 2024, which was double the figure it posted in November 2023. The latest milestone for the AI assistant comes after a huge uproar over new rival platform DeepSeek earlier in the year, which raised questions about whether the current crop of leading AI tools was about to be dethroned. OpenAI is on the verge of a move to simplify its ChatGPT offerings so that users won't have to select which reasoning model will respond to an input, and it will make its GPT-4.5 and GPT-5 models available soon in the chat and API clients.
Why OpenAI is trying to untangle its 'bespoke' corporate structure
On the Friday after Christmas, OpenAI published a blog post titled "Why OpenAI's structure must evolve to advance our mission." In it, the company detailed a plan to reorganize its for-profit arm into a public benefit corporation (PBC). In the weeks since that announcement, I've spoken to some of the country's leading corporate law experts to gain a better understanding of OpenAI's plan, and, more importantly, what it might mean for its mission to build safe artificial general intelligence (AGI). "Public benefit corporations are a relatively recent addition to the universe of business entity types," says Jens Dammann, professor of corporate law at the University of Texas School of Law. Depending on who you ask, you may get a different history of PBCs, but in the dominant narrative, they came out of a certification program created by a nonprofit called B Lab. Companies that complete a self-assessment and pay an annual fee to B Lab can carry the B Lab logo on their products and websites and call themselves B-Corps.
I'm Not Convinced Ethical Generative AI Currently Exists
Are there generative AI tools I can use that are perhaps slightly more ethical than others? No, I don't think any one generative AI tool from the major players is more ethical than any other. For me, the ethics of generative AI use can be broken down to issues with how the models are developed--specifically, how the data used to train them was accessed--as well as ongoing concerns about their environmental impact. In order to power a chatbot or image generator, an obscene amount of data is required, and the decisions developers have made in the past--and continue to make--to obtain this repository of data are questionable and shrouded in secrecy. Even what people in Silicon Valley call "open source" models hide the training datasets inside.
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
This paper explores the intersection of artificial intelligence and higher education administration, focusing on liberal arts colleges (LACs). It examines AI's opportunities and challenges in academic and student affairs, legal compliance, and accreditation processes, while also addressing the ethical considerations of AI deployment in mission-driven institutions. Considering AI's value pluralism and potential allocative or representational harms caused by algorithmic bias, LACs must ensure AI aligns with its mission and principles. The study highlights other strategies for responsible AI integration, balancing innovation with institutional values.
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model
Liu, Yi, Xu, Changran, Zhou, Yunhao, Li, Zeju, Xu, Qiang
Recent advancements in large language models (LLMs) have shown significant potential for automating hardware description language (HDL) code generation from high-level natural language instructions. While fine-tuning has improved LLMs' performance in hardware design tasks, prior efforts have largely focused on Verilog generation, overlooking the equally critical task of Verilog understanding. Furthermore, existing models suffer from weak alignment between natural language descriptions and Verilog code, hindering the generation of high-quality, synthesizable designs. To address these issues, we present DeepRTL, a unified representation model that excels in both Verilog understanding and generation. Based on CodeT5+, DeepRTL is fine-tuned on a comprehensive dataset that aligns Verilog code with rich, multi-level natural language descriptions. We also introduce the first benchmark for Verilog understanding and take the initiative to apply embedding similarity and GPT Score to evaluate the models' understanding capabilities. These metrics capture semantic similarity more accurately than traditional methods like BLEU and ROUGE, which are limited to surface-level n-gram overlaps. By adapting curriculum learning to train DeepRTL, we enable it to significantly outperform GPT-4 in Verilog understanding tasks, while achieving performance on par with OpenAI's o1-preview model in Verilog generation tasks.
Methods and Trends in Detecting Generated Images: A Comprehensive Review
Mahara, Arpan, Rishe, Naphtali
The proliferation of generative models, such as Generative Adversarial Networks (GANs), Diffusion Models, and Variational Autoencoders (VAEs), has enabled the synthesis of high-quality multimedia data. However, these advancements have also raised significant concerns regarding adversarial attacks, unethical usage, and societal harm. Recognizing these challenges, researchers have increasingly focused on developing methodologies to detect synthesized data effectively, aiming to mitigate potential risks. Prior reviews have primarily focused on deepfake detection and often lack coverage of recent advancements in synthetic image detection, particularly methods leveraging multimodal frameworks for improved forensic analysis. To address this gap, the present survey provides a comprehensive review of state-of-the-art methods for detecting and classifying synthetic images generated by advanced generative AI models. This review systematically examines core detection methodologies, identifies commonalities among approaches, and categorizes them into meaningful taxonomies. Furthermore, given the crucial role of large-scale datasets in this field, we present an overview of publicly available datasets that facilitate further research and benchmarking in synthetic data detection.
Hardware-Friendly Static Quantization Method for Video Diffusion Transformers
Yi, Sanghyun, Liu, Qingfeng, El-Khamy, Mostafa
Diffusion Transformers for video generation have gained significant research interest since the impressive performance of SORA. Efficient deployment of such generative-AI models on GPUs has been demonstrated with dynamic quantization. However, resource-constrained devices cannot support dynamic quantization, and need static quantization of the models for their efficient deployment on AI processors. In this paper, we propose a novel method for the post-training quantization of OpenSora\cite{opensora}, a Video Diffusion Transformer, without relying on dynamic quantization techniques. Our approach employs static quantization, achieving video quality comparable to FP16 and dynamically quantized ViDiT-Q methods, as measured by CLIP, and VQA metrics. In particular, we utilize per-step calibration data to adequately provide a post-training statically quantized model for each time step, incorporating channel-wise quantization for weights and tensor-wise quantization for activations. By further applying the smooth-quantization technique, we can obtain high-quality video outputs with the statically quantized models. Extensive experimental results demonstrate that static quantization can be a viable alternative to dynamic quantization for video diffusion transformers, offering a more efficient approach without sacrificing performance.