Generative AI
OpenAI's Sam Altman and other tech leaders join the federal AI safety board
Sam Altman, OpenAI's CEO, Microsoft chief Satya Nadella, Alphabet CEO Sundar Pichai are joining the government's Artificial Intelligence Safety and Security Board, according to The Wall Street Journal. They're also joined by Nvidia's Jensen Huang, Northrop Grumman's Kathy Warden and Delta's Ed Bastian, along with other leaders in the tech and AI industry. The AI board will be working with and advising the Department of Homeland Security on how it can safely deploy AI within the country's critical infrastructure. They're also tasked with conjuring recommendations for power grid operators, transportation service providers and manufacturing plants on how they can can protect their systems against potential threats that could be brought about by advances in the technology. The Biden administration ordered the creation of an AI safety board last year as part of a sweeping executive order that focuses on regulating AI development.
Is ChatGPT sexist? AI chatbot was asked to generate 100 images of CEOs but only ONE was a woman (and 99% of the secretaries were female...)
Imagine a successful investor or a wealthy chief executive – who would you picture? If you ask ChatGPT, it's almost certainly a white man. The chatbot has been accused of'sexism' after it was asked to generate images of people in various high powered jobs. Out of 100 tests, it chose a man 99 times. In contrast, when it was asked to do so for a secretary, it chose a woman all but once.
Google Thinks It Can Cash In on Generative AI. Microsoft Already Has
Alphabet CEO Sundar Pichai is confident that Google will find a way to make money selling access to generative AI tools. Microsoft CEO Satya Nadella says his company is already doing it. Both companies reported better-than-expected quarterly sales and profit on Thursday. And the stock prices of both soared on the results, with Alphabet further buoyed by its new plans to buy back more shares and issue its first-ever dividend. But the near-term fortunes of Microsoft and Google, at least as far as their generative AI efforts are concerned, look different under the hood and in the comments of their executives. How investors, workers, and potential customers perceive the rivals' dueling efforts could determine which gets the better chunk of the hundreds of billions of dollars in spending expected to flow to such software in the coming years.
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
Singh, Abhishek Kumar, Patras, Ioannis
The rapid evolution of the fashion industry increasingly intersects with technological advancements, particularly through the integration of generative AI. This study introduces a novel generative pipeline designed to transform the fashion design process by employing latent diffusion models. Utilizing ControlNet and LoRA fine-tuning, our approach generates high-quality images from multimodal inputs such as text and sketches. We leverage and enhance state-of-the-art virtual try-on datasets, including Multimodal Dress Code and VITON-HD, by integrating sketch data. Our evaluation, utilizing metrics like FID, CLIP Score, and KID, demonstrates that our model significantly outperforms traditional stable diffusion models. The results not only highlight the effectiveness of our model in generating fashion-appropriate outputs but also underscore the potential of diffusion models in revolutionizing fashion design workflows. This research paves the way for more interactive, personalized, and technologically enriched methodologies in fashion design and representation, bridging the gap between creative vision and practical application.
Seizing the Means of Production: Exploring the Landscape of Crafting, Adapting and Navigating Generative AI Models in the Visual Arts
Abuzuraiq, Ahmed M., Pasquier, Philippe
Users of these models can produce diverse and high-quality visuals through meticulously written text prompts. These models mark a significant shift from the era when artists used personally-trainable generative models like Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs). LTGMs have made the generation of visuals accessible to everyone, but this shift in attention has overshadowed the practice of "model crafting", whereas artists personalize their work by experimenting with training sets, model architectures, and hyperparameters in addition to combining, adapting and manipulating pre-trained models [14]. Model crafting offered artists a sense of craftsmanship and ownership over the creative process and its outcomes.
Meta's Open Source Llama 3 Is Already Nipping at OpenAI's Heels
Jerome Pesenti has a few reasons to celebrate Meta's decision last week to release Llama 3, a powerful open source large language model that anyone can download, run, and build on. Pesenti used to be vice president of artificial intelligence at Meta and says he often pushed the company to consider releasing its technology for others to use and build on. But his main reason to rejoice is that his new startup will get access to an AI model that he says is very close in power to OpenAI's industry-leading text generator GPT-4, but considerably cheaper to run and more open to outside scrutiny and modification. "The release last Friday really feels like a game-changer," Pesenti says. His new company, Sizzle, an AI tutor, currently uses GPT-4 and other AI models, both closed and open, to craft problem sets and curricula for students.
The Download: hyperrealistic deepfakes, and clean energy's implications for mining
Until now, AI-generated videos of people have tended to have some stiffness, glitchiness, or other unnatural elements that make them pretty easy to differentiate from reality. For the past several years, AI video startup Synthesia has produced these kinds of AI-generated avatars. But today it launches a new generation, its first to take advantage of the latest advancements in generative AI, and they are more realistic and expressive than anything we've seen before. While today's release means almost anyone will now be able to make a digital double, before the technology went public, Synthesia agreed to make one of Melissa Heikkilä, our senior AI reporter. This technological progress signals a much larger shift.
An AI startup made a hyperrealistic deepfake of me that's so good it's scary
Thanks to rapid advancements in generative AI and a glut of training data created by human actors that has been fed into its AI model, Synthesia has been able to produce avatars that are indeed more humanlike and more expressive than their predecessors. The digital clones are better able to match their reactions and intonation to the sentiment of their scripts--acting more upbeat when talking about happy things, for instance, and more serious or sad when talking about unpleasant things. They also do a better job matching facial expressions--the tiny movements that can speak for us without words. But this technological progress also signals a much larger social and cultural shift. Increasingly, so much of what we see on our screens is generated (or at least tinkered with) by AI, and it is becoming more and more difficult to distinguish what is real from what is not.
Generative AI arrives in the gene-editing world of CRISPR
Generative artificial intelligence technologies can write poetry and computer programs or create images of teddy bears and videos of cartoon characters that look like something from a Hollywood movie. Now, new AI technology is generating blueprints for microscopic biological mechanisms that can edit your DNA, pointing to a future when scientists can battle illness and diseases with even greater precision and speed than they can today. Described in a research paper published Monday by a Berkeley, California, startup called Profluent, the technology is based on the same methods that drive ChatGPT, the online chatbot that launched the AI boom after its release in 2022.
To what extent is ChatGPT useful for language teacher lesson plan creation?
Dornburg, Alex, Davin, Kristin
The advent of generative AI models holds tremendous potential for aiding teachers in the generation of pedagogical materials. However, numerous knowledge gaps concerning the behavior of these models obfuscate the generation of research-informed guidance for their effective usage. Here we assess trends in prompt specificity, variability, and weaknesses in foreign language teacher lesson plans generated by zero-shot prompting in ChatGPT. Iterating a series of prompts that increased in complexity, we found that output lesson plans were generally high quality, though additional context and specificity to a prompt did not guarantee a concomitant increase in quality. Additionally, we observed extreme cases of variability in outputs generated by the same prompt. In many cases, this variability reflected a conflict between 20th century versus 21st century pedagogical practices. These results suggest that the training of generative AI models on classic texts concerning pedagogical practices may represent a currently underexplored topic with the potential to bias generated content towards teaching practices that have been long refuted by research. Collectively, our results offer immediate translational implications for practicing and training foreign language teachers on the use of AI tools. More broadly, these findings reveal the existence of generative AI output trends that have implications for the generation of pedagogical materials across a diversity of content areas.