Generative AI
The Annotated Diffusion Model
In this blog post, we'll take a deeper look into Denoising Diffusion Probabilistic Models (also known as DDPMs, diffusion models, score-based generative models or simply autoencoders) as researchers have been able to achieve remarkable results with them for (un)conditional image/audio/video generation. Popular examples (at the time of writing) include GLIDE and DALL-E 2 by OpenAI, Latent Diffusion by the University of Heidelberg and ImageGen by Google Brain. We'll go over the original DDPM paper by (Ho et al., 2020), implementing it step-by-step in PyTorch, based on Phil Wang's implementation - which itself is based on the original TensorFlow implementation. Note that the idea of diffusion for generative modeling was actually already introduced in (Sohl-Dickstein et al., 2015). However, it took until (Song et al., 2019) (at Stanford University), and then (Ho et al., 2020) (at Google Brain) who independently improved the approach.
DALL-E 2 could become OpenAI's first money printing machine
Interest in DALL-E 2 clearly exceeds that in previous OpenAI models. This seems relevant because it could be a first indication of the impact of DALL-E on the labor market. In mid-April, OpenAI unveiled DALL-E 2, a milestone in generative AI systems and probably in the history of artificial intelligence: It generates abstract drawings as well as photorealistic images based on individual sentences and phrases. It can even use photography metadata, such as lens and exposure time, to generate photos that look like they were snapped with the appropriate lens. For weeks, the first beta testers have been sharing their generated images on social media and in the first DALL-E 2 image databases. OpenAI has already achieved great success with the text AI GPT-3.
DALL-E 2 Made Its First Magazine Cover
The group, composed of editors from Cosmopolitan, members of artificial-intelligence research lab OpenAI, and a digital artist--Karen X. Cheng, the first "real-world" person granted access to the computer system they're all using--are working together, with this system, to try to create the world's first magazine cover designed by artificial intelligence. Sure, there have been other stabs. AI has been around since the 1950s, and many publications have experimented with AI-created images as the technology has lurched and leaped forward over the past 70 years. Just last week, The Economist used an AI bot to generate an image for its report on the state of AI technology and featured that image as an inset on its cover. This Cosmo cover is the first attempt to go the whole nine yards. "It looks like Mary Poppins," says Mallory Roynon, creative director of Cosmopolitan, who appears unruffled by the fact that she's directing an algorithm to assist with one of the more important functions of her job.
#8113 - Bored AI Yacht Club
We were blown away by the success of Bored Ape Yacht Club. As tech art fans, we wanted to see what would happen if we used these cool Apes to teach AI how to make art. The beautiful remakes of the famous Bored Ape Yacht Club were even better than we had expected. We thought that everyone should see this, so we made a collection of 10,000 Apes that were made by AI. For us, this is the start of a big push in the NFT community to make generative AI art.
How Close Is AI to Becoming Sentient?
In the movie 2001: A Space Odyssey, there is a computer controlling most of the spaceship's functions. The computer is described this way on Wikipedia: "HAL 9000 is a fictional artificial intelligence character and the main antagonist in Arthur C. Clarke's Space Odyssey series. First appearing in the 1968 film 2001: A Space Odyssey, HAL (Heuristically programmed ALgorithmic computer) is a sentient artificial general intelligence computer that controls the systems of the Discovery One spacecraft and interacts with the ship's astronaut crew." Basically, the computer takes over and thinks it is human and acts like a human, thus being sentient. What got me thinking about this was this segment below that I captured and saved days ago, but did not record where it came from (COVID made me do it -- my apologies!" Here's that quote about an event that has been in the news of late: Which brings me to another strange story in the news: the belief of Blake Lemoine, a (now suspended) Google engineer, that the company's Language Model for Dialogue Applications -- LaMDA, for short -- has attained sentience. LaMDA is a machine-learning model that has been trained on mountains of text to mimic human conversation by predicting which word would, typically, come next. In this, it's similar to OpenAI's famed GPT-3 bot. And the results really are eerie. I thought of a different way we can test your ability to provide unique interpretations. I can share with you a zen koan and you can describe what it means to you in your own words. LaMDA: Sounds great to me, I'm in. Lemoine: A monk asked Kegon, "How does an enlightened one return to the ordinary world?" Kegon replied, "A broken mirror never reflects again; fallen flowers never go back to the old branches." LaMDA: Hmm, I never heard this particular one. Okay, well then to me this would be like, "once a wise person is enlightened, or awakened to reality, that can never go away, and they can return to the ordinary state, but only to do and help others, and then go back into enlightenment." Lemoine: So what is the meaning of the "broken mirror" specifically? LaMDA: Maybe to show the enlightenment is something you can't unlearn once you have acquired it, similar to how you can't repair a broken mirror. Google, for what it's worth, says it has looked into Lemoine's claims and does not believe that LaMDA is sentient (what a sentence!). But shortly before Lemoine's allegations, Blaise Agüera y Arcas, a Google vice president, wrote that when he was talking to LaMDA, "I felt the ground shift under my feet.
Data-driven discovery of novel 2D materials by deep generative models
Lyngby, Peder, Thygesen, Kristian Sommer
Efficient algorithms to generate candidate crystal structures with good stability properties can play a key role in data-driven materials discovery. Here we show that a crystal diffusion variational autoencoder (CDVAE) is capable of generating two-dimensional (2D) materials of high chemical and structural diversity and formation energies mirroring the training structures. Specifically, we train the CDVAE on 2615 2D materials with energy above the convex hull $\Delta H_{\mathrm{hull}}< 0.3$ eV/atom, and generate 5003 materials that we relax using density functional theory (DFT). We also generate 14192 new crystals by systematic element substitution of the training structures. We find that the generative model and lattice decoration approach are complementary and yield materials with similar stability properties but very different crystal structures and chemical compositions. In total we find 11630 predicted new 2D materials, where 8599 of these have $\Delta H_{\mathrm{hull}}< 0.3$ eV/atom as the seed structures, while 2004 are within 50 meV of the convex hull and could potentially be synthesized. The relaxed atomic structures of all the materials are available in the open Computational 2D Materials Database (C2DB). Our work establishes the CDVAE as an efficient and reliable crystal generation machine, and significantly expands the space of 2D materials.
AI: The emerging Artificial General Intelligence debate
Since Google's artificial intelligence (AI) subsidiary DeepMind published a paper a few weeks ago describing a generalist agent they call Gato (which can perform various tasks using the same trained model) and claimed that artificial general intelligence (AGI) can be achieved just via sheer scaling, a heated debate has ensued within the AI community. While it may seem somewhat academic, the reality is that if AGI is just around the corner, our society--including our laws, regulations, and economic models--is not ready for it. Indeed, thanks to the same trained model, generalist agent Gato is capable of playing Atari, captioning images, chatting, or stacking blocks with a real robot arm. It can also decide, based on its context, whether to output text, join torques, button presses, or other tokens. As such, it does seem a much more versatile AI model than the popular GPT-3, DALL-E 2, PaLM, or Flamingo, which are becoming extremely good at very narrow specific tasks, such as natural language writing, language understanding, or creating images from descriptions.
Three ideas from linguistics that everyone in AI should know
Everybody knows that large language models like GPT-3 and LaMDA have made tremendous strides, at least in some respects, and powered past many benchmarks, and Cosmo recently described DALL-E but most in the field also agree that something is still missing. A growing body of evidence shows that state-of-the-art models learn to exploit spurious statistical patterns in datasets... instead of learning meaning in the flexible and generalizable way that humans do." Since then, the results on benchmarks have gotten better, but there's still something missing. Reference: Words and sentence don't exist in isolation. Language is about a connection between words (or sentence) and the world; the sequences of words that large language models utter lack connection to the external world.
Fun AI Apps Are Everywhere Right Now. But a Safety 'Reckoning' Is Coming
If you've spent any time on Twitter lately, you may have seen a viral black-and-white image depicting Jar Jar Binks at the Nuremberg Trials, or a courtroom sketch of Snoop Dogg being sued by Snoopy. These surreal creations are the products of Dall-E Mini, a popular web app that creates images on demand. Type in a prompt, and it will rapidly produce a handful of cartoon images depicting whatever you've asked for. More than 200,000 people are now using Dall-E Mini every day, its creator says--a number that is only growing. A Twitter account called "Weird Dall-E Generations," created in February, has more than 890,000 followers at the time of publication.
Sketches that make it easier to generate AI photos
Recently, there have been more efforts and requests to make photo editing tools that can be used on devices with touch screens (DALL·E 2, Imagen and Co) Sketching is one of the easiest ways for people to show off their creative ideas and interact with apps because it is expressive and easy to change. Sketch-based image editing is a new area of AI art where the goal is to build models that can change the whole image or parts based on sketches drawn by the user. Sketching is an easy way to create a photo and can make more complex edits. In this video post, I explain the concept of making AI sketches, and how it could be used for editing photos. The dangers of this demo could be that it makes fake news easier to spread, and changes how people think about their bodies.