Generative AI
Can AI Learn Better without Learning Anything at All?
The human mind can get really complicated at times. That's when we turn to meditation, taking deep breaths to forget all the chaos. It gives us certain fulfilment--by bringing out new traits of understanding and empathy--sometimes by not doing anything at all. What if machines could also meditate or do nothing for a day or two--to learn better? As absurd as it may sound, a group of researchers have now discovered how artificial neural networks can mimic sleep patterns of the human brain, boosting their utility across a spectrum of research areas.
See Samuel L. Jackson As R2-D2 In Creepy AI Generated Art
Samuel L Jackon has been made into R2D2 using the DALL-E 2 artificial intelligence program. By now, many of us have seen hilarious or ridiculous images gracing our social media feed courtesy of DALL-E 2, a complex ai art generator that allows for users to create expansive images using only a basic text prompt. DALL-E 2 can create brand new images from scratch or edit existing images, often leading to disturbing or creepy distortions of people or things that seem familiar but… not quite right. For instance, Samuel L Jackson as R2-D2 provides enough visual context to make anyone aware of who is being depicted at a passing glance, but still doesn't look quite like the Mace Windu actor, or even quite human at all. Images such as this, or Kanye Light Year, provide a glimpse into the uncanny valley.
A bot that watched 70,000 hours of Minecraft could unlock AI's next big thing
The result is a breakthrough for a technique known as imitation learning, in which neural networks are trained how to perform tasks by watching humans do them. Imitation learning can be used to train AI to control robot arms, drive cars or navigate webpages. There is a vast amount of video online showing people doing different tasks. By tapping into this resource, the researchers hope to do for imitation learning what GPT-3 did for large language models. "In the last few years we've seen the rise of this GPT-3 paradigm where we see amazing capabilities come from big models trained on enormous swathes of the internet," says Bowen Baker at OpenAI, one of the team behind the new Minecraft bot.
Is generative AI really a threat to creative professionals?
When the concept artist and illustrator RJ Palmer first witnessed the fine-tuned photorealism of compositions produced by the AI image generator Dall-E 2, his feeling was one of unease. The tool, released by the AI research company OpenAI, showed a marked improvement on 2021's Dall-E, and was quickly followed by rivals such as Stable Diffusion and Midjourney. Type in any surreal prompt, from Kermit the frog in the style of Edvard Munch, to Gollum from The Lord of the Rings feasting on a slice of watermelon, and these tools will return a startlingly accurate depiction moments later. Cosmopolitan trumpeted the world's first AI-generated magazine cover, and technology investors fell over themselves to wave in the new era of "generative AI". The image-generation capabilities have already spread to video, with the release of Google's Imagen Video and Meta's Make-A-Video.
Stable Diffusion made copying artists and generating porn harder and users are mad
Changes to Stable Diffusion are notable, as the software is hugely influential and helps set norms in the fast-moving generative AI scene. Unlike rival models like OpenAI's DALL-E, Stable Diffusion is open source. This allows the community to quickly improve on the tool and for developers to integrate it into their products free of charge. But it also means Stable Diffusion has fewer constraints in how it's used and, as a consequence, has attracted significant criticism. In particular, many artists, like Rutkowski, are annoyed that Stable Diffusion and other image generating models were trained on their artwork without their consent and can now reproduce their styles.
Google has a secret new project that is teaching artificial intelligence to write and fix code. It could reduce the need for human engineers in the future.
Google is working on a secretive project that uses machine learning to train code to write, fix, and update itself. This project is part of a broader push by Google into so-called generative artificial intelligence, which uses algorithms to create images, videos, code, and more. It could have profound implications for the company's future and developers who write code. The project, which began life inside Alphabet's X research unit and was codenamed Pitchfork, moved into Google's Labs group this summer, according to people familiar with the matter. By moving into Google, it signaled its increased importance to leaders.
Harvey, which uses AI to answer legal questions, lands cash from OpenAI
Harvey, a startup building what it describes as a "copilot for lawyers," today emerged from stealth with $5 million in funding led by the OpenAI Startup Fund, the tranche through which OpenAI and its partners are investing in early-stage AI companies tackling major problems. Also participating in the round was Jeff Dean, the lead of Google AI, Google's AI research division. Harvey was founded by Winston Weinberg, a former securities and antitrust litigator at law firm O'Melveny & Myers, and Gabriel Pereyra, previously a research scientist at DeepMind, Google Brain (another of Google's AI groups) and Meta AI. Weinberg and Pereyra are roomates -- Pereyra showed Weinberg OpenAI's GPT-3 text-generating system and Weinberg realized that it could be used to improve legal workflows. "Our product provides lawyers with a natural language interface for their existing legal workflows," Pereyra told TechCrunch in an email interview. "Instead of manually editing legal documents or performing legal research, Harvey enables lawyers to describe the task they wish to accomplish in simple instructions and receive the generated result.
Improving dermatology classifiers across populations using images generated by large diffusion models
Sagers, Luke W., Diao, James A., Groh, Matthew, Rajpurkar, Pranav, Adamson, Adewole S., Manrai, Arjun K.
Dermatological classification algorithms developed without sufficiently diverse training data may generalize poorly across populations. While intentional data collection and annotation offer the best means for improving representation, new computational approaches for generating training data may also aid in mitigating the effects of sampling bias. In this paper, we show that DALL$\cdot$E 2, a large-scale text-to-image diffusion model, can produce photorealistic images of skin disease across skin types. Using the Fitzpatrick 17k dataset as a benchmark, we demonstrate that augmenting training data with DALL$\cdot$E 2-generated synthetic images improves classification of skin disease overall and especially for underrepresented groups.
Schr\"{o}dinger's Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition
White, Jennifer C., Cotterell, Ryan
Recent work has shown that despite their impressive capabilities, text-to-image diffusion models such as DALL-E 2 (Ramesh et al., 2022) can display strange behaviours when a prompt contains a word with multiple possible meanings, often generating images containing both senses of the word (Rassin et al., 2022). In this work we seek to put forward a possible explanation of this phenomenon. Using the similar Stable Diffusion model (Rombach et al., 2022), we first show that when given an input that is the sum of encodings of two distinct words, the model can produce an image containing both concepts represented in the sum. We then demonstrate that the CLIP encoder used to encode prompts (Radford et al., 2021) encodes polysemous words as a superposition of meanings, and that using linear algebraic techniques we can edit these representations to influence the senses represented in the generated images. Combining these two findings, we suggest that the homonym duplication phenomenon described by Rassin et al. (2022) is caused by diffusion models producing images representing both of the meanings that are present in superposition in the encoding of a polysemous word.
Robustness Analysis of Deep Learning Models for Population Synthesis
Mensah, Daniel Opoku, Badu-Marfo, Godwin, Farooq, Bilal
Deep generative models have become useful for synthetic data generation, particularly population synthesis. The models implicitly learn the probability distribution of a dataset and can draw samples from a distribution. Several models have been proposed, but their performance is only tested on a single cross-sectional sample. The implementation of population synthesis on single datasets is seen as a drawback that needs further studies to explore the robustness of the models on multiple datasets. While comparing with the real data can increase trust and interpretability of the models, techniques to evaluate deep generative models' robustness for population synthesis remain underexplored. In this study, we present bootstrap confidence interval for the deep generative models, an approach that computes efficient confidence intervals for mean errors predictions to evaluate the robustness of the models to multiple datasets. Specifically, we adopt the tabular-based Composite Travel Generative Adversarial Network (CTGAN) and Variational Autoencoder (VAE), to estimate the distribution of the population, by generating agents that have tabular data using several samples over time from the same study area. The models are implemented on multiple travel diaries of Montreal Origin- Destination Survey of 2008, 2013, and 2018 and compare the predictive performance under varying sample sizes from multiple surveys. Results show that the predictive errors of CTGAN have narrower confidence intervals indicating its robustness to multiple datasets of the varying sample sizes when compared to VAE. Again, the evaluation of model robustness against varying sample size shows a minimal decrease in model performance with decrease in sample size. This study directly supports agent-based modelling by enabling finer synthetic generation of populations in a reliable environment.