Generative AI
Generally Intelligent secures cash from OpenAI vets to build capable AI systems
A new AI research company is launching out of stealth today with an ambitious goal: to research the fundamentals of human intelligence that machines currently lack. Called Generally Intelligent, it plans to do this by turning these fundamentals into an array of tasks to be solved and by designing and testing different systems' ability to learn to solve them in highly complex 3D worlds built by their team. "We believe that generally intelligent computers will someday unlock extraordinary potential for human creativity and insight," CEO Kanjun Qiu told TechCrunch in an email interview. "However, today's AI models are missing several key elements of human intelligence, which inhibits the development of general-purpose AI systems that can be deployed safely โฆ Generally Intelligent's work aims to understand the fundamentals of human intelligence in order to engineer safe AI systems that can learn and understand the way humans do." Qiu, the former chief of staff at Dropbox and the co-founder of Ember Hardware, which designed laser displays for VR headsets, co-founded Generally Intelligent in 2021 after shutting down her previous startup, Sourceress, a recruiting company that used AI to scour the web.
OpenAI offers early look at DALL-E API, showcases text-to-image use case
Did you miss a session from MetaBeat 2022? Head over to the on-demand library for all of our featured sessions here. The DALL-E API won't be officially announced until later this fall, according to OpenAI, but today the company shared details about a customer already leveraging the DALL-E API for a specific enterprise use case. New York City-based Cala, a startup that bills itself as the "world's first operating system for fashion," offers a digital platform (including a mobile app launched in March) that allows creators to design and produce clothing lines, unifying the process from product ideation through order fulfillment. With the addition of DALL-E-powered text-to-image generating tools, users can generate new visual design ideas from natural text descriptions or uploaded reference images โ which the company says are first-of-its-kind capabilities for the fashion industry.
Efficient Diffusion Models for Vision: A Survey
Ulhaq, Anwaar, Akhtar, Naveed, Pogrebna, Ganna
Diffusion Models (DMs) have demonstrated state-of-the-art performance in content generation without requiring adversarial training. These models are trained using a two-step process. First, a forward - diffusion - process gradually adds noise to a datum (usually an image). Then, a backward - reverse diffusion - process gradually removes the noise to turn it into a sample of the target distribution being modelled. DMs are inspired by non-equilibrium thermodynamics and have inherent high computational complexity. Due to the frequent function evaluations and gradient calculations in high-dimensional spaces, these models incur considerable computational overhead during both training and inference stages. This can not only preclude the democratization of diffusion-based modelling, but also hinder the adaption of diffusion models in real-life applications. Not to mention, the efficiency of computational models is fast becoming a significant concern due to excessive energy consumption and environmental scares. These factors have led to multiple contributions in the literature that focus on devising computationally efficient DMs. In this review, we present the most recent advances in diffusion models for vision, specifically focusing on the important design aspects that affect the computational efficiency of DMs. In particular, we emphasize the recently proposed design choices that have led to more efficient DMs. Unlike the other recent reviews, which discuss diffusion models from a broad perspective, this survey is aimed at pushing this research direction forward by highlighting the design strategies in the literature that are resulting in practicable models for the broader research community. We also provide a future outlook of diffusion models in vision from their computational efficiency viewpoint. EEP generative modelling has emerged as one of the most exciting computational tools that is even challenging human creativity [1].
Composing Ensembles of Pre-trained Models via Iterative Consensus
Li, Shuang, Du, Yilun, Tenenbaum, Joshua B., Torralba, Antonio, Mordatch, Igor
Large pre-trained models exhibit distinct and complementary capabilities dependent on the data they are trained on. Language models such as GPT-3 are capable of textual reasoning but cannot understand visual information, while vision models such as DALL-E can generate photorealistic photos but fail to understand complex language descriptions. In this work, we propose a unified framework for composing ensembles of different pre-trained models - combining the strengths of each individual model to solve various multimodal problems in a zero-shot manner. We use pre-trained models as "generators" or "scorers" and compose them via closed-loop iterative consensus optimization. The generator constructs proposals and the scorers iteratively provide feedback to refine the generated result. Such closed-loop communication enables models to correct errors caused by other models, significantly boosting performance on downstream tasks, e.g. We demonstrate that consensus achieved by an ensemble of scorers outperforms the feedback of a single scorer, by leveraging the strengths of each expert model. Results show that the proposed method can be used as a general purpose framework for a wide range of zero-shot multimodal tasks, such as image generation, video question answering, mathematical reasoning, and robotic manipulation. Large pre-trained models have shown remarkable zero-shot generalization abilities, ranging from zero-shot image generation and natural language processing to machine reasoning and action planning. Such models are trained on large datasets scoured from the internet, often consisting of billions of datapoints. Individual pre-trained models capture different aspects of knowledge on the internet, with language models (LMs) capturing textual information in news, articles, and Wikipedia pages, and visual-language models (VLMs) modeling the alignments between visual and textual information. While it is desirable to have a single sizable pre-trained model capturing all possible modalities of data on the internet, such a comprehensive model is challenging to obtain and maintain, requiring intensive memory, an enormous amount of energy, months of training time, and millions of dollars. A more scalable alternative approach is to compose different pre-trained models together, leveraging the knowledge from different expert models to solve complex multimodal tasks. Building a unified framework for composing multiple models is challenging.
Make-A-Video: Text-to-Video Generation's Next... Generation? - NAB Amplify
The inevitable has happened, albeit a little sooner than expected. After all the hoopla surrounding text-to-image AI generators in recent months, Meta is first out of the gate with a text-to-video version. Perhaps Meta wanted to establish some headline leadership in this space, since the results aren't ready for primetime. But as developments in text-to-image generation has shown, by the time you read this the technology will already have advanced. Meta is only giving a glimpse to the public at the tech it calls Make-A-Video.
Focus on Whisper, OpenAI's automatic speech recognition system - Actu IA
OpenAI recently released Whisper, a 1.6 billion parameter AI model capable of transcribing and translating speech audio from 97 different languages, showing robust performance on a wide range of automated speech recognition (ASR) tasks. The model trained on 680,000 hours of audio data collected from the web was soon published as open source on GitHub. Whisper uses a transform-encoder-decoder architecture, the input audio is split into 30-second chunks, converted to a log-Mel spectrogram, and then passed through an encoder. Unlike most state-of-the-art ASR models, it has not been fitted to a specific data set, but instead has been trained using weak supervision on a large-scale noisy data set collected from the Internet. Although it did not beat the specialized LibriSpeech performance models, in zero-shot evaluations on a diverse dataset, Whisper proved to be more robust and made 50% fewer errors than those models.
David O. Houwen on LinkedIn: #AI #LLMs #OpenAI
Do not keep calm and carry on, girls!'' Do we really care more about Van Gogh's sunflowers than real ones? Gedorfge Monbiot The Guardian The response by the media and government to the two Just Stop Oil activists who threw soup at Vincent van Gogh's Sunflowers in the National Gallery in London speaks volumes. Decorating the glass protecting the painting with tomato soup (the painting itself was, as the protesters calculated, undamaged) appears to horrify some people more than the collapse of our planet, which these campaigners are seeking to prevent. Everywhere I see claims that the "extreme" tactics of environmental campaigners will prompt people to "stop listening". But how could we listen any less to the warnings of scientists and campaigners and eminent committees?
AI delivers a hilarious upgrade to '90s video game characters
AI image generators have been responsible for so much in the last few months. Barely a day goes by without us discovering some new experiment, whether it's resurrecting late celebs or creating a tool so we can turn anyone into a Pokรฉmon. Now someone's used them to enhance what 30 years ago was the cutting edge in video game graphics. An artist has fed the characters of Sega's nineties fighting game Virtua Fighter through an AI image generator, turning the original 3D polygon graphics into photorealistic images (well, almost). To see how this technology works, take a look at our piece on how to use DALL-E 2 (opens in new tab).
Adobe commits to transparency in use of generative AI
Did you miss a session from MetaBeat 2022? Head over to the on-demand library for all of our featured sessions here. Today, at Adobe MAX, billed as the world's largest creativity conference, Adobe announced its commitment to support creatives by ensuring transparency in the use of generative AI tools. In a year dominated by the rise of generative AI tools โ such as OpenAI's DALL-E 2, Google's Imagen, Stable Diffusion and MidJourney โ Adobe, the world's leading computer graphics software company, said its approach to developing creator-centric generative AI offerings would leverage its Content Authenticity Initiative (CAI) standards and invest in new research to support creatives' control over their style and work. The CAI is an Adobe-led initiative that enables creators to securely attach provenance data to digital content, helping ensure creators get credit for their work and audiences understand who made a piece of content and how it was created.
Bizarre Halloween Candy Courtesy of AI Tool Dall-E: Farte Cats, Anyone?
The DIY art world hasn't been the same since the passing of Bob Ross (rest in peace in a forest of happy little trees, king), but AI art creation tool Dall-E at least offers an entertaining and quicker way to generate masterpieces that seem appropriate for the bizarro timeline we all now share. AI researcher Janelle Shane has made a hobby of prompting machine learning systems to engage in the 2022 equivalent of some very weird improv comedy. Her interactions with AI prompt plenty of hilarity -- including a charming set of Valentine's Day cards that almost work. Her latest bit is simply asking Dall-E to paint a picture of the most popular Halloween candies in each US state. The results are filled with lots of odd gibberish and overflowing with candy corn like any child's basket on Nov. 1 when all the good stuff from the previous night's haul has already been devoured. The first thing that becomes clear as Shane starts to work her way through all 50 states alphabetically is that Dall-E associates Halloween treats very strongly with candy corn, which is pretty fair given its abundance this time of year.