Goto

Collaborating Authors

 Generative AI


MAUVE Scores for Generative Models: Theory and Practice

arXiv.org Artificial Intelligence

Generative artificial intelligence has made significant strides, producing text indistinguishable from human prose and remarkably photorealistic images. Automatically measuring how close the generated data distribution is to the target distribution is central to diagnosing existing models and developing better ones. We present MAUVE, a family of comparison measures between pairs of distributions such as those encountered in the generative modeling of text or images. These scores are statistical summaries of divergence frontiers capturing two types of errors in generative modeling. We explore three approaches to statistically estimate these scores: vector quantization, non-parametric estimation, and classifier-based estimation. We provide statistical bounds for the vector quantization approach. Empirically, we find that the proposed scores paired with a range of $f$-divergences and statistical estimation methods can quantify the gaps between the distributions of human-written text and those of modern neural language models by correlating with human judgments and identifying known properties of the generated texts. We demonstrate in the vision domain that MAUVE can identify known properties of generated images on par with or better than existing metrics. In conclusion, we present practical recommendations for using MAUVE effectively with language and image modalities.


Generative AI's iPhone Moment

The Atlantic - Technology

After nearly seven months of rumors and delays, Google has finally released its most advanced generative-AI model to date: Gemini 1.0, a program the company is advertising as one of the most capable pieces of software ever. It can purportedly solve calculus problems, explain memes, write code, and--in a real example offered by the company--provide feedback on cooking photos to help you decide when your omelet is done. Google is even billing Gemini as "a first step toward a truly universal AI model," one that is designed from the ground up to engage with images, video, text, audio, and computer code in a range of contexts. And, somehow, it all feels a bit underwhelming. Perhaps that is because today's announcement feels like any other Silicon Valley product launch.


Meta's AI image generator is available as a standalone website

Engadget

Meta has launched a standalone version of its image generator as it tests dozens of new generative AI features across Facebook, Instagram and WhatsApp. The image generator, called Imagine, was first previewed at the company's Connect event in November and has been available as part of Meta's AI chatbot. Now, with its own dedicated website at imagine.meta.com, the tool will be available outside of the company's messaging apps. Like other generative AI tools, Imagine allows users to create images from simple text prompts. Imagine, which relies on Meta's Emu model, will generate four images for each prompt.


Google CEO Sundar Pichai on Gemini and the coming age of AI

MIT Technology Review

Pichai, who previously oversaw Chrome and Android, is famously product obsessed. In his first founder's letter as CEO in 2016, he predicted that "[w]e will move from mobile first to an AI first world." In the years since, Pichai has infused AI deeply into all of Google's products, from Android devices all the way up to the cloud. Despite that, the last year has largely been defined by the AI releases from another company, OpenAI. The rollout of DALL-E and GPT-3.5 last year, followed by GPT-4 this year, dominated the sector and kicked off an arms race between startups and tech giants alike.


Google DeepMind Unveils Its Most Powerful AI Offering Yet

TIME - Tech

Google DeepMind has announced its much-anticipated family of artificial intelligence chatbots, Gemini, which will compete with OpenAI's GPT series. According to Google, Gemini Ultra, its largest and most capable new model, outperforms OpenAI's most capable model, GPT-4, at a number of text-based, image-based, coding, and reasoning tasks. Gemini Ultra will be available through a new AI chat feature called Bard Advanced from early next year, the company said. It is currently being refined and is undergoing "trust and safety checks, including red-teaming by trusted external parties," according to the announcement. Google DeepMind also announced the launch of Gemini Pro, which is now available to the public through Google's Bard chat interface, and the smaller Gemini Nano, which will run on Google's Pixel 8 Pro smartphone.


Google claims new Gemini AI 'thinks more carefully'

BBC News

However, a new, more powerful version of the OpenAI software is due to be released next year, with chief executive Sam Altman saying the firm's new products would make its current ones look like "a quaint relative".


Google's 'Gemini' is the latest AI software entering fierce competition

Washington Post - Technology News

Facebook owner Meta has been an AI player for years, hiring some of the field's smartest researchers and using the tech to help decide which of its users should see certain advertisements. In July, it doubled down on a very different approach to AI than its Big Tech rivals. It announced that Llama 2, its GPT4 competitor, would be "open source" -- available for anyone to download, modify and add to their own products for free. The approach won Meta plaudits from tech start-ups who were worried that Google, Microsoft and OpenAI would try to corner the market for advanced AI and squeeze out any competitors. But it's also been criticized for making it easier for people to use AI for malicious purposes.


Google's answer to GPT-4 is Gemini: 'the most capable model we've ever built'

Engadget

OpenAI's spot atop the generative AI heap may be coming to an end as Google officially introduced its most capable large language model to date on Wednesday, dubbed Gemini 1.0. It's the first of "a new generation of AI models, inspired by the way people understand and interact with the world," CEO Sundar Pichai wrote in a Google blog post. "Ever since programming AI for computer games as a teenager, and throughout my years as a neuroscience researcher trying to understand the workings of the brain, I've always believed that if we could build smarter machines, we could harness them to benefit humanity in incredible ways," Pichai continued. The result of extensive collaboration between Google's DeepMind and Research divisions, Gemini has all the bells and whistles cutting-edge genAIs have to offer. "Its capabilities are state-of-the-art in nearly every domain," Pichai declared.


Google announces new AI processing chips and a cloud 'hypercomputer'

Engadget

Undoubtedly, 2023 has been the year of generative AI, and Google is marking its end with even more AI developments. The company has announced the creation of its most powerful TPU (formally known as Tensor Processing Units) yet, Cloud TPU v5p, and an AI Hypercomputer from Google Cloud. "The growth in [generative] AI models -- with a tenfold increase in parameters annually over the past five years -- brings heightened requirements for training, tuning, and inference," Amin Vahdat, Google's Engineering Fellow and Vice President for the Machine Leaning, Systems, and Cloud AI team, said in a release. The Cloud TPU v5p is an AI accelerator, training and serving models. Google designed Cloud TPUs to work with models that are large, have long training periods, are mostly made of matrix computations and have no custom operations inside its main training loop, such as TensorFlow or JAX.


Google Just Launched Gemini, Its Long-Awaited Answer to ChatGPT

WIRED

Increasing talk of artificial intelligence developing with potentially dangerous speed is hardly slowing things down. A year after OpenAI launched ChatGPT and triggered a new race to develop AI technology, Google today revealed an AI project intended to reestablish the search giant as the world leader in AI. Gemini, a new type of AI model that can work with text, images, and video, could be the most important algorithm in Google's history after PageRank, which vaulted the search engine into the public psyche and created a corporate giant. An initial version of Gemini starts to roll out today inside Google's chatbot Bard for the English language setting. It will be available in more than 170 countries and territories.