Media
Score-Based Generative Modeling with Critically-Damped Langevin Diffusion
Dockhorn, Tim, Vahdat, Arash, Kreis, Karsten
Score-based generative models (SGMs) have demonstrated remarkable synthesis quality. SGMs rely on a diffusion process that gradually perturbs the data towards a tractable distribution, while the generative model learns to denoise. The complexity of this denoising task is, apart from the data distribution itself, uniquely determined by the diffusion process. We argue that current SGMs employ overly simplistic diffusions, leading to unnecessarily complex denoising processes, which limit generative modeling performance. Based on connections to statistical mechanics, we propose a novel critically-damped Langevin diffusion (CLD) and show that CLD-based SGMs achieve superior performance. CLD can be interpreted as running a joint diffusion in an extended space, where the auxiliary variables can be considered "velocities" that are coupled to the data variables as in Hamiltonian dynamics. We derive a novel score matching objective for CLD and show that the model only needs to learn the score function of the conditional distribution of the velocity given data, an easier task than learning scores of the data directly. We also derive a new sampling scheme for efficient synthesis from CLD-based diffusion models. We find that CLD outperforms previous SGMs in synthesis quality for similar network architectures and sampling compute budgets. We show that our novel sampler for CLD significantly outperforms solvers such as Euler--Maruyama. Our framework provides new insights into score-based denoising diffusion models and can be readily used for high-resolution image synthesis. Project page and code: https://nv-tlabs.github.io/CLD-SGM.
Pose Estimation of Specific Rigid Objects
In this thesis, we address the problem of estimating the 6D pose of rigid objects from a single RGB or RGB-D input image, assuming that 3D models of the objects are available. This problem is of great importance to many application fields such as robotic manipulation, augmented reality, and autonomous driving. First, we propose EPOS, a method for 6D object pose estimation from an RGB image. The key idea is to represent an object by compact surface fragments and predict the probability distribution of corresponding fragments at each pixel of the input image by a neural network. Each pixel is linked with a data-dependent number of fragments, which allows systematic handling of symmetries, and the 6D poses are estimated from the links by a RANSAC-based fitting method. EPOS outperformed all RGB and most RGB-D and D methods on several standard datasets. Second, we present HashMatch, an RGB-D method that slides a window over the input image and searches for a match against templates, which are pre-generated by rendering 3D object models in different orientations. The method applies a cascade of evaluation stages to each window location, which avoids exhaustive matching against all templates. Third, we propose ObjectSynth, an approach to synthesize photorealistic images of 3D object models for training methods based on neural networks. The images yield substantial improvements compared to commonly used images of objects rendered on top of random photographs. Fourth, we introduce T-LESS, the first dataset for 6D object pose estimation that includes 3D models and RGB-D images of industry-relevant objects. Fifth, we define BOP, a benchmark that captures the status quo in the field. BOP comprises eleven datasets in a unified format, an evaluation methodology, an online evaluation system, and public challenges held at international workshops organized at the ICCV and ECCV conferences.
RheFrameDetect: A Text Classification System for Automatic Detection of Rhetorical Frames in AI from Open Sources
Ghosh, Saurav, Loustaunau, Philippe
Rhetorical Frames in AI can be thought of as expressions that describe AI development as a competition between two or more actors, such as governments or companies. Examples of such Frames include robotic arms race, AI rivalry, technological supremacy, cyberwarfare dominance and 5G race. Detection of Rhetorical Frames from open sources can help us track the attitudes of governments or companies towards AI, specifically whether attitudes are becoming more cooperative or competitive over time. Given the rapidly increasing volumes of open sources (online news media, twitter, blogs), it is difficult for subject matter experts to identify Rhetorical Frames in (near) real-time. Moreover, these sources are in general unstructured (noisy) and therefore, detecting Frames from these sources will require state-of-the-art text classification techniques. In this paper, we develop RheFrameDetect, a text classification system for (near) real-time capture of Rhetorical Frames from open sources. Given an input document, RheFrameDetect employs text classification techniques at multiple levels (document level and paragraph level) to identify all occurrences of Frames used in the discussion of AI. We performed extensive evaluation of the text classification techniques used in RheFrameDetect against human annotated Frames from multiple news sources. To further demonstrate the effectiveness of RheFrameDetect, we show multiple case studies depicting the Frames identified by RheFrameDetect compared against human annotated Frames.
Mom claims Amazon Alexa suggested dangerous online challenge to child
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A mom's Amazon Echo reportedly recommended a dangerous online challenge to her child earlier this week. On Sunday, Kristin Livdahl tweeted a screengrab writing, "My 10 year old just asked Alexa on our Echo for a challenge and this is what she said." In the screengrab, which appears to show Livdahl's Echo's activity, the apparent command says, "Tell me a challenge to do" followed by the Echo's response.
If you need your music everywhere, Apple Music's "voice plan" isn't for you
And while the idea of a discounted streaming service you have to talk to may seem a little odd, even that isn't all that weird. Two years ago, Amazon launched a cheaper version of Music Unlimited service that only runs on Echo speakers, and Apple Music's voice plan seemed tailor-made to compete with it on affordable smart speakers like HomePod minis. And if the promise of cheaper music access gets more people talking to Siri, that could mean more training data Apple could use to improve its voice assistant's performance down the road.
6 Photography Trends 2022: AI Modes, Post-Instagram, Camera Shortages, More
The start of a new year implies one thing: new and exciting aesthetic photography trends to embrace. Even the quickest telephoto lenses can't keep up with the fast-paced world of photography. However, it's also an exciting period for photographers, and it's set to get even more interesting to experience the photography trends in 2022. This year has delivered some ground-breaking equipment that continues to alter both cameras and photography, whether users shoot with a smartphone, film camera, or mirrorless powerhouse; these trends are sought to be seen in 2022. Desktop photo editors like Photoshop and Luminar AI are trying to outdo one other with new AI tactics, similar to the computational photography battles being conducted between smartphones.