Technology
Contextual bandits with surrogate losses: Margin bounds and efficient algorithms
We use surrogate losses to obtain several new regret bounds and new algorithms for contextual bandit learning. Using the ramp loss, we derive a new margin-based regret bound in terms of standard sequential complexity measures of a benchmark class of real-valued regression functions. Using the hinge loss, we derive an efficient algorithm with a $\sqrt{dT}$-type mistake bound against benchmark policies induced by $d$-dimensional regressors. Under realizability assumptions, our results also yield classical regret bounds.
Contour location via entropy reduction leveraging multiple information sources
We introduce an algorithm to locate contours of functions that are expensive to evaluate. The problem of locating contours arises in many applications, including classification, constrained optimization, and performance analysis of mechanical and dynamical systems (reliability, probability of failure, stability, etc.). Our algorithm locates contours using information from multiple sources, which are available in the form of relatively inexpensive, biased, and possibly noisy approximations to the original function. Considering multiple information sources can lead to significant cost savings. We also introduce the concept of contour entropy, a formal measure of uncertainty about the location of the zero contour of a function approximated by a statistical surrogate model.
KDGAN: Knowledge Distillation with Generative Adversarial Networks
Knowledge distillation (KD) aims to train a lightweight classifier suitable to provide accurate inference with constrained resources in multi-label learning. Instead of directly consuming feature-label pairs, the classifier is trained by a teacher, i.e., a high-capacity model whose training may be resource-hungry. The accuracy of the classifier trained this way is usually suboptimal because it is difficult to learn the true data distribution from the teacher. An alternative method is to adversarially train the classifier against a discriminator in a two-player game akin to generative adversarial networks (GAN), which can ensure the classifier to learn the true data distribution at the equilibrium of this game. However, it may take excessively long time for such a two-player game to reach equilibrium due to high-variance gradient updates.
Hierarchical Reinforcement Learning for Zero-shot Generalization with Subtask Dependencies
We introduce a new RL problem where the agent is required to generalize to a previously-unseen environment characterized by a subtask graph which describes a set of subtasks and their dependencies. Unlike existing hierarchical multitask RL approaches that explicitly describe what the agent should do at a high level, our problem only describes properties of subtasks and relationships among them, which requires the agent to perform complex reasoning to find the optimal subtask to execute. To solve this problem, we propose a neural subtask graph solver (NSGS) which encodes the subtask graph using a recursive neural network embedding. To overcome the difficulty of training, we propose a novel non-parametric gradient-based policy, graph reward propagation, to pre-train our NSGS agent and further finetune it through actor-critic method. The experimental results on two 2D visual domains show that our agent can perform complex reasoning to find a near-optimal way of executing the subtask graph and generalize well to the unseen subtask graphs. In addition, we compare our agent with a Monte-Carlo tree search (MCTS) method showing that our method is much more efficient than MCTS, and the performance of NSGS can be further improved by combining it with MCTS.
Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks
Real-world image recognition is often challenged by the variability of visual styles including object textures, lighting conditions, filter effects, etc. Although these variations have been deemed to be implicitly handled by more training data and deeper networks, recent advances in image style transfer suggest that it is also possible to explicitly manipulate the style information. Extending this idea to general visual recognition problems, we present Batch-Instance Normalization (BIN) to explicitly normalize unnecessary styles from images. Considering certain style features play an essential role in discriminative tasks, BIN learns to selectively normalize only disturbing styles while preserving useful styles. The proposed normalization module is easily incorporated into existing network architectures such as Residual Networks, and surprisingly improves the recognition performance in various scenarios. Furthermore, experiments verify that BIN effectively adapts to completely different tasks like object classification and style transfer, by controlling the trade-off between preserving and removing style variations. BIN can be implemented with only a few lines of code using popular deep learning frameworks.
On GANs and GMMs
A longstanding problem in machine learning is to find unsupervised methods that can learn the statistical structure of high dimensional signals. In recent years, GANs have gained much attention as a possible solution to the problem, and in particular have shown the ability to generate remarkably realistic high resolution sampled images. At the same time, many authors have pointed out that GANs may fail to model the full distribution (mode collapse) and that using the learned models for anything other than generating samples may be very difficult. In this paper, we examine the utility of GANs in learning statistical models of images by comparing them to perhaps the simplest statistical model, the Gaussian Mixture Model. First, we present a simple method to evaluate generative models based on relative proportions of samples that fall into predetermined bins.
Self-Supervised Generation of Spatial Audio for 360 Video
We introduce an approach to convert mono audio recorded by a 360 video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360 video viewing, but spatial audio microphones are still rare in current 360 video production. Our system consists of end-to-end trainable neural networks that separate individual sound sources and localize them on the viewing sphere, conditioned on multi-modal analysis from the audio and 360 video frames. We introduce several datasets, including one filmed ourselves, and one collected in-the-wild from YouTube, consisting of 360 videos uploaded with spatial audio. During training, ground truth spatial audio serves as self-supervision and a mixed down mono track forms the input to our network. Using our approach we show that it is possible to infer the spatial localization of sounds based only on a synchronized 360 video and the mono audio track.
Synthesized Policies for Transfer and Adaptation across Tasks and Environments
The ability to transfer in reinforcement learning is key towards building an agent of general artificial intelligence. In this paper, we consider the problem of learning to simultaneously transfer across both environments and tasks, probably more importantly, by learning from only sparse (environment, task) pairs out of all the possible combinations. We propose a novel compositional neural network architecture which depicts a meta rule for composing policies from environment and task embeddings. Notably, one of the main challenges is to learn the embeddings jointly with the meta rule. We further propose new training methods to disentangle the embeddings, making them both distinctive signatures of the environments and tasks and effective building blocks for composing the policies. Experiments on GridWorld and THOR, of which the agent takes as input an egocentric view, show that our approach gives rise to high success rates on all the (environment, task) pairs after learning from only 40% of them.
Where OpenAI's technology could show up in Iran
Where OpenAI's technology could show up in Iran Three places to watch, from the margins of war to the center of combat. It's been just over two weeks since OpenAI reached a controversial agreement to allow the Pentagon to use its AI in classified environments. There are still pressing questions about what exactly OpenAI's agreement allows for; Sam Altman said the military can't use his company's technology to build autonomous weapons, but the agreement really just demands that the military follow its own (quite permissive) guidelines about such weapons. OpenAI's other main claim, that the agreement will prevent use of its technology for domestic surveillance, appears equally dubious . It's not the first tech giant to embrace military contracts it had once vowed never to enter into, but the speed of the pivot was notable. Perhaps it's just about money; OpenAI is spending lots on AI training and is on the hunt for more revenue (from sources including ads).
Billionaire Peter Thiel holds secret 'Antichrist' meetings on the Vatican's doorstep
Trump announces White House Chief of Staff Susie Wiles diagnosed with'early stage' breast cancer Trump's billionaire adviser publicly rebukes Iran war as JD Vance camp erupts over Israel nuke threat Kristi Noem referred for criminal investigation after'lying under oath' about $220M vanity scheme You don't have to fly to Turkey or Thailand... and can do it on your lunch break! Diet that cures pain and inflammation, devised by experts: Constant sickness and aching joints are the first signs of problems that left unchecked can turn deadly. Timothee Chalamet and Kylie Jenner's strained Oscars chat decoded by lip-reader as he gets snubbed and mocked The snubbed A-lister, drunken pics and C-List stars who plagued the most'exclusive' party: All the Oscars gossip Hollywood didn't want you to see at very messy afterparty Proof Leonardo DiCaprio sent a CLONE to the Oscars... alarming truth about Teyana Taylor's blow up... and a very dirty Barbra Streisand rumor: KENNEDY's most brutal review yet NYC's smiling socialist mayor is VERY different behind the scenes, as progressives who crossed him allege tyrannical and ruthless behavior Awful Timothee Chalamet's ego is bigger than Kylie's inflated butt... but it's so clear what's really going on here. Trump stunned by lurid rumor about Iran's new'gay' ayatollah Chilling new details of dismembered Emily Pike's final hours after she was snatched in Arizona desert and man detectives now believe murdered her'It's like he was possessed': Terrifying moment Alexander brother turned into a'monster' and raped me... and the four chilling words he said after horror attack - alleged victim claims After Oscars 2026, the whispered fear among Hollywood doctors is now massive... this is so much bigger than Ozempic. A-list stars ditch formal Oscars red carpet dresses for sexy party looks - with Jeff Goldblum's wife Emilie Livingston, Heidi Klum, Amelia Gray Hamlin and Kate Hudson turning up the heat at Vanity Fair bash Shock as man begs for death penalty for HIMSELF after pinning dead pastor's hands to wall and targeting other religious leaders How Oscars 2026 proved Hollywood has overdosed on Ozempic: Leading doctors name stars now at'extreme' risk... and reveal terrifying new side effects Billionaire Peter Thiel holds secret'Antichrist' meetings on the Vatican's doorstep READ MORE: Catholic priest warns'the stage is set' for the rise of the Antichrist US billionaire Peter Thiel is hosting a series of closed-door lectures in Rome on the doorstep of the Vatican, focused on the concept of the Antichrist.