Technology
Learning to Communicate with Deep Multi-Agent Reinforcement Learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, Shimon Whiteson
We consider the problem of multiple agents sensing and acting in environments with the goal of maximising their shared utility. In these environments, agents must learn communication protocols in order to share information that is needed to solve the tasks. By embracing deep neural networks, we are able to demonstrate endto-end learning of protocols in complex environments inspired by communication riddles and multi-agent computer vision problems with partial observability. We propose two approaches for learning in these domains: Reinforced Inter-Agent Learning (RIAL) and Differentiable Inter-Agent Learning (DIAL). The former uses deep Q-learning, while the latter exploits the fact that, during learning, agents can backpropagate error derivatives through (noisy) communication channels. Hence, this approach uses centralised learning but decentralised execution. Our experiments introduce new environments for studying the learning of communication protocols and present a set of engineering innovations that are essential for success in these domains.
Without-Replacement Sampling for Stochastic Gradient Methods Ohad Shamir Department of Computer Science and Applied Mathematics Weizmann Institute of Science Rehovot, Israel ohad.shamir@weizmann.ac.il
Stochastic gradient methods for machine learning and optimization problems are usually analyzed assuming data points are sampled with replacement. In contrast, sampling without replacement is far less understood, yet in practice it is very common, often easier to implement, and usually performs better. In this paper, we provide competitive convergence guarantees for without-replacement sampling under several scenarios, focusing on the natural regime of few passes over the data. Moreover, we describe a useful application of these results in the context of distributed optimization with randomly-partitioned data, yielding a nearly-optimal algorithm for regularized least squares (in terms of both communication complexity and runtime complexity) under broad parameter regimes. Our proof techniques combine ideas from stochastic optimization, adversarial online learning and transductive learning theory, and can potentially be applied to other stochastic optimization and learning problems.
Sample Complexity of Automated Mechanism Design
Maria-Florina F. Balcan, Tuomas Sandholm, Ellen Vitercik
The design of revenue-maximizing combinatorial auctions, i.e. multi-item auctions over bundles of goods, is one of the most fundamental problems in computational economics, unsolved even for two bidders and two items for sale. In the traditional economic models, it is assumed that the bidders' valuations are drawn from an underlying distribution and that the auction designer has perfect knowledge of this distribution. Despite this strong and oftentimes unrealistic assumption, it is remarkable that the revenue-maximizing combinatorial auction remains unknown. In recent years, automated mechanism design has emerged as one of the most practical and promising approaches to designing high-revenue combinatorial auctions. The most scalable automated mechanism design algorithms take as input samples from the bidders' valuation distribution and then search for a high-revenue auction in a rich auction class. In this work, we provide the first sample complexity analysis for the standard hierarchy of deterministic combinatorial auction classes used in automated mechanism design. In particular, we provide tight sample complexity bounds on the number of samples needed to guarantee that the empirical revenue of the designed mechanism on the samples is close to its expected revenue on the underlying, unknown distribution over bidder valuations, for each of the auction classes in the hierarchy. In addition to helping set automated mechanism design on firm foundations, our results also push the boundaries of learning theory. In particular, the hypothesis functions used in our contexts are defined through multi-stage combinatorial optimization procedures, rather than simple decision boundaries, as are common in machine learning.
Optimizing affinity-based binary hashing using auxiliary coordinates
Ramin Raziperchikolaei, Miguel A. Carreira-Perpinan
In supervised binary hashing, one wants to learn a function that maps a highdimensional feature vector to a vector of binary codes, for application to fast image retrieval. This typically results in a difficult optimization problem, nonconvex and nonsmooth, because of the discrete variables involved. Much work has simply relaxed the problem during training, solving a continuous optimization, and truncating the codes a posteriori. This gives reasonable results but is quite suboptimal. Recent work has tried to optimize the objective directly over the binary codes and achieved better results, but the hash function was still learned a posteriori, which remains suboptimal. We propose a general framework for learning hash functions using affinity-based loss functions that uses auxiliary coordinates. This closes the loop and optimizes jointly over the hash functions and the binary codes so that they gradually match each other. The resulting algorithm can be seen as an iterated version of the procedure of optimizing first over the codes and then learning the hash function. Compared to this, our optimization is guaranteed to obtain better hash functions while being not much slower, as demonstrated experimentally in various supervised datasets.
Stochastic Online AUC Maximization
Yiming Ying, Longyin Wen, Siwei Lyu
Area under ROC (AUC) is a metric which is widely used for measuring the classification performance for imbalanced data. It is of theoretical and practical interest to develop online learning algorithms that maximizes AUC for large-scale data. A specific challenge in developing online AUC maximization algorithm is that the learning objective function is usually defined over a pair of training examples of opposite classes, and existing methods achieves on-line processing with higher space and time complexity. In this work, we propose a new stochastic online algorithm for AUC maximization. In particular, we show that AUC optimization can be equivalently formulated as a convex-concave saddle point problem. From this saddle representation, a stochastic online algorithm (SOLAM) is proposed which has time and space complexity of one datum. We establish theoretical convergence of SOLAM with high probability and demonstrate its effectiveness on standard benchmark datasets.
Young Chinese use AI to launch one-person firms over job anxiety
One-person company SoloNest sounder Karen Dai preparing for a coffee chat at a conference room in Shanghai on April 12. | AFP-JIJI Shanghai - Young Chinese, many who fear age discrimination in their workplace after turning 35, are increasingly starting one-person companies that have artificial intelligence do most of the work. Smaller startups are already in vogue in Silicon Valley and elsewhere, with rapidly advancing AI tools seen as a welcome teammate even as they threaten layoffs at existing firms. More young people in China are subscribing to the model, as cities pledge millions of dollars in funding and rent subsidies for such ventures, in alignment with Beijing's political goal of technological self-reliance. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right.
Pentagon seeks 75 billion for drones in record budget ask
A soldier carries a drone during a military parade in Washington on June 14, 2025. The Pentagon's largest-ever budget request earmarks $75 billion for drones and technologies to counter them, mainly for a massive increase for a little-known office working with U.S. commandos to test and evaluate various systems, according to defense officials. The drone-funding proposal includes $54.6 billion for the Defense Autonomous Working Group, or DAWG, from just $225.9 million this year. That would appear to be the largest single year-over-year boost of any defense program or office, meaning it's likely to draw particular congressional and public scrutiny in an already eye-catching $1.5 trillion request that's 42% larger than this year's budget. The big boost for the Pentagon's little-known drone unit comes as the U.S. and Israeli war against Iran illustrates how drones can help level the playing field against even the world's most well-funded armed forces.
A drone delivered her lethal dose of fentanyl in a church parking lot. Now her dealer is going to prison
Things to Do in L.A. Tap to enable a layout that focuses on the article. A drone delivered her lethal dose of fentanyl in a church parking lot. The Drug Enforcement Administration was among agencies involved in the investigation. This is read by an automated voice. Please report any issues or inconsistencies here .