Goto

Collaborating Authors

 xpo


Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

arXiv.org Machine Learning

Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment. We consider online exploration in RLHF, which exploits interactive access to human or AI feedback by deliberately encouraging the model to produce diverse, maximally informative responses. By allowing RLHF to confidently stray from the pre-trained model, online exploration offers the possibility of novel, potentially super-human capabilities, but its full potential as a paradigm for language model training has yet to be realized, owing to computational and statistical bottlenecks in directly adapting existing reinforcement learning techniques. We propose a new algorithm for online exploration in RLHF, Exploratory Preference Optimization (XPO), which is simple and practical -- a one-line change to (online) Direct Preference Optimization (DPO; Rafailov et al., 2023) -- yet enjoys the strongest known provable guarantees and promising empirical performance. XPO augments the DPO objective with a novel and principled exploration bonus, empowering the algorithm to explore outside the support of the initial model and human feedback data. In theory, we show that XPO is provably sample-efficient and converges to a near-optimal language model policy under natural exploration conditions, irrespective of whether the initial model has good coverage. Our analysis, which builds on the observation that DPO implicitly performs a form of $Q^{\star}$-approximation (or, Bellman error minimization), combines previously disparate techniques from language modeling and theoretical reinforcement learning in a serendipitous fashion through the perspective of KL-regularized Markov decision processes. Empirically, we find that XPO is more sample-efficient than non-exploratory DPO variants in a preliminary evaluation.


From Reindeer to Robots, Automation Set to Deliver This Holiday Season

WSJ.com: WSJD - Technology

"It's a fight for talent…It's like'Game of Thrones' out there," Erik Caldwell, chief operating officer for supply chain in the Americas and Asia Pacific at XPO Logistics Inc., XPO 2.83% said at an industry conference earlier this year, discussing the company's use of robots to fulfill online orders. The use of robotics and other automation technology in industrial operations is growing, although the vast majority of warehouse work remains largely manual. About 16.5% of organizations across several industries including warehousing are now using commercial service robots, and 21.5% have them in pilot programs, according to a 2018 survey of 600 respondents by research firm IDC. The holiday shopping season highlights a warehouse-worker squeeze that is driving more logistics operators to embrace automation, as the growth of online commerce pushes more retail sales from storefronts to distribution centers. Online fulfillment centers--where companies like Amazon.com Inc. AMZN -0.94% pick, pack and ship consumer orders--require two to three times as many workers as traditional warehouses.


XPO to deploy 5,000 robots in its warehouses

#artificialintelligence

XPO Logistics is to deploy 5,000 "intelligent robots" throughout its logistics sites in North America and Europe. The robots, designed to "collaborate with humans", will supplement XPO's existing workforce and support future growth, said the company in statement. XPO has a strategic partnership with robotics manufacturer GreyOrange that makes XPO the exclusive logistics provider for use of its robots in North America, the UK and eight other European countries. Bradley Jacobs, chief executive officer of XPO Logistics, said, "We've developed our logistics technology to integrate the latest intelligent automation and adapt it at lightning speed. This allows us to dramatically improve fulfillment time and cut costs. "The addition of 5,000 collaborative robots will make our logistics operations safer and more productive in picking, packing and sortation.


Solving Bongard Problems with a Visual Language and Pragmatic Reasoning

arXiv.org Artificial Intelligence

More than 50 years ago Bongard introduced 100 visual concept learning problems as a testbed for intelligent vision systems. These problems are now known as Bongard problems. Although they are well known in the cognitive science and AI communities only moderate progress has been made towards building systems that can solve a substantial subset of them. In the system presented here, visual features are extracted through image processing and then translated into a symbolic visual vocabulary. We introduce a formal language that allows representing complex visual concepts based on this vocabulary. Using this language and Bayesian inference, complex visual concepts can be induced from the examples that are provided in each Bongard problem. Contrary to other concept learning problems the examples from which concepts are induced are not random in Bongard problems, instead they are carefully chosen to communicate the concept, hence requiring pragmatic reasoning. Taking pragmatic reasoning into account we find good agreement between the concepts with high posterior probability and the solutions formulated by Bongard himself. While this approach is far from solving all Bongard problems, it solves the biggest fraction yet.