sipo
AI dives into a sea of data, from plankton to pollution
When asked why Jean-Olivier Irisson, a scientist at Sorbonne Université in Paris, decided to dedicate his life to studying microscopic creatures in the sea, his answer was simple: "They are beautiful." Beauty may not be the first thing that comes to mind when we think of plankton - organisms that drift in water and come in an extraordinary variety of shapes and sizes. But images by Irisson's team tell a different story. Shown in striking blues and oranges, as well as black and white, they reveal an unfamiliar and strangely beautiful world. "This one served as the model for the head of the creature in the Alien movie franchise," Irisson said, pointing to one particularly unusual specimen.
Iteratively Learn Diverse Strategies with State Distance Information
In complex reinforcement learning (RL) problems, policies with similar rewards may have substantially different behaviors. It remains a fundamental challenge to optimize rewards while also discovering as many strategies as possible, which can be crucial in many practical applications. Our study examines two design choices for tackling this challenge, i.e., and . First, we find that with existing diversity measures, visually indistinguishable policies can still yield high diversity scores. To accurately capture the behavioral difference, we propose to incorporate the state-space distance information into the diversity measure.
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
Li, Moxin, Zhang, Yuantao, Wang, Wenjie, Shi, Wentao, Liu, Zhuo, Feng, Fuli, Chua, Tat-Seng
Multi-Objective Alignment (MOA) aims to align LLMs' responses with multiple human preference objectives, with Direct Preference Optimization (DPO) emerging as a prominent approach. However, we find that DPO-based MOA approaches suffer from widespread preference conflicts in the data, where different objectives favor different responses. This results in conflicting optimization directions, hindering the optimization on the Pareto Front. To address this, we propose to construct Pareto-optimal responses to resolve preference conflicts. To efficiently obtain and utilize such responses, we propose a self-improving DPO framework that enables LLMs to self-generate and select Pareto-optimal responses for self-supervised preference alignment. Extensive experiments on two datasets demonstrate the superior Pareto Front achieved by our framework compared to various baselines. Code is available at \url{https://github.com/zyttt-coder/SIPO}.
Iteratively Learn Diverse Strategies with State Distance Information
In complex reinforcement learning (RL) problems, policies with similar rewards may have substantially different behaviors. It remains a fundamental challenge to optimize rewards while also discovering as many diverse strategies as possible, which can be crucial in many practical applications. Our study examines two design choices for tackling this challenge, i.e., diversity measure and computation framework. First, we find that with existing diversity measures, visually indistinguishable policies can still yield high diversity scores. To accurately capture the behavioral difference, we propose to incorporate the state-space distance information into the diversity measure.
China's research institutes file more AI patents than businesses
Chinese academic institutions are more prolific patent filers in the artificial intelligence (AI) area than domestic companies, according to China's State Intellectual Property Office (SIPO). SIPO shared the statement, based on a release from China IP News, on Wednesday, August 1. The release is based on "China's AI Development Report 2018", which was recently published by Tsinghua University, in Beijing. The university's report revealed that the most prolific filers in AI tend to come from research institutions, such as universities. Unlike in other countries, industry players in China file fewer patents in the AI sphere than those in research institutions. The country's "top IT giants" such as Alibaba and Tencent are "overwhelmed" by the filings of foreign companies, such as IBM and Microsoft, SIPO said.