Goto

Collaborating Authors

 Personal


AI 'godfather' predicts another revolution in the tech in next five years

The Guardian

One of the "godfathers" of modern artificial intelligence has predicted a further revolution in the technology by the end of the decade, and says current systems are too limited to create domestic robots and fully automated cars. Yann LeCun, the chief AI scientist at Mark Zuckerberg's Meta, said new breakthroughs are needed in order for the systems to understand and interact with the physical world. LeCun spoke as one of seven engineers who were awarded the 500,000 Queen Elizabeth prize for engineering on Tuesday for their contributions to machine learning, a cornerstone of AI. Recent breakthroughs in the sector, led by the launch of OpenAI's ChatGPT chatbot, have heightened expectations – and fears – of systems gaining human levels of intelligence. However, LeCun said there was some way to go before AIs matched humans or animals, with the current cutting-edge technology excelling at "manipulating language" but not at understanding the physical world.


Stuart J. Russell wins 2025 AAAI Award for Artificial Intelligence for the Benefit of Humanity

AIHub

The AAAI Award for Artificial Intelligence for the Benefit of Humanity recognizes positive impacts of artificial intelligence to protect, enhance, and improve human life in meaningful ways with long-lived effects. The award is given annually at the conference for the Association for the Advancement of Artificial Intelligence (AAAI). This year, the AAAI Awards Committee has announced that the 2025 recipient of the award and 25,000 prize is Stuart J. Russell, "for his work on the conceptual and theoretical foundations of provably beneficial AI and his leadership in creating the field of AI safety". Stuart will give an invited talk at AAAI 2025 entitled "Can AI Benefit Humanity?" Stuart J. Russell is a Distinguished Professor of Computer Science at the University of California, Berkeley, and holds the Michael H. Smith and Lotfi A. Zadeh Chair in Engineering.


France pitches AI summit as 'wake-up call' for Europe

The Japan Times

France hosts top tech players next week at an artificial intelligence summit meant as a "wake-up call" for Europe as it struggles with AI challenges from the United States and China. Players from across the sector and representatives from 80 nations will gather in the French capital on Feb. 10 and 11 in the sumptuous Grand Palais, built for the 1900 Universal Exhibition.


Consistent Client Simulation for Motivational Interviewing-based Counseling

arXiv.org Artificial Intelligence

Simulating human clients in mental health counseling is crucial for training and evaluating counselors (both human or simulated) in a scalable manner. Nevertheless, past research on client simulation did not focus on complex conversation tasks such as mental health counseling. In these tasks, the challenge is to ensure that the client's actions (i.e., interactions with the counselor) are consistent with with its stipulated profiles and negative behavior settings. In this paper, we propose a novel framework that supports consistent client simulation for mental health counseling. Our framework tracks the mental state of a simulated client, controls its state transitions, and generates for each state behaviors consistent with the client's motivation, beliefs, preferred plan to change, and receptivity. By varying the client profile and receptivity, we demonstrate that consistent simulated clients for different counseling scenarios can be effectively created. Both our automatic and expert evaluations on the generated counseling sessions also show that our client simulation method achieves higher consistency than previous methods.


Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose Estimation

arXiv.org Artificial Intelligence

Nine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization to intra-class unknown objects. However, these methods require manual collection and labeling of large-scale real-world training data. To address this problem, we introduce a diffusion-based paradigm for domain-generalized category-level 9-DoF object pose estimation. Our motivation is to leverage the latent generalization ability of the diffusion model to address the domain generalization challenge in object pose estimation. This entails training the model exclusively on rendered synthetic data to achieve generalization to real-world scenes. We propose an effective diffusion model to redefine 9-DoF object pose estimation from a generative perspective. Our model does not require any 3D shape priors during training or inference. By employing the Denoising Diffusion Implicit Model, we demonstrate that the reverse diffusion process can be executed in as few as 3 steps, achieving near real-time performance. Finally, we design a robotic grasping system comprising both hardware and software components. Through comprehensive experiments on two benchmark datasets and the real-world robotic system, we show that our method achieves state-of-the-art domain generalization performance. Our code will be made public at https://github.com/CNJianLiu/Diff9D.


Source Code by Bill Gates review – growing pains of a computer geek

The Guardian

The enduring mystery about William Henry Gates III is this: how did a precocious and sometimes obnoxious kid evolve into a billionaire tech lord and then into an elder statesman and philanthropist? This book gives us only the first part of the story, tracing Gates's evolution from birth in 1955 to the founding of Microsoft in 1975. For the next part of the story, we will just have to wait for the sequel. In a way, the volume's title describes it well. In the era before machine learning and AI, when computer programs were exclusively written by humans, the term "source code" meant something.


Wizard of Shopping: Target-Oriented E-commerce Dialogue Generation with Decision Tree Branching

arXiv.org Artificial Intelligence

The goal of conversational product search (CPS) is to develop an intelligent, chat-based shopping assistant that can directly interact with customers to understand shopping intents, ask clarification questions, and find relevant products. However, training such assistants is hindered mainly due to the lack of reliable and large-scale datasets. Prior human-annotated CPS datasets are extremely small in size and lack integration with real-world product search systems. We propose a novel approach, TRACER, which leverages large language models (LLMs) to generate realistic and natural conversations for different shopping domains. TRACER's novelty lies in grounding the generation to dialogue plans, which are product search trajectories predicted from a decision tree model, that guarantees relevant product discovery in the shortest number of search conditions. We also release the first target-oriented CPS dataset Wizard of Shopping (WoS), containing highly natural and coherent conversations (3.6k) from three shopping domains. Finally, we demonstrate the quality and effectiveness of WoS via human evaluations and downstream tasks.


AIhub monthly digest: January 2025 – artists' perspectives on GenAI, biomedical knowledge graphs, and ML for studying greenhouse gas emissions

AIHub

Welcome to our monthly digest, where you can catch up with any AIhub stories you may have missed, peruse the latest news, recap recent events, and more. This month, we hear about artists' perspectives on generative AI, learn how to explain neural networks using logic, and find out about using machine learning for studying greenhouse gas emissions. We caught up with Erica Kimei to find out about her research studying gas emissions from agriculture, specifically ruminant livestock. Erica combines machine learning and remote sensing technology to monitor and forecast such emissions. This interview is the latest in our series highlighting members of the AfriClimate AI community.


International AI Safety Report

arXiv.org Artificial Intelligence

I am honoured to present the International AI Safety Report. It is the work of 96 international AI experts who collaborated in an unprecedented effort to establish an internationally shared scientific understanding of risks from advanced AI and methods for managing them. We embarked on this journey just over a year ago, shortly after the countries present at the Bletchley Park AI Safety Summit agreed to support the creation of this report. Since then, we published an Interim Report in May 2024, which was presented at the AI Seoul Summit. We are now pleased to publish the present, full report ahead of the AI Action Summit in Paris in February 2025. Since the Bletchley Summit, the capabilities of general-purpose AI, the type of AI this report focuses on, have increased further. For example, new models have shown markedly better performance at tests of Professor Yoshua Bengio programming and scientific reasoning.


2SSP: A Two-Stage Framework for Structured Pruning of LLMs

arXiv.org Artificial Intelligence

We propose a novel Two-Stage framework for Structured Pruning (2SSP) for pruning Large Language Models (LLMs), which combines two different strategies of pruning, namely Width and Depth Pruning. The first stage (Width Pruning) removes entire neurons, hence their corresponding rows and columns, aiming to preserve the connectivity among the pruned structures in the intermediate state of the Feed-Forward Networks in each Transformer block. This is done based on an importance score measuring the impact of each neuron over the output magnitude. The second stage (Depth Pruning), instead, removes entire Attention submodules. This is done by applying an iterative process that removes the Attention submodules with the minimum impact on a given metric of interest (in our case, perplexity). We also propose a novel mechanism to balance the sparsity rate of the two stages w.r.t. to the desired global sparsity. We test 2SSP on four LLM families and three sparsity rates (25\%, 37.5\%, and 50\%), measuring the resulting perplexity over three language modeling datasets as well as the performance over six downstream tasks. Our method consistently outperforms five state-of-the-art competitors over three language modeling and six downstream tasks, with an up to two-order-of-magnitude gain in terms of pruning time. The code is available at available at \url{https://github.com/FabrizioSandri/2SSP}.