Deep Learning
PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
Lykov, Artem, Sam, Jeffrin, Nguyen, Hung Khang, Kozlovskiy, Vladislav, Mahmoud, Yara, Serpiva, Valerii, Cabrera, Miguel Altamirano, Konenkov, Mikhail, Tsetserukou, Dzmitry
Abstract-- We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video demonstrations of candidate trajectories, executes them on the robot, and iteratively re-plans in response to failures. This approach enables robust recovery from execution errors. We evaluate PhysicalAgent across multiple perceptual modalities (egocentric, third-person, and simulated) and robotic embodiments (bimanual UR3, Unitree G1 humanoid, simulated GR1), comparing against state-of-the-art task-specific baselines. Experiments demonstrate that our method consistently outperforms prior approaches, achieving up to 83% success on human-familiar tasks. Physical trials reveal that first-attempt success is limited (20-30%), yet iterative correction increases overall success to 80% across platforms. These results highlight the potential of video-based generative reasoning for general-purpose robotic manipulation and underscore the importance of iterative execution for recovering from initial failures. Our framework paves the way for scalable, adaptable, and robust robot control. The rapid progress of large foundation models has transformed the design of AI agents.
Synthetic Data Generation for Screen Time and App Usage
Kruger, Gustavo, Sachdeva, Nikhil, Sobolev, Michael
Smartphone usage data can provide valuable insights for understanding interaction with technology and human behavior. However, collecting large-scale, in-the-wild smartphone usage logs is challenging due to high costs, privacy concerns, under representative user samples and biases like non-response that can skew results. These challenges call for exploring alternative approaches to obtain smartphone usage datasets. In this context, large language models (LLMs) such as Open AI's ChatGPT present a novel approach for synthetic smartphone usage data generation, addressing limitations of real-world data collection. We describe a case study on how four prompt strategies influenced the quality of generated smartphone usage data. We contribute with insights on prompt design and measures of data quality, reporting a prompting strategy comparison combining two factors, prompt level of detail (describing a user persona, describing the expected results characteristics) and seed data inclusion (with versus without an initial real usage example). Our findings suggest that using LLMs to generate structured and behaviorally plausible smartphone use datasets is feasible for some use cases, especially when using detailed prompts. Challenges remain in capturing diverse nuances of human behavioral patterns in a single synthetic dataset, and evaluating tradeoffs between data fidelity and diversity, suggesting the need for use-case-specific evaluation metrics and future research with more diverse seed data and different LLM models.
Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation
Vaish, Puru, Meister, Felix, Heimann, Tobias, Brune, Christoph, Wolterink, Jelmer M.
Many recent approaches in representation learning implicitly assume that uncorrelated views of a data point are sufficient to learn meaningful representations for various downstream tasks. In this work, we challenge this assumption and demonstrate that meaningful structure in the latent space does not emerge naturally. Instead, it must be explicitly induced. W e propose a method that aligns representations from different views of the data to align complementary information without inducing false positives. Our experiments show that our proposed self-supervised learning method, Consistent View Alignment, improves performance for downstream tasks, highlighting the critical role of structured view alignment in learning effective representations. Our method achieved first and second place in the MICCAI 2025 SSL3D challenge when using a Primus vision transformer and ResEnc convolutional neural network, respectively.
Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
Kambara, Motonari, Sugiura, Komei
In this work, we address the problem of predicting the future success of open-vocabulary object manipulation tasks. Conventional approaches typically determine success or failure after the action has been carried out. However, they make it difficult to prevent potential hazards and rely on failures to trigger replanning, thereby reducing the efficiency of object manipulation sequences. To overcome these challenges, we propose a model, which predicts the alignment between a pre-manipulation egocentric image with the planned trajectory and a given natural language instruction. We introduce a Multi-Level Trajectory Fusion module, which employs a state-of-the-art deep state-space model and a transformer encoder in parallel to capture multi-level time-series self-correlation within the end effector trajectory. Our experimental results indicate that the proposed method outperformed existing methods, including foundation models.
Large Language Models Discriminate Against Speakers of German Dialects
Bui, Minh Duc, Holtermann, Carolin, Hofmann, Valentin, Lauscher, Anne, von der Wense, Katharina
Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022). However, despite cultural importance, individuals speaking dialects often face negative societal stereotypes. We examine whether such stereotypes are mirrored by large language models (LLMs). We draw on the sociolinguistic literature on dialect perception to analyze traits commonly associated with dialect speakers. Based on these traits, we assess the dialect naming bias and dialect usage bias expressed by LLMs in two tasks: an association task and a decision task. To assess a model's dialect usage bias, we construct a novel evaluation corpus that pairs sentences from seven regional German dialects (e.g., Alemannic and Bavarian) with their standard German counterparts. We find that: (1) in the association task, all evaluated LLMs exhibit significant dialect naming and dialect usage bias against German dialect speakers, reflected in negative adjective associations; (2) all models reproduce these dialect naming and dialect usage biases in their decision making; and (3) contrary to prior work showing minimal bias with explicit demographic mentions, we find that explicitly labeling linguistic demographics--German dialect speakers--amplifies bias more than implicit cues like dialect usage.
Findings of the Third Automatic Minuting (AutoMin) Challenge
Shinde, Kartik, Besacier, Laurent, Bojar, Ondrej, Thonet, Thibaut, Ghosal, Tirthankar
This paper presents the third edition of AutoMin, a shared task on automatic meeting summarization into minutes. In 2025, AutoMin featured the main task of minuting, the creation of structured meeting minutes, as well as a new task: question answering (QA) based on meeting transcripts. The minuting task covered two languages, English and Czech, and two domains: project meetings and European Parliament sessions. The QA task focused solely on project meetings and was available in two settings: monolingual QA in English, and cross-lingual QA, where questions were asked and answered in Czech based on English meetings. Participation in 2025 was more limited compared to previous years, with only one team joining the minuting task and two teams participating in QA. However, as organizers, we included multiple baseline systems to enable a comprehensive evaluation of current (2025) large language models (LLMs) on both tasks.
Circuit realization and hardware linearization of monotone operator equilibrium networks
--It is shown that the port behavior of a resistor-diode network corresponds to the solution of a ReLU monotone operator equilibrium network (a neural network in the limit of infinite depth), giving a parsimonious construction of a neural network in analog hardware. We furthermore show that the gradient of such a circuit can be computed directly in hardware, using a procedure we call hardware linearization . This allows the network to be trained in hardware, which we demonstrate with a device-level circuit simulation. We extend the results to cascades of resistor-diode networks, which can be used to implement feedforward and other asymmetric networks. We finally show that different nonlinear elements give rise to different activation functions, and introduce the novel diode ReLU which is induced by a non-ideal diode model. The idea of building a neural network in analog hardware is classical [1]-[5]. Since the discovery of semiconductor devices with memristive properties [6], and in light of the growing energy intensiveness of machine learning systems, there has been a resurgence of interest in building devices which incorporate analog memristive components and are specially suited for deep learning applications [7], [8]. One of the primary advantages of such devices is that memristors, and similar elements such as phase change memory, act as both memory and computational units. This allows the transport delay between memory and computation to be circumvented. A particularly successful design is to arrange a number of memristors in a crossbar array, which can be used to perform matrix-vector calculation in a single operation [9]-[12].
Bridging the Synthetic-Real Gap: Supervised Domain Adaptation for Robust Spacecraft 6-DoF Pose Estimation
Singh, Inder Pal, Chenni, Nidhal Eddine, Shabayek, Abd El Rahman, Rathinam, Arunkumar, Aouada, Djamila
The monocular vision-based 6-DoF pose estimation involves deducing the rotation and translation data of a target spacecraft from 2D images captured from the chaser spacecraft. Accurate estimation of the 6-DoF pose from 2D images is particularly challenging in the space environment due to illumination changes, specular reflections, limited texture, and significant variations in the apparent size of the target caused by changes in range during approach [5]. Early spacecraft pose estimation methods relied primarily on geometric computer vision techniques such as edge matching, template alignment, and photogrammetry-based measurements [6]. Although effective in controlled conditions, these methods are sensitive to noise and environmental variations, which limit their robustness in real on-orbit imagery. With the rise of deep learning (DL), fully end-to-end networks have been explored, mapping raw images directly to pose parameters [7, 8]. Such approaches can implicitly learn complex visual cues, but often require vast amounts of labelled data and can be less interpretable or adaptable to new spacecraft geometries. Alternate approaches to end-to-end approaches are hybrid modular approaches to estimate spacecraft poses (Figure 1). These methods combine data-driven feature extraction with geometric model-based solvers, exploiting the strengths of both paradigms [9, 10]. In a typical pipeline, the process begins with spacecraft localization, where a deep learning (DL) object detection model predicts a bounding box enclosing the target spacecraft (see Figure 1(a)).
Floating-Body Hydrodynamic Neural Networks
Zhang, Tianshuo, Zhai, Wenzhe, Yann, Rui, Gao, Jia, Cao, He, Xing, Xianglei
Fluid-structure interaction is common in engineering and natural systems, where floating-body motion is governed by added mass, drag, and background flows. Modeling these dissipative dynamics is difficult: black-box neural models regress state derivatives with limited interpretability and unstable long-horizon predictions. We propose Floating-Body Hydrodynamic Neural Networks (FHNN), a physics-structured framework that predicts interpretable hydrodynamic parameters such as directional added masses, drag coefficients, and a streamfunction-based flow, and couples them with analytic equations of motion. This design constrains the hypothesis space, enhances interpretability, and stabilizes integration. On synthetic vortex datasets, FHNN achieves up to an order-of-magnitude lower error than Neural ODEs, recovers physically consistent flow fields. Compared with Hamiltonian and Lagrangian neural networks, FHNN more effectively handles dissipative dynamics while preserving interpretability, which bridges the gap between black-box learning and transparent system identification.
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
Bian, Zhipeng, Zhu, Jieming, Xie, Xuyang, Dai, Quanyu, Zhao, Zhou, Dong, Zhenhua
The rapid advancement of generative AI technologies is driving the integration of diverse AI-powered services into smartphones, transforming how users interact with their devices. To simplify access to predefined AI services, this paper introduces MIRA, a pioneering framework for task instruction recommendation that enables intuitive one-touch AI tasking on smartphones. With MIRA, users can long-press on images or text objects to receive contextually relevant instruction recommendations for executing AI tasks. Our work introduces three key innovations: 1) A multimodal large language model (MLLM)-based recommendation pipeline with structured reasoning to extract key entities, infer user intent, and generate precise instructions; 2) A template-augmented reasoning mechanism that integrates high-level reasoning templates, enhancing task inference accuracy; 3) A prefix-tree-based constrained decoding strategy that restricts outputs to predefined instruction candidates, ensuring coherent and intent-aligned suggestions. Through evaluation using a real-world annotated datasets and a user study, MIRA has demonstrated substantial improvements in the accuracy of instruction recommendation. The encouraging results highlight MIRA's potential to revolutionize the way users engage with AI services on their smartphones, offering a more seamless and efficient experience.