Africa
Single-grasp deformable object discrimination: the effect of gripper morphology, sensing modalities, and action parameters
Pliska, Michal, Patni, Shubhan, Mares, Michal, Stoudek, Pavel, Straka, Zdenek, Stepanova, Karla, Hoffmann, Matej
In haptic object discrimination, the effect of gripper embodiment, action parameters, and sensory channels has not been systematically studied. We used two anthropomorphic hands and two 2-finger grippers to grasp two sets of deformable objects. On the object classification task, we found: (i) among classifiers, SVM on sensory features and LSTM on raw time series performed best across all grippers; (ii) faster compression speeds degraded performance; (iii) generalization to different grasping configurations was limited; transfer to different compression speeds worked well for the Barrett Hand only. Visualization of the feature spaces using PCA showed that the gripper morphology and the action parameters were the main source of variance, rendering generalization across embodiment or grasp configurations very hard. On the highly challenging dataset consisting of polyurethane foams alone, only the Barrett Hand achieved excellent performance. Tactile sensors can thus provide a key advantage even if recognition is based on stiffness rather than shape. The dataset with 24000 measurements is publicly available.
Biden repeats dubious claim about son's death in call to fallen service member's family: 'The nerve'
During a call with the parents of fallen service member Spc. Kennedy Ladon Sanders, Biden claimed he "lost" his son, Beau Biden, to the war in Iraq. President Biden repeated a dubious claim about the death of his son, Beau Biden, during a call with the parents of a U.S. service member who was recently killed in an attack on a base in Jordan near the border with Syria. While speaking on Tuesday to the parents of 24-year-old Specialist Kennedy Ladon Sanders, who lost her life in an Iran-backed drone strike this month in northeast Jordan that killed three service members total and injured 25 others, Biden said he lost his son to the war in Iraq. During the call, which was first shared by the Atlanta Journal Constitution, Biden told Shawn Sanders and Oneida Oliver-Sanders that their daughter was being posthumously promoted to sergeant.
Google reveals another text-to-image generative AI tool, ImageFX
Google is rolling out a swathe of updates on the generative AI front, including a new text-to-image tool. What's different about ImageFX is that it has an interface that features "expressive chips." The idea here is that these will help you "quickly experiment with adjacent dimensions of your creation and ideas." Alongside the debut of ImageFX, Google says it has improved MusicFX and TextFX. The company's claims that it's made upgrades to the MusicLM model that include faster generation of music and higher-quality audio, along with new features.
US military targets 10 Houthi drones in new Yemen strikes
The United States military has carried out new strikes against 10 drones belonging to the Iran-aligned Houthi rebels in Yemen as well as a ground control centre. On Thursday, US forces targeted a "Houthi UAV ground control station and 10 Houthi one-way UAVs" that "presented an imminent threat to merchant vessels and the US Navy ships in the region", the US military's Central Command (CENTCOM) said in a statement referring to unmanned aerial vehicles. "This action will protect freedom of navigation and make international waters safer and more secure for US Navy vessels and merchant vessels," it added. The group said on Wednesday that all US and British warships participating in "aggression" against Yemen are targets, heightening concerns over the escalating tensions in the region as well as the increased disruption to world trade. CENTCOM said earlier that the USS Carney had shot down an antiship ballistic missile fired by the Houthis and downed three Iranian drones less than an hour later.
Universal Imitation Games
Alan Turing proposed in 1950 a framework called an imitation game to decide if a machine could think. Using mathematics developed largely after Turing -- category theory -- we analyze a broader class of universal imitation games (UIGs), which includes static, dynamic, and evolutionary games. In static games, the participants are in a steady state. In dynamic UIGs, "learner" participants are trying to imitate "teacher" participants over the long run. In evolutionary UIGs, the participants are competing against each other in an evolutionary game, and participants can go extinct and be replaced by others with higher fitness. We use the framework of category theory -- in particular, two influential results by Yoneda -- to characterize each type of imitation game. Universal properties in categories are defined by initial and final objects. We characterize dynamic UIGs where participants are learning by inductive inference as initial algebras over well-founded sets, and contrast them with participants learning by conductive inference over the final coalgebra of non-well-founded sets. We briefly discuss the extension of our categorical framework for UIGs to imitation games on quantum computers.
Hierarchical hybrid modeling for flexible tool use
Priorelli, Matteo, Stoianov, Ivilin Peev
In a recent computational framework called active inference, discrete models can be linked to their continuous counterparts to perform decision-making in changing environments. From another perspective, simple agents can be combined to better capture the causal relationships of the world. How can we use these two features together to achieve efficient goal-directed behavior? We present an architecture composed of several hybrid -- continuous and discrete -- units replicating the agent's configuration, controlled by a high-level discrete model that achieves dynamic planning and synchronized behavior. Additional factorizations within each level allow to represent hierarchically other agents and objects in relation to the self. We evaluate this hierarchical hybrid model on a non-trivial task: reaching a moving object after having picked a moving tool. This study extends past work on control as inference and proposes an alternative direction to deep reinforcement learning.
Character-based Outfit Generation with Vision-augmented Style Extraction via LLMs
Forouzandehmehr, Najmeh, Cao, Yijie, Thakurdesai, Nikhil, Giahi, Ramin, Ma, Luyi, Farrokhsiar, Nima, Xu, Jianpeng, Korpeoglu, Evren, Achan, Kannan
The outfit generation problem involves recommending a complete outfit to a user based on their interests. Existing approaches focus on recommending items based on anchor items or specific query styles but do not consider customer interests in famous characters from movie, social media, etc. In this paper, we define a new Character-based Outfit Generation (COG) problem, designed to accurately interpret character information and generate complete outfit sets according to customer specifications such as age and gender. To tackle this problem, we propose a novel framework LVA-COG that leverages Large Language Models (LLMs) to extract insights from customer interests (e.g., character information) and employ prompt engineering techniques for accurate understanding of customer preferences. Additionally, we incorporate text-to-image models to enhance the visual understanding and generation (factual or counterfactual) of cohesive outfits. Our framework integrates LLMs with text-to-image models and improves the customer's approach to fashion by generating personalized recommendations. With experiments and case studies, we demonstrate the effectiveness of our solution from multiple dimensions.
The Information of Large Language Model Geometry
Tan, Zhiquan, Li, Chenghai, Huang, Weiran
This paper investigates the information encoded in the embeddings of large language models (LLMs). We conduct simulations to analyze the representation entropy and discover a power law relationship with model sizes. Building upon this observation, we propose a theory based on (conditional) entropy to elucidate the scaling law phenomenon. Furthermore, we delve into the auto-regressive structure of LLMs and examine the relationship between the last token and previous context tokens using information theory and regression techniques. Specifically, we establish a theoretical connection between the information gain of new tokens and ridge regression. Additionally, we explore the effectiveness of Lasso regression in selecting meaningful tokens, which sometimes outperforms the closely related attention weights. Finally, we conduct controlled experiments, and find that information is distributed across tokens, rather than being concentrated in specific "meaningful" tokens alone.
Harm Amplification in Text-to-Image Models
Hao, Susan, Shelby, Renee, Liu, Yuchi, Srinivasan, Hansa, Bhutani, Mukul, Ayan, Burcu Karagol, Poddar, Shivani, Laszlo, Sarah
Warning: The content of this paper as well as some blurred images shown may include references to nudity, sexualization, violence, and gore. Text-to-image (T2I) models have emerged as a significant advancement in generative AI; however, there exist safety concerns regarding their potential to produce harmful image outputs even when users input seemingly safe prompts. This phenomenon, where T2I models generate harmful representations that were not explicit in the input, poses a potentially greater risk than adversarial prompts, leaving users unintentionally exposed to harms. Our paper addresses this issue by first introducing a formal definition for this phenomenon, termed harm amplification. We further contribute to the field by developing methodologies to quantify harm amplification in which we consider the harm of the model output in the context of user input. We then empirically examine how to apply these different methodologies to simulate real-world deployment scenarios including a quantification of disparate impacts across genders resulting from harm amplification. Together, our work aims to offer researchers tools to comprehensively address safety challenges in T2I systems and contribute to the responsible deployment of generative AI models.