Government
The Download: inside the US defense tech aid package, and how AI is improving vegan cheese
After weeks of drawn-out congressional debate over how much the United States should spend on conflicts abroad, President Joe Biden signed a 95 billion aid package into law last week. The bill will send a significant quantity of supplies to Ukraine and Israel, while also supporting Taiwan with submarine technology to aid its defenses against China. It's also sparked renewed calls for stronger crackdowns on Iranian-produced drones. James O'Donnell, our AI reporter, spoke to Andrew Metrick, a fellow with the defense program at the Center for a New American Security, a think tank, to discuss how the spending bill provides a window into US strategies around four key defense technologies with the power to reshape how today's major conflicts are being fought. This piece is part of MIT Technology Review Explains: a series delving into the complex, messy world of technology to help you understand what's coming next.
Tesla clears key regulatory hurdles for self-driving in China
Tesla has cleared some key regulatory hurdles that have long hindered it from rolling out its self-driving software in China, paving the way for a favorable result from Elon Musk's surprise visit to the U.S. automaker's second-largest market. Tesla CEO Musk arrived in the Chinese capital on Sunday, where he was expected to discuss the rollout of Full Self-Driving (FSD) software and permission to transfer driving data overseas, according to a person with knowledge of the matter. The billionaire's whirlwind visit, during which he met with Chinese Premier Li Qiang, came just over a week after he scrapped a planned trip to India to meet with Prime Minister Narendra Modi, citing "very heavy Tesla obligations."
Why China Is So Bad at Disinformation
"China will use AI to disrupt elections in the US, South Korea and India, Microsoft warns" one read. "China Is Using AI to Sow Disinformation and Stoke Discord Across Asia and the US," another claimed. The headlines were based on a report published earlier this month by Microsoft's Threat Analysis Center which outlined how a Chinese disinformation campaign was now utilizing artificial technology to inflame divisions and disrupt elections in the US and around the world. The campaign, which has already targeted Taiwan's elections, uses AI-generated audio and memes designed to grab user attention and boost engagement. But what these headlines and Microsoft itself failed to adequately convey is that the Chinese government-linked disinformation campaign, known as Spamouflage Dragon or Dragonbridge, has so far been virtually ineffective.
Russia-Ukraine war: List of key events, day 795
Ukraine's top commander Colonel General Oleksandr Syrskii said Kyiv's troops fell back to new positions west of three villages on the eastern front, as the situation on the front line worsened. Syrskii said the "most difficult" areas were west of Russian-occupied Maryinka and northwest of Avdiivka, the town captured by Russian forces in February. Syrskii also said his forces were closely monitoring an increase in the number of Russian troops in the area of Kharkiv, Ukraine's second-largest city and just 30km (19 miles) from the Russian border. "In the most threatening directions, our troops have been reinforced by artillery and tank units," he said. Russia's Ministry of Defence said its troops had captured the village of Novobakhmutivka in the Donetsk region, about 10km (six miles) north of Avdiivka.
Bridging Data Barriers among Participants: Assessing the Potential of Geoenergy through Federated Learning
Peng, Weike, Gao, Jiaxin, Chen, Yuntian, Wang, Shengwei
Machine learning algorithms emerge as a promising approach in energy fields, but its practical is hindered by data barriers, stemming from high collection costs and privacy concerns. This study introduces a novel federated learning (FL) framework based on XGBoost models, enabling safe collaborative modeling with accessible yet concealed data from multiple parties. Hyperparameter tuning of the models is achieved through Bayesian Optimization. To ascertain the merits of the proposed FL-XGBoost method, a comparative analysis is conducted between separate and centralized models to address a classical binary classification problem in geoenergy sector. The results reveal that the proposed FL framework strikes an optimal balance between privacy and accuracy. FL models demonstrate superior accuracy and generalization capabilities compared to separate models, particularly for participants with limited data or low correlation features and offers significant privacy benefits compared to centralized model. The aggregated optimization approach within the FL agreement proves effective in tuning hyperparameters. This study opens new avenues for assessing unconventional reservoirs through collaborative and privacy-preserving FL techniques.
Enhancing Traffic Incident Management with Large Language Models: A Hybrid Machine Learning Approach for Severity Classification
Grigorev, Artur, Saleh, Khaled, Ou, Yuming, Mihaita, Adriana-Simona
This research showcases the innovative integration of Large Language Models into machine learning workflows for traffic incident management, focusing on the classification of incident severity using accident reports. By leveraging features generated by modern language models alongside conventional data extracted from incident reports, our research demonstrates improvements in the accuracy of severity classification across several machine learning algorithms. Our contributions are threefold. First, we present an extensive comparison of various machine learning models paired with multiple large language models for feature extraction, aiming to identify the optimal combinations for accurate incident severity classification. Second, we contrast traditional feature engineering pipelines with those enhanced by language models, showcasing the superiority of language-based feature engineering in processing unstructured text. Third, our study illustrates how merging baseline features from accident reports with language-based features can improve the severity classification accuracy. This comprehensive approach not only advances the field of incident management but also highlights the cross-domain application potential of our methodology, particularly in contexts requiring the prediction of event outcomes from unstructured textual data or features translated into textual representation. Specifically, our novel methodology was applied to three distinct datasets originating from the United States, the United Kingdom, and Queensland, Australia. This cross-continental application underlines the robustness of our approach, suggesting its potential for widespread adoption in improving incident management processes globally.
Predicting Safety Misbehaviours in Autonomous Driving Systems using Uncertainty Quantification
Grewal, Ruben, Tonella, Paolo, Stocco, Andrea
The automated real-time recognition of unexpected situations plays a crucial role in the safety of autonomous vehicles, especially in unsupported and unpredictable scenarios. This paper evaluates different Bayesian uncertainty quantification methods from the deep learning domain for the anticipatory testing of safety-critical misbehaviours during system-level simulation-based testing. Specifically, we compute uncertainty scores as the vehicle executes, following the intuition that high uncertainty scores are indicative of unsupported runtime conditions that can be used to distinguish safe from failure-inducing driving behaviors. In our study, we conducted an evaluation of the effectiveness and computational overhead associated with two Bayesian uncertainty quantification methods, namely MC- Dropout and Deep Ensembles, for misbehaviour avoidance. Overall, for three benchmarks from the Udacity simulator comprising both out-of-distribution and unsafe conditions introduced via mutation testing, both methods successfully detected a high number of out-of-bounds episodes providing early warnings several seconds in advance, outperforming two state-of-the-art misbehaviour prediction methods based on autoencoders and attention maps in terms of effectiveness and efficiency. Notably, Deep Ensembles detected most misbehaviours without any false alarms and did so even when employing a relatively small number of models, making them computationally feasible for real-time detection. Our findings suggest that incorporating uncertainty quantification methods is a viable approach for building fail-safe mechanisms in deep neural network-based autonomous vehicles.
Iconic Gesture Semantics
Lücking, Andy, Henlein, Alexander, Mehler, Alexander
The "meaning" of an iconic gesture is conditioned on its informational evaluation. Only informational evaluation lifts a gesture to a quasi-linguistic level that can interact with verbal content. Interaction is either vacuous or regimented by usual lexicon-driven inferences. Informational evaluation is spelled out as extended exemplification (extemplification) in terms of perceptual classification of a gesture's visual iconic model. The iconic model is derived from Frege/Montague-like truth-functional evaluation of a gesture's form within spatially extended domains. We further argue that the perceptual classification of instances of visual communication requires a notion of meaning different from Frege/Montague frameworks. Therefore, a heuristic for gesture interpretation is provided that can guide the working semanticist. In sum, an iconic gesture semantics is introduced which covers the full range from kinematic gesture representations over model-theoretic evaluation to inferential interpretation in dynamic semantic frameworks.
Foundations of Multisensory Artificial Intelligence
Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits, such as in supporting human health and well-being, enabling multimedia content processing, and enhancing real-world autonomous agents. By synthesizing a range of theoretical frameworks and application domains, this thesis aims to advance the machine learning foundations of multisensory AI. In the first part, we present a theoretical framework formalizing how modalities interact with each other to give rise to new information for a task. These interactions are the basic building blocks in all multimodal problems, and their quantification enables users to understand their multimodal datasets, design principled approaches to learn these interactions, and analyze whether their model has succeeded in learning. In the second part, we study the design of practical multimodal foundation models that generalize over many modalities and tasks, which presents a step toward grounding large language models to real-world sensory modalities. We introduce MultiBench, a unified large-scale benchmark across a wide range of modalities, tasks, and research areas, followed by the cross-modal attention and multimodal transformer architectures that now underpin many of today's multimodal foundation models. Scaling these architectures on MultiBench enables the creation of general-purpose multisensory AI systems, and we discuss our collaborative efforts in applying these models for real-world impact in affective computing, mental health, cancer prognosis, and robotics. Finally, we conclude this thesis by discussing how future work can leverage these ideas toward more general, interactive, and safe multisensory AI.
Clio: Real-time Task-Driven Open-Set 3D Scene Graphs
Maggio, Dominic, Chang, Yun, Hughes, Nathan, Trang, Matthew, Griffith, Dan, Dougherty, Carlyn, Cristofalo, Eric, Schmid, Lukas, Carlone, Luca
Modern tools for class-agnostic image segmentation (e.g., SegmentAnything) and open-set semantic understanding (e.g., CLIP) provide unprecedented opportunities for robot perception and mapping. While traditional closed-set metric-semantic maps were restricted to tens or hundreds of semantic classes, we can now build maps with a plethora of objects and countless semantic variations. This leaves us with a fundamental question: what is the right granularity for the objects (and, more generally, for the semantic concepts) the robot has to include in its map representation? While related work implicitly chooses a level of granularity by tuning thresholds for object detection, we argue that such a choice is intrinsically task-dependent. The first contribution of this paper is to propose a task-driven 3D scene understanding problem, where the robot is given a list of tasks in natural language and has to select the granularity and the subset of objects and scene structure to retain in its map that is sufficient to complete the tasks. We show that this problem can be naturally formulated using the Information Bottleneck (IB), an established information-theoretic framework. The second contribution is an algorithm for task-driven 3D scene understanding based on an Agglomerative IB approach, that is able to cluster 3D primitives in the environment into task-relevant objects and regions and executes incrementally. The third contribution is to integrate our task-driven clustering algorithm into a real-time pipeline, named Clio, that constructs a hierarchical 3D scene graph of the environment online using only onboard compute, as the robot explores it. Our final contribution is an extensive experimental campaign showing that Clio not only allows real-time construction of compact open-set 3D scene graphs, but also improves the accuracy of task execution by limiting the map to relevant semantic concepts.