Government
Delving into the Openness of CLIP
Ren, Shuhuai, Li, Lei, Ren, Xuancheng, Zhao, Guangxiang, Sun, Xu
Contrastive Language-Image Pre-training (CLIP) formulates image classification as an image-to-text matching task, i.e., matching images to the corresponding natural language descriptions instead of discrete category IDs. This allows for open-vocabulary visual recognition, where the model can recognize images from an open class set (also known as an open vocabulary) in a zero-shot manner. However, evaluating the openness of CLIP-like models is challenging, as the models are open to arbitrary vocabulary in theory, but their accuracy varies in practice. To address this, we resort to an incremental perspective to assess the openness through vocabulary expansions, and define extensibility to measure a model's ability to handle novel classes. Our evaluation shows that CLIP-like models are not truly open, and their performance deteriorates as the vocabulary expands. We further dissect the feature space of CLIP from the perspectives of representation alignment and uniformity. Our investigation reveals that the overestimation of openness is due to confusion among competing text features, rather than a failure to capture the similarity between image features and text features of novel classes. We hope that our investigation and analysis will facilitate future research on the CLIP openness issue.
Data Models for Dataset Drift Controls in Machine Learning With Optical Images
Oala, Luis, Aversa, Marco, Nobis, Gabriel, Willis, Kurt, Neuenschwander, Yoan, Buck, Michèle, Matek, Christian, Extermann, Jerome, Pomarico, Enrico, Samek, Wojciech, Murray-Smith, Roderick, Clausen, Christoph, Sanguinetti, Bruno
Camera images are ubiquitous in machine learning research. They also play a central role in the delivery of important services spanning medicine and environmental surveying. However, the application of machine learning models in these domains has been limited because of robustness concerns. A primary failure mode are performance drops due to differences between the training and deployment data. While there are methods to prospectively validate the robustness of machine learning models to such dataset drifts, existing approaches do not account for explicit models of the primary object of interest: the data. This limits our ability to study and understand the relationship between data generation and downstream machine learning model performance in a physically accurate manner. In this study, we demonstrate how to overcome this limitation by pairing traditional machine learning with physical optics to obtain explicit and differentiable data models. We demonstrate how such data models can be constructed for image data and used to control downstream machine learning model performance related to dataset drift. The findings are distilled into three applications. First, drift synthesis enables the controlled generation of physically faithful drift test cases to power model selection and targeted generalization. Second, the gradient connection between machine learning task model and data model allows advanced, precise tolerancing of task model sensitivity to changes in the data generation. These drift forensics can be used to precisely specify the acceptable data environments in which a task model may be run. Third, drift optimization opens up the possibility to create drifts that can help the task model learn better faster, effectively optimizing the data generating process itself. A guide to access the open code and datasets is available at https://github.com/aiaudit-org/raw2logit.
Are Synonym Substitution Attacks Really Synonym Substitution Attacks?
Chiang, Cheng-Han, Lee, Hung-yi
In this paper, we explore the following question: Are synonym substitution attacks really synonym substitution attacks (SSAs)? We approach this question by examining how SSAs replace words in the original sentence and show that there are still unresolved obstacles that make current SSAs generate invalid adversarial samples. We reveal that four widely used word substitution methods generate a large fraction of invalid substitution words that are ungrammatical or do not preserve the original sentence's semantics. Next, we show that the semantic and grammatical constraints used in SSAs for detecting invalid word replacements are highly insufficient in detecting invalid adversarial samples.
Integrated Space Domain Awareness and Communication System
Cetin, Selen Gecgel, Ozbek, Berna, Kurt, Gunes Karabulut
Space has been reforming and this evolution brings new threats that, together with technological developments and malicious intent, can pose a major challenge. Space domain awareness (SDA), a new conceptual idea, has come to the forefront. It aims sensing, detection, identification and countermeasures by providing autonomy, intelligence and flexibility against potential threats in space. In this study, we first present an insightful and clear view of the new space. Secondly, we propose an integrated SDA and communication (ISDAC) system for attacker detection. We assume that the attacker has beam-steering antennas and is capable to vary attack scenarios, such as random attacks on some receiver antennas. To track random patterns and meet SDA requirements, a lightweight convolutional neural network architecture is developed. The proposed ISDAC system shows superior and robust performance under 12 different attacker configurations with a detection accuracy of over 97.8%.
GRAPE for Fast and Scalable Graph Processing and random walk-based Embedding
Cappelletti, Luca, Fontana, Tommaso, Casiraghi, Elena, Ravanmehr, Vida, Callahan, Tiffany J., Cano, Carlos, Joachimiak, Marcin P., Mungall, Christopher J., Robinson, Peter N., Reese, Justin, Valentini, Giorgio
Graph Representation Learning (GRL) methods opened new avenues for addressing complex, real-world problems represented by graphs. However, many graphs used in these applications comprise millions of nodes and billions of edges and are beyond the capabilities of current methods and software implementations. We present GRAPE, a software resource for graph processing and embedding that can scale with big graphs by using specialized and smart data structures, algorithms, and a fast parallel implementation of random walk-based methods. Compared with state-of-the-art software resources, GRAPE shows an improvement of orders of magnitude in empirical space and time complexity, as well as a competitive edge and node label prediction performance. GRAPE comprises about 1.7 million well-documented lines of Python and Rust code and provides 69 node embedding methods, 25 inference models, a collection of efficient graph processing utilities and over 80,000 graphs from the literature and other sources. Standardized interfaces allow seamless integration of third-party libraries, while ready-to-use and modular pipelines permit an easy-to-use evaluation of GRL methods, therefore also positioning GRAPE as a software resource to perform a fair comparison between methods and libraries for graph processing and embedding.
Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
Crothers, Evan, Japkowicz, Nathalie, Viktor, Herna
Machine generated text is increasingly difficult to distinguish from human authored text. Powerful open-source models are freely available, and user-friendly tools that democratize access to generative models are proliferating. ChatGPT, which was released shortly after the first edition of this survey, epitomizes these trends. The great potential of state-of-the-art natural language generation (NLG) systems is tempered by the multitude of avenues for abuse. Detection of machine generated text is a key countermeasure for reducing abuse of NLG models, with significant technical challenges and numerous open problems. We provide a survey that includes both 1) an extensive analysis of threat models posed by contemporary NLG systems, and 2) the most complete review of machine generated text detection methods to date. This survey places machine generated text within its cybersecurity and social context, and provides strong guidance for future work addressing the most critical threat models, and ensuring detection systems themselves demonstrate trustworthiness through fairness, robustness, and accountability.
The 2000s Video Game With an Unexpected Lesson for Today's Transportation Debates
In the spring of 2021, just months before Congress passed the Infrastructure Investment and Jobs Act--heralded by the Biden administration as the largest-ever federal investment in public transit, bridge repair, and clean energy--I found myself playing a lot of Mass Effect Legendary Edition. This was a happy coincidence, because never in my life had the nation been so embroiled in wonky debates about infrastructure priorities and spending. And as it turns out, Mass Effect was the perfect 100-hour video game for that particular moment in history: It's absolutely obsessed with transportation technologies and their social, cultural, and political implications. Despite its revolutionary capacity, we often conceptualize transportation in mundane, frustrating terms: long commutes and congested highways, spotty bus service and increasingly crowded sidewalks littered with e-scooters. That's what makes fiction centered around these questions so important--especially when it comes to thinking through the big investments we want to make in infrastructure, what we hope to accomplish, and the challenges we should anticipate.
The next fear on AI: Hollywood's killer robots become the military's tools
Washington – When U.S. President Joe Biden announced sharp restrictions in October on selling the most advanced computer chips to China, he sold it, in part, as a way of giving American industry a chance to restore its competitiveness. But at the Pentagon and the National Security Council, there was a second agenda: arms control. If the Chinese military cannot get the chips, the theory goes, it may slow its effort to develop weapons driven by artificial intelligence. That would give the White House, and the world, time to figure out some rules for the use of AI in everything from sensors, missiles and cyberweapons, and ultimately to guard against some of the nightmares conjured by Hollywood -- autonomous killer robots and computers that lock out their human creators. This could be due to a conflict with your ad-blocking or security software.
U.S. sees a new era of nuclear risk dawning in China-Russia cooperation
The deepening cooperation between China and Russia threatens to overturn decades of international stability in nuclear arms control, according to a top adviser to U.S. President Joe Biden. To avert miscalculations, nuclear-weapons states must engage on existing and potential threats, from Iran's atomic ambitions to the use of artificial intelligence for decision-making during crises, Pranay Vaddi, the National Security Council's senior director for arms control, said in an interview in Vienna. "We're entering a different period," Vaddi said after talks at the International Atomic Energy Agency. "It requires a little bit of experimentation." This could be due to a conflict with your ad-blocking or security software.
Japan and Singapore leaders affirm alignment on rules-based global order
SINGAPORE – Prime Minister Fumio Kishida and his Singaporean counterpart, Lee Hsien Loong, have reaffirmed their commitment to uphold the rules-based international order amid Russia's aggression against Ukraine and China's growing military and economic clout. During talks Friday at Singapore's Changi Airport following a six-day visit to Africa, Kishida told Lee that negotiations on a deal that would allow the transfer of defense equipment and technology between the two countries are making progress, according to the Japanese Foreign Ministry. "We want to strengthen security and defense cooperation," Kishida was quoted as saying, while also calling for deepening cooperation in areas such as start-ups and building resilient supply chains. This could be due to a conflict with your ad-blocking or security software. Please add japantimes.co.jp and piano.io to your list of allowed sites.