Government
Areas of Strategic Visibility: Disability Bias in Biometrics
Mankoff, Jennifer, Kasnitz, Devva, Studies, Disability, Camp, L Jean, Lazar, Jonathan, Hochheiser, Harry
Yet many of these systems are not accessible to people who experience different kinds of disability exclusion. Different personal characteristics may impact any or all of the physical (DNA, fingerprints, face or retina) and behavioral (gesture, gait, voice) characteristics listed in the RFI as examples of biometric signals. We define disability here in terms of the discriminatory and often systemic problems with available infrastructure's ability to meet the needs of all people [UN 2017, Oliver, 2013). Using this definition, "[biometrics] could either mitigate or amplify disability depending on how they are designed." (Guo, 2019). As Whittaker and colleauges (2019) state, this is not simply a matter of algorithmic accuracy: "...discrimination against people of color, women, and other historically marginalized groups has often been justified by representing these groups as disabled . Thus disability is entwined with, and serves to justify, practices of marginalization." It is critical that we look beyond inclusion to full and fully accommodated participation.
Developing a Series of AI Challenges for the United States Department of the Air Force
Gadepally, Vijay, Angelides, Gregory, Barbu, Andrei, Bowne, Andrew, Brattain, Laura J., Broderick, Tamara, Cabrera, Armando, Carl, Glenn, Carter, Ronisha, Cha, Miriam, Cowen, Emilie, Cummings, Jesse, Freeman, Bill, Glass, James, Goldberg, Sam, Hamilton, Mark, Heldt, Thomas, Huang, Kuan Wei, Isola, Phillip, Katz, Boris, Koerner, Jamie, Lin, Yen-Chen, Mayo, David, McAlpin, Kyle, Perron, Taylor, Piou, Jean, Rao, Hrishikesh M., Reynolds, Hayley, Samuel, Kaira, Samsi, Siddharth, Schmidt, Morgan, Shing, Leslie, Simek, Olga, Swenson, Brandon, Sze, Vivienne, Taylor, Jonathan, Tylkin, Paul, Veillette, Mark, Weiss, Matthew L, Wollaber, Allan, Yuditskaya, Sophia, Kepner, Jeremy
Through a series of federal initiatives and orders, the U.S. Government has been making a concerted effort to ensure American leadership in AI. These broad strategy documents have influenced organizations such as the United States Department of the Air Force (DAF). The DAF-MIT AI Accelerator is an initiative between the DAF and MIT to bridge the gap between AI researchers and DAF mission requirements. Several projects supported by the DAF-MIT AI Accelerator are developing public challenge problems that address numerous Federal AI research priorities. These challenges target priorities by making large, AI-ready datasets publicly available, incentivizing open-source solutions, and creating a demand signal for dual use technologies that can stimulate further research. In this article, we describe these public challenges being developed and how their application contributes to scientific advances.
Unified 2D and 3D Pre-Training of Molecular Representations
Zhu, Jinhua, Xia, Yingce, Wu, Lijun, Xie, Shufang, Qin, Tao, Zhou, Wengang, Li, Houqiang, Liu, Tie-Yan
Molecular representation learning has attracted much attention recently. A molecule can be viewed as a 2D graph with nodes/atoms connected by edges/bonds, and can also be represented by a 3D conformation with 3-dimensional coordinates of all atoms. We note that most previous work handles 2D and 3D information separately, while jointly leveraging these two sources may foster a more informative representation. In this work, we explore this appealing idea and propose a new representation learning method based on a unified 2D and 3D pre-training. Atom coordinates and interatomic distances are encoded and then fused with atomic representations through graph neural networks. The model is pre-trained on three tasks: reconstruction of masked atoms and coordinates, 3D conformation generation conditioned on 2D graph, and 2D graph generation conditioned on 3D conformation. We evaluate our method on 11 downstream molecular property prediction tasks: 7 with 2D information only and 4 with both 2D and 3D information. Our method achieves state-of-the-art results on 10 tasks, and the average improvement on 2D-only tasks is 8.3%. Our method also achieves significant improvement on two 3D conformation generation tasks.
Neural Data-to-Text Generation Based on Small Datasets: Comparing the Added Value of Two Semi-Supervised Learning Approaches on Top of a Large Language Model
van der Lee, Chris, Ferreira, Thiago Castro, Emmery, Chris, Wiltshire, Travis, Krahmer, Emiel
This study discusses the effect of semi-supervised learning in combination with pretrained language models for data-to-text generation. It is not known whether semi-supervised learning is still helpful when a large-scale language model is also supplemented. This study aims to answer this question by comparing a data-to-text system only supplemented with a language model, to two data-to-text systems that are additionally enriched by a data augmentation or a pseudo-labeling semi-supervised learning approach. Results show that semi-supervised learning results in higher scores on diversity metrics. In terms of output quality, extending the training set of a data-to-text system with a language model using the pseudo-labeling approach did increase text quality scores, but the data augmentation approach yielded similar scores to the system without training set extension. These results indicate that semi-supervised learning approaches can bolster output quality and diversity, even when a language model is also present.
Detecting Volunteer Cotton Plants in a Corn Field with Deep Learning on UAV Remote-Sensing Imagery
Yadav, Pappu Kumar, Thomasson, J. Alex, Hardin, Robert, Searcy, Stephen W., Braga-Neto, Ulisses, Popescu, Sorin C., Martin, Daniel E., Rodriguez, Roberto, Meza, Karem, Enciso, Juan, Diaz, Jorge Solorzano, Wang, Tianyi
The cotton boll weevil, Anthonomus grandis Boheman is a serious pest to the U.S. cotton industry that has cost more than 16 billion USD in damages since it entered the United States from Mexico in the late 1800s. This pest has been nearly eradicated; however, southern part of Texas still faces this issue and is always prone to the pest reinfestation each year due to its sub-tropical climate where cotton plants can grow year-round. Volunteer cotton (VC) plants growing in the fields of inter-seasonal crops, like corn, can serve as hosts to these pests once they reach pin-head square stage (5-6 leaf stage) and therefore need to be detected, located, and destroyed or sprayed . In this paper, we present a study to detect VC plants in a corn field using YOLOv3 on three band aerial images collected by unmanned aircraft system (UAS). The two-fold objectives of this paper were : (i) to determine whether YOLOv3 can be used for VC detection in a corn field using RGB (red, green, and blue) aerial images collected by UAS and (ii) to investigate the behavior of YOLOv3 on images at three different scales (320 x 320, S1; 416 x 416, S2; and 512 x 512, S3 pixels) based on average precision (AP), mean average precision (mAP) and F1-score at 95% confidence level. No significant differences existed for mAP among the three scales, while a significant difference was found for AP between S1 and S3 (p = 0.04) and S2 and S3 (p = 0.02). A significant difference was also found for F1-score between S2 and S3 (p = 0.02). The lack of significant differences of mAP at all the three scales indicated that the trained YOLOv3 model can be used on a computer vision-based remotely piloted aerial application system (RPAAS) for VC detection and spray application in near real-time.
Combining Diverse Feature Priors
Jain, Saachi, Tsipras, Dimitris, Madry, Aleksander
The driving force behind deep learning's success is its ability to automatically discover predictive features in complex high-dimensional datasets. These features can generalize beyond the specific task at hand, thus enabling models to transfer to other (similar) tasks [DJV+14]. At the same time, the set of features that the model learns has a large impact on the model's performance on unseen inputs, especially in the presence of distribution shift [PBE+06; TE11; SKH+20] or spurious correlations [HM17; BVP18; Mei18]. Motivated by this, recent work focuses on encouraging specific modes of behavior by preventing the models from relying on certain features.
Overview of Abusive and Threatening Language Detection in Urdu at FIRE 2021
Amjad, Maaz, Zhila, Alisa, Sidorov, Grigori, Labunets, Andrey, Butta, Sabur, Amjad, Hamza Imam, Vitman, Oxana, Gelbukh, Alexander
With the growth of social media platform influence, the effect of their misuse becomes more and more impactful. The importance of automatic detection of threatening and abusive language can not be overestimated. However, most of the existing studies and state-of-the-art methods focus on English as the target language, with limited work on low- and medium-resource languages. In this paper, we present two shared tasks of abusive and threatening language detection for the Urdu language which has more than 170 million speakers worldwide. Both are posed as binary classification tasks where participating systems are required to classify tweets in Urdu into two classes, namely: (i) Abusive and Non-Abusive for the first task, and (ii) Threatening and Non-Threatening for the second. We present two manually annotated datasets containing tweets labelled as (i) Abusive and Non-Abusive, and (ii) Threatening and Non-Threatening. The abusive dataset contains 2400 annotated tweets in the train part and 1100 annotated tweets in the test part. The threatening dataset contains 6000 annotated tweets in the train part and 3950 annotated tweets in the test part. We also provide logistic regression and BERT-based baseline classifiers for both tasks. In this shared task, 21 teams from six countries registered for participation (India, Pakistan, China, Malaysia, United Arab Emirates, and Taiwan), 10 teams submitted their runs for Subtask A, which is Abusive Language Detection and 9 teams submitted their runs for Subtask B, which is Threatening Language detection, and seven teams submitted their technical reports. The best performing system achieved an F1-score value of 0.880 for Subtask A and 0.545 for Subtask B. For both subtasks, m-Bert based transformer model showed the best performance.
A Theoretically Grounded Benchmark for Evaluating Machine Commonsense
Santos, Henrique, Shen, Ke, Mulvehill, Alice M., Razeghi, Yasaman, McGuinness, Deborah L., Kejriwal, Mayank
Programming machines with commonsense reasoning (CSR) abilities is a longstanding challenge in the Artificial Intelligence community. Current CSR benchmarks use multiple-choice (and in relatively fewer cases, generative) question-answering instances to evaluate machine commonsense. Recent progress in transformer-based language representation models suggest that considerable progress has been made on existing benchmarks. However, although tens of CSR benchmarks currently exist, and are growing, it is not evident that the full suite of commonsense capabilities have been systematically evaluated. Furthermore, there are doubts about whether language models are 'fitting' to a benchmark dataset's training partition by picking up on subtle, but normatively irrelevant (at least for CSR), statistical features to achieve good performance on the testing partition. To address these challenges, we propose a benchmark called Theoretically-Grounded Commonsense Reasoning (TG-CSR) that is also based on discriminative question answering, but with questions designed to evaluate diverse aspects of commonsense, such as space, time, and world states. TG-CSR is based on a subset of commonsense categories first proposed as a viable theory of commonsense by Gordon and Hobbs. The benchmark is also designed to be few-shot (and in the future, zero-shot), with only a few training and validation examples provided. This report discusses the structure and construction of the benchmark. Preliminary results suggest that the benchmark is challenging even for advanced language representation models designed for discriminative CSR question answering tasks. Benchmark access and leaderboard: https://codalab.lisn.upsaclay.fr/competitions/3080 Benchmark website: https://usc-isi-i2.github.io/TGCSR/
ESA Fully Cuts Mars Mission Ties With Russia, Angering Moscow
The European Space Agency has officially terminated cooperation with Russia on a mission to put a rover on Mars, with Russia's space chief furiously responding by banning cosmonauts on the ISS from using a Europe-made robotic arm. The ESA had previously suspended ties on the joint ExoMars mission, which had planned to use Russian rockets to put Europe's Rosalind Franklin rover on the red planet to drill for signs of life, due to Russia's invasion of Ukraine. ESA Director-General Josef Aschbacher tweeted on Tuesday that because the war and resulting sanctions "continue to prevail", the agency would "officially terminate" ties with Russia on ExoMars and its landing platform. The firebrand head of Russian space agency Roscosmos Dmitry Rogozin issued an angry response. "Has the head of the European Space Agency thought about the work of thousands of scientists and engineers in Europe and Russia which has been ended by this decision? Is he prepared to answer for sabotaging a joint Mars mission?" Rogozin said on Telegram.
Autonomous flight startup Merlin Labs lands $120M and U.S. Air Force partnership โ TechCrunch
Autonomous flight is a grand challenge in aviation -- and a gold mine. The first company to crack it at scale stands to reap handsome profits from transportation and logistics alone. In 2020, the size of the global cargo airline industry was $110.8 billion, according to Statista, and one source estimates that it'll generate hundreds of billions in revenue by 2027. Xwing is one of the startups chasing after self-flying planes, as is Reliable Robotics, Pyka and the unicorn Volocopter. Roughly a year ago, Boston-based Merlin Labs emerged from stealth with an autonomous flight system designed to be installed in existing aircraft.