Government
TAPS Responsibility Matrix: A tool for responsible data science by design
Urovi, Visara, Celebi, Remzi, Sun, Chang, Rieswijk, Linda, Erard, Michael, Yilmaz, Arif, Moodley, Kody, Kumar, Parveen, Dumontier, Michel
Data science is an interdisciplinary research area where scientists are typically working with data coming from different fields. When using and analyzing data, the scientists implicitly agree to follow standards, procedures, and rules set in these fields. However, guidance on the responsibilities of the data scientists and the other involved actors in a data science project is typically missing. While literature shows that novel frameworks and tools are being proposed in support of open-science, data reuse, and research data management, there are currently no frameworks that can fully express responsibilities of a data science project. In this paper, we describe the Transparency, Accountability, Privacy, and Societal Responsibility Matrix (TAPS-RM) as framework to explore social, legal, and ethical aspects of data science projects. TAPS-RM acts as a tool to provide users with a holistic view of their project beyond key outcomes and clarifies the responsibilities of actors. We map the developed model of TAPS-RM with well-known initiatives for open data (such as FACT, FAIR and Datasheets for datasets). We conclude that TAPS-RM is a tool to reflect on responsibilities at a data science project level and can be used to advance responsible data science by design.
A Light Recipe to Train Robust Vision Transformers
Debenedetti, Edoardo, Sehwag, Vikash, Mittal, Prateek
In this paper, we ask whether Vision Transformers (ViTs) can serve as an underlying architecture for improving the adversarial robustness of machine learning models against evasion attacks. While earlier works have focused on improving Convolutional Neural Networks, we show that also ViTs are highly suitable for adversarial training to achieve competitive performance. We achieve this objective using a custom adversarial training recipe, discovered using rigorous ablation studies on a subset of the ImageNet dataset. The canonical training recipe for ViTs recommends strong data augmentation, in part to compensate for the lack of vision inductive bias of attention modules, when compared to convolutions. We show that this recipe achieves suboptimal performance when used for adversarial training. In contrast, we find that omitting all heavy data augmentation, and adding some additional bag-of-tricks ($\varepsilon$-warmup and larger weight decay), significantly boosts the performance of robust ViTs. We show that our recipe generalizes to different classes of ViT architectures and large-scale models on full ImageNet-1k. Additionally, investigating the reasons for the robustness of our models, we show that it is easier to generate strong attacks during training when using our recipe and that this leads to better robustness at test time. Finally, we further study one consequence of adversarial training by proposing a way to quantify the semantic nature of adversarial perturbations and highlight its correlation with the robustness of the model. Overall, we recommend that the community should avoid translating the canonical training recipes in ViTs to robust training and rethink common training choices in the context of adversarial training.
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
Lewicki, Kornel, Lee, Michelle Seng Ah, Cobbe, Jennifer, Singh, Jatinder
"AI as a Service" (AIaaS) is a rapidly growing market, offering various plug-and-play AI services and tools. AIaaS enables its customers (users) - who may lack the expertise, data, and/or resources to develop their own systems - to easily build and integrate AI capabilities into their applications. Yet, it is known that AI systems can encapsulate biases and inequalities that can have societal impact. This paper argues that the context-sensitive nature of fairness is often incompatible with AIaaS' 'one-size-fits-all' approach, leading to issues and tensions. Specifically, we review and systematise the AIaaS space by proposing a taxonomy of AI services based on the levels of autonomy afforded to the user. We then critically examine the different categories of AIaaS, outlining how these services can lead to biases or be otherwise harmful in the context of end-user applications. In doing so, we seek to draw research attention to the challenges of this emerging area.
Combining Deep Neural Reranking and Unsupervised Extraction for Multi-Query Focused Summarization
Seeberger, Philipp, Riedhammer, Korbinian
The CrisisFACTS Track aims to tackle challenges such as multi-stream fact-finding in the domain of event tracking; participants' systems extract important facts from several disaster-related events while incorporating the temporal order. We propose a combination of retrieval, reranking, and the well-known Integer Linear Programming (ILP) and Maximal Marginal Relevance (MMR) frameworks. In the former two modules, we explore various methods including an entity-based baseline, pre-trained and fine-tuned Question Answering systems, and ColBERT. We then use the latter module as an extractive summarization component by taking diversity and novelty criteria into account. The automatic scoring runs show strong results across the evaluation setups but also reveal shortcomings and challenges.
PiC: A Phrase-in-Context Dataset for Phrase Understanding and Semantic Search
Pham, Thang M., Yoon, Seunghyun, Bui, Trung, Nguyen, Anh
While contextualized word embeddings have been a de-facto standard, learning contextualized phrase embeddings is less explored and being hindered by the lack of a human-annotated benchmark that tests machine understanding of phrase semantics given a context sentence or paragraph (instead of phrases alone). To fill this gap, we propose PiC -- a dataset of ~28K of noun phrases accompanied by their contextual Wikipedia pages and a suite of three tasks for training and evaluating phrase embeddings. Training on PiC improves ranking models' accuracy and remarkably pushes span-selection (SS) models (i.e., predicting the start and end index of the target phrase) near-human accuracy, which is 95% Exact Match (EM) on semantic search given a query phrase and a passage. Interestingly, we find evidence that such impressive performance is because the SS models learn to better capture the common meaning of a phrase regardless of its actual context. SotA models perform poorly in distinguishing two senses of the same phrase in two contexts (~60% EM) and in estimating the similarity between two different phrases in the same context (~70% EM).
Keyword Assisted Topic Models
Eshima, Shusei, Imai, Kosuke, Sasaki, Tomoya
The unsupervised nature of the models makes them suitable for exploring topics in a corpus without prior knowledge. However, researchers find that these models often fail to measure specific concepts of substantive interest by inadvertently creating multiple topics with similar content and combining distinct themes into a single topic. In this paper, we empirically demonstrate that providing a small number of keywords can substantially enhance the measurement performance of topic models. An important advantage of the proposed keyword assisted topic model (keyATM) is that the specification of keywords requires researchers to label topics prior to fitting a model to the data. This contrasts with a widespread practice of post-hoc topic interpretation and adjustments that compromises the objectivity of empirical findings. In our application, we find that keyATM provides more interpretable results, has better document classification performance, and is less sensitive to the number of topics than the standard topic models. Finally, we show that keyATM can also incorporate covariates and model time trends. An open-source software package is available for implementing the proposed methodology. Verification Materials: The data and materials required to verify the computational reproducibility of the results, procedures and analyses in this article are available on the American Journal of Political Science Dataverse within the Harvard Dataverse Network, at: https://doi.org/10.7910/DVN/RKNNVL
Taiwan Parliament speaker says country is a 'beacon of democracy for Chinese-speaking peoples'
Hudson Institute senior fellow Michael Pillsbury tells "Fox News @ Night" that war games conducted by the Center for Strategic and International Studies on a Chinese invasion of Taiwan is "all the more reason to try to deter" an invasion. You Si-kun, the speaker of Taiwan's Parliament, spoke at the International Religious Freedom Summit on Wednesday and voiced why it is important that free nations protect Taiwan. "If Taiwan falls into the sphere of influence of CCP (Chinese Communist Party), then the beacon of democracy will be destroyed. And China may invade the first island chain and will cause a threat to the entire world," he said. You Si-kun, the speaker of Taiwan's Parliament, addresses the International Religious Freedom Summit in Washington, D.C. (IRF Summit / Matt Rybczynski) Freedom House's 2022 Freedom in the World report ranked Taiwan a perfect score of 4 concerning religious freedom.
The Supreme Court Considers the Algorithm
When the Ninth Circuit Court of Appeals considered a lawsuit against Google in 2020, Judge Ronald M. Gould stated his view of the tech giant's most significant asset bluntly: "So-called'neutral' algorithms," he wrote, can be "transformed into deadly missiles of destruction by ISIS." According to Gould, it was time to challenge the boundaries of a little snippet of the 1996 Communications Decency Act known as Section 230, which protects online platforms from liability for the things their users post. The plaintiffs in this case, the family of a young woman who was killed during a 2015 Islamic State attack in Paris, alleged that Google had violated the Anti-terrorism Act by allowing YouTube's recommendation system to promote terrorist content. The algorithms that amplified ISIS videos were a danger in and of themselves, they argued. Gould was in the minority, and the case was decided in Google's favor.
Cybersecurity Will Shift in 2023 Thanks to AI - RTInsights
AI will form a key component of cyber defense strategies in 2023, allowing companies to move to an entirely new approach to cybersecurity. Because of this, companies look to innovative tools to respond to threats and--even better--prevent them in the first place. Previously, Gartner outlined its top seven cybersecurity trends for last year. With each one, it becomes more apparent that humans will need the support of artificial intelligence and machine learning tools to stay ahead of the curve. These predictions for 2022 are becoming even more potent for this year.