Africa
In the Shadowy, Hard-to-Track Poaching Industry, Governments Hope a New Tool Can Solve an Old Problem
In August 2021, forest range officer Remya Raghavan caught three people carrying wild boar meat in the Wayanad forest of Kerala, a state in southern India. Possessing wild animal meat is a crime under the country's 1972 Wildlife Protection Act, so Raghavan entered all the details of the crime--location, witnesses, names of the accused, items seized, and section of the forest--in a mobile application. Just like that, the case was officially registered in the app-based system, which signaled that it needed to be taken to court. The app Raghavan used is called HAWK, or Hostile Activity Watch Kernel, and it appears to be the first such digital intelligence gathering system for wildlife crime in India. It helps officers like Raghavan centralize and share information on forest and wildlife crimes in real time.
Contrastive Decoding: Open-ended Text Generation as Optimization
Li, Xiang Lisa, Holtzman, Ari, Fried, Daniel, Liang, Percy, Eisner, Jason, Hashimoto, Tatsunori, Zettlemoyer, Luke, Lewis, Mike
Given a language model (LM), maximum probability is a poor decoding objective for open-ended generation, because it produces short and repetitive text. On the other hand, sampling can often produce incoherent text that drifts from the original topics. We propose contrastive decoding (CD), a reliable decoding approach that optimizes a contrastive objective subject to a plausibility constraint. The contrastive objective returns the difference between the likelihood under a large LM (called the expert, e.g. OPT-13B) and a small LM (called the amateur, e.g. OPT-125M), and the constraint ensures that the outputs are plausible. CD is inspired by the fact that the failures of larger LMs (e.g., repetition, incoherence) are even more prevalent in smaller LMs, and that this difference signals which texts should be preferred. CD requires zero additional training, and produces higher quality text than decoding from the larger LM alone. It also works across model scales (OPT-13B and GPT2-1.5B) and significantly outperforms four strong decoding algorithms (e.g., nucleus, top-k) in automatic and human evaluations across wikipedia, news and story domains.
Latent Space Perspicacity and Interpretation Enhancement (LS-PIE) Framework
Stevens, Jesse, Wilke, Daniel N., Setshedi, Itumeleng
Linear latent variable models such as principal component analysis (PCA), independent component analysis (ICA), canonical correlation analysis (CCA), and factor analysis (FA) identify latent directions (or loadings) either ordered or unordered. The data is then projected onto the latent directions to obtain their projected representations (or scores). For example, PCA solvers usually rank the principal directions by explaining the most to least variance, while ICA solvers usually return independent directions unordered and often with single sources spread across multiple directions as multiple sub-sources, which is of severe detriment to their usability and interpretability. This paper proposes a general framework to enhance latent space representations for improving the interpretability of linear latent spaces. Although the concepts in this paper are language agnostic, the framework is written in Python. This framework automates the clustering and ranking of latent vectors to enhance the latent information per latent vector, as well as, the interpretation of latent vectors. Several innovative enhancements are incorporated including latent ranking (LR), latent scaling (LS), latent clustering (LC), and latent condensing (LCON). For a specified linear latent variable model, LR ranks latent directions according to a specified metric, LS scales latent directions according to a specified metric, LC automatically clusters latent directions into a specified number of clusters, while, LCON automatically determines an appropriate number of clusters into which to condense the latent directions for a given metric. Additional functionality of the framework includes single-channel and multi-channel data sources, data preprocessing strategies such as Hankelisation to seamlessly expand the applicability of linear latent variable models (LLVMs) to a wider variety of data. The effectiveness of LR, LS, and LCON are showcased on two crafted foundational problems with two applied latent variable models, namely, PCA and ICA.
SITTA: A Semantic Image-Text Alignment for Image Captioning
Paischer, Fabian, Adler, Thomas, Hofmarcher, Markus, Hochreiter, Sepp
Textual and semantic comprehension of images is essential for generating proper captions. The comprehension requires detection of objects, modeling of relations between them, an assessment of the semantics of the scene and, finally, representing the extracted knowledge in a language space. To achieve rich language capabilities while ensuring good image-language mappings, pretrained language models (LMs) were conditioned on pretrained multi-modal (image-text) models that allow for image inputs. This requires an alignment of the image representation of the multi-modal model with the language representations of a generative LM. However, it is not clear how to best transfer semantics detected by the vision encoder of the multi-modal model to the LM. We introduce two novel ways of constructing a linear mapping that successfully transfers semantics between the embedding spaces of the two pretrained models. The first aligns the embedding space of the multi-modal language encoder with the embedding space of the pretrained LM via token correspondences. The latter leverages additional data that consists of image-text pairs to construct the mapping directly from vision to language space. Using our semantic mappings, we unlock image captioning for LMs without access to gradient information. By using different sources of data we achieve strong captioning performance on MS-COCO and Flickr30k datasets. Even in the face of limited data, our method partly exceeds the performance of other zero-shot and even finetuned competitors. Our ablation studies show that even LMs at a scale of merely 250M parameters can generate decent captions employing our semantic mappings. Our approach makes image captioning more accessible for institutions with restricted computational resources.
Online Tensor Learning: Computational and Statistical Trade-offs, Adaptivity and Optimal Regret
Cai, Jian-Feng, Li, Jingyang, Xia, Dong
We investigate a generalized framework for estimating latent low-rank tensors in an online setting, encompassing both linear and generalized linear models. This framework offers a flexible approach for handling continuous or categorical variables. Additionally, we investigate two specific applications: online tensor completion and online binary tensor learning. To address these challenges, we propose the online Riemannian gradient descent algorithm, which demonstrates linear convergence and the ability to recover the low-rank component under appropriate conditions in all applications. Furthermore, we establish a precise entry-wise error bound for online tensor completion. Notably, our work represents the first attempt to incorporate noise in the online low-rank tensor recovery task. Intriguingly, we observe a surprising trade-off between computational and statistical aspects in the presence of noise. Increasing the step size accelerates convergence but leads to higher statistical error, whereas a smaller step size yields a statistically optimal estimator at the expense of slower convergence. Moreover, we conduct regret analysis for online tensor regression. Under the fixed step size regime, a fascinating trilemma concerning the convergence rate, statistical error rate, and regret is observed. With an optimal choice of step size we achieve an optimal regret of $O(\sqrt{T})$. Furthermore, we extend our analysis to the adaptive setting where the horizon T is unknown. In this case, we demonstrate that by employing different step sizes, we can attain a statistically optimal error rate along with a regret of $O(\log T)$. To validate our theoretical claims, we provide numerical results that corroborate our findings and support our assertions.
QI2 -- an Interactive Tool for Data Quality Assurance
Geerkens, Simon, Sieberichs, Christian, Braun, Alexander, Waschulzik, Thomas
The importance of high data quality is increasing with the growing impact and distribution of ML systems and big data. Also the planned AI Act from the European commission defines challenging legal requirements for data quality especially for the market introduction of safety relevant ML systems. In this paper we introduce a novel approach that supports the data quality assurance process of multiple data quality aspects. This approach enables the verification of quantitative data quality requirements. The concept and benefits are introduced and explained on small example data sets. How the method is applied is demonstrated on the well known MNIST data set based an handwritten digits.
Exploring Antitrust and Platform Power in Generative AI
The concentration of power in a few digital technology companies has become a subject of increasing interest in both academic and non-academic discussions. One of the most noteworthy contributions to the debate is Lina Khan's Amazon's Antitrust Paradox. In this work, Khan contends that Amazon has systematically exerted its dominance in online retail to eliminate competitors and subsequently charge above-market prices. This work contributed to Khan's appointment as the chair of the US Federal Trade Commission (FTC), one of the most influential antitrust organisations. Today, several ongoing antitrust lawsuits in the US and Europe involve major technology companies like Apple, Google/Alphabet, and Facebook/Meta. In the realm of generative AI, we are once again witnessing the same companies taking the lead in technological advancements, leaving little room for others to compete. This article examines the market dominance of these corporations in the technology stack behind generative AI from an antitrust law perspective.
Cross-Lingual Retrieval Augmented Prompt for Low-Resource Languages
Nie, Ercong, Liang, Sheng, Schmid, Helmut, Schütze, Hinrich
Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent empirical cross-lingual transfer studies. In this paper, we propose the Prompts Augmented by Retrieval Crosslingually (PARC) pipeline to improve the zero-shot performance on low-resource languages (LRLs) by augmenting the context with semantically similar sentences retrieved from a high-resource language (HRL) as prompts. PARC improves the zero-shot performance on three downstream tasks (binary sentiment classification, topic categorization and natural language inference) with multilingual parallel test sets across 10 LRLs covering 6 language families in both unlabeled settings (+5.1%) and labeled settings (+16.3%). PARC-labeled also outperforms the finetuning baseline by 3.7%. We find a significant positive correlation between cross-lingual transfer performance on one side, and the similarity between the high- and low-resource languages as well as the amount of low-resource pretraining data on the other side. A robustness analysis suggests that PARC has the potential to achieve even stronger performance with more powerful MPLMs.
Most Women Ignore Their "Reply Guys." Then There Are These People.
In May, Sydney Leathers confessed to her tens of thousands of Twitter followers that she was smitten. Where'd she meet the guy? Not on a dating app, or through friends, but in the last place she ever expected to find a real connection: her mentions. "Still can't believe I fell in love with one of my reply guys. Apparently, things had progressed since December, when she last posted about him: "I had sex with someone who started as my reply guy and I hope this doesn't inspire confidence in the rest of you because frankly your replies are not that good," she wrote. Leathers is a writer, adult performer, and startup employee whose name you may recognize from her part in the Anthony Weiner sexting scandal--this wasn't exactly her first brush with online flirtation. But it was her first time falling for a reply guy, or someone who was, effectively, a fan. The term "reply guy" emerged on Twitter about five years ago to describe the behavior of a certain subset of people, usually with very few social media followers of their own, who stake out space in the mentions of prominent users. They can be counted on to reply promptly and frequently to the tweets of whomever they've chosen as their object of devotion, and they often seek attention by nitpicking, mansplaining, joke one-upping, and harassing them. Because of this, reply guys--who can also be girls, or people of any gender--are generally understood to be pathetic creatures, without a chance in hell of getting said person to like their replies, much less return their affections. So the revelation that this gambit actually worked for someone is … pretty noteworthy. Reply guy success stories may be happening more than we realize. Abby, a 25-year-old in Brooklyn who runs a meme page on Instagram with several thousand followers, told me that she got frisky with one of her reply guys last year. "I'm not the only person that I know that has hooked up with reply guys," she said. "It's not as uncommon as you might think." Now, Leathers' Twitter feed is a monument to her relationship, by turns adorable and lewd. "This definitely caught me by surprise," she told me. "But it's been the best, happiest relationship I've had." To attain this goal, a reply guy's first challenge is to stand out from the crowd. The meme account Abby is the admin for is about politics, so she likes when a guy can show not just that he's hot, but that they share a political sensibility. "I have to be attracted to them," she said. "And they have to have some sort of compelling thing to say." "I feel like I've never more than mildly acknowledged a reply guy before now," she said. "I generally don't even follow them back." But when her now-boyfriend started responding to her tweets last year after discovering her through a winding path that involved the singer of the band Eve 6, she took notice. "I'd seen him reply to my stuff a few times.
Mixed integer linear optimization formulations for learning optimal binary classification trees
Alston, Brandon, Validi, Hamidreza, Hicks, Illya V.
Decision trees are powerful tools for classification and regression that attract many researchers working in the burgeoning area of machine learning. One advantage of decision trees over other methods is their interpretability, which is often preferred over other higher accuracy methods that are relatively uninterpretable. A binary classification tree has two types of vertices: (i) branching vertices which have exactly two children and where datapoints are assessed on a set of discrete features; and (ii) leaf vertices at which datapoints are given a discrete prediction. An optimal binary classification tree can be obtained by solving a biobjective optimization problem that seeks to (i) maximize the number of correctly classified datapoints and (ii) minimize the number of branching vertices. In this paper, we propose four mixed integer linear optimization (MILO) formulations for designing optimal binary classification trees: two flow-based formulations and two-cut based formulations. We provide theoretical comparisons between our proposed formulations and the strongest flow-based MILO formulation of Aghaei et al. (2021). We conduct experiments on 13 publicly available datasets to show the models' ability to scale and the strength of a biobjective approach using Pareto frontiers. Our code and data are available on GitHub.