Goto

Collaborating Authors

 Education


Do Language Models Plagiarize?

arXiv.org Artificial Intelligence

In this work, therefore, we study three types of plagiarism (i.e., verbatim, paraphrase, and idea) among GPT-2 generated texts, Language Models (LMs) have become core elements of Natural in comparison to its training data, and further analyze the plagiarism Language Processing (NLP) solutions, excelling in a wide range of patterns of fine-tuned LMs with domain-specific corpora which are tasks such as natural language generation (NLG), speech recognition, extensively used in practice. Our results suggest that (1) three types machine translation, and question answering. The development of plagiarism widely exist in LMs beyond memorization, (2) both of large-scale text corpora (generally scraped from the Web) has size and decoding methods of LMs are strongly associated with the enabled researchers to train increasingly large-scale LMs. Especially, degrees of plagiarism they exhibit, and (3) fine-tuned LMs' plagiarism large-scale LMs have demonstrated unprecedented performance on patterns vary based on their corpus similarity and homogeneity. NLG such that LM-generated texts routinely show more novel and Given that a majority of LMs' training data is scraped from the Web interesting stories than human writings do [35], and the distinction without informing content owners, their reiteration of words, phrases, between machine-authored and human-written texts has become and even core ideas from training sets into generated texts has ethical non-trivial [52, 53]. As a result, there has been a significant increase implications. Their patterns are likely to exacerbate as both in the use of LMs in user-facing products and critical applications.


Understanding Transformer Memorization Recall Through Idioms

arXiv.org Artificial Intelligence

To produce accurate predictions, language models (LMs) must balance between generalization and memorization. Yet, little is known about the mechanism by which transformer LMs employ their memorization capacity. When does a model decide to output a memorized phrase, and how is this phrase then retrieved from memory? In this work, we offer the first methodological framework for probing and characterizing recall of memorized sequences in transformer LMs. First, we lay out criteria for detecting model inputs that trigger memory recall, and propose idioms as inputs that typically fulfill these criteria. Next, we construct a dataset of English idioms and use it to compare model behavior on memorized vs. non-memorized inputs. Specifically, we analyze the internal prediction construction process by interpreting the model's hidden representations as a gradual refinement of the output probability distribution. We find that across different model sizes and architectures, memorized predictions are a two-step process: early layers promote the predicted token to the top of the output distribution, and upper layers increase model confidence. This suggests that memorized information is stored and retrieved in the early layers of the network. Last, we demonstrate the utility of our methodology beyond idioms in memorized factual statements. Overall, our work makes a first step towards understanding memory recall, and provides a methodological basis for future studies of transformer memorization.


Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment

arXiv.org Artificial Intelligence

The presence of decision-making algorithms in society is rapidly increasing nowadays, while concerns about their transparency and the possibility of these algorithms becoming new sources of discrimination are arising. There is a certain consensus about the need to develop AI applications with a Human-Centric approach. Human-Centric Machine Learning needs to be developed based on four main requirements: (i) utility and social good; (ii) privacy and data ownership; (iii) transparency and accountability; and (iv) fairness in AI-driven decision-making processes. All these four Human-Centric requirements are closely related to each other. With the aim of studying how current multimodal algorithms based on heterogeneous sources of information are affected by sensitive elements and inner biases in the data, we propose a fictitious case study focused on automated recruitment: FairCVtest. We train automatic recruitment algorithms using a set of multimodal synthetic profiles including image, text, and structured data, which are consciously scored with gender and racial biases. FairCVtest shows the capacity of the Artificial Intelligence (AI) behind automatic recruitment tools built this way (a common practice in many other application scenarios beyond recruitment) to extract sensitive information from unstructured data and exploit it in combination to data biases in undesirable (unfair) ways. We present an overview of recent works developing techniques capable of removing sensitive information and biases from the decision-making process of deep learning architectures, as well as commonly used databases for fairness research in AI. We demonstrate how learning approaches developed to guarantee privacy in latent spaces can lead to unbiased and fair automatic decision-making process.


Book review: 'Quantum in Pictures'

Oxford Comp Sci

The latest work by computer scientists Bob Coecke and Stefano Gogioso, 'Quantum in Pictures', aims to make the quantum world more accessible and inclusive. So, whether you're a high school student or a science enthusiast, the authors are confident that anyone mastering the tools in the book will gain an understanding equivalent to that of a quantum mechanics graduate at university. But what if a complete novice in quantum computing, i.e., this reviewer, could gain a genuine understanding of the field by simply reading this book? Let's test this out, shall we? Full disclosure from the get-go, I have absolutely no prior knowledge or expertise in quantum computing, therefore Coecke and Gogioso's latest research and book is not only worthy of a review but also a lesson for someone who barely scraped a C in GSCE Maths โ€“ a learning curve, if you will. For context, 'Quantum in Pictures' is the brainchild of Quantinuum's chief scientist Professor Bob Coecke and Dr Stefano Gogioso of Oxford University. The book introduces a formalism for quantum mechanics based on using'ZX-calculus' (or'ZX'), to describe quantum processes.


Building a Career in Data Science

#artificialintelligence

I currently work at Rebaie Analytics Group to develop algorithms in computer vision, natural language processing, and other deep learning fields. In college, I started reading about the impact of data science in transforming business and even in the way humans interact with machines in our daily lives. Further inspired by the AI influencer and keynote speaker Ali Rebaie, I wanted to apply an anthropological perspective to solve current AI challenges. Like I do with any subject I'm interested in, I jumped right into learning everything I could, starting with taking machine learning courses online. I was glad to find Coursera -- it's really the most effective and interactive e-learning platform out there.


iiot machinelearning, Twitter, 2/10/2023 12:29:00 PM, 289068

#artificialintelligence

The graph represents a network of 1,389 Twitter users whose recent tweets contained "iiot machinelearning", or who were replied to, mentioned, retweeted or quoted in those tweets, taken from a data set limited to a maximum of 5,000 tweets, tweeted between 3/26/2006 12:00:00 AM and 2/9/2023 5:00:35 PM. The network was obtained from Twitter on Friday, 10 February 2023 at 12:24 UTC. The tweets in the network were tweeted over the 1474-day, 5-hour, 23-minute period from Sunday, 27 January 2019 at 19:36 UTC to Friday, 10 February 2023 at 01:00 UTC. There is an edge for each "replies-to" relationship in a tweet, an edge for each "mentions" relationship in a tweet, an edge for each "retweet" relationship in a tweet, an edge for each "quote" relationship in a tweet, an edge for each "mention in retweet" relationship in a tweet, an edge for each "mention in reply-to" relationship in a tweet, an edge for each "mention in quote" relationship in a tweet, an edge for each "mention in quote reply-to" relationship in a tweet, and a self-loop edge for each tweet that is not from above. The graph's vertices were grouped by cluster using the Clauset-Newman-Moore cluster algorithm.


AI and the future of Teaching and learning

#artificialintelligence

In Minnesota, a start-up company recently created a 19-lesson, fully online, three-hour multimedia course in just 10 hours using ChatGPT, the artificial intelligence tool launched in November 2022. ChatGPT found images and relevant video materials and developed a quiz to assess learning. Subsequent courses created by this same team are being created in less time -- just one hour to create a three-hour learning module. Elsewhere, ChatGPT is used to create multimedia webpages that can be quickly inserted into websites, and to create code in python (and other computer languages) that can be incorporated into apps or web spaces. ChatGPT is one of many similar AI services that use natural language to respond to user questions or requirements.


From high-dimensional & mean-field dynamics to dimensionless ODEs: A unifying approach to SGD in two-layers networks

arXiv.org Artificial Intelligence

This manuscript investigates the one-pass stochastic gradient descent (SGD) dynamics of a two-layer neural network trained on Gaussian data and labels generated by a similar, though not necessarily identical, target function. We rigorously analyse the limiting dynamics via a deterministic and low-dimensional description in terms of the sufficient statistics for the population risk. Our unifying analysis bridges different regimes of interest, such as the classical gradient-flow regime of vanishing learning rate, the high-dimensional regime of large input dimension, and the overparameterised "mean-field" regime of large network width, covering as well the intermediate regimes where the limiting dynamics is determined by the interplay between these behaviours. In particular, in the high-dimensional limit, the infinite-width dynamics is found to remain close to a low-dimensional subspace spanned by the target principal directions. Our results therefore provide a unifying picture of the limiting SGD dynamics with synthetic data.


Physics informed WNO

arXiv.org Artificial Intelligence

Deep neural operators are recognized as an effective tool for learning solution operators of complex partial differential equations (PDEs). As compared to laborious analytical and computational tools, a single neural operator can predict solutions of PDEs for varying initial or boundary conditions and different inputs. A recently proposed Wavelet Neural Operator (WNO) is one such operator that harnesses the advantage of time-frequency localization of wavelets to capture the manifolds in the spatial domain effectively. While WNO has proven to be a promising method for operator learning, the data-hungry nature of the framework is a major shortcoming. In this work, we propose a physics-informed WNO for learning the solution operators of families of parametric PDEs without labeled training data. The efficacy of the framework is validated and illustrated with four nonlinear spatiotemporal systems relevant to various fields of engineering and science.


AIDA: Legal Judgment Predictions for Non-Professional Fact Descriptions via Partial-and-Imbalanced Domain Adaptation

arXiv.org Artificial Intelligence

In this paper, we study the problem of legal domain adaptation problem from an imbalanced source domain to a partial target domain. The task aims to improve legal judgment predictions for non-professional fact descriptions. We formulate this task as a partial-and-imbalanced domain adaptation problem. Though deep domain adaptation has achieved cutting-edge performance in many unsupervised domain adaptation tasks. However, due to the negative transfer of samples in non-shared classes, it is hard for current domain adaptation model to solve the partial-and-imbalanced transfer problem. In this work, we explore large-scale non-shared but related classes data in the source domain with a hierarchy weighting adaptation to tackle this limitation. We propose to embed a novel pArtial Imbalanced Domain Adaptation technique (AIDA) in the deep learning model, which can jointly borrow sibling knowledge from non-shared classes to shared classes in the source domain and further transfer the shared classes knowledge from the source domain to the target domain. Experimental results show that our model outperforms the state-of-the-art algorithms.