Goto

Collaborating Authors

 Media


Exploring Graph Representation of Chorales

arXiv.org Artificial Intelligence

This work explores areas overlapping music, graph theory, and machine learning. An embedding representation of a node, in a weighted undirected graph $\mathcal{G}$, is a representation that captures the meaning of nodes in an embedding space. In this work, 383 Bach chorales were compiled and represented as a graph. Two application cases were investigated in this paper (i) learning node embedding representation using \emph{Continuous Bag of Words (CBOW), skip-gram}, and \emph{node2vec} algorithms, and (ii) learning node labels from neighboring nodes based on a collective classification approach. The results of this exploratory study ascertains many salient features of the graph-based representation approach applicable to music applications.


Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation

arXiv.org Artificial Intelligence

Large pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", without a sufficiently large training set. In this work, we propose a novel conditional neural process-based approach for few-shot text classification that learns to transfer from other diverse tasks with rich annotation. Our key idea is to represent each task using gradient information from a base model and to train an adaptation network that modulates a text classifier conditioned on the task representation. While previous task-aware few-shot learners represent tasks by input encoding, our novel task representation is more powerful, as the gradient captures input-output relationships of a task. Experimental results show that our approach outperforms traditional fine-tuning, sequential transfer learning, and state-of-the-art meta learning approaches on a collection of diverse few-shot tasks. We further conducted analysis and ablations to justify our design choices.


Top 60 Data Science Interview Questions and Answers 2022

#artificialintelligence

Harvard Business Review referred to data scientist as the "Sexiest Job of the 21st Century." Glassdoor placed it #1 on the 25 Best Jobs in America list. According to IBM, demand for this role will soar 28 percent by 2020. It should come as no surprise that in the new era of big data and machine learning, data scientists are becoming rock stars. Companies that are able to leverage massive amounts of data to improve the way they serve customers, build products, and run their operations will be positioned to thrive in this economy. And if you're moving down the path to becoming a data scientist, you must be prepared to impress prospective employers with your knowledge. And to do that you must be able to crack your next data science interview in one go! We have clubbed a list of the most popular data science interview questions you can expect in your next interview!


Team Yao at Factify 2022: Utilizing Pre-trained Models and Co-attention Networks for Multi-Modal Fact Verification

arXiv.org Artificial Intelligence

In recent years, social media has enabled users to get exposed to a myriad of misinformation and disinformation; thus, misinformation has attracted a great deal of attention in research fields and as a social issue. To address the problem, we propose a framework, Pre-CoFact, composed of two pre-trained models for extracting features from text and images, and multiple co-attention networks for fusing the same modality but different sources and different modalities. Besides, we adopt the ensemble method by using different pre-trained models in Pre-CoFact to achieve better performance. We further illustrate the effectiveness from the ablation study and examine different pre-trained models for comparison. Our team, Yao, won the fifth prize (F1-score: 74.585\%) in the Factify challenge hosted by De-Factify @ AAAI 2022, which demonstrates that our model achieved competitive performance without using auxiliary tasks or extra information. The source code of our work is publicly available at https://github.com/wywyWang/Multi-Modal-Fact-Verification-2021


Explainable Patterns for Distinction and Prediction of Moral Judgement on Reddit

arXiv.org Artificial Intelligence

The forum r/AmITheAsshole in Reddit hosts discussion on moral issues based on concrete narratives presented by users. Existing analysis of the forum focuses on its comments, and does not make the underlying data publicly available. In this paper we build a new dataset of comments and also investigate the classification of the posts in the forum. Further, we identify textual patterns associated with the provocation of moral judgement by posts, with the expression of moral stance in comments, and with the decisions of trained classifiers of posts and comments.


Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

arXiv.org Artificial Intelligence

Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles -- written by students from across the country -- we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban ZIP codes are more likely to be classified as high quality. We then demonstrate that the filter's measurement of quality is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.


All the Star Wars games currently in development

Washington Post - Technology News

While details on the three new games remain sparse, EA did offer some personnel info. Development on the new "Jedi" game will be headed by Stig Asmussen, who helmed the previous entry and, before that, Sony's "God of War III." The new first-person shooter, meanwhile, will be directed by Peter Hirschmann, who previously worked on numerous Star Wars games including "Star Wars: Battlefront" and "Star Wars: The Force Unleashed." The strategy game will be designed by a new studio formed by Greg Foertsch, a developer on the revered XCOM series of turn-based, sci-fi strategy games. Respawn Entertainment, creator of Titanfall and "Apex Legends," will lead development on the "Jedi" sequel and the shooter while handling production for the strategy game.


James Cameron warns of the dangers of deepfakes

#artificialintelligence

Legendary director James Cameron has warned of the dangers that deepfakes pose to society. Deepfakes leverage machine learning and AI techniques to convincingly manipulate or generate visual and audio content. Their high potential to deceive makes them a powerful tool for spreading disinformation, committing fraud, trolling, and more. "Every time we improve these tools, we're actually in a sense building a toolset to create fake media -- and we're seeing it happening now," said Cameron in a BBC video interview. "Right now the tools are -- the people just playing around on apps aren't that great. But over time, those limitations will go away. Things that you see and fully believe you're seeing could be faked."


Preserving Integrity in Online Social Networks

Communications of the ACM

The goal of online social networks is to help create connections between people (online and offline), to connect people to communities of interest, and to provide a forum for advancing culture. Social networks advance these causes by providing a platform for free expression by anyone, whether they are well-known figures or your next-door neighbor. Unfortunately, open platforms for free expression can be used for malicious purposes. People and organizations can distribute misinformation and hate speech and can use the platform to commit crimes such as selling illegal drugs, coordinating sex trafficking, or child exploitation. All these violations existed much before the advent of social networks, but social networks exacerbate the scale and sophistication with which these activities can be carried out. Naturally, fighting these violations, which we collectively refer to as the problem of preserving integrity in online networks (or simply, integrity), has become a huge priority for the companies running them and for society at large. Setting policies for what content and behavior are allowed on social networks is an area fraught with debate because it involves striking a balance between free expression and removing offending content. In addition, the policies must be sensitive to a variety of cultures and political climates all over the world. While we touch on the policy backdrop, this survey focuses on the technical challenges that arise in enforcing the policies.


How Topic Novelty Impacts the Effectiveness of News Veracity Interventions

Communications of the ACM

As previously mentioned, we employed a 2x2 design, in which we studied two news conditions (familiar news and novel news) and two AI conditions (No AI and AI Intervention). The familiar news condition contained news articles on vaccination and climate change, both of which had been widely reported prior to this study, while the novel news condition contained news articles on COVID-19. The No AI condition provided just the article, while the AI Intervention condition displayed one of two statements at the top of the article: either "Our smart AI system rates this article as accurate and reliable" or "Our smart AI system rates this article as inaccurate and unreliable." Each participant read one randomly chosen article and answered two questions: 1. Do you believe the information in this news article?