Media
Identifying Misinformation from Website Screenshots
Abdali, Sara, Gurav, Rutuja, Menon, Siddharth, Fonseca, Daniel, Entezari, Negin, Shah, Neil, Papalexakis, Evangelos E.
Can the look and the feel of a website give information about the trustworthiness of an article? In this paper, we propose to use a promising, yet neglected aspect in detecting the misinformativeness: the overall look of the domain webpage. To capture this overall look, we take screenshots of news articles served by either misinformative or trustworthy web domains and leverage a tensor decomposition based semi-supervised classification technique. The proposed approach i.e., VizFake is insensitive to a number of image transformations such as converting the image to grayscale, vectorizing the image and losing some parts of the screenshots. VizFake leverages a very small amount of known labels, mirroring realistic and practical scenarios, where labels (especially for known misinformative articles), are scarce and quickly become dated. The F1 score of VizFake on a dataset of 50k screenshots of news articles spanning more than 500 domains is roughly 85% using only 5% of ground truth labels. Furthermore, tensor representations of VizFake, obtained in an unsupervised manner, allow for exploratory analysis of the data that provides valuable insights into the problem. Finally, we compare VizFake with deep transfer learning, since it is a very popular black-box approach for image classification and also well-known text text-based methods. VizFake achieves competitive accuracy with deep transfer learning models while being two orders of magnitude faster and not requiring laborious hyper-parameter tuning.
Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon
Veisi, Hadi, Hosseini, Hawre, Mohammadamini, Mohammad, Fathy, Wirya, Mahmudi, Aso
In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several countries, but due to the lack of speech and text resources, there is no speech recognition system for this language. To fill this gap, we introduce the first speech corpus and pronunciation lexicon for the Kurdish language. Regarding speech corpus, we designed a sentence collection in which the ratio of di-phones in the collection resembles the real data of the Central Kurdish language. The designed sentences are uttered by 576 speakers in a controlled environment with noise-free microphones (called AsoSoft Speech-Office) and in Telegram social network environment using mobile phones (denoted as AsoSoft Speech-Crowdsourcing), resulted in 43.68 hours of speech. Besides, a test set including 11 different document topics is designed and recorded in two corresponding speech conditions (i.e., Office and Crowdsourcing). Furthermore, a 60K pronunciation lexicon is prepared in this research in which we faced several challenges and proposed solutions for them. The Kurdish language has several dialects and sub-dialects that results in many lexical variations. Our methods for script standardization of lexical variations and automatic pronunciation of the lexicon tokens are presented in detail. To setup the recognition engine, we used the Kaldi toolkit. A statistical tri-gram language model that is extracted from the AsoSoft text corpus is used in the system. Several standard recipes including HMM-based models (i.e., mono, tri1, tr2, tri2, tri3), SGMM, and DNN methods are used to generate the acoustic model. These methods are trained with AsoSoft Speech-Office and AsoSoft Speech-Crowdsourcing and a combination of them. The best performance achieved by the SGMM acoustic model which results in 13.9% of the average word error rate (on different document topics) and 4.9% for the general topic.
Us vs. Them: A Dataset of Populist Attitudes, News Bias and Emotions
Huguet-Cabot, Pere-Lluís, Abadi, David, Fischer, Agneta, Shutova, Ekaterina
Computational modelling of political discourse tasks has become an increasingly important area of research in natural language processing. Populist rhetoric has risen across the political sphere in recent years; however, computational approaches to it have been scarce due to its complex nature. In this paper, we present the new $\textit{Us vs. Them}$ dataset, consisting of 6861 Reddit comments annotated for populist attitudes and the first large-scale computational models of this phenomenon. We investigate the relationship between populist mindsets and social groups, as well as a range of emotions typically associated with these. We set a baseline for two tasks related to populist attitudes and present a set of multi-task learning models that leverage and demonstrate the importance of emotion and group identification as auxiliary tasks.
Video shows melting snowflakes freezing back into original form
Capturing snowflakes on film can be quite the feat, as photographers have mere before the tiny ice crystal's intricate details melt – but a new video shows the event in reverse. Photographer Jens recently shared a stunning video showing already melted snowflakes freezing back to their original form. Each shot begins with a small droplet of water that begins to sprout icicles until it returns to the unique design. The movie was done using highly detailed macro photography, which is capable of making very small object look larger than life size. Capturing snowflakes on film can be quite the feat, as photographers have mere before the tiny ice crystal's intricate details melt – but a new video shows the event in reverse.
What UFOs and Joe McCarthy Have to Do With the Assault on the Capitol
On a cold December night in 1950, red-baiting Sen. Joe McCarthy spent a charity dinner at Washington's Sulgrave Club trading insults with liberal journalist Drew Pearson. McCarthy had attacked Pearson on the floor of the Senate, calling for a boycott of his radio show. Pearson had attacked McCarthy on air and in his newspaper column, accusing the senator of lying about communist infiltration of the American government. McCarthy had recklessly accused the State Department of harboring hundreds of communists, sparking a massive investigation and an ongoing purge. After dinner, the two ran into each other in the cloakroom and their conflict turned physical.
In Science Fiction, We Are Never Home - Issue 95: Escape
This essay first appeared in our "Home" issue way back in 2013. But somehow feels so timely today. Halfway through director Alfonso Cuarón's Gravity, Sandra Bullock suffers the most cosmic case of homesick blues since Keir Dullea was hurled toward the infinite in 2001: A Space Odyssey nearly half a century ago. For Bullock, home is (as it was for Dullea) the Earth, looming below so huge it would seem she couldn't miss it, if she could somehow just fall from her shattered spacecraft. She cares about nothing more than getting back to where she came from, even as 2001's Dullea is in flight, accepting his exile and even embracing it.
Modeling Dynamic User Interests: A Neural Matrix Factorization Approach
Dhillon, Paramveer, Aral, Sinan
In recent years, there has been significant interest in understanding users' online content consumption patterns. But, the unstructured, high-dimensional, and dynamic nature of such data makes extracting valuable insights challenging. Here we propose a model that combines the simplicity of matrix factorization with the flexibility of neural networks to efficiently extract nonlinear patterns from massive text data collections relevant to consumers' online consumption patterns. Our model decomposes a user's content consumption journey into nonlinear user and content factors that are used to model their dynamic interests. This natural decomposition allows us to summarize each user's content consumption journey with a dynamic probabilistic weighting over a set of underlying content attributes. The model is fast to estimate, easy to interpret and can harness external data sources as an empirical prior. These advantages make our method well suited to the challenges posed by modern datasets. We use our model to understand the dynamic news consumption interests of Boston Globe readers over five years. Thorough qualitative studies, including a crowdsourced evaluation, highlight our model's ability to accurately identify nuanced and coherent consumption patterns. These results are supported by our model's superior and robust predictive performance over several competitive baseline methods.