Goto

Collaborating Authors

 Media


Fake News Detection Using Machine Learning Ensemble Methods

#artificialintelligence

The advent of the World Wide Web and the rapid adoption of social media platforms (such as Facebook and Twitter) paved the way for information dissemination that has never been witnessed in the human history before. With the current usage of social media platforms, consumers are creating and sharing more information than ever before, some of which are misleading with no relevance to reality. Automated classification of a text article as misinformation or disinformation is a challenging task. Even an expert in a particular domain has to explore multiple aspects before giving a verdict on the truthfulness of an article. In this work, we propose to use machine learning ensemble approach for automated classification of news articles. Our study explores different textual properties that can be used to distinguish fake contents from real. By using those properties, we train a combination of different machine learning algorithms using various ensemble methods and evaluate their performance on 4 real world datasets. Experimental evaluation confirms the superior performance of our proposed ensemble learner approach in comparison to individual learners. The advent of the World Wide Web and the rapid adoption of social media platforms (such as Facebook and Twitter) paved the way for information dissemination that has never been witnessed in the human history before. Besides other use cases, news outlets benefitted from the widespread use of social media platforms by providing updated news in near real time to its subscribers. The news media evolved from newspapers, tabloids, and magazines to a digital form such as online news platforms, blogs, social media feeds, and other digital media formats [1]. It became easier for consumers to acquire the latest news at their fingertips.


Assisting the Human Fact-Checkers: Detecting All Previously Fact-Checked Claims in a Document

arXiv.org Artificial Intelligence

Given the recent proliferation of false claims online, there has been a lot of manual fact-checking effort. As this is very time-consuming, human fact-checkers can benefit from tools that can support them and make them more efficient. Here, we focus on building a system that could provide such support. Given an input document, it aims to detect all sentences that contain a claim that can be verified by some previously fact-checked claims (from a given database). The output is a re-ranked list of the document sentences, so that those that can be verified are ranked as high as possible, together with corresponding evidence. Unlike previous work, which has looked into claim retrieval, here we take a document-level perspective. We create a new manually annotated dataset for the task, and we propose suitable evaluation measures. We further experiment with a learning-to-rank approach, achieving sizable performance gains over several strong baselines. Our analysis demonstrates the importance of modeling text similarity and stance, while also taking into account the veracity of the retrieved previously fact-checked claims. We believe that this research would be of interest to fact-checkers, journalists, media, and regulatory authorities.


Controllable Dialogue Generation with Disentangled Multi-grained Style Specification and Attribute Consistency Reward

arXiv.org Artificial Intelligence

Controllable text generation is an appealing but challenging task, which allows users to specify particular attributes of the generated outputs. In this paper, we propose a controllable dialogue generation model to steer response generation under multi-attribute constraints. Specifically, we define and categorize the commonly used control attributes into global and local ones, which possess different granularities of effects on response generation. Then, we significantly extend the conventional seq2seq framework by introducing a novel two-stage decoder, which first uses a multi-grained style specification layer to impose the stylistic constraints and determine word-level control states of responses based on the attributes, and then employs a response generation layer to generate final responses maintaining both semantic relevancy to the contexts and fidelity to the attributes. Furthermore, we train our model with an attribute consistency reward to promote response control with explicit supervision signals. Extensive experiments and in-depth analyses on two datasets indicate that our model can significantly outperform competitive baselines in terms of response quality, content diversity and controllability.


How AI in Video Will Enhance Work in the Modern-Day Work Environment - ReadWrite

#artificialintelligence

Shaking off the dust from what could be described as the longest year known to man -- remote work is a hot topic in the world of employment. By establishing both its benefits, as well as its challenges, remote work has people talking about its permanence. What is more, employees have become accustomed to remote working, in fact, many of them actually prefer it to the office. According to a FlexJobs survey, 65% of employee respondents reported wanting to be full-time remote post-pandemic, and 31% want a hybrid remote work environment -- that's 96% who desire some form of remote work. These numbers inevitably mean that the methods in which we worked during the pandemic, primarily via the screen and through video calls, will have some longevity.


'National Artificial Intelligence Advisory Committee' to guide Biden on tech affecting society …

#artificialintelligence

"We must be sure that these [artificial intelligence] advances are matched by similar progress in ensuring that AI is trustworthy, and that it ensures …


You, Me, and My AI-Generated Alternate Identity

#artificialintelligence

There, she posts pictures of herself in a biker shirt, posing in front of her gleaming red-and-blue Yamaha Telkor on dirt roads and hilltops and misty beaches. But one day, she accidentally posted a picture of her bike on Twitter that captured her reflection in the rear-view mirror. The reflection was of a middle-aged man–because the woman in the photo was actually a 50-year-old man named Soya who transformed his face using a machine-learning-powered face tune app. "No-one will read what a normal middle-aged man, taking care of his motorcycle and taking pictures outside, posts on his account," Soya told the Japanese TV program Getsuyou Kara Yofukashi. That said, happily, his fans responded mostly positively to his late-in-life, accidental gender reveal.


AI study reveals the secret of an artistic 'hot streak'

Daily Mail - Science & tech

Whether an artist, scientist, or film director, trailblazers in particular fields often have a critically-acclaimed'hot streak' where they produce a series of outstanding work in short succession. Now, scientists at Northwestern University in Illinois claim to have pinpointed the secret formula that often triggers a pioneer's best work. Using a form of artificial intelligence (AI) called deep learning, they mined data related to thousands of artists, film directors and scientists to identify a magical formula for success. Hot streaks directly result from years of'exploration' (studying diverse styles or topics), immediately followed by years of'exploitation' (focusing on a narrow area to develop deep expertise), they claim. They define a hot streak as a burst of high-impact works clustered together in close succession – as achieved by artists such as Vincent Van Gogh and Jackson Pollock, or film directors like Peter Jackson or Alfred Hitchcock.


Artificial Intelligence (AI) in Cybersecurity 2021

#artificialintelligence

Organizations across industries are turning to artificial intelligence (AI) in cybersecurity to protect their networks and relieve often …


Online Learning of Optimally Diverse Rankings

arXiv.org Machine Learning

Search engines answer users' queries by listing relevant items (e.g. documents, songs, products, web pages, ...). These engines rely on algorithms that learn to rank items so as to present an ordered list maximizing the probability that it contains relevant item. The main challenge in the design of learning-to-rank algorithms stems from the fact that queries often have different meanings for different users. In absence of any contextual information about the query, one often has to adhere to the {\it diversity} principle, i.e., to return a list covering the various possible topics or meanings of the query. To formalize this learning-to-rank problem, we propose a natural model where (i) items are categorized into topics, (ii) users find items relevant only if they match the topic of their query, and (iii) the engine is not aware of the topic of an arriving query, nor of the frequency at which queries related to various topics arrive, nor of the topic-dependent click-through-rates of the items. For this problem, we devise LDR (Learning Diverse Rankings), an algorithm that efficiently learns the optimal list based on users' feedback only. We show that after $T$ queries, the regret of LDR scales as $O((N-L)\log(T))$ where $N$ is the number of all items. We further establish that this scaling cannot be improved, i.e., LDR is order optimal. Finally, using numerical experiments on both artificial and real-world data, we illustrate the superiority of LDR compared to existing learning-to-rank algorithms.


Hetero-SCAN: Towards Social Context Aware Fake News Detection via Heterogeneous Graph Neural Network

arXiv.org Artificial Intelligence

Fake news, false or misleading information presented as news, has a great impact on many aspects of society, such as politics and healthcare. To handle this emerging problem, many fake news detection methods have been proposed, applying Natural Language Processing (NLP) techniques on the article text. Considering that even people cannot easily distinguish fake news by news content, these text-based solutions are insufficient. To further improve fake news detection, researchers suggested graph-based solutions, utilizing the social context information such as user engagement or publishers information. However, existing graph-based methods still suffer from the following four major drawbacks: 1) expensive computational cost due to a large number of user nodes in the graph, 2) the error in sub-tasks, such as textual encoding or stance detection, 3) loss of rich social context due to homogeneous representation of news graphs, and 4) the absence of temporal information utilization. In order to overcome the aforementioned issues, we propose a novel social context aware fake news detection method, Hetero-SCAN, based on a heterogeneous graph neural network. Hetero-SCAN learns the news representation from the heterogeneous graph of news in an end-to-end manner. We demonstrate that Hetero-SCAN yields significant improvement over state-of-the-art text-based and graph-based fake news detection methods in terms of performance and efficiency.