Goto

Collaborating Authors

 Genre


A Longitudinal Study of Topic Classification on Twitter

AAAI Conferences

Twitter represents a massively distributed information source over a kaleidoscope of topics ranging from social and political events to entertainment and sports news. While recent work has suggested that variations on standard classifiers can be effectively trained as topical filters (Lin, Snow, and Morgan 2011; Yang et al. 2014; Magdy and Elsayed 2014), there remain many open questions about the efficacy of such classification-based filtering approaches. For example, over a year or more after training, how well do such classifiers generalize to future novel topical content, and are such results stable across a range of topics? Furthermore, what features and feature classes are most critical for long-term classifier performance? To answer these questions, we collected a corpus of over 800 million English Tweets via the Twitter streaming API during 2013 and 2014 and learned topic classifiers for 10 diverse themes ranging from social issues to celebrity deaths to the โ€œIran nuclear dealโ€. The results of this long-term study of topic classifier performance provide a number of important insights, among them that (1) such classifiers can indeed generalize to novel topical content with high precision over a year or more after training and (2) simple terms and locations are the most informative feature classes (despite training on classes labeled via hashtags).


Spatio-Temporal Analysis of Reverted Wikipedia Edits

AAAI Conferences

Little is known about what causes anti-social behavior online. The paper at hand analyzes vandalism and damage in Wikipedia with regard to the time it is conducted and the country it originates from. First, we identify vandalism and damaging edits via ex post facto evidence by mining Wikipediaโ€™s revert graph. Second, we geolocate the cohort of edits from anonymous Wikipedia editors using their associated IP addresses and edit times, showing the feasibility of reliable historic geolocation with respect to country and time zone, even under limited geolocation data. Third, we conduct the first spatio-temporal analysis of vandalism on Wikipedia. Our analysis reveals significant differences for vandalism activities during the day, and for different days of the week, seasons, countries of origin, as well as Wikipediaโ€™s languages. For the analyzed countries, the ratio is typically highest at non-summer workday mornings, with additional peaks after break times. We hence assume that Wikipedia vandalism is linked to labor, perhaps serving as relief from stress or boredom, whereas cultural differences have a large effect. Our results open up avenues for new research on collaborative writing at scale, and advanced technologies to identify and handle antisocial behavior in online communities.


A rational analysis of curiosity

arXiv.org Artificial Intelligence

We present a rational analysis of curiosity, proposing that people's curiosity is driven by seeking stimuli that maximize their ability to make appropriate responses in the future. This perspective offers a way to unify previous theories of curiosity into a single framework. Experimental results confirm our model's predictions, showing how the relationship between curiosity and confidence can change significantly depending on the nature of the environment.


Learning activation functions from data using cubic spline interpolation

arXiv.org Machine Learning

Neural networks require a careful design in order to perform properly on a given task. In particular, selecting a good activation function (possibly in a data-dependent fashion) is a crucial step, which remains an open problem in the research community. Despite a large amount of investigations, most current implementations simply select one fixed function from a small set of candidates, which is not adapted during training, and is shared among all neurons throughout the different layers. However, neither two of these assumptions can be supposed optimal in practice. In this paper, we present a principled way to have data-dependent adaptation of the activation functions, which is performed independently for each neuron. This is achieved by leveraging over past and present advances on cubic spline interpolation, allowing for local adaptation of the functions around their regions of use. The resulting algorithm is relatively cheap to implement, and overfitting is counterbalanced by the inclusion of a novel damping criterion, which penalizes unwanted oscillations from a predefined shape. Experimental results validate the proposal over two well-known benchmarks.


How artificial intelligence is helping detect tuberculosis in remote areas

#artificialintelligence

Researchers are training artificial intelligence to identify tuberculosis on chest X-rays, an initiative that could help screening and evaluation efforts in TB-prevalent areas lacking access to radiologists. The findings are part of a study published online in the journal Radiology. "An artificial intelligence solution that could interpret radiographs for the presence of TB in a cost-effective way could expand the reach of early identification and treatment in developing nations," study co-author Paras Lakhani, MD, from Thomas Jefferson University Hospital in Philadelphia, wrote in the journal. For the study, Lakhani and his colleague, Baskaran Sundaram, MD, obtained 1,007 X-rays of patients with and without active TB. The cases consisted of multiple chest X-ray datasets from the National Institutes of Health, the Belarus Tuberculosis Portal, and Thomas Jefferson University Hospital.



AI detective analyses police data to learn how to crack cases

New Scientist

UK police are trialling a computer system that can piece together what might have happened at a crime scene. The idea is that the system, called VALCRI, will be able to do the laborious parts of a crime analyst's job in seconds, freeing them to focus on the case, while also provoking new lines of enquiry and possible narratives that may have been missed. "Everyone thinks policing is about connecting the dots, but that's the easy bit," says William Wong, who leads the project at Middlesex University London. "The hard part is working out which dots need to be connected." VALCRI's main job is to help generate plausible ideas about how, when and why a crime was committed as well as who did it.


Automated Machine Learning -- A Paradigm Shift That Accelerates Data Scientist Productivity @ Airbnb

#artificialintelligence

At Airbnb, we are always searching for ways to improve our data science workflow. A fair amount of our data science projects involve machine learning, and many parts of this workflow are repetitive. Exploratory Data Analysis: Visualizing data before embarking on a modeling exercise is a crucial step in machine learning. Automating tasks such as plotting all your variables against the target variable being predicted as well as computing summary statistics can save lots of time. Feature Transformations: There are many choices in how you can encode categorical variables, impute missing values, encode sequences and text, etc.


AI-augmented government

#artificialintelligence

For decades, artificial intelligence (AI) researchers have sought to enable computers to perform a wide range of tasks once thought to be reserved for humans. In recent years, the technology has moved from science fiction into real life: AI programs can play games, recognize faces and speech, learn, and make informed decisions. As striking as AI programs may be (and as potentially unsettling to filmgoers suffering periodic nightmares about robots becoming self-aware and malevolent), the cognitive technologies behind artificial intelligence are already having a real impact on many people's lives and work. AI-based technologies include machine learning, computer vision, speech recognition, natural language processing, and robotics;1 they are powerful, scalable, and improving at an exponential rate. Developers are working on implementing AI solutions in everything from self-driving cars to swarms of autonomous drones, from "intelligent" robots to stunningly accurate speech translation.2 And the public sector is seeking--and finding--applications to improve services; indeed, cognitive technologies could eventually revolutionize every facet of government operations. For instance, the Department of Homeland Security's Citizenship and Immigration and Services has created a virtual assistant, EMMA, that can respond accurately to human language. EMMA uses its intelligence simply, showing relevant answers to questions--almost a half-million questions per month at present. Learning from her own experiences, the virtual assistant gets smarter as she answers more questions. Customer feedback tells EMMA which answers helped, honing her grasp of the data in a process called "supervised learning."3 While EMMA is a relatively simple application, developers are thinking bigger as well: Today's cognitive technologies can track the course, speed, and destination of nearly 2,000 airliners at a time, allowing them to fly safely.4


Study finds our taste in movies is highly idiosyncratic

Daily Mail - Science & tech

Taste in movies is idiosyncratic, and not linked to the demographic traits that film studios target, a study has found. It also shows that moviegoers' ratings don't always match those of film critics. The survey of more more than 3,000 people found that the best predictor of a non-critics' response to a film was the aggregated evaluations of other non-critics, such as those on the Internet Movie Database (IMDb). The results of a study revealed that there were generally low levels of correlation in movie preferences among non-film critics - in other words, their movie tastes were highly individualistic. 'What we find enjoyable in movies is strikingly subjective - so much so that the industry's targeting of film goers by broad demographic categories seems off the mark,' says Dr Pascal Wallisch, a clinical assistant professor in New York University's Department of Psychology and the senior author of the study.