Genre
Toward Finding Malicious Cyber Discussions in Social Media
Lippman, Richard P. (MIT Lincoln Laboratory) | Weller-Fahy, David J. (MIT Lincoln Laboratory) | Mensch, Alyssa C. (MIT Lincoln Laboratory) | Campbell, William M. (MIT Lincoln Laboratory) | Campbell, Joseph P. (MIT Lincoln Laboratory) | Streilein, William W. (MIT Lincoln Laboratory) | Carter, Kevin M. (MIT Lincoln Laboratory)
Security analysts gather essential information about cyber attacks, exploits, vulnerabilities, and victims by manually searching social media sites. This effort can be dramatically reduced using natural language machine learning techniques. Using a new English text corpus containing more than 250K discussions from Stack Exchange, Reddit, and Twitter on cyber and non-cyber topics, we demonstrate the ability to detect more than 90% of the cyber discussions with fewer than 1% false alarms. If an original searched document corpus includes only 5% cyber documents, then our processing provides an enriched corpus for analysts where 83% to 95% of the documents are on cyber topics. Good performance was obtained using term frequency (TF) โ inverse document frequency (IDF) (TFโIDF) features and either logistic regression or linear support vector machine (SVM) classifiers. A classifier trained using prior historical data accurately detected 86% of emergent Heartbleed discussions and retrospective experiments demonstrate that classifier performance remains stable up to a year without retraining.
A Game-Theoretic Approach for Alert Prioritization
Laszka, Aron (Vanderbilt University) | Vorobeychik, Yevgeniy (Vanderbilt University) | Fabbri, Daniel (Vanderbilt University) | Yan, Chao (Vanderbilt University) | Malin, Bradley (Vanderbilt University)
The quantity of information that is collected and stored in computer systems continues to grow rapidly. At the same time, the sensitivity of such information (e.g., detailed medical records) often makes such information valuable to both external attackers, who may obtain information by compromising a system, and malicious insiders, who may misuse information by exercising their authorization. To mitigate compromises and deter misuse, the security administrators of these resources often deploy various types of intrusion and misuse detection systems, which provide alerts of suspicious events that are worthy of follow-up review. However, in practice, these systems may generate a large number of false alerts, wasting the time of investigators. Given that security administrators have limited budget for investigating alerts, they must prioritize certain types of alerts over others. An important challenge in alert prioritization is that adversaries may take advantage of such behavior to evade detection โ specifically by mounting attacks that trigger alerts that are less likely to be investigated. In this paper, we model alert prioritization with adaptive adversaries using a Stackelberg game and introduce an approach to compute the optimal prioritization of alert types. We evaluate our approach using both synthetic data and a real-world dataset of alerts generated from the audit logs of an electronic medical record system in use at a large academic medical center.
Towards A Multi-Tiered Knowledge-Based System for Autonomous Cloud Security Auditing
Khan, Saad Ullah (University of Huddersfield) | Parkinson, Simon (University of Huddersfield)
Every cloud platform has a large number of software components, making it difficult to manage the security of the entire system. This paper discusses the requirement for an intelligent cloud security auditing solution, and an expert system architecture is presented. The solution can identify data confidentiality threats in the OpenStack cloud platform, as well as propose solutions to remove vulnerabilities before an attack occurs. Data confidentiality threats cover a wide range of security risks where attackers usually try to steal/corrupt personal data and are a major concern of users. For this reason, cloud infrastructures need frequent security auditing. The key features of the proposed expert system architecture include: acquisition of information detailing the latest cloud security threats and solutions, the conversion of acquired raw data into usable format, the application of a forward chaining inference algorithm, and the ability for the user to add/modify knowledge, which is then utilised to provide feasible solutions in ranked order. These components provide an automated mechanism to generate human-readable audit reports, improving the overall security status without the need for expert knowledge.
The Meta-Turing Test
Walsh, Toby (University of New South Wales and Data61)
We propose an alternative to the Turing test that removes the inherent asymmetry between humans and machines in Turingโs original imitation game. In this new test, both humans and machines judge each other. We argue that this makes the test more robust against simple deceptions. We also propose a small number of refinements to improve further the test. These refinements could be applied also to Turingโs original imitation game.
Homelessness Service Provision: A Data Science Perspective
Gao, Yuan (Washington University in St. Louis) | Das, Sanmay (Washington University in St. Louis) | Fowler, Patrick (Washington University in St. Louis)
We study homeless service provision in the United States from a data science perspective, with the goal of informing homelessness prevention efforts. We use machine learning techniques to predict household reentry into a homeless system using an administrative dataset containing both demographic and service information. This data recorded all publicly funded services provided in a Midwestern US community from 2007 through 2014. We find that several techniques can provide useful lift in the prediction task, with random forests achieving an AUC around 0.7. Prediction improves significantly when conducted within calendar years, compared to across years, suggesting that changing dynamics drive repeated need for homeless services. We also analyze key service usage patterns that are associated with lower probabilities for reentry. Counterintuitively, individuals receiving the least intensive services provided through the homelessness system exhibit significantly lower likelihoods for further system involvement compared to individuals who received more intensive services, even after accounting for initial differences through propensity score and nearest neighbor matching. These result provide intriguing insights into homelessness service delivery that need to be further probed. In particular, it is unclear whether these less intensive services sustainably address housing needs, or whether, in contrast, frustration with inadequate services drives clients away from the homelessness system. Our results provide a proof-of-concept for how data science approaches can drive interesting, socially important research in the provision of public services.
Cluster-based Kriging Approximation Algorithms for Complexity Reduction
van Stein, Bas, Wang, Hao, Kowalczyk, Wojtek, Emmerich, Michael, Bรคck, Thomas
Kriging or Gaussian Process Regression is applied in many fields as a nonlinear regression model as well as a surrogate model in the field of evolutionary computation. However, the computational and space complexity of Kriging, that is cubic and quadratic in the number of data points respectively, becomes a major bottleneck with more and more data available nowadays. In this paper, we propose a general methodology for the complexity reduction, called cluster Kriging, where the whole data set is partitioned into smaller clusters and multiple Kriging models are built on top of them. In addition, four Kriging approximation algorithms are proposed as candidate algorithms within the new framework. Each of these algorithms can be applied to much larger data sets while maintaining the advantages and power of Kriging. The proposed algorithms are explained in detail and compared empirically against a broad set of existing state-of-the-art Kriging approximation methods on a well-defined testing framework. According to the empirical study, the proposed algorithms consistently outperform the existing algorithms. Moreover, some practical suggestions are provided for using the proposed algorithms. Kriging, or Gaussian Process Regression [1] is a popular and elegant kernel based regression model capable of modeling very complex functions. Kriging is used in many fields e.g. Many other regression models exist, such as parametric models, which are easy to interpret but may lack expressive power to model complex functions.
Network-based methods for outcome prediction in the "sample space"
In this thesis we present the novel semi-supervised network-based algorithm P-Net, which is able to rank and classify patients with respect to a specific phenotype or clinical outcome under study. The peculiar and innovative characteristic of this method is that it builds a network of samples/patients, where the nodes represent the samples and the edges are functional or genetic relationships between individuals (e.g. similarity of expression profiles), to predict the phenotype under study. In other words, it constructs the network in the "sample space" and not in the "biomarker space" (where nodes represent biomolecules (e.g. genes, proteins) and edges represent functional or genetic relationships between nodes), as usual in state-of-the-art methods. To assess the performances of P-Net, we apply it on three different publicly available datasets from patients afflicted with a specific type of tumor: pancreatic cancer, melanoma and ovarian cancer dataset, by using the data and following the experimental set-up proposed in two recently published papers [Barter et al., 2014, Winter et al., 2012]. We show that network-based methods in the "sample space" can achieve results competitive with classical supervised inductive systems. Moreover, the graph representation of the samples can be easily visualized through networks and can be used to gain visual clues about the relationships between samples, taking into account the phenotype associated or predicted for each sample. To our knowledge this is one of the first works that proposes graph-based algorithms working in the "sample space" of the biomolecular profiles of the patients to predict their phenotype or outcome, thus contributing to a novel research line in the framework of the Network Medicine.
Google's Diane Greene: AI will cost jobs, so skills training is critical - SiliconANGLE
Machine learning will cost us jobs, a prominent technology executive acknowledged today, but she said job disruption isn't the insurmountable problem that many observers fear. Diane Greene, senior vice president in charge of Google Inc.'s cloud business, said at the Women in Data Science conference at Stanford University today that there's "no question" that machine learning, a branch of artificial intelligence that uses data to help computers learn rather than explicitly programming them, is replacing jobs. SiliconANGLE Media's mobile live video studio, theCUBE, is doing live interviews at the conference. Already, Greene said, "machines are better than humans" at some tasks. Recently they've started to do better at some kinds of image and speech recognition, and they're performing tasks such as finding signs of disease in photos better than humans.
'Bat Bot' Flying Robot Mimics 'Ridiculously Stupid' Complexity Of Bat Flight
One of the problems with bats, if you're a robotics expert, is that they have so many joints. That's what robotics researchers at the University of Illinois Urbana-Champaign and Caltech quickly learned when they set out to build a robot version of the flying mammal. "Bats use more than 40 active and passive joints, [along with] the flexible membranes of their wings," Soon-Jo Chung of Caltech told Popular Mechanics. "It's impractical, or impossible, to incorporate [all 40] of these joints in the robot's design." Or as biologist Dan Riskin of the University of Toronto put it to PBS, "bats are ridiculously stupid in terms of how complex they are."
Python Machine Learning: Scikit-Learn Tutorial
Machine learning is a branch in computer science that studies the design of algorithms that can learn. Typical tasks are concept learning, function learning or "predictive modeling", clustering and finding predictive patterns. These tasks are learned through available data that were observed through experiences or instructions, for example. The hope that comes with this discipline is that including the experience into its tasks will eventually improve the learning. But this improvement needs to happen in such a way that the learning itself becomes automatic so that humans like ourselves don't need to interfere anymore is the ultimate goal. There are close ties between this discipline and Knowledge Discovery, Data Mining, Artificial Intelligence (AI) and Statistics. Typical applications can be classified into scientific knowledge discovery and more commercial ones, ranging from the "Robot Scientist" to anti-spam filtering and recommender systems. But above all, you will know this discipline because it's one of the topics that you need to master if you want to excel in data science. Today's scikit-learn tutorial will introduce you to the basics of Python machine learning: step-by-step, it will show you how to use Python and its libraries to explore your data with the help of matplotlib, work with the well-known algorithms KMeans and Support Vector Machines (SVM) to construct models, to fit the data to these models, to predict values and to validate the models that you have build. The first step to about anything in data science is loading in your data. This is also the starting point of this scikit-learn tutorial.