Country
Maximum Entropy Based Significance of Itemsets
We consider the problem of defining the significance of an itemset. We say that the itemset is significant if we are surprised by its frequency when compared to the frequencies of its sub-itemsets. In other words, we estimate the frequency of the itemset from the frequencies of its sub-itemsets and compute the deviation between the real value and the estimate. For the estimation we use Maximum Entropy and for measuring the deviation we use Kullback-Leibler divergence. A major advantage compared to the previous methods is that we are able to use richer models whereas the previous approaches only measure the deviation from the independence model. We show that our measure of significance goes to zero for derivable itemsets and that we can use the rank as a statistical test. Our empirical results demonstrate that for our real datasets the independence assumption is too strong but applying more flexible models leads to good results.
Compositional generalization in a deep seq2seq model by separating syntax and semantics
Russin, Jake, Jo, Jason, O'Reilly, Randall C., Bengio, Yoshua
Standard methods in deep learning for natural language processing fail to capture the compositional structure of human language that allows for systematic generalization outside of the training distribution. However, human learners readily generalize in this way, e.g. by applying known grammatical rules to novel words. Inspired by work in neuroscience suggesting separate brain systems for syntactic and semantic processing, we implement a modification to standard approaches in neural machine translation, imposing an analogous separation. The novel model, which we call Syntactic Attention, substantially outperforms standard methods in deep learning on the SCAN dataset, a compositional generalization task, without any hand-engineered features or additional supervision. Our work suggests that separating syntactic from semantic learning may be a useful heuristic for capturing compositional structure.
Gender specific and Age dependent classification model for improved diagnosis in Parkinson's disease
Gupta, Ujjwal, Bansal, Hritik, Joshi, Deepak
Accurate diagnosis is crucial for preventing the progression of Parkinson's, as well as improving the quality of life with individuals with Parkinson's disease. In this paper, we develop a gender specific and age dependent classification method to diagnose the Parkinson's disease using the handwriting based measurements. The gender specific and age dependent classifier was observed significantly outperforming the generalized classifier. An improved accuracy of 83.75% (SD=1.63) with the female specific classifier, and 79.55% (SD=1.58) with the old age dependent classifier was observed in comparison to 75.76% (SD=1.17) accuracy with the generalized classifier. Finally, combining the age and gender information proved to be encouraging in classification. We performed a rigorous analysis to observe the dominance of gender specific and age dependent features for Parkinson's detection and ranked them using the support vector machine(SVM) ranking method. Distinct set of features were observed to be dominating for higher classification accuracy in different category of classification.
Capturing human categorization of natural images at scale by combining deep networks and cognitive models
Battleday, Ruairidh M., Peterson, Joshua C., Griffiths, Thomas L.
Human categorization is one of the most important and successful targets of cognitive modeling in psychology, yet decades of development and assessment of competing models have been contingent on small sets of simple, artificial experimental stimuli. Here we extend this modeling paradigm to the domain of natural images, revealing the crucial role that stimulus representation plays in categorization and its implications for conclusions about how people form categories. Applying psychological models of categorization to natural images required two significant advances. First, we conducted the first large-scale experimental study of human categorization, involving over 500,000 human categorization judgments of 10,000 natural images from ten non-overlapping object categories. Second, we addressed the traditional bottleneck of representing high-dimensional images in cognitive models by exploring the best of current supervised and unsupervised deep and shallow machine learning methods. We find that selecting sufficiently expressive, data-driven representations is crucial to capturing human categorization, and using these representations allows simple models that represent categories with abstract prototypes to outperform the more complex memory-based exemplar accounts of categorization that have dominated in studies using less naturalistic stimuli.
Graph Neural Reasoning for 2-Quantified Boolean Formula Solvers
Yang, Zhanfu, Wang, Fei, Chen, Ziliang, Wei, Guannan, Rompf, Tiark
In this paper, we investigate the feasibility of learning GNN (Graph Neural Network) based solvers and GNN-based heuristics for specified QBF (Quantified Boolean Formula) problems. We design and evaluate several GNN architectures for 2QBF formulae, and conjecture that GNN has limitations in learning 2QBF solvers. Then we show how to learn a heuristic CEGAR 2QBF solver. We further explore generalizing GNN-based heuristics to larger unseen instances, and uncover some interesting challenges. In summary, this paper provides a comprehensive surveying view of applying GNN-embeddings to specified QBF solvers, and aims to offer guidance in applying ML to more complicated symbolic reasoning problems.
ARCHANGEL: Tamper-proofing Video Archives using Temporal Content Hashes on the Blockchain
Bui, Tu, Cooper, Daniel, Collomosse, John, Bell, Mark, Green, Alex, Sheridan, John, Higgins, Jez, Das, Arindra, Keller, Jared, Thereaux, Olivier, Brown, Alan
We present ARCHANGEL; a novel distributed ledger based system for assuring the long-term integrity of digital video archives. First, we describe a novel deep network architecture for computing compact temporal content hashes (TCHs) from audio-visual streams with durations of minutes or hours. Our TCHs are sensitive to accidental or malicious content modification (tampering) but invariant to the codec used to encode the video. This is necessary due to the curatorial requirement for archives to format shift video over time to ensure future accessibility. Second, we describe how the TCHs (and the models used to derive them) are secured via a proof-of-authority blockchain distributed across multiple independent archives. We report on the efficacy of ARCHANGEL within the context of a trial deployment in which the national government archives of the United Kingdom, Estonia and Norway participated.
Towards Data Poisoning Attack against Knowledge Graph Embedding
Zhang, Hengtong, Zheng, Tianhang, Gao, Jing, Miao, Chenglin, Su, Lu, Li, Yaliang, Ren, Kui
Knowledge graph embedding (KGE) is a technique for learning continuous embeddings for entities and relations in the knowledge graph.Due to its benefit to a variety of downstream tasks such as knowledge graph completion, question answering and recommendation, KGE has gained significant attention recently. Despite its effectiveness in a benign environment, KGE' robustness to adversarial attacks is not well-studied. Existing attack methods on graph data cannot be directly applied to attack the embeddings of knowledge graph due to its heterogeneity. To fill this gap, we propose a collection of data poisoning attack strategies, which can effectively manipulate the plausibility of arbitrary targeted facts in a knowledge graph by adding or deleting facts on the graph. The effectiveness and efficiency of the proposed attack strategies are verified by extensive evaluations on two widely-used benchmarks.
CoachAI: A Conversational Agent Assisted Health Coaching Platform
Fadhil, Ahmed, Schiavo, Gianluca, Wang, Yunlong
Poor lifestyle represents a health risk factor and is the leading cause of morbidity and chronic conditions. The impact of poor lifestyle can be significantly altered by individual behavior change. Although the current shift in healthcare towards a long-lasting modifiable behavior, however, with increasing caregiver workload and individuals' continuous needs of care, there is a need to ease caregiver's work while ensuring continuous interaction with users. This paper describes the design and validation of CoachAI, a conversational agent-assisted health coaching system to support health intervention delivery to individuals and groups. This research provides three main contributions to the preventive healthcare & healthy lifestyle promotion: (1) it presents the conversational agent to aid the caregiver; (2) it aims to decrease caregiver's workload and enhance care given to users, by handling (automating) repetitive caregiver tasks; and (3) it presents a domain-independent mobile health conversational agent for health intervention delivery. We will discuss our approach and analyze the results of a one-month validation study on physical activity, healthy diet and stress management. Introduction Adhering to a healthy lifestyle is among the most contributor to health promotion and disease prevention [27,28]. A varied diet and regular physical activity have significant benefits for individuals' overall health [27,28,6]. Similarly, mental wellness is associated with social competence and coping skills that lead to positive outcomes in adulthood and later stages of individuals life [29,9]. Although the benefit of pursuing a healthy lifestyle, several barriers exist in the process of health promotion. For instance, individuals' motivation to change, their demographics and preparedness are all factors that contribute to their intention to follow a healthy lifestyle. Several studies tackled the issue of poor lifestyle through mobile technologies. Approaches [31,32] developed mobile applications to mitigate the risk of poor diet, sedentary lifestyle and anxiety. That said, the learning curve associated with mobile apps is still an issue, especially for individuals with low digital literacy.
Factored Contextual Policy Search with Bayesian Optimization
Pinsler, Robert, Karkus, Peter, Kupcsik, Andras, Hsu, David, Lee, Wee Sun
Scarce data is a major challenge to scaling robot learning to truly complex tasks, as we need to generalize locally learned policies over different task contexts. Contextual policy search offers data-efficient learning and generalization by explicitly conditioning the policy on a parametric context space. In this paper, we further structure the contextual policy representation. We propose to factor contexts into two components: target contexts that describe the task objectives, e.g. target position for throwing a ball; and environment contexts that characterize the environment, e.g. initial position or mass of the ball. Our key observation is that experience can be directly generalized over target contexts. We show that this can be easily exploited in contextual policy search algorithms. In particular, we apply factorization to a Bayesian optimization approach to contextual policy search both in sampling-based and active learning settings. Our simulation results show faster learning and better generalization in various robotic domains. See our supplementary video: https://youtu.be/MNTbBAOufDY.
The Potential of Restarts for ProbSAT
Lorenz, Jan-Hendrik, Nickerl, Julian
This work analyses the potential of restarts for probSAT, a quite successful algorithm for k-SAT, by estimating its runtime distributions on random 3-SAT instances that are close to the phase transition. We estimate an optimal restart time from empirical data, reaching a potential speedup factor of 1.39. Calculating restart times from fitted probability distributions reduces this factor to a maximum of 1.30. A spin-off result is that the Weibull distribution approximates the runtime distribution for over 93% of the used instances well. A machine learning pipeline is presented to compute a restart time for a fixed-cutoff strategy to exploit this potential. The main components of the pipeline are a random forest for determining the distribution type and a neural network for the distribution's parameters. ProbSAT performs statistically significantly better than Luby's restart strategy and the policy without restarts when using the presented approach. The structure is particularly advantageous on hard problems.