Oceania
Differential Privacy in Natural Language Processing: The Story So Far
Klymenko, Oleksandra, Meisenbacher, Stephen, Matthes, Florian
In an age where a vast amount of data is being Alas, in the field of NLP, where the core unit of produced daily, the opportunities created by this data is unstructured, fuzzy text rather than a structured proliferation increase concurrently. The availability data point, an initial attempt to apply Differential of big data enables countless downstream tasks Privacy poses some challenges. Chief among whose accuracy and utility seem to increase with these is the challenge of how to transfer the core the amount of data used. Specifically, the fields concepts of Differential Privacy, namely the "individual" of Machine Learning (ML) and Deep Learning and adjacency, to the textual domain where (DL) have profited from such data. Particularly in these concepts are not easily perceivable. Thus, it the case of Natural Language Processing (NLP), becomes the goal to find new ways of reasoning the tasks at hand more often than not concern the about Differential Privacy in order to adapt it to handling of unstructured data, meaning data that the unstructured data domain of NLP. Through the is not neatly organized into a traditional row-like course of this paper, the foundations of Differential database structure, and furthermore, data that is Privacy in the lens of NLP will be investigated, motivated not necessarily static. In fact, it is estimated that by some privacy vulnerabilities that surface data on the order of zettabytes (ZB) is being produced from NLP techniques. Afterwards, the limitations every day (Begum and Nausheen, 2018), and and open questions of Differential Privacy with within this amount, roughly 80% is unstructured, NLP will be analyzed with an in-depth discussion.
BERTifying Sinhala -- A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification
Dhananjaya, Vinura, Demotte, Piyumal, Ranathunga, Surangika, Jayasena, Sanath
This research provides the first comprehensive analysis of the performance of pre-trained language models for Sinhala text classification. We test on a set of different Sinhala text classification tasks and our analysis shows that out of the pre-trained multilingual models that include Sinhala (XLM-R, LaBSE, and LASER), XLM-R is the best model by far for Sinhala text classification. We also pre-train two RoBERTa-based monolingual Sinhala models, which are far superior to the existing pre-trained language models for Sinhala. We show that when fine-tuned, these pre-trained language models set a very strong baseline for Sinhala text classification and are robust in situations where labeled data is insufficient for fine-tuning. We further provide a set of recommendations for using pre-trained models for Sinhala text classification. We also introduce new annotated datasets useful for future research in Sinhala text classification and publicly release our pre-trained models.
'Ask all the time: why do I need this?' How to stop your vacuum from spying on you
This month, Amazon inked a deal to acquire smart vacuum company iRobot โ the makers of Roomba โ for a tidy US$1.7bn. As some see it, if the purchase goes through, that should worry us. "It's all about the data," says David Vaile from the Australian Privacy Foundation. Privacy advocates such as Vaile are concerned the robot vacuum cleaner will give Amazon access to floor plans of users' homes, using mapping features some iRobot products already offer. Amazon are yet to release details about what existing and future iRobot data will be used for; and the company told Reuters that they safeguard customer privacy and do not sell their data.
A Sequence Tagging based Framework for Few-Shot Relation Extraction
Relation Extraction (RE) refers to extracting the relation triples in the input text. Existing neural work based systems for RE rely heavily on manually labeled training data, but there are still a lot of domains where sufficient labeled data does not exist. Inspired by the distance-based few-shot named entity recognition methods, we put forward the definition of the few-shot RE task based on the sequence tagging joint extraction approaches, and propose a few-shot RE framework for the task. Besides, we apply two actual sequence tagging models to our framework (called Few-shot TPLinker and Few-shot BiTT), and achieves solid results on two few-shot RE tasks constructed from a public dataset.
Fast Heterogeneous Federated Learning with Hybrid Client Selection
Shen, Guangyuan, Gao, Dehong, Song, Duanxiao, Yang, Libin, Zhou, Xukai, Pan, Shirui, Lou, Wei, Zhou, Fang
Client selection schemes are widely adopted to handle the communication-efficient problems in recent studies of Federated Learning (FL). However, the large variance of the model updates aggregated from the randomly-selected unrepresentative subsets directly slows the FL convergence. We present a novel clustering-based client selection scheme to accelerate the FL convergence by variance reduction. Simple yet effective schemes are designed to improve the clustering effect and control the effect fluctuation, therefore, generating the client subset with certain representativeness of sampling. Theoretically, we demonstrate the improvement of the proposed scheme in variance reduction. We also present the tighter convergence guarantee of the proposed method thanks to the variance reduction. Experimental results confirm the exceed efficiency of our scheme compared to alternatives.
Deception for Cyber Defence: Challenges and Opportunities
Liebowitz, David, Nepal, Surya, Moore, Kristen, Christopher, Cody J., Kanhere, Salil S., Nguyen, David, Timmer, Roelien C., Longland, Michael, Rathakumar, Keerth
Deception is rapidly growing as an important tool for cyber defence, complementing existing perimeter security measures to rapidly detect breaches and data theft. One of the factors limiting the use of deception has been the cost of generating realistic artefacts by hand. Recent advances in Machine Learning have, however, created opportunities for scalable, automated generation of realistic deceptions. This vision paper describes the opportunities and challenges involved in developing models to mimic many common elements of the IT stack for deception effects.
Online Target Localization using Adaptive Belief Propagation in the HMM Framework
This paper proposes a novel adaptive sample space-based Viterbi algorithm for target localization in an online manner. The method relies on discretizing the target's motion space into cells representing a finite number of hidden states. Then, the most probable trajectory of the tracked target is computed via dynamic programming in a Hidden Markov Model (HMM) framework. The proposed method uses a Bayesian estimation framework which is neither limited to Gaussian noise models nor requires a linearized target motion model or sensor measurement models. However, an HMM-based approach to localization can suffer from poor computational complexity in scenarios where the number of hidden states increases due to high-resolution modeling or target localization in a large space. To improve this poor computational complexity, this paper proposes a belief propagation in the most probable belief space with a low to high-resolution sequentially, reducing the required resources significantly. The proposed method is inspired by the k-d Tree algorithm (e.g., quadtree) commonly used in the computer vision field. Experimental tests using an ultra-wideband (UWB) sensor network demonstrate our results.
WiFi Based Distance Estimation Using Supervised Machine Learning
Kostas, Kahraman, Kostas, Rabia Yasa, Zampella, Francisco, Alsehly, Firas
In recent years WiFi became the primary source of information to locate a person or device indoor. Collecting RSSI values as reference measurements with known positions, known as WiFi fingerprinting, is commonly used in various positioning methods and algorithms that appear in literature. However, measuring the spatial distance between given set of WiFi fingerprints is heavily affected by the selection of the signal distance function used to model signal space as geospatial distance. In this study, the authors proposed utilization of machine learning to improve the estimation of geospatial distance between fingerprints. This research examined data collected from 13 different open datasets to provide a broad representation aiming for general model that can be used in any indoor environment. The proposed novel approach extracted data features by examining a set of commonly used signal distance metrics via feature selection process that includes feature analysis and genetic algorithm. To demonstrate that the output of this research is venue independent, all models were tested on datasets previously excluded during the training and validation phase. Finally, various machine learning algorithms were compared using wide variety of evaluation metrics including ability to scale out the test bed to real world unsolicited datasets.
DSR: Direct Simultaneous Registration for Multiple 3D Images
Mao, Zhehua, Zhao, Liang, Huang, Shoudong, Fan, Yiting, Lee, Alex Pui-Wai
This paper presents a novel algorithm named Direct Simultaneous Registration (DSR) that registers a collection of 3D images in a simultaneous fashion without specifying any reference image, feature extraction and matching, or information loss or reuse. The algorithm optimizes the global poses of local image frames by maximizing the similarity between a predefined panoramic image and local images. Although we formulate the problem as a Direct Bundle Adjustment (DBA) that jointly optimizes the poses of local frames and the intensities of the panoramic image, by investigating the independence of pose estimation from the panoramic image in the solving process, DSR is proposed to solve the poses only and proved to be able to obtain the same optimal poses as DBA. The proposed method is particularly suitable for the scenarios where distinct features are not available, such as Transesophageal Echocardiography (TEE) images. DSR is evaluated by comparing it with four widely used methods via simulated and in-vivo 3D TEE images. It is shown that the proposed method outperforms these four methods in terms of accuracy and requires much fewer computational resources than the state-of-the-art accumulated pairwise estimates (APE). Codes of DSR are available at https://github.com/ZH-Mao/DSR.
BenchPress: A Deep Active Benchmark Generator
Tsimpourlas, Foivos, Petoumenos, Pavlos, Xu, Min, Cummins, Chris, Hazelwood, Kim, Rajan, Ajitha, Leather, Hugh
We develop BenchPress, the first ML benchmark generator for compilers that is steerable within feature space representations of source code. BenchPress synthesizes compiling functions by adding new code in any part of an empty or existing sequence by jointly observing its left and right context, achieving excellent compilation rate. BenchPress steers benchmark generation towards desired target features that has been impossible for state of the art synthesizers (or indeed humans) to reach. It performs better in targeting the features of Rodinia benchmarks in 3 different feature spaces compared with (a) CLgen - a state of the art ML synthesizer, (b) CLSmith fuzzer, (c) SRCIROR mutator or even (d) human-written code from GitHub. BenchPress is the first generator to search the feature space with active learning in order to generate benchmarks that will improve a downstream task. We show how using BenchPress, Grewe's et al. CPU vs GPU heuristic model can obtain a higher speedup when trained on BenchPress's benchmarks compared to other techniques. BenchPress is a powerful code generator: Its generated samples compile at a rate of 86%, compared to CLgen's 2.33%. Starting from an empty fixed input, BenchPress produces 10x more unique, compiling OpenCL benchmarks than CLgen, which are significantly larger and more feature diverse.