Oceania
Transformer-Based Language Models for Software Vulnerability Detection
Thapa, Chandra, Jang, Seung Ick, Ahmed, Muhammad Ejaz, Camtepe, Seyit, Pieprzyk, Josef, Nepal, Surya
The large transformer-based language models demonstrate excellent performance in natural language processing. By considering the transferability of the knowledge gained by these models in one domain to other related domains, and the closeness of natural languages to high-level programming languages, such as C/C++, this work studies how to leverage (large) transformer-based language models in detecting software vulnerabilities and how good are these models for vulnerability detection tasks. In this regard, firstly, a systematic (cohesive) framework that details source code translation, model preparation, and inference is presented. Then, an empirical analysis is performed with software vulnerability datasets with C/C++ source codes having multiple vulnerabilities corresponding to the library function call, pointer usage, array usage, and arithmetic expression. Our empirical results demonstrate the good performance of the language models in vulnerability detection. Moreover, these language models have better performance metrics, such as F1-score, than the contemporary models, namely bidirectional long short-term memory and bidirectional gated recurrent unit. Experimenting with the language models is always challenging due to the requirement of computing resources, platforms, libraries, and dependencies. Thus, this paper also analyses the popular platforms to efficiently fine-tune these models and present recommendations while choosing the platforms.
Opportunities for Data Science Innovation in the Policing Sector
According to Peter K. Manning, in Anglo-American societies, the purpose of the police is to "sustain politically defined order and ordering via tracking, surveillance, coercion and arrest" (2014: p.6). Consisting of several authoritatively coordinated and legitimate organizations (ibid.), the policing sector serves governments in protecting their communities, preventing crime and disorder, and ensuring justice (The Policy Circle, 2022). The police's position as acting in the communities' interest suggests that their functions are heavily dependent on public trust and societal consensus concerning social justice and fairness (Manning, 2014). While there are large numbers of police officers employed in Australia (67,200 in 2021), a number which is expected to increase in the future (Australian Industry and Skills Committee, 2022), Ransley & Mazerolle (2009) have argued that trends in public governance and regulation have caused the increased pluralization and privatisation of policing efforts. Nowadays, the policing sector thus constitutes a large network of private, public and welfare organizations geared at controlling and preventing crimes (ibid.). In this essay, I will thus focus on data science opportunities for a variety of stakeholders involved in ensuring public security and order.
How 'Lord of the Rings' Used AI to Change Big-Screen Battles Forever
An invading force, 10,000 strong, marches through a storm toward a fortress built into the side of a mountain. From a distance, the combatants look like ants -- menacing and alarmingly well organized. They rattle their spears and snarl through teeth that have never known modern dentistry, and when lightning strikes, it reveals their sheer numbers. Volleys of arrows fly, swords find their way to the weak spots around breast plates. Bodies on both sides hit the ground. This bloody affair is the Battle of Helm's Deep, from The Lord of the Rings: The Two Towers.
The super-rich 'preppers' planning to save themselves from the apocalypse
As a humanist who writes about the impact of digital technology on our lives, I am often mistaken for a futurist. The people most interested in hiring me for my opinions about technology are usually less concerned with building tools that help people live better lives in the present than they are in identifying the Next Big Thing through which to dominate them in the future. I don't usually respond to their inquiries. Why help these guys ruin what's left of the internet, much less civilisation? Still, sometimes a combination of morbid curiosity and cold hard cash is enough to get me on a stage in front of the tech elite, where I try to talk some sense into them about how their businesses are affecting our lives out here in the real world. That's how I found myself accepting an invitation to address a group mysteriously described as "ultra-wealthy stakeholders", out in the middle of the desert. A limo was waiting for me at the airport.
Flinders University Is Testing a Driverless Shuttle Bus On Campus
An autonomous shuttle bus is currently being tested at Flinders University and has now entered the second stage of its trial. Dubbed the "Flinders University Express Shuttle" (FLEX), the bus can carry 11 seated passengers. It operates on a 2.8km route and is described as a "test bed" for the future of autonomous vehicles in South Australia. In what continues to be one of Australia's only public autonomous vehicle testing programs, the Flinders University autonomous shuttle bus travels around the Tonsely innovation district, between the train station, the residential village, the university and the TAFE. It's a walking distance route, but keep in mind that it's only a trial at the moment.
PhishClone: Measuring the Efficacy of Cloning Evasion Attacks
Wong, Arthur, Abuadbba, Alsharif, Almashor, Mahathir, Kanhere, Salil
Web-based phishing accounts for over 90% of data breaches, and most web-browsers and security vendors rely on machine-learning (ML) models as mitigation. Despite this, links posted regularly on anti-phishing aggregators such as PhishTank and VirusTotal are shown to easily bypass existing detectors. Prior art suggests that automated website cloning, with light mutations, is gaining traction with attackers. This has limited exposure in current literature and leads to sub-optimal ML-based countermeasures. The work herein conducts the first empirical study that compiles and evaluates a variety of state-of-the-art cloning techniques in wide circulation. We collected 13,394 samples and found 8,566 confirmed phishing pages targeting 4 popular websites using 7 distinct cloning mechanisms. These samples were replicated with malicious code removed within a controlled platform fortified with precautions that prevent accidental access. We then reported our sites to VirusTotal and other platforms, with regular polling of results for 7 days, to ascertain the efficacy of each cloning technique. Results show that no security vendor detected our clones, proving the urgent need for more effective detectors. Finally, we posit 4 recommendations to aid web developers and ML-based defences to alleviate the risks of cloning attacks.
Towards Understanding the Overfitting Phenomenon of Deep Click-Through Rate Prediction Models
Zhang, Zhao-Yu, Sheng, Xiang-Rong, Zhang, Yujing, Jiang, Biye, Han, Shuguang, Deng, Hongbo, Zheng, Bo
Deep learning techniques have been applied widely in industrial recommendation systems. However, far less attention has been paid to the overfitting problem of models in recommendation systems, which, on the contrary, is recognized as a critical issue for deep neural networks. In the context of Click-Through Rate (CTR) prediction, we observe an interesting one-epoch overfitting problem: the model performance exhibits a dramatic degradation at the beginning of the second epoch. Such a phenomenon has been witnessed widely in real-world applications of CTR models. Thereby, the best performance is usually achieved by training with only one epoch. To understand the underlying factors behind the one-epoch phenomenon, we conduct extensive experiments on the production data set collected from the display advertising system of Alibaba. The results show that the model structure, the optimization algorithm with a fast convergence rate, and the feature sparsity are closely related to the one-epoch phenomenon. We also provide a likely hypothesis for explaining such a phenomenon and conduct a set of proof-of-concept experiments. We hope this work can shed light on future research on training more epochs for better performance.
Cross-Network Social User Embedding with Hybrid Differential Privacy Guarantees
Ren, Jiaqian, Jiang, Lei, Peng, Hao, Lyu, Lingjuan, Liu, Zhiwei, Chen, Chaochao, Wu, Jia, Bai, Xu, Yu, Philip S.
Integrating multiple online social networks (OSNs) has important implications for many downstream social mining tasks, such as user preference modelling, recommendation, and link prediction. However, it is unfortunately accompanied by growing privacy concerns about leaking sensitive user information. How to fully utilize the data from different online social networks while preserving user privacy remains largely unsolved. To this end, we propose a Cross-network Social User Embedding framework, namely DP-CroSUE, to learn the comprehensive representations of users in a privacy-preserving way. We jointly consider information from partially aligned social networks with differential privacy guarantees. In particular, for each heterogeneous social network, we first introduce a hybrid differential privacy notion to capture the variation of privacy expectations for heterogeneous data types. Next, to find user linkages across social networks, we make unsupervised user embedding-based alignment in which the user embeddings are achieved by the heterogeneous network embedding technology. To further enhance user embeddings, a novel cross-network GCN embedding model is designed to transfer knowledge across networks through those aligned users. Extensive experiments on three real-world datasets demonstrate that our approach makes a significant improvement on user interest prediction tasks as well as defending user attribute inference attacks from embedding.
Autonomous Cross Domain Adaptation under Extreme Label Scarcity
Weng, Weiwei, Pratama, Mahardhika, Za'in, Choiru, De Carvalho, Marcus, Appan, Rakaraddi, Ashfahani, Andri, Yee, Edward Yapp Kien
A cross domain multistream classification is a challenging problem calling for fast domain adaptations to handle different but related streams in never-ending and rapidly changing environments. Notwithstanding that existing multistream classifiers assume no labelled samples in the target stream, they still incur expensive labelling cost since they require fully labelled samples of the source stream. This paper aims to attack the problem of extreme label shortage in the cross domain multistream classification problems where only very few labelled samples of the source stream are provided before process runs. Our solution, namely Learning Streaming Process from Partial Ground Truth (LEOPARD), is built upon a flexible deep clustering network where its hidden nodes, layers and clusters are added and removed dynamically in respect to varying data distributions. A deep clustering strategy is underpinned by a simultaneous feature learning and clustering technique leading to clustering-friendly latent spaces. A domain adaptation strategy relies on the adversarial domain adaptation technique where a feature extractor is trained to fool a domain classifier classifying source and target streams. Our numerical study demonstrates the efficacy of LEOPARD where it delivers improved performances compared to prominent algorithms in 15 of 24 cases. Source codes of LEOPARD are shared in \url{https://github.com/wengweng001/LEOPARD.git} to enable further study.
Reinforced Continual Learning for Graphs
Rakaraddi, Appan, Lam, Siew Kei, Pratama, Mahardhika, De Carvalho, Marcus
Graph Neural Networks (GNNs) have become the backbone for a myriad of tasks pertaining to graphs and similar topological data structures. While many works have been established in domains related to node and graph classification/regression tasks, they mostly deal with a single task. Continual learning on graphs is largely unexplored and existing graph continual learning approaches are limited to the task-incremental learning scenarios. This paper proposes a graph continual learning strategy that combines the architecture-based and memory-based approaches. The structural learning strategy is driven by reinforcement learning, where a controller network is trained in such a way to determine an optimal number of nodes to be added/pruned from the base network when new tasks are observed, thus assuring sufficient network capacities. The parameter learning strategy is underpinned by the concept of Dark Experience replay method to cope with the catastrophic forgetting problem. Our approach is numerically validated with several graph continual learning benchmark problems in both task-incremental learning and class-incremental learning settings. Compared to recently published works, our approach demonstrates improved performance in both the settings. The implementation code can be found at \url{https://github.com/codexhammer/gcl}.