Oceania
Improving Multilayer-Perceptron(MLP)-based Network Anomaly Detection with Birch Clustering on CICIDS-2017 Dataset
Yin, Yuhua, Jang-Jaccard, Julian, Sabrina, Fariza, Kwak, Jin
Machine learning algorithms have been widely used in intrusion detection systems, including Multi-layer Perceptron (MLP). In this study, we proposed a two-stage model that combines the Birch clustering algorithm and MLP classifier to improve the performance of network anomaly multi-classification. In our proposed method, we first apply Birch or Kmeans as an unsupervised clustering algorithm to the CICIDS-2017 dataset to pre-group the data. The generated pseudo-label is then added as an additional feature to the training of the MLP-based classifier. The experimental results show that using Birch and K-Means clustering for data pre-grouping can improve intrusion detection system performance. Our method can achieve 99.73% accuracy in multi-classification using Birch clustering, which is better than similar researches using a stand-alone MLP model.
Leveraging Locality in Abstractive Text Summarization
Liu, Yixin, Ni, Ansong, Nan, Linyong, Deb, Budhaditya, Zhu, Chenguang, Awadallah, Ahmed H., Radev, Dragomir
Neural attention models have achieved significant improvements on many natural language processing tasks. However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization. Instead of designing more efficient attention modules, we approach this problem by investigating if models with a restricted context can have competitive performance compared with the memory-efficient attention models that maintain a global context by treating the input as a single sequence. Our model is applied to individual pages which contain parts of inputs grouped by the principle of locality during both encoding and decoding. We empirically investigated three kinds of locality in text summarization at different levels of granularity, ranging from sentences to documents. Our experimental results show that our model has a better performance compared with strong baselines with efficient attention modules, and our analysis provides further insights into our locality-aware modeling strategy.
Explainable Predictive Decision Mining for Operational Support
Park, Gyunam, Kรผsters, Aaron, Tews, Mara, Pitsch, Cameron, Schneider, Jonathan, van der Aalst, Wil M. P.
Several decision points exist in business processes (e.g., whether a purchase order needs a manager's approval or not), and different decisions are made for different process instances based on their characteristics (e.g., a purchase order higher than e500 needs a manager approval). Decision mining in process mining aims to describe/predict the routing of a process instance at a decision point of the process. By predicting the decision, one can take proactive actions to improve the process. For instance, when a bottleneck is developing in one of the possible decisions, one can predict the decision and bypass the bottleneck. However, despite its huge potential for such operational support, existing techniques for decision mining have focused largely on describing decisions but not on predicting them, deploying decision trees to produce logical expressions to explain the decision. In this work, we aim to enhance the predictive capability of decision mining to enable proactive operational support by deploying more advanced machine learning algorithms. Our proposed approach provides explanations of the predicted decisions using SHAP values to support the elicitation of proactive actions. We have implemented a Web application to support the proposed approach and evaluated the approach using the implementation.
A review of machine learning concepts and methods for addressing challenges in probabilistic hydrological post-processing and forecasting
Papacharalampous, Georgia, Tyralis, Hristos
"Prediction" is a broad and generic term that describes any process for obtaining guesses of unseen variables based on any available information, as well as each of these guesses. On the other hand, "forecasting" is a more specific term that describes any process for issuing predictions for future variables based on information (which most commonly takes the form of time series) about the present and the past, with these particular predictions being broadly called "forecasts". Forecasting is a key theme and topic for this study. Therefore, in what follows, the general focus will be on it and not on prediction in general, although many of the statements and methods that will be referring to it are equally relevant and applicable to other prediction types. The origins of forecasting trace back to the early humans and their pronounced need for certainty in the practical endeavour of supporting their various everyday life decisions (Petropoulos et al. 2022). Thus, forecasting has met until today and still meets numerous implementations, formal and informal. Independently of their exact categorization and features, the formal implementations of forecasting rely, in principal, on concepts, theory and practice that originate from or can be attributed to the predictive branch of statistical modelling, although forecasting is also considered as an entire field on its own because of the major role that the temporal dependence plays in the formulation of its methods. The predictive branch of statistical modelling exhibits profound and fundamental differences with respect to the descriptive and explanatory ones, as it is thoroughly explained in Shmueli (2010).
DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation
Chen, Wei, Gong, Yeyun, Wang, Song, Yao, Bolun, Qi, Weizhen, Wei, Zhongyu, Hu, Xiaowu, Zhou, Bartuer, Mao, Yi, Chen, Weizhu, Cheng, Biao, Duan, Nan
Dialog response generation in open domain is an important research topic where the main challenge is to generate relevant and diverse responses. In this paper, we propose a new dialog pre-training framework called DialogVED, which introduces continuous latent variables into the enhanced encoder-decoder pre-training framework to increase the relevance and diversity of responses. With the help of a large dialog corpus (Reddit), we pre-train the model using the following 4 tasks adopted in language models (LMs) and variational autoencoders (VAEs): 1) masked language model; 2) response generation; 3) bag-of-words prediction; and 4) KL divergence reduction. We also add additional parameters to model the turn structure in dialogs to improve the performance of the pre-trained model. We conduct experiments on PersonaChat, DailyDialog, and DSTC7-AVSD benchmarks for response generation. Experimental results show that our model achieves the new state-of-the-art results on all these datasets.
Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning Framework
Chen, Yiming, Zhang, Yan, Wang, Bin, Liu, Zuozhu, Li, Haizhou
Most sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals. Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks. In this work, we propose a semi-supervised sentence embedding framework, GenSE, that effectively leverages large-scale unlabeled data. Our method include three parts: 1) Generate: A generator/discriminator model is jointly trained to synthesize sentence pairs from open-domain unlabeled corpus; 2) Discriminate: Noisy sentence pairs are filtered out by the discriminator to acquire high-quality positive and negative sentence pairs; 3) Contrast: A prompt-based contrastive approach is presented for sentence representation learning with both annotated and synthesized data. Comprehensive experiments show that GenSE achieves an average correlation score of 85.19 on the STS datasets and consistent performance improvement on four domain adaptation tasks, significantly surpassing the state-of-the-art methods and convincingly corroborating its effectiveness and generalization ability.Code, Synthetic data and Models available at https://github.com/MatthewCYM/GenSE.
Shoring up drones with artificial intelligence helps surf lifesavers spot sharks at the beach
Australian surf lifesavers are increasingly using drones to spot sharks at the beach before they get too close to swimmers. But just how reliable are they? Discerning whether that dark splodge in the water is a shark or just, say, seaweed isn't always straightforward and, in reasonable conditions, drone pilots generally make the right call only 60% of the time. While this has implications for public safety, it can also lead to unnecessary beach closures and public alarm. Engineers are trying to boost the accuracy of these shark-spotting drones with artificial intelligence (AI).
Role of Artificial Intelligence in Military Aviation - Indian Defence Review
The discovery of gunpowder in the ninth century and the invention of the atomic bomb in the twentieth century may be considered the first two revolutions in warfare. The third revolution in warfare is Artificial Intelligence (AI), the branch of computer sciences that is engaged in the development of intelligence machines i.e. those that could think and function like human beings. AI has gained enough prominence in military spheres by way of autonomous weaponry on land, sea, air, space and cyber domains to be considered as a breakthrough that militaries around the world are scampering to exploit so as to dominate, or at least gain an advantage over, potential or existing adversaries. Air power, from the days of Douhet, concerns air supremacy; that is to say, it aims at possessing the capability to use the medium of air to own advantage while denying its use to the adversary. However, concepts of air power thought have evolved remarkably since Douhet on account of technological innovations. From gladiatorial dogfights between knights of the air, the instruments of air power have progressed astoundingly with the advent of Beyond Visual Range (BVR) missiles, air-to-surface weapons launched from long distances without visually sighting the targets they are aimed at, stealth and speed enhancements and aircraft performance in terms of manoeuver ability and agility.
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Zhang, Chen, D'Haro, Luis Fernando, Zhang, Qiquan, Friedrichs, Thomas, Li, Haizhou
Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality dimensions at the dialogue level. To this end, we are motivated to propose a multi-dimensional dialogue-level metric, which consists of three sub-metrics with each targeting a specific dimension. The sub-metrics are trained with novel self-supervised objectives and exhibit strong correlations with human judgment for their respective dimensions. Moreover, we explore two approaches to combine the sub-metrics: metric ensemble and multitask learning. Both approaches yield a holistic metric that significantly outperforms individual sub-metrics. Compared to the existing state-of-the-art metric, the combined metrics achieve around 16% relative improvement on average across three high-quality dialogue-level evaluation benchmarks.
Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation
Li, Zhuang, Qu, Lizhen, Xu, Qiongkai, Wu, Tongtong, Zhan, Tianyang, Haffari, Gholamreza
In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of task-specific labeled examples. In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space. Both types of priors satisfy a novel property called $\epsilon$-disentangled. We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work. The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings. Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.