Oceania
CoHS-CQG: Context and History Selection for Conversational Question Generation
Do, Xuan Long, Zou, Bowei, Pan, Liangming, Chen, Nancy F., Joty, Shafiq, Aw, Ai Ti
Conversational question generation (CQG) serves as a vital task for machines to assist humans, such as interactive reading comprehension, through conversations. Compared to traditional single-turn question generation (SQG), CQG is more challenging in the sense that the generated question is required not only to be meaningful, but also to align with the occurred conversation history. While previous studies mainly focus on how to model the flow and alignment of the conversation, there has been no thorough study to date on which parts of the context and history are necessary for the model. We argue that shortening the context and history is crucial as it can help the model to optimise more on the conversational alignment property. To this end, we propose CoHS-CQG, a two-stage CQG framework, which adopts a CoHS module to shorten the context and history of the input. In particular, CoHS selects contiguous sentences and history turns according to their relevance scores by a top-p strategy. Our model achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings.
BanglaParaphrase: A High-Quality Bangla Paraphrase Dataset
Akil, Ajwad, Sultana, Najrin, Bhattacharjee, Abhik, Shahriyar, Rifat
In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the low resource status of the Bangla language in the NLP domain through the introduction of BanglaParaphrase, which ensures quality by preserving both semantics and diversity, making it particularly useful to enhance other Bangla datasets. We show a detailed comparative analysis between our dataset and models trained on it with other existing works to establish the viability of our synthetic paraphrase data generation pipeline. We are making the dataset and models publicly available at https://github.com/csebuetnlp/banglaparaphrase to further the state of Bangla NLP.
Modular Multi-Copter Structure Control for Cooperative Aerial Cargo Transportation
Chaikalis, Dimitris, Evangeliou, Nikolaos, Tzes, Anthony, Khorrami, Farshad
The control problem of a multi-copter swarm, mechanically coupled through a modular lattice structure of connecting rods, is considered in this article. The system's structural elasticity is considered in deriving the system's dynamics. The devised controller is robust against the induced flexibilities, while an inherent adaptation scheme allows for the control of asymmetrical configurations and the transportation of unknown payloads. Certain optimization metrics are introduced for solving the individual agent thrust allocation problem while achieving maximum system flight time, resulting in a platform-independent control implementation. Experimental studies are offered to illustrate the efficiency of the suggested controller under typical flight conditions, increased rod elasticities and payload transportation.
HYCEDIS: HYbrid Confidence Engine for Deep Document Intelligence System
Nguyen, Bao-Sinh, Tran, Quang-Bach, Dang, Tuan-Anh Nguyen, Nguyen, Duc, Le, Hung
In this paper, we introduce a novel neural architecture that can judge the result of extracted structured information from documents Measuring the confidence of AI models is critical for safely deploying provided by the information extracting neural networks AI in real-world industrial systems. One important application (hereafter referred to as the IE Networks). Our architecture is hybrid, of confidence measurement is information extraction from scanned consisting of two models, which are a Multi-modal Conformal documents. However, there exists no solution to provide reliable Predictor (MCP) and an Variational Cluster-oriented Anomaly Detector confidence score for current state-of-the-art deep-learning-based (VCAD). The former aims to combine the neural signals from information extractors. In this paper, we propose a complete and 3 main stages of information extraction processes including textbox novel architecture to measure confidence of current deep learning localization, OCR, and key-value recognition to predict the models in document information extraction task. Our architecture confidence level for each extracted key-value output. The later computes consists of a Multi-modal Conformal Predictor and a Variational anomaly scores for the raw input document image, providing Cluster-oriented Anomaly Detector, trained to faithfully estimate the MCP with additional features to produce better confidence estimation.
Instance-Based Uncertainty Estimation for Gradient-Boosted Regression Trees
Brophy, Jonathan, Lowd, Daniel
Gradient-boosted regression trees (GBRTs) are hugely popular for solving tabular regression problems, but provide no estimate of uncertainty. We propose Instance-Based Uncertainty estimation for Gradient-boosted regression trees (IBUG), a simple method for extending any GBRT point predictor to produce probabilistic predictions. IBUG computes a non-parametric distribution around a prediction using the $k$-nearest training instances, where distance is measured with a tree-ensemble kernel. The runtime of IBUG depends on the number of training examples at each leaf in the ensemble, and can be improved by sampling trees or training instances. Empirically, we find that IBUG achieves similar or better performance than the previous state-of-the-art across 22 benchmark regression datasets. We also find that IBUG can achieve improved probabilistic performance by using different base GBRT models, and can more flexibly model the posterior distribution of a prediction than competing methods. We also find that previous methods suffer from poor probabilistic calibration on some datasets, which can be mitigated using a scalar factor tuned on the validation data. Source code is available at https://www.github.com/jjbrophy47/ibug.
Embedding-Enhanced Giza++: Improving Alignment in Low- and High- Resource Scenarios Using Embedding Space Geometry
Marchisio, Kelly, Xiong, Conghao, Koehn, Philipp
A popular natural language processing task decades ago, word alignment has been dominated until recently by GIZA++, a statistical method based on the 30-year-old IBM models. New methods that outperform GIZA++ primarily rely on large machine translation models, massively multilingual language models, or supervision from GIZA++ alignments itself. We introduce Embedding-Enhanced GIZA++, and outperform GIZA++ without any of the aforementioned factors. Taking advantage of monolingual embedding spaces of source and target language only, we exceed GIZA++'s performance in every tested scenario for three languages pairs. In the lowest-resource setting, we outperform GIZA++ by 8.5, 10.9, and 12 AER for Ro-En, De-En, and En-Fr, respectively. We release our code at https://github.com/kellymarchisio/ee-giza.
Is artificial intelligence capable of attacking humanity? - study
Will AI be capable of overpowering humanity? Any sci-fi fan can tell you that a good apocalypse begins and ends with artificial intelligence and a god complex. But the main question is how close we really are to bringing aspects like Skynet from the Terminator franchise or Brainiac from DC Comics into the world, and if it is possible to control a being whose intelligence far exceeds that of the brightest and smartest humanity has to offer. A number of scientists, philosophers and technology experts at international educational institutions published an article in which they analyzed the danger, stating among other things that if such an entity were to arise, it would not be possible to stop it. The authors of the study, who come from academic institutions and technological bodies in the US, Australia and Madrid, published their findings in the Journal of Artificial Intelligence Research.
QAScore -- An Unsupervised Unreferenced Metric for the Question Generation Evaluation
Ji, Tianbo, Lyu, Chenyang, Jones, Gareth, Zhou, Liting, Graham, Yvette
Question Generation (QG) aims to automate the task of composing questions for a passage with a set of chosen answers found within the passage. In recent years, the introduction of neural generation models has resulted in substantial improvements of automatically generated questions in terms of quality, especially compared to traditional approaches that employ manually crafted heuristics. However, the metrics commonly applied in QG evaluations have been criticized for their low agreement with human judgement. We therefore propose a new reference-free evaluation metric that has the potential to provide a better mechanism for evaluating QG systems, called QAScore. Instead of fine-tuning a language model to maximize its correlation with human judgements, QAScore evaluates a question by computing the cross entropy according to the probability that the language model can correctly generate the masked words in the answer to that question. Furthermore, we conduct a new crowd-sourcing human evaluation experiment for the QG evaluation to investigate how QAScore and other metrics can correlate with human judgements. Experiments show that QAScore obtains a stronger correlation with the results of our proposed human evaluation method compared to existing traditional word-overlap-based metrics such as BLEU and ROUGE, as well as the existing pretrained-model-based metric BERTScore.
Modeling and Mining Multi-Aspect Graphs With Scalable Streaming Tensor Decomposition
Graphs emerge in almost every real-world application domain, ranging from online social networks all the way to health data and movie viewership patterns. Typically, such real-world graphs are big and dynamic, in the sense that they evolve over time. Furthermore, graphs usually contain multi-aspect information i.e. in a social network, we can have the "means of communication" between nodes, such as who messages whom, who calls whom, and who comments on whose timeline and so on. How can we model and mine useful patterns, such as communities of nodes in that graph, from such multi-aspect graphs? How can we identify dynamic patterns in those graphs, and how can we deal with streaming data, when the volume of data to be processed is very large? In order to answer those questions, in this thesis, we propose novel tensor-based methods for mining static and dynamic multi-aspect graphs. In general, a tensor is a higher-order generalization of a matrix that can represent high-dimensional multi-aspect data such as time-evolving networks, collaboration networks, and spatio-temporal data like Electroencephalography (EEG) brain measurements. The thesis is organized in two synergistic thrusts: First, we focus on static multi-aspect graphs, where the goal is to identify coherent communities and patterns between nodes by leveraging the tensor structure in the data. Second, as our graphs evolve dynamically, we focus on handling such streaming updates in the data without having to re-compute the decomposition, but incrementally update the existing results.
A Differentiable Distance Approximation for Fairer Image Classification
Rosa, Nicholas, Drummond, Tom, Harandi, Mehrtash
Naïvely trained AI models can be heavily biased. This can be particularly problematic when the biases involve legally or morally protected attributes such as ethnic background, age or gender. Existing solutions to this problem come at the cost of extra computation, unstable adversarial optimisation or have losses on the feature space structure that are disconnected from fairness measures and only loosely generalise to fairness. In this work we propose a differentiable approximation of the variance of demographics, a metric that can be used to measure the bias, or unfairness, in an AI model. Our approximation can be optimised alongside the regular training objective which eliminates the need for any extra models during training and directly improves the fairness of the regularised models. We demonstrate that our approach improves the fairness of AI models in varied task and dataset scenarios, whilst still maintaining a high level of classification accuracy.