Oceania
Gender Bias in Fake News: An Analysis
Data science research into fake news has gathered much momentum in recent years, arguably facilitated by the emergence of large public benchmark datasets. While it has been well-established within media studies that gender bias is an issue that pervades news media, there has been very little exploration into the relationship between gender bias and fake news. In this work, we provide the first empirical analysis of gender bias vis-a-vis fake news, leveraging simple and transparent lexicon-based methods over public benchmark datasets. Our analysis establishes the increased prevalance of gender bias in fake news across three facets viz., abundance, affect and proximal words. The insights from our analysis provide a strong argument that gender bias needs to be an important consideration in research into fake news.
A New cross-domain strategy based XAI models for fake news detection
A New cross-domain strategy based XAI models for fake news detection v0.1.1 ABSTRACT The Advancement in technology and rapid usage of social media has made communication easier and faster than ever before. Fake news threatens the community, democracy, egalitarianism and people's trust. Cross-domain text classification is a task of a model adopting a target domain by using the knowledge of the source domain. Natural Language Processing and Deep Learning models are used to identify misleading information. Explainability is crucial in understanding the behaviour of these complex models. In this study, we propose a four-level cross-domain strategy to study the impact of explainability on cross-domain models. The latest findings in the natural language process, the "Bidirectional Encoder Representations from Transformers" (BERT) model published by Devlin et al. (2018) google used to implement the concept of transfer learning. A fine-tune BERT model is used to perform cross-domain classification. Using this model, we conducted four experiments using datasets from different domains. Explanatory models like Anchor, ELI5, LIME and SHAP are used to design a novel explainable approach to cross-domain levels. The experimental analysis has given an ideal pair of XAI models on different levels of cross-domain. INTRODUCTION Nowadays, social media has become a potential influencing tool. According to the statistics published by Datareportal in July 2022, there is exponential growth in social media platforms, declaring that more than half of the world's population (59 per cent) is using them. Consequently, these platforms have deterministic effects on people's lives and the integrity of societies and local communities. Groups of people forming social media clusters use, unfortunately, these tools to spread speculation - so-called "fake news". In 2008, a journalist posted a report about Steve jobs medical condition. It has created massive confusion and controversy within societies and led to fluctuations in the stock price of Apple Inc. Rubin (2017). During the Covid-19 pandemic, fake news was largely spread among people and has created panic within societies. Recent statistics published by the United States support receiving reports from 80 per cent of consumers about the fake news outbreak. Insufficient data is one of the reasons behind unreliable communication, making it difficult to distinguish fake from real news. In 2016, fake news was popular mainly during the United States elections. They have created a great source of influence on people's opinions about two constants.
Inferencing the earth moving equipment-environment interaction in open pit mining
In mining, grade control generally focuses on blast hole sampling and the estimation of ore control block models with little or no attention given to how the materials are being excavated from the ground. In the process of loading trucks, the underlying variability of the individual bucket load will determine the variability of truck payload. Hence, accurate material movement demands a good knowledge of the excavation process and the buckets interaction with the environment. However, equipment frequently goes into off nominal states due to unexpected delays, disturbances or faults. The large amount of such disturbances causes information loss that reduces the statistical power and biases estimates, leading to increased uncertainty in the production. A reliable method that inferences the missing knowledge about the interaction between the machine and the environment from the available data sources, is vital to accurately model the material movement. In this study, a twostep method was implemented that performed unsupervised clustering and then predicted the missing information. The first method is DBSCAN based spatial clustering which divides the diggers and buckets positional data into connected loading segments. Clear patterns of segmented bucket dig positions were observed. The second model utilized Gaussian process regression which was trained with the clustered data and the model was then used to infer the mean locations of the test clusters. Bucket dig locations were then simulated at the inferred mean locations for different durations and compared against the known bucket dig locations. This method was tested at an open pit mine in the Pilbara of Western Australia. The results demonstrate the advantage of the proposed method in inferencing the missing information of bucket environment interactions and therefore enables miners to continuously track the material movement.
GRANDE: a neural model over directed multigraphs with application to anti-money laundering
Wu, Ruofan, Ma, Boqun, Jin, Hong, Zhao, Wenlong, Wang, Weiqiang, Zhang, Tianyi
The application of graph representation learning techniques to the area of financial risk management (FRM) has attracted significant attention recently. However, directly modeling transaction networks using graph neural models remains challenging: Firstly, transaction networks are directed multigraphs by nature, which could not be properly handled with most of the current off-the-shelf graph neural networks (GNN). Secondly, a crucial problem in FRM scenarios like anti-money laundering (AML) is to identify risky transactions and is most naturally cast into an edge classification problem with rich edge-level features, which are not fully exploited by the prevailing GNN design that follows node-centric message passing protocols. In this paper, we present a systematic investigation of design aspects of neural models over directed multigraphs and develop a novel GNN protocol that overcomes the above challenges via efficiently incorporating directional information, as well as proposing an enhancement that targets edge-related tasks using a novel message passing scheme over an extension of edge-to-node dual graph. A concrete GNN architecture called GRANDE is derived using the proposed protocol, with several further improvements and generalizations to temporal dynamic graphs. We apply the GRANDE model to both a real-world anti-money laundering task and public datasets. Experimental evaluations show the superiority of the proposed GRANDE architecture over recent state-of-the-art models on dynamic graph modeling and directed graph modeling.
Loss-Controlling Calibration for Predictive Models
Wang, Di, Shi, Junzhi, Wang, Pingping, Zhuang, Shuo, Li, Hongyue
We propose a learning framework for calibrating predictive models to make loss-controlling prediction for exchangeable data, which extends our recently proposed conformal loss-controlling prediction for more general cases. By comparison, the predictors built by the proposed loss-controlling approach are not limited to set predictors, and the loss function can be any measurable function without the monotone assumption. To control the loss values in an efficient way, we introduce transformations preserving exchangeability to prove finite-sample controlling guarantee when the test label is obtained, and then develop an approximation approach to construct predictors. The transformations can be built on any predefined function, which include using optimization algorithms for parameter searching. This approach is a natural extension of conformal loss-controlling prediction, since it can be reduced to the latter when the set predictors have the nesting property and the loss functions are monotone. Our proposed method is applied to selective regression and high-impact weather forecasting problems, which demonstrates its effectiveness for general loss-controlling prediction.
Self-supervised Multi-view Disentanglement for Expansion of Visual Collections
Jain, Nihal, Vaddamanu, Praneetha, Maheshwari, Paridhi, Vinay, Vishwa, Kulkarni, Kuldeep
Image search engines enable the retrieval of images relevant to a query image. In this work, we consider the setting where a query for similar images is derived from a collection of images. For visual search, the similarity measurements may be made along multiple axes, or views, such as style and color. We assume access to a set of feature extractors, each of which computes representations for a specific view. Our objective is to design a retrieval algorithm that effectively combines similarities computed over representations from multiple views. To this end, we propose a self-supervised learning method for extracting disentangled view-specific representations for images such that the inter-view overlap is minimized. We show how this allows us to compute the intent of a collection as a distribution over views. We show how effective retrieval can be performed by prioritizing candidate expansion images that match the intent of a query collection. Finally, we present a new querying mechanism for image search enabled by composing multiple collections and perform retrieval under this setting using the techniques presented in this paper.
Monitoring the risk of a tailings dam collapse through spectral analysis of satellite InSAR time-series data
Das, Sourav, Priyadarshana, Anuradha, Grebby, Stephen
Slope failures possess destructive power that can cause significant damage to both life and infrastructure. Monitoring slopes prone to instabilities is therefore critical in mitigating the risk posed by their failure. The purpose of slope monitoring is to detect precursory signs of stability issues, such as changes in the rate of displacement with which a slope is deforming. This information can then be used to predict the timing or probability of an imminent failure in order to provide an early warning. In this study, a more objective, statistical-learning algorithm is proposed to detect and characterise the risk of a slope failure, based on spectral analysis of serially correlated displacement time series data. The algorithm is applied to satellite-based interferometric synthetic radar (InSAR) displacement time series data to retrospectively analyse the risk of the 2019 Brumadinho tailings dam collapse in Brazil. Two potential risk milestones are identified and signs of a definitive but emergent risk (27 February 2018 to 26 August 2018) and imminent risk of collapse of the tailings dam (27 June 2018 to 24 December 2018) are detected by the algorithm. Importantly, this precursory indication of risk of failure is detected as early as at least five months prior to the dam collapse on 25 January 2019. The results of this study demonstrate that the combination of spectral methods and second order statistical properties of InSAR displacement time series data can reveal signs of a transition into an unstable deformation regime, and that this algorithm can provide sufficient early warning that could help mitigate catastrophic slope failures.
The Construction of Reality in an AI: A Review
AI constructivism as inspired by Jean Piaget, described and surveyed by Frank Guerin, and representatively implemented by Gary Drescher seeks to create algorithms and knowledge structures that enable agents to acquire, maintain, and apply a deep understanding of the environment through sensorimotor interactions. This paper aims to increase awareness of constructivist AI implementations to encourage greater progress toward enabling lifelong learning by machines. It builds on Guerin's 2008 "Learning Like a Baby: A Survey of AI approaches." After briefly recapitulating that survey, it summarizes subsequent progress by the Guerin referents, numerous works not covered by Guerin (or found in other surveys), and relevant efforts in related areas. The focus is on knowledge representations and learning algorithms that have been used in practice viewed through lenses of Piaget's schemas, adaptation processes, and staged development. The paper concludes with a preview of a simple framework for constructive AI being developed by the author that parses concepts from sensory input and stores them in a semantic memory network linked to episodic data.
The Heritage Digital Twin: a bicycle made for two. The integration of digital methodologies into cultural heritage research
Niccolucci, Franco, Markhoff, Béatrice, Theodoridou, Maria, Felicetti, Achille, Hermon, Sorin
According to the authors, such integration is like riding a bicycle made for two, also known as a tandem. This kind of vehicle requires a strong collaboration between the two riders to pedal synchronically and the one in front must be able and willing to drive the tandem towards a common destination, on which both riders agree. The structure of the bicycle should suit a diversity of users: tall and short; married couples and perfect strangers; sportspeople and lazy ones. The way it can be used must adapt to any kind of road, dirt trails and urban well-paved streets alike. Cycling metaphors aside, the convergence and integration of two different disciplines puts requirements to the method and the attitude of both and of all participants.
Towards a responsible machine learning approach to identify forced labor in fisheries
Joo, Rocío, McDonald, Gavin, Miller, Nathan, Kroodsma, David, Farthing, Courtney, Belhabib, Dyhia, Hochberg, Timothy
Many fishing vessels use forced labor, but identifying vessels that engage in this practice is challenging because few are regularly inspected. We developed a positive-unlabeled learning algorithm using vessel characteristics and movement patterns to estimate an upper bound of the number of positive cases of forced labor, with the goal of helping make accurate, responsible, and fair decisions. 89% of the reported cases of forced labor were correctly classified as positive (recall) while 98% of the vessels certified as having decent working conditions were correctly classified as negative. The recall was high for vessels from different regions using different gears, except for trawlers. We found that as much as ~28% of vessels may operate using forced labor, with the fraction much higher in squid jiggers and longlines. This model could inform risk-based port inspections as part of a broader monitoring, control, and surveillance regime to reduce forced labor. * Translated versions of the English title and abstract are available in five languages in S1 Text: Spanish, French, Simplified Chinese, Traditional Chinese, and Indonesian.