Performance Analysis
Emotion Detection From Social Media Posts
Rahman, Md Mahbubur, Shova, Shaila
Over the last few years, social media has evolved into a medium for expressing personal views, emotions, and even business and political proposals, recommendations, and advertisements. We address the topic of identifying emotions from text data obtained from social media posts like Twitter in this research. We have deployed different traditional machine learning techniques such as Support Vector Machines (SVM), Naive Bayes, Decision Trees, and Random Forest, as well as deep neural network models such as LSTM, CNN, GRU, BiLSTM, BiGRU to classify these tweets into four emotion categories (Fear, Anger, Joy, and Sadness). Furthermore, we have constructed a BiLSTM and BiGRU ensemble model. The evaluation result shows that the deep neural network models(BiGRU, to be specific) produce the most promising results compared to traditional machine learning models, with an 87.53 % accuracy rate. The ensemble model performs even better (87.66 %), albeit the difference is not significant. This result will aid in the development of a decision-making tool that visualizes emotional fluctuations.
TPE-Net: Track Point Extraction and Association Network for Rail Path Proposal Generation
Kang, Jungwon, Ghorbanalivakili, Mohammadjavad, Sohn, Gunho, Beach, David, Marin, Veronica
One essential feature of an autonomous train is minimizing collision risks with third-party objects. To estimate the risk, the control system must identify topological information of all the rail routes ahead on which the train can possibly move, especially within merging or diverging rails. This way, the train can figure out the status of potential obstacles with respect to its route and hence, make a timely decision. Numerous studies have successfully extracted all rail tracks as a whole within forward-looking images without considering element instances. Still, some image-based methods have employed hard-coded prior knowledge of railway geometry on 3D data to associate left-right rails and generate rail route instances. However, we propose a rail path extraction pipeline in which left-right rail pixels of each rail route instance are extracted and associated through a fully convolutional encoder-decoder architecture called TPE-Net. Two different regression branches for TPE-Net are proposed to regress the locations of center points of each rail route, along with their corresponding left-right pixels. Extracted rail pixels are then spatially clustered to generate topological information of all the possible train routes (ego-paths), discarding non-ego-path ones. Experimental results on a challenging, publicly released benchmark show true-positive-pixel level average precision and recall of 0.9207 and 0.8721, respectively, at about 12 frames per second. Even though our evaluation results are not higher than the SOTA, the proposed regression pipeline performs remarkably in extracting the correspondences by looking once at the image. It generates strong rail route hypotheses without reliance on camera parameters, 3D data, and geometrical constraints.
Bootstrapping Multilingual Semantic Parsers using Large Language Models
Awasthi, Abhijeet, Gupta, Nitish, Samanta, Bidisha, Dave, Shachi, Sarawagi, Sunita, Talukdar, Partha
Despite cross-lingual generalization demonstrated by pre-trained multilingual models, the translate-train paradigm of transferring English datasets across multiple languages remains to be a key mechanism for training task-specific multilingual models. However, for many low-resource languages, the availability of a reliable translation service entails significant amounts of costly human-annotated translation pairs. Further, translation services may continue to be brittle due to domain mismatch between task-specific input text and general-purpose text used for training translation models. For multilingual semantic parsing, we demonstrate the effectiveness and flexibility offered by large language models (LLMs) for translating English datasets into several languages via few-shot prompting. Through extensive comparisons on two public datasets, MTOP and MASSIVE, spanning 50 languages and several domains, we show that our method of translating data using LLMs outperforms a strong translate-train baseline on 41 out of 50 languages. We study the key design choices that enable more effective multilingual data translation via prompted LLMs.
Evaluation of Data Augmentation and Loss Functions in Semantic Image Segmentation for Drilling Tool Wear Detection
Schlager, Elke, Windisch, Andreas, Hanna, Lukas, Klünsner, Thomas, Hagendorfer, Elias Jan, Teppernegg, Tamara
Tool wear monitoring is crucial for quality control and cost reduction in manufacturing processes, of which drilling applications are one example. In this paper, we present a U-Net based semantic image segmentation pipeline, deployed on microscopy images of cutting inserts, for the purpose of wear detection. The wear area is differentiated in two different types, resulting in a multiclass classification problem. Joining the two wear types in one general wear class, on the other hand, allows the problem to be formulated as a binary classification task. Apart from the comparison of the binary and multiclass problem, also different loss functions, i. e., Cross Entropy, Focal Cross Entropy, and a loss based on the Intersection over Union (IoU), are investigated. Furthermore, models are trained on image tiles of different sizes, and augmentation techniques of varying intensities are deployed. We find, that the best performing models are binary models, trained on data with moderate augmentation and an IoU-based loss function.
The out-of-sample $R^2$: estimation and inference
Hawinkel, Stijn, Waegeman, Willem, Maere, Steven
Out-of-sample prediction is the acid test of predictive models, yet an independent test dataset is often not available for assessment of the prediction error. For this reason, out-of-sample performance is commonly estimated using data splitting algorithms such as cross-validation or the bootstrap. For quantitative outcomes, the ratio of variance explained to total variance can be summarized by the coefficient of determination or in-sample $R^2$, which is easy to interpret and to compare across different outcome variables. As opposed to the in-sample $R^2$, the out-of-sample $R^2$ has not been well defined and the variability on the out-of-sample $\hat{R}^2$ has been largely ignored. Usually only its point estimate is reported, hampering formal comparison of predictability of different outcome variables. Here we explicitly define the out-of-sample $R^2$ as a comparison of two predictive models, provide an unbiased estimator and exploit recent theoretical advances on uncertainty of data splitting estimates to provide a standard error for the $\hat{R}^2$. The performance of the estimators for the $R^2$ and its standard error are investigated in a simulation study. We demonstrate our new method by constructing confidence intervals and comparing models for prediction of quantitative $\text{Brassica napus}$ and $\text{Zea mays}$ phenotypes based on gene expression data.
CCDN: Checkerboard Corner Detection Network for Robust Camera Calibration
Chen, Ben, Xiong, Caihua, Zhang, Qi
Aiming to improve the checkerboard corner detection robustness against the images with poor quality, such as lens distortion, extreme poses, and noise, we propose a novel detection algorithm which can maintain high accuracy on inputs under multiply scenarios without any prior knowledge of the checkerboard pattern. This whole algorithm includes a checkerboard corner detection network and some post-processing techniques. The network model is a fully convolutional network with improvements of loss function and learning rate, which can deal with the images of arbitrary size and produce correspondingly-sized output with a corner score on each pixel by efficient inference and learning. Besides, in order to remove the false positives, we employ three post-processing techniques including threshold related to maximum response, non-maximum suppression, and clustering. Evaluations on two different datasets show its superior robustness, accuracy and wide applicability in quantitative comparisons with the state-of-the-art methods, like MATE, ChESS, ROCHADE and OCamCalib.
Artificial Intelligence System for Detection and Screening of Cardiac Abnormalities using Electrocardiogram Images
Zhang, Deyun, Geng, Shijia, Zhou, Yang, Xu, Weilun, Wei, Guodong, Wang, Kai, Yu, Jie, Zhu, Qiang, Li, Yongkui, Zhao, Yonghong, Chen, Xingyue, Zhang, Rui, Fu, Zhaoji, Zhou, Rongbo, E, Yanqi, Fan, Sumei, Zhao, Qinghao, Cheng, Chuandong, Peng, Nan, Zhang, Liang, Zheng, Linlin, Chu, Jianjun, Xu, Hongbin, Tan, Chen, Liu, Jian, Tao, Huayue, Liu, Tong, Chen, Kangyin, Jiang, Chenyang, Liu, Xingpeng, Hong, Shenda
The artificial intelligence (AI) system has achieved expert-level performance in electrocardiogram (ECG) signal analysis. However, in underdeveloped countries or regions where the healthcare information system is imperfect, only paper ECGs can be provided. Analysis of real-world ECG images (photos or scans of paper ECGs) remains challenging due to complex environments or interference. In this study, we present an AI system developed to detect and screen cardiac abnormalities (CAs) from real-world ECG images. The system was evaluated on a large dataset of 52,357 patients from multiple regions and populations across the world. On the detection task, the AI system obtained area under the receiver operating curve (AUC) of 0.996 (hold-out test), 0.994 (external test 1), 0.984 (external test 2), and 0.979 (external test 3), respectively. Meanwhile, the detection results of AI system showed a strong correlation with the diagnosis of cardiologists (cardiologist 1 (R=0.794, p<1e-3), cardiologist 2 (R=0.812, p<1e-3)). On the screening task, the AI system achieved AUCs of 0.894 (hold-out test) and 0.850 (external test). The screening performance of the AI system was better than that of the cardiologists (AI system (0.846) vs. cardiologist 1 (0.520) vs. cardiologist 2 (0.480)). Our study demonstrates the feasibility of an accurate, objective, easy-to-use, fast, and low-cost AI system for CA detection and screening. The system has the potential to be used by healthcare professionals, caregivers, and general users to assess CAs based on real-world ECG images.
Sequential Strategic Screening
Cohen, Lee, Sharifi-Malvajerdi, Saeed, Stangl, Kevin, Vakilian, Ali, Ziani, Juba
We initiate the study of strategic behavior in screening processes with multiple classifiers. We focus on two contrasting settings: a conjunctive setting in which an individual must satisfy all classifiers simultaneously, and a sequential setting in which an individual to succeed must satisfy classifiers one at a time. In other words, we introduce the combination of strategic classification with screening processes. We show that sequential screening pipelines exhibit new and surprising behavior where individuals can exploit the sequential ordering of the tests to zig-zag between classifiers without having to simultaneously satisfy all of them. We demonstrate an individual can obtain a positive outcome using a limited manipulation budget even when far from the intersection of the positive regions of every classifier. Finally, we consider a learner whose goal is to design a sequential screening process that is robust to such manipulations, and provide a construction for the learner that optimizes a natural objective.
Two-step counterfactual generation for OOD examples
Keshtmand, Nawid, Santos-Rodriguez, Raul, Lawry, Jonathan
However, they still make erroneous predictions when exposed to inputs from an unfamiliar distribution. This poses a significant obstacle to the deployment of ML models in safety-critical applications such as healthcare and autonomous vehicles. Consequently, for applications in these domains, two fundamental requirements for the deployment of ML models are; 1) being able to identify data that is from a different distribution from the data on which the model was trained, which is referred to as out-of-distribution (OOD) detection, outlier detection, or anomaly detection [30]; 2) being able to explain the prediction of the model [24]. There has been significant work on improving the accuracy of OOD detectors although, there has not been much work on explaining why a data point is OOD [20]. As OOD detection algorithms are increasingly used in safety-critical domains, providing explanations for high-stakes decisions has become an ethical and regulatory requirement [26]. Therefore, it is important to develop methods that provide both accurate OOD scores and also provide an explanation of why specific data points are detected as OOD. OOD detection can be considered a binary classification problem, where a data point can belong either to the in-distribution (ID) class or to the OOD class [4]. Additionally, there are different versions of the OOD detection problem, which are referred to as near-OOD and far-OOD detection [23, 29]. OOD data points that have neither non-discriminative (class-irrelevant) nor discriminative (class-relevant) features are referred to as far-OOD data and are therefore very dissimilar to the ID data.
AIROGS: Artificial Intelligence for RObust Glaucoma Screening Challenge
de Vente, Coen, Vermeer, Koenraad A., Jaccard, Nicolas, Wang, He, Sun, Hongyi, Khader, Firas, Truhn, Daniel, Aimyshev, Temirgali, Zhanibekuly, Yerkebulan, Le, Tien-Dung, Galdran, Adrian, Ballester, Miguel Ángel González, Carneiro, Gustavo, G, Devika R, S, Hrishikesh P, Puthussery, Densen, Liu, Hong, Yang, Zekang, Kondo, Satoshi, Kasai, Satoshi, Wang, Edward, Durvasula, Ashritha, Heras, Jónathan, Zapata, Miguel Ángel, Araújo, Teresa, Aresta, Guilherme, Bogunović, Hrvoje, Arikan, Mustafa, Lee, Yeong Chan, Cho, Hyun Bin, Choi, Yoon Ho, Qayyum, Abdul, Razzak, Imran, van Ginneken, Bram, Lemij, Hans G., Sánchez, Clara I.
The early detection of glaucoma is essential in preventing visual impairment. Artificial intelligence (AI) can be used to analyze color fundus photographs (CFPs) in a cost-effective manner, making glaucoma screening more accessible. While AI models for glaucoma screening from CFPs have shown promising results in laboratory settings, their performance decreases significantly in real-world scenarios due to the presence of out-of-distribution and low-quality images. To address this issue, we propose the Artificial Intelligence for Robust Glaucoma Screening (AIROGS) challenge. This challenge includes a large dataset of around 113,000 images from about 60,000 patients and 500 different screening centers, and encourages the development of algorithms that are robust to ungradable and unexpected input data. We evaluated solutions from 14 teams in this paper, and found that the best teams performed similarly to a set of 20 expert ophthalmologists and optometrists. The highest-scoring team achieved an area under the receiver operating characteristic curve of 0.99 (95% CI: 0.98-0.99) for detecting ungradable images on-the-fly. Additionally, many of the algorithms showed robust performance when tested on three other publicly available datasets. These results demonstrate the feasibility of robust AI-enabled glaucoma screening.