Statistical Learning
Analysis of child development facts and myths using text mining techniques and classification models
Tajrian, Mehedi, Rahman, Azizur, Kabir, Muhammad Ashad, Islam, Md Rafiqul
The rapid dissemination of misinformation on the internet complicates the decision-making process for individuals seeking reliable information, particularly parents researching child development topics. This misinformation can lead to adverse consequences, such as inappropriate treatment of children based on myths. While previous research has utilized text-mining techniques to predict child abuse cases, there has been a gap in the analysis of child development myths and facts. This study addresses this gap by applying text mining techniques and classification models to distinguish between myths and facts about child development, leveraging newly gathered data from publicly available websites. The research methodology involved several stages. First, text mining techniques were employed to pre-process the data, ensuring enhanced accuracy. Subsequently, the structured data was analysed using six robust Machine Learning (ML) classifiers and one Deep Learning (DL) model, with two feature extraction techniques applied to assess their performance across three different training-testing splits. To ensure the reliability of the results, cross-validation was performed using both k-fold and leave-one-out methods. Among the classification models tested, Logistic Regression (LR) demonstrated the highest accuracy, achieving a 90% accuracy with the Bag-of-Words (BoW) feature extraction technique. LR stands out for its exceptional speed and efficiency, maintaining low testing time per statement (0.97 microseconds). These findings suggest that LR, when combined with BoW, is effective in accurately classifying child development information, thus providing a valuable tool for combating misinformation and assisting parents in making informed decisions.
Amortized Bayesian Multilevel Models
Habermann, Daniel, Schmitt, Marvin, Kühmichel, Lars, Bulling, Andreas, Radev, Stefan T., Bürkner, Paul-Christian
Obtaining accurate inference and faithful uncertainty quantification in reasonable time is a frontier of today's statistical research (Cranmer et al., 2020). One major difficulty arising in most experimental and almost all observational data is the presence of complex dependency structures, for example, due to natural groupings (e.g., data gathered in different countries) or repeated measurements of the same observational units over time (e.g., particles, bacteria, or people; Gelman and Hill, 2006). To leverage these dependency structures, multilevel models (MLMs), also referred to as latent variable, hierarchical, random, or mixed effects models, have become an integral part of modern Bayesian statistics (Goldstein, 2011; Gelman et al., 2013; McGlothlin and Viele, 2018; Finch et al., 2019; Yao et al., 2022). Despite the wide success of Bayesian MLMs across the quantitative sciences, a major challenge is their limited efficiency and scalability when dealing with large and complex data. This is because estimating the full posterior distribution of all parameters of interest can be very costly (Gelman et al., 2013).
Recent advances in Meta-model of Optimal Prognosis
In real case applications within the virtual prototyping process, it is not always possible to reduce the complexity of the physical models and to obtain numerical models which can be solved quickly. Usually, every single numerical simulation takes hours or even days. Although the progresses in numerical methods and high performance computing, in such cases, it is not possible to explore various model configurations, hence efficient surrogate models are required. Generally the available meta-model techniques show several advantages and disadvantages depending on the investigated problem. In this paper we present an automatic approach for the selection of the optimal suitable meta-model for the actual problem. Together with an automatic reduction of the variable space using advanced filter techniques an efficient approximation is enabled also for high dimensional problems.
Predicting Affective States from Screen Text Sentiment
Teng, Songyan, Zhang, Tianyi, D'Alfonso, Simon, Kostakos, Vassilis
The proliferation of mobile sensing technologies has enabled the Mobile sensing technologies have been widely used in wellbeing study of various physiological and behavioural phenomena through studies and applications, and the significant advancements in sensing unobtrusive data collection from smartphone sensors. This approach over the last decade have spurred heightened interest in this offers real-time insights into individuals' physical and mental field, often referred to as "digital phenotyping". This approach often states, creating opportunities for personalised treatment and involves the use of smartphone sensors to continuously and unobtrusively interventions. However, the potential of analysing the textual content collect data on various physiological and behavioural viewed on smartphones to predict affective states remains phenomena [9]. Data from a range of smartphone sensors can be underexplored. To better understand how the screen text that users integrated to obtain a comprehensive understanding of a person's are exposed to and interact with can influence their affects, we surroundings, activities, and behaviours [1]. This approach allows investigated a subset of data obtained from a digital phenotyping for real-time monitoring and analysis of individuals' physical and study of Australian university students conducted in 2023. We employed mental states, providing valuable insights into their overall wellbeing linear regression, zero-shot, and multi-shot prompting using and creating opportunities for delivering recommendations and a large language model (LLM) to analyse relationships between interventions based on the user's context.
The News Comment Gap and Algorithmic Agenda Setting in Online Forums
Böwing, Flora, Gildersleve, Patrick
The disparity between news stories valued by journalists and those preferred by readers, known as the "News Gap", is well-documented. However, the difference in expectations regarding news related user-generated content is less studied. Comment sections, hosted by news websites, are popular venues for reader engagement, yet still subject to editorial decisions. It is thus important to understand journalist vs reader comment preferences and how these are served by various comment ranking algorithms that represent discussions differently. We analyse 1.2 million comments from Austrian newspaper Der Standard to understand the "News Comment Gap" and the effects of different ranking algorithms. We find that journalists prefer positive, timely, complex, direct responses, while readers favour comments similar to article content from elite authors. We introduce the versatile Feature-Oriented Ranking Utility Metric (FORUM) to assess the impact of different ranking algorithms and find dramatic differences in how they prioritise the display of comments by sentiment, topical relevance, lexical diversity, and readability. Journalists can exert substantial influence over the discourse through both curatorial and algorithmic means. Understanding these choices' implications is vital in fostering engaging and civil discussions while aligning with journalistic objectives, especially given the increasing legal scrutiny and societal importance of online discourse.
COVID-19 Probability Prediction Using Machine Learning: An Infectious Approach
Ilani, Mohsen Asghari, Tehran, Saba Moftakhar, Kavei, Ashkan, Radmehr, Arian
The ongoing COVID-19 pandemic continues to pose significant challenges to global public health, despite the widespread availability of vaccines. Early detection of the disease remains paramount in curbing its transmission and mitigating its impact on public health systems. In response, this study delves into the application of advanced machine learning (ML) techniques for predicting COVID-19 infection probability. We conducted a rigorous investigation into the efficacy of various ML models, including XGBoost, LGBM, AdaBoost, Logistic Regression, Decision Tree, RandomForest, CatBoost, KNN, and Deep Neural Networks (DNN). Leveraging a dataset comprising 4000 samples, with 3200 allocated for training and 800 for testing, our experiment offers comprehensive insights into the performance of these models in COVID-19 prediction. Our findings reveal that Deep Neural Networks (DNN) emerge as the top-performing model, exhibiting superior accuracy and recall metrics. With an impressive accuracy rate of 89%, DNN demonstrates remarkable potential in early COVID-19 detection. This underscores the efficacy of deep learning approaches in leveraging complex data patterns to identify COVID-19 infections accurately. This study underscores the critical role of machine learning, particularly deep learning methodologies, in augmenting early detection efforts amidst the ongoing pandemic. The success of DNN in accurately predicting COVID-19 infection probability highlights the importance of continued research and development in leveraging advanced technologies to combat infectious diseases. I. INTRODUCTION Health represents the cornerstone of any society, yet the world currently grapples with a profound health crisis due to the widespread dissemination of the coronavirus. The global COVID-19 pandemic has inflicted extensive loss of life and has profoundly impacted individuals, both directly and indirectly.
A Comparison of Deep Learning and Established Methods for Calf Behaviour Monitoring
Dissanayake, Oshana, Riaboff, Lucile, McPherson, Sarah E., Kennedy, Emer, Cunningham, Pádraig
In recent years, there has been considerable progress in research on human activity recognition using data from wearable sensors. This technology also has potential in the context of animal welfare in livestock science. In this paper, we report on research on animal activity recognition in support of welfare monitoring. The data comes from collar-mounted accelerometer sensors worn by Holstein and Jersey calves, the objective being to detect changes in behaviour indicating sickness or stress. A key requirement in detecting changes in behaviour is to be able to classify activities into classes, such as drinking, running or walking. In Machine Learning terms, this is a time-series classification task, and in recent years, the Rocket family of methods have emerged as the state-of-the-art in this area. We have over 27 hours of labelled time-series data from 30 calves for our analysis. Using this data as a baseline, we present Rocket's performance on a 6-class classification task. Then, we compare this against the performance of 11 Deep Learning (DL) methods that have been proposed as promising methods for time-series classification. Given the success of DL in related areas, it is reasonable to expect that these methods will perform well here as well. Surprisingly, despite taking care to ensure that the DL methods are configured correctly, none of them match Rocket's performance. A possible explanation for the impressive success of Rocket is that it has the data encoding benefits of DL models in a much simpler classification framework.
SHEDAD: SNN-Enhanced District Heating Anomaly Detection for Urban Substations
van Dreven, Jonne, Cheddad, Abbas, Alawadi, Sadi, Ghazi, Ahmad Nauman, Koussa, Jad Al, Vanhoudt, Dirk
District Heating (DH) systems are essential for energy-efficient urban heating. However, despite the advancements in automated fault detection and diagnosis (FDD), DH still faces challenges in operational faults that impact efficiency. This study introduces the Shared Nearest Neighbor Enhanced District Heating Anomaly Detection (SHEDAD) approach, designed to approximate the DH network topology and allow for local anomaly detection without disclosing sensitive information, such as substation locations. The approach leverages a multi-adaptive k-Nearest Neighbor (k-NN) graph to improve the initial neighborhood creation. Moreover, it introduces a merging technique that reduces noise and eliminates trivial edges. We use the Median Absolute Deviation (MAD) and modified z-scores to flag anomalous substations. The results reveal that SHEDAD outperforms traditional clustering methods, achieving significantly lower intra-cluster variance and distance. Additionally, SHEDAD effectively isolates and identifies two distinct categories of anomalies: supply temperatures and substation performance. We identified 30 anomalous substations and reached a sensitivity of approximately 65\% and specificity of approximately 97\%. By focusing on this subset of poor-performing substations in the network, SHEDAD enables more targeted and effective maintenance interventions, which can reduce energy usage while optimizing network performance.
Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
Rusci, Manuele, Paci, Francesco, Fariselli, Marco, Flamand, Eric, Tuytelaars, Tinne
This paper proposes a self-learning framework to incrementally train (fine-tune) a personalized Keyword Spotting (KWS) model after the deployment on ultra-low power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% vs. the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient Microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training 10x lower than the labeling energy if sampling a new utterance every 5 s or 16.4 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.
Efficient Learning for Linear Properties of Bounded-Gate Quantum Circuits
Du, Yuxuan, Hsieh, Min-Hsiu, Tao, Dacheng
The vast and complicated large-qubit state space forbids us to comprehensively capture the dynamics of modern quantum computers via classical simulations or quantum tomography. However, recent progress in quantum learning theory invokes a crucial question: given a quantum circuit containing d tunable RZ gates and G-d Clifford gates, can a learner perform purely classical inference to efficiently predict its linear properties using new classical inputs, after learning from data obtained by incoherently measuring states generated by the same circuit but with different classical inputs? In this work, we prove that the sample complexity scaling linearly in d is necessary and sufficient to achieve a small prediction error, while the corresponding computational complexity may scale exponentially in d. Building upon these derived complexity bounds, we further harness the concept of classical shadow and truncated trigonometric expansion to devise a kernel-based learning model capable of trading off prediction error and computational complexity, transitioning from exponential to polynomial scaling in many practical settings. Our results advance two crucial realms in quantum computation: the exploration of quantum algorithms with practical utilities and learning-based quantum system certification. We conduct numerical simulations to validate our proposals across diverse scenarios, encompassing quantum information processing protocols, Hamiltonian simulation, and variational quantum algorithms up to 60 qubits.