Education
Predicting Affective States from Screen Text Sentiment
Teng, Songyan, Zhang, Tianyi, D'Alfonso, Simon, Kostakos, Vassilis
The proliferation of mobile sensing technologies has enabled the Mobile sensing technologies have been widely used in wellbeing study of various physiological and behavioural phenomena through studies and applications, and the significant advancements in sensing unobtrusive data collection from smartphone sensors. This approach over the last decade have spurred heightened interest in this offers real-time insights into individuals' physical and mental field, often referred to as "digital phenotyping". This approach often states, creating opportunities for personalised treatment and involves the use of smartphone sensors to continuously and unobtrusively interventions. However, the potential of analysing the textual content collect data on various physiological and behavioural viewed on smartphones to predict affective states remains phenomena [9]. Data from a range of smartphone sensors can be underexplored. To better understand how the screen text that users integrated to obtain a comprehensive understanding of a person's are exposed to and interact with can influence their affects, we surroundings, activities, and behaviours [1]. This approach allows investigated a subset of data obtained from a digital phenotyping for real-time monitoring and analysis of individuals' physical and study of Australian university students conducted in 2023. We employed mental states, providing valuable insights into their overall wellbeing linear regression, zero-shot, and multi-shot prompting using and creating opportunities for delivering recommendations and a large language model (LLM) to analyse relationships between interventions based on the user's context.
Reconciling Different Theories of Learning with an Agent-based Model of Procedural Learning
Rismanchian, Sina, Doroudi, Shayan
Computational models of human learning can play a significant role in enhancing our knowledge about nuances in theoretical and qualitative learning theories and frameworks. There are many existing frameworks in educational settings that have shown to be verified using empirical studies, but at times we find these theories make conflicting claims or recommendations for instruction. In this study, we propose a new computational model of human learning, Procedural ABICAP, that reconciles the ICAP, Knowledge-Learning-Instruction (KLI), and cognitive load theory (CLT) frameworks for learning procedural knowledge. ICAP assumes that constructive learning generally yields better learning outcomes, while theories such as KLI and CLT claim that this is not always true. We suppose that one reason for this may be that ICAP is primarily used for conceptual learning and is underspecified as a framework for thinking about procedural learning. We show how our computational model, both by design and through simulations, can be used to reconcile different results in the literature. More generally, we position our computational model as an executable theory of learning that can be used to simulate various educational settings.
Avatar Visual Similarity for Social HCI: Increasing Self-Awareness
Hilpert, Bernhard, da Silva, Claudio Alves, Christidis, Leon, Bhuvaneshwara, Chirag, Gebhard, Patrick, Nunnari, Fabrizio, Tsovaltzi, Dimitra
Self-awareness is a critical factor in social human-human interaction and, hence, in social HCI interaction. Increasing self-awareness through mirrors or video recordings is common in face-to-face trainings, since it influences antecedents of self-awareness like explicit identification and implicit affective identification (affinity). However, increasing self-awareness has been scarcely examined in virtual trainings with virtual avatars, which allow for adjusting the similarity, e.g. to avoid negative effects of self-consciousness. Automatic visual similarity in avatars is an open issue related to high costs. It is important to understand which features need to be manipulated and which degree of similarity is necessary for self-awareness to leverage the added value of using avatars for self-awareness. This article examines the relationship between avatar visual similarity and increasing self-awareness in virtual training environments. We define visual similarity based on perceptually important facial features for human-human identification and develop a theory-based methodology to systematically manipulate visual similarity of virtual avatars and support self-awareness. Three personalized versions of virtual avatars with varying degrees of visual similarity to participants were created (weak, medium and strong facial features manipulation). In a within-subject study (N=33), we tested effects of degree of similarity on perceived similarity, explicit identification and implicit affective identification (affinity). Results show significant differences between the weak similarity manipulation, and both the strong manipulation and the random avatar for all three antecedents of self-awareness. An increasing degree of avatar visual similarity influences antecedents of self-awareness in virtual environments.
Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese
Doan, Khang T., Huynh, Bao G., Hoang, Dung T., Pham, Thuc D., Pham, Nhat H., Nguyen, Quan T. M., Vo, Bang Q., Hoang, Suong N.
In this report, we introduce Vintern-1B, a reliable 1-billion-parameters multimodal large language model (MLLM) for Vietnamese language tasks. By integrating the Qwen2-0.5B-Instruct language model with the InternViT-300M-448px visual model, Vintern-1B is optimized for a range of applications, including optical character recognition (OCR), document extraction, and general question-answering in Vietnamese context. The model is fine-tuned on an extensive dataset of over 3 million image-question-answer pairs, achieving robust performance and reliable results across multiple Vietnamese language benchmarks like OpenViVQA and ViTextVQA. Vintern-1B is small enough to fit into various on-device applications easily. Additionally, we have open-sourced several Vietnamese vision question answering (VQA) datasets for text and diagrams, created with Gemini 1.5 Flash. Our models are available at: https://huggingface.co/5CD-AI/Vintern-1B-v2.
MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark
Zhou, Minxuan, Liang, Hao, Li, Tianpeng, Wu, Zhiyu, Lin, Mingan, Sun, Linzhuang, Zhou, Yaqi, Zhang, Yan, Huang, Xiaoqin, Chen, Yicong, Qiao, Yujing, Chen, Weipeng, Cui, Bin, Zhang, Wentao, Zhou, Zenan
With the development of Multimodal Large Language Models (MLLMs), the evaluation of multimodal models in the context of mathematical problems has become a valuable research field. Multimodal visual-textual mathematical reasoning serves as a critical indicator for evaluating the comprehension and complex multi-step quantitative reasoning abilities of MLLMs. However, previous multimodal math benchmarks have not sufficiently integrated visual and textual information. To address this gap, we proposed MathScape, a new benchmark that emphasizes the understanding and application of combined visual and textual information. MathScape is designed to evaluate photo-based math problem scenarios, assessing the theoretical understanding and application ability of MLLMs through a categorical hierarchical approach. We conduct a multi-dimensional evaluation on 11 advanced MLLMs, revealing that our benchmark is challenging even for the most sophisticated models. By analyzing the evaluation results, we identify the limitations of MLLMs, offering valuable insights for enhancing model performance.
How ancient tech is thwarting AI cheating in the classroom
Nearly two years ago, ChatGPT's AI writing powers set off a firestorm in classrooms. How would teachers be able to determine which assignments were actually authored by the student? A host of AI-powered services answered the call. Today, there are even more services promising to catch AI cheaters. "My hand cramped up so much," my eldest son complained about his AP World History course he took last year, and the requirement to handwrite all papers and tests because of AI concerns.
California high school principal placed on leave after video surfaces of inappropriate dance with mascot
A viral video shows a high school principal engaging in a seemingly risqué dance with the school's mascot during the back-to-school rally. A high school principal in central California has been placed on administrative leave as an investigation is underway into a video showing him dancing in what some have called an inappropriate manner with the school's mascot during a back-to-school rally. The Merced Union High School District shared a statement with Fox News Digital that said Robert Nunes, principal of Buhach Colony High School in Atwater, was on administrative leave effective Aug. 19. The district said this action is in response to an incident at the back-to-school rally on Aug. 16. "The District is conducting a comprehensive review of the situation. While the investigation is ongoing, Mr. Nunes will not be participating in any school-related responsibilities or activities," Viviana Fuentes, director of communications for the school district, said in the statement.
A Language-agnostic Model of Child Language Acquisition
Mahon, Louis, Abend, Omri, Berger, Uri, Demuth, Katherine, Johnson, Mark, Steedman, Mark
This work reimplements a recent semantic bootstrapping child-language acquisition model, which was originally designed for English, and trains it to learn a new language: Hebrew. The model learns from pairs of utterances and logical forms as meaning representations, and acquires both syntax and word meanings simultaneously. The results show that the model mostly transfers to Hebrew, but that a number of factors, including the richer morphology in Hebrew, makes the learning slower and less robust. This suggests that a clear direction for future work is to enable the model to leverage the similarities between different word forms.
Weight Scope Alignment: A Frustratingly Easy Method for Model Merging
Xu, Yichu, Li, Xin-Chun, Gan, Le, Zhan, De-Chuan
Merging models becomes a fundamental procedure in some applications that consider model efficiency and robustness. The training randomness or Non-I.I.D. data poses a huge challenge for averaging-based model fusion. Previous research efforts focus on element-wise regularization or neural permutations to enhance model averaging while overlooking weight scope variations among models, which can significantly affect merging effectiveness. In this paper, we reveal variations in weight scope under different training conditions, shedding light on its influence on model merging. Fortunately, the parameters in each layer basically follow the Gaussian distribution, which inspires a novel and simple regularization approach named Weight Scope Alignment (WSA). It contains two key components: 1) leveraging a target weight scope to guide the model training process for ensuring weight scope matching in the subsequent model merging. 2) fusing the weight scope of two or more models into a unified one for multi-stage model fusion. We extend the WSA regularization to two different scenarios, including Mode Connectivity and Federated Learning. Abundant experimental studies validate the effectiveness of our approach.
Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
Rusci, Manuele, Paci, Francesco, Fariselli, Marco, Flamand, Eric, Tuytelaars, Tinne
This paper proposes a self-learning framework to incrementally train (fine-tune) a personalized Keyword Spotting (KWS) model after the deployment on ultra-low power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% vs. the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient Microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training 10x lower than the labeling energy if sampling a new utterance every 5 s or 16.4 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.