Education
Investigating learning-independent abstract reasoning in artificial neural networks
Barak, Tomer, Loewenstein, Yonatan
Humans are capable of solving complex abstract reasoning tests. Whether this ability reflects a learning-independent inference mechanism applicable to any novel unlearned problem or whether it is a manifestation of extensive training throughout life is an open question. Addressing this question in humans is challenging because it is impossible to control their prior training. However, assuming a similarity between the cognitive processing of Artificial Neural Networks (ANNs) and humans, the extent to which training is required for ANNs' abstract reasoning is informative about this question in humans. Previous studies demonstrated that ANNs can solve abstract reasoning tests. However, this success required extensive training. In this study, we examined the learning-independent abstract reasoning of ANNs. Specifically, we evaluated their performance without any pretraining, with the ANNs' weights being randomly-initialized, and only change in the process of problem solving. We found that naive ANN models can solve non-trivial visual reasoning tests, similar to those used to evaluate human learning-independent reasoning. We further studied the mechanisms that support this ability. Our results suggest the possibility of learning-independent abstract reasoning that does not require extensive training.
A Model for Combinatorial Dictionary Learning and Inference
Blum, Avrim, Ravichandran, Kavya
We are often interested in decomposing complex, structured data into simple components that explain the data. The linear version of this problem is well-studied as dictionary learning and factor analysis. In this work, we propose a combinatorial model in which to study this question, motivated by the way objects occlude each other in a scene to form an image. First, we identify a property we call "well-structuredness" of a set of low-dimensional components which ensures that no two components in the set are too similar. We show how well-structuredness is sufficient for learning the set of latent components comprising a set of sample instances. We then consider the problem: given a set of components and an instance generated from some unknown subset of them, identify which parts of the instance arise from which components. We consider two variants: (1) determine the minimal number of components required to explain the instance; (2) determine the correct explanation for as many locations as possible. For the latter goal, we also devise a version that is robust to adversarial corruptions, with just a slightly stronger assumption on the components. Finally, we show that the learning problem is computationally infeasible in the absence of any assumptions.
Difficulty Estimation and Simplification of French Text Using LLMs
Jamet, Henri, Shrestha, Yash Raj, Vlachos, Michalis
We frame both tasks as prediction problems and develop a difficulty classification model using labeled examples, transfer learning, and large language models, demonstrating superior accuracy compared to previous approaches. For simplification, we evaluate the trade-off between simplification quality and meaning preservation, comparing zero-shot and fine-tuned performances of large language models. We show that meaningful text simplifications can be obtained with limited fine-tuning. Our experiments are conducted on French texts, but our methods are language-agnostic and directly applicable to other foreign languages.
Online Learning for Autonomous Management of Intent-based 6G Networks
Karakaya, Erciyes, Ercetin, Ozgur, Ozkan, Huseyin, Karaca, Mehmet, Biyar, Elham Dehghan, Palaios, Alexandros
The growing complexity of networks and the variety of future scenarios with diverse and often stringent performance requirements call for a higher level of automation. Intent-based management emerges as a solution to attain high level of automation, enabling human operators to solely communicate with the network through high-level intents. The intents consist of the targets in the form of expectations (i.e., latency expectation) from a service and based on the expectations the required network configurations should be done accordingly. It is almost inevitable that when a network action is taken to fulfill one intent, it can cause negative impacts on the performance of another intent, which results in a conflict. In this paper, we aim to address the conflict issue and autonomous management of intent-based networking, and propose an online learning method based on the hierarchical multi-armed bandits approach for an effective management. Thanks to this hierarchical structure, it performs an efficient exploration and exploitation of network configurations with respect to the dynamic network conditions. We show that our algorithm is an effective approach regarding resource allocation and satisfaction of intent expectations.
Statistical Batch-Based Bearing Fault Detection
Jorry, Victoria, Duma, Zina-Sabrina, Sihvonen, Tuomas, Reinikainen, Satu-Pia, Roininen, Lassi
In the domain of rotating machinery, bearings are vulnerable to different mechanical faults, including ball, inner, and outer race faults. Various techniques can be used in condition-based monitoring, from classical signal analysis to deep learning methods. Based on the complex working conditions of rotary machines, multivariate statistical process control charts such as Hotelling's $T^2$ and Squared Prediction Error are useful for providing early warnings. However, these methods are rarely applied to condition monitoring of rotating machinery due to the univariate nature of the datasets. In the present paper, we propose a multivariate statistical process control-based fault detection method that utilizes multivariate data composed of Fourier transform features extracted for fixed-time batches. Our approach makes use of the multidimensional nature of Fourier transform characteristics, which record more detailed information about the machine's status, in an effort to enhance early defect detection and diagnosis. Experiments with varying vibration measurement locations (Fan End, Drive End), fault types (ball, inner, and outer race faults), and motor loads (0-3 horsepower) are used to validate the suggested approach. The outcomes illustrate our method's effectiveness in fault detection and point to possible broader uses in industrial maintenance.
How Do Students Interact with an LLM-powered Virtual Teaching Assistant in Different Educational Settings?
Maiti, Pratyusha, Goel, Ashok K.
In Jill Watson has been equipped with OpenAI's GPT-this paper, we analyze student interactions with Jill across 3.5 Turbo model, accessed via the OpenAI API, and coupled multiple courses and colleges, focusing on the types and with several other technologies to facilitate more nuanced, complexity of student questions based on Bloom's Revised context-aware, and safe interactions with students. Jill has Taxonomy and tool usage patterns. We find that, by supporting been deployed in both online and offline classrooms[10] across a wide range of cognitive demands, Jill encourages different educational institutes and courses. This paper examines students to engage in sophisticated, higher-order cognitive student interactions with Jill Watson, to understand questions. However, the frequency of usage varies significantly how AI-based educational tools may engage students in meaningful across deployments, and the types of questions asked and deeper learning experiences.
Welcome: 2024 Regional Special Section, Latin America
It is with great pleasure that we introduce the second edition of the Communications of the ACM Latin American Regional Special Section. In this edition, we use this opportunity to showcase some of the region's most interesting as well as impactful advancements in computer science. Latin America is a highly heterogeneous continent with great diversity in culture, geography, demography, ethnicities, languages, and science and technology. When it comes to computer science, Latin American researchers have made significant contributions to multiple areas, such as software engineering, databases, networking and distributed systems, artificial intelligence, computer theory, and computer science education. In this Regional Special Section, we present only a small portion of the work researchers in Latin America are currently conducting. To create this edition, we put out a general call for contributions, receiving submissions from all regions of the continent.
Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model
Petković, Uroš, Frenkel, Jonas, Hellwich, Olaf, Lazarides, Rebecca
This paper introduces a novel computational approach for analyzing nonverbal social behavior in educational settings. Integrating multimodal behavioral cues, including facial expressions, gesture intensity, and spatial dynamics, the model assesses the nonverbal immediacy (NVI) of teachers from RGB classroom videos. A dataset of 400 30-second video segments from German classrooms was constructed for model training and validation. The gesture intensity regressor achieved a correlation of 0.84, the perceived distance regressor 0.55, and the NVI model 0.44 with median human ratings. The model demonstrates the potential to provide a valuable support in nonverbal behavior assessment, approximating the accuracy of individual human raters. Validated against both questionnaire data and trained observer ratings, our models show moderate to strong correlations with relevant educational outcomes, indicating their efficacy in reflecting effective teaching behaviors. This research advances the objective assessment of nonverbal communication behaviors, opening new pathways for educational research.
Amman City, Jordan: Toward a Sustainable City from the Ground Up
The idea of smart cities (SCs) has gained substantial attention in recent years. The SC paradigm aims to improve citizens' quality of life and protect the city's environment. As we enter the age of next-generation SCs, it is important to explore all relevant aspects of the SC paradigm. In recent years, the advancement of Information and Communication Technologies (ICT) has produced a trend of supporting daily objects with smartness, targeting to make human life easier and more comfortable. The paradigm of SCs appears as a response to the purpose of building the city of the future with advanced features. SCs still face many challenges in their implementation, but increasingly more studies regarding SCs are implemented. Nowadays, different cities are employing SC features to enhance services or the residents quality of life. This work provides readers with useful and important information about Amman Smart City.
Building a Domain-specific Guardrail Model in Production
Niknazar, Mohammad, Haley, Paul V, Ramanan, Latha, Truong, Sang T., Shrinivasan, Yedendra, Bhowmick, Ayan Kumar, Dey, Prasenjit, Jagmohan, Ashish, Maheshwari, Hema, Ponoth, Shom, Smith, Robert, Vempaty, Aditya, Haber, Nick, Koyejo, Sanmi, Sundararajan, Sharad
Generative AI holds the promise of enabling a range of sought-after capabilities and revolutionizing workflows in various consumer and enterprise verticals. However, putting a model in production involves much more than just generating an output. It involves ensuring the model is reliable, safe, performant and also adheres to the policy of operation in a particular domain. Guardrails as a necessity for models has evolved around the need to enforce appropriate behavior of models, especially when they are in production. In this paper, we use education as a use case, given its stringent requirements of the appropriateness of content in the domain, to demonstrate how a guardrail model can be trained and deployed in production. Specifically, we describe our experience in building a production-grade guardrail model for a K-12 educational platform. We begin by formulating the requirements for deployment to this sensitive domain. We then describe the training and benchmarking of our domain-specific guardrail model, which outperforms competing open- and closed- instruction-tuned models of similar and larger size, on proprietary education-related benchmarks and public benchmarks related to general aspects of safety. Finally, we detail the choices we made on architecture and the optimizations for deploying this service in production; these range across the stack from the hardware infrastructure to the serving layer to language model inference optimizations. We hope this paper will be instructive to other practitioners looking to create production-grade domain-specific services based on generative AI and large language models.