Education
On conceptualisation and an overview of learning path recommender systems in e-learning
Fuster-López, A., Cruz, J. M., Guerrero-García, P., Hendrix, E. M. T., Košir, A., Nowak, I., Oneto, L., Sirmakessis, S., Pacheco, M. F., Fernandes, F. P., Pereira, A. I.
In recent years, the landscape of e-learning has witnessed exceptional advancements, providing students with tools to improve their performance. In the pursuit of optimizing the e-learning experience, one emerging area of focus is the integration of recommender systems. By leveraging sophisticated algorithms, recommender systems aim to personalize the learning path by tailoring recommendations based on individual student performance, preferences, learning style and other factors.
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
Song, Zheshu, Zhuo, Jianheng, Yang, Yifan, Ma, Ziyang, Zhang, Shixiong, Chen, Xie
When new languages need to be integrated into a multilingual ASR system, a naive Recent years have witnessed significant progress in multilingual approach is to fine-tune the ASR model using data from these automatic speech recognition (ASR), driven by the emergence new languages. Unfortunately, this often results in catastrophic of end-to-end (E2E) models and the scaling of multilingual forgetting, referring to the phenomenon that the recognition performance datasets. Despite that, two main challenges persist in multilingual of base languages tends to decline. To solve the above ASR: language interference and the incorporation of problem, Li et al. [26] proposes lifelong learning [27] solution new languages without degrading the performance of the existing which remedies the language interference problem by mixing ones. This paper proposes LoRA-Whisper, which incorporates base language data and new language data. However, this approach LoRA matrix into Whisper for multilingual ASR, is inefficient and time-consuming. Libera et al. [28] explores effectively mitigating language interference. Furthermore, by various continual learning methods [29-34] to address leveraging LoRA and the similarities between languages, we the issue of catastrophic forgetting. While these approaches can achieve better performance on new languages while upholding have helped alleviate the problem, it still persists.
Lean Workbook: A large-scale Lean problem set formalized from natural language math problems
Ying, Huaiyuan, Wu, Zijian, Geng, Yihan, Wang, Jiayu, Lin, Dahua, Chen, Kai
Large language models have demonstrated impressive capabilities across various natural language processing tasks, especially in solving mathematical problems. However, large language models are not good at math theorem proving using formal languages like Lean. A significant challenge in this area is the scarcity of training data available in these formal languages. To address this issue, we propose a novel pipeline that iteratively generates and filters synthetic data to translate natural language mathematical problems into Lean 4 statements, and vice versa. Our results indicate that the synthetic data pipeline can provide useful training data and improve the performance of LLMs in translating and understanding complex mathematical problems and proofs. Our final dataset contains about 57K formal-informal question pairs along with searched proof from the math contest forum and 21 new IMO questions.
LogiCode: an LLM-Driven Framework for Logical Anomaly Detection
Zhang, Yiheng, Cao, Yunkang, Xu, Xiaohao, Shen, Weiming
This paper presents LogiCode, a novel framework that leverages Large Language Models (LLMs) for identifying logical anomalies in industrial settings, moving beyond traditional focus on structural inconsistencies. By harnessing LLMs for logical reasoning, LogiCode autonomously generates Python codes to pinpoint anomalies such as incorrect component quantities or missing elements, marking a significant leap forward in anomaly detection technologies. A custom dataset "LOCO-Annotations" and a benchmark "LogiBench" are introduced to evaluate the LogiCode's performance across various metrics including binary classification accuracy, code generation success rate, and precision in reasoning. Findings demonstrate LogiCode's enhanced interpretability, significantly improving the accuracy of logical anomaly detection and offering detailed explanations for identified anomalies. This represents a notable shift towards more intelligent, LLM-driven approaches in industrial anomaly detection, promising substantial impacts on industry-specific applications.
Language Guided Skill Discovery
Rho, Seungeun, Smith, Laura, Li, Tianyu, Levine, Sergey, Peng, Xue Bin, Ha, Sehoon
Skill discovery methods enable agents to learn diverse emergent behaviors without explicit rewards. To make learned skills useful for unknown downstream tasks, obtaining a semantically diverse repertoire of skills is essential. While some approaches introduce a discriminator to distinguish skills and others aim to increase state coverage, no existing work directly addresses the "semantic diversity" of skills. We hypothesize that leveraging the semantic knowledge of large language models (LLMs) can lead us to improve semantic diversity of resulting behaviors. In this sense, we introduce Language Guided Skill Discovery (LGSD), a skill discovery framework that aims to directly maximize the semantic diversity between skills. LGSD takes user prompts as input and outputs a set of semantically distinctive skills. The prompts serve as a means to constrain the search space into a semantically desired subspace, and the generated LLM outputs guide the agent to visit semantically diverse states within the subspace. We demonstrate that LGSD enables legged robots to visit different user-intended areas on a plane by simply changing the prompt. Furthermore, we show that language guidance aids in discovering more diverse skills compared to five existing skill discovery methods in robot-arm manipulation environments. Lastly, LGSD provides a simple way of utilizing learned skills via natural language.
A model of early word acquisition based on realistic-scale audiovisual naming events
Khorrami, Khazar, Räsänen, Okko
As they grow, infants gradually acquire understanding of their native language without direct supervision. By the age of 6 months, infants' perception has already attuned to native language phonetic contrasts [1] and they show first signs of word comprehension [2-4] and familiar word identification [5-7]. By 12 months, they already recognize dozens of words [8]. During this learning process, the infants must learn to parse the speech stream into words and to associate the words with their referential meanings in the external world. From a cognitive perspective, the discovery of words and word-meaning mappings is a task far from trivial: acoustic speech is a continuous and complex signal without transparent linguistic structure (see, e.g., [9, 10]), and there is substantial ambiguity in how individual words embedded in larger utterances are related to specific objects and events in the visual scene, also known as referential ambiguity [11]. Figure 1 illustrates referential ambiguity in everyday life situations between speech and its visual references through some examples.
SLOPE: Search with Learned Optimal Pruning-based Expansion
Bokan, Davor, Ajanovic, Zlatan, Lacevic, Bakir
Heuristic search is often used for motion planning and pathfinding problems, for finding the shortest path in a graph while also promising completeness and optimal efficiency. The drawback is it's space complexity, specifically storing all expanded child nodes in memory and sorting large lists of active nodes, which can be a problem in real-time scenarios with limited on-board computation. To combat this, we present the Search with Learned Optimal Pruning-based Expansion (SLOPE), which, learns the distance of a node from a possible optimal path, unlike other approaches that learn a cost-to-go value. The unfavored nodes are then pruned according to the said distance, which in turn reduces the size of the open list. This ensures that the search explores only the region close to optimal paths while lowering memory and computational costs. Unlike traditional learning methods, our approach is orthogonal to estimating cost-to-go heuristics, offering a complementary strategy for improving search efficiency. We demonstrate the effectiveness of our approach evaluating it as a standalone search method and in conjunction with learned heuristic functions, achieving comparable-or-better node expansion metrics, while lowering the number of child nodes in the open list. Our code is available at https://github.com/dbokan1/SLOPE.
Are We Done with MMLU?
Gema, Aryo Pradipta, Leang, Joshua Ong Jun, Hong, Giwon, Devoto, Alessio, Mancino, Alberto Carlo Maria, Saxena, Rohit, He, Xuanli, Zhao, Yu, Du, Xiaotang, Madani, Mohammad Reza Ghasemi, Barale, Claire, McHardy, Robert, Harris, Joshua, Kaddour, Jean, van Krieken, Emile, Minervini, Pasquale
We identify and analyse errors in the popular Massive Multitask Language Understanding (MMLU) benchmark. Even though MMLU is widely adopted, our analysis demonstrates numerous ground truth errors that obscure the true capabilities of LLMs. For example, we find that 57% of the analysed questions in the Virology subset contain errors. To address this issue, we introduce a comprehensive framework for identifying dataset errors using a novel error taxonomy. Then, we create MMLU-Redux, which is a subset of 3,000 manually re-annotated questions across 30 MMLU subjects. Using MMLU-Redux, we demonstrate significant discrepancies with the model performance metrics that were originally reported. Our results strongly advocate for revising MMLU's error-ridden questions to enhance its future utility and reliability as a benchmark.
TLEX: An Efficient Method for Extracting Exact Timelines from TimeML Temporal Graphs
Ocal, Mustafa, Xie, Ning, Finlayson, Mark
A timeline provides a total ordering of events and times, and is useful for a number of natural language understanding tasks. However, qualitative temporal graphs that can be derived directly from text -- such as TimeML annotations -- usually explicitly reveal only partial orderings of events and times. In this work, we apply prior work on solving point algebra problems to the task of extracting timelines from TimeML annotated texts, and develop an exact, end-to-end solution which we call TLEX (TimeLine EXtraction). TLEX transforms TimeML annotations into a collection of timelines arranged in a trunk-and-branch structure. Like what has been done in prior work, TLEX checks the consistency of the temporal graph and solves it; however, it adds two novel functionalities. First, it identifies specific relations involved in an inconsistency (which could then be manually corrected) and, second, TLEX performs a novel identification of sections of the timelines that have indeterminate order, information critical for downstream tasks such as aligning events from different timelines. We provide detailed descriptions and analysis of the algorithmic components in TLEX, and conduct experimental evaluations by applying TLEX to 385 TimeML annotated texts from four corpora. We show that 123 of the texts are inconsistent, 181 of them have more than one ``real world'' or main timeline, and there are 2,541 indeterminate sections across all four corpora. A sampling evaluation showed that TLEX is 98--100% accurate with 95% confidence along five dimensions: the ordering of time-points, the number of main timelines, the placement of time-points on main versus subordinate timelines, the connecting point of branch timelines, and the location of the indeterminate sections. We provide a reference implementation of TLEX, the extracted timelines for all texts, and the manual corrections of the inconsistent texts.
Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios
There is increasing interest in distilling task-specific knowledge from large language models (LLM) to smaller student models. Nonetheless, LLM distillation presents a dual challenge: 1) there is a high cost associated with querying the teacher LLM, such as GPT-4, for gathering an ample number of demonstrations; 2) the teacher LLM might provide imperfect outputs with a negative impact on the student's learning process. To enhance sample efficiency within resource-constrained, imperfect teacher scenarios, we propose a three-component framework leveraging three signal types. The first signal is the student's self-consistency (consistency of student multiple outputs), which is a proxy of the student's confidence. Specifically, we introduce a ``teaching assistant'' (TA) model to assess the uncertainty of both the student's and the teacher's outputs via confidence scoring, which serves as another two signals for student training. Furthermore, we propose a two-stage training schema to first warm up the student with a small proportion of data to better utilize student's signal. Experiments have shown the superiority of our proposed framework for four complex reasoning tasks. On average, our proposed two-stage framework brings a relative improvement of up to 20.79% compared to fine-tuning without any signals across datasets.