Goto

Collaborating Authors

 Education


AceMap: Knowledge Discovery through Academic Graph

arXiv.org Artificial Intelligence

The exponential growth of scientific literature requires effective management and extraction of valuable insights. While existing scientific search engines excel at delivering search results based on relational databases, they often neglect the analysis of collaborations between scientific entities and the evolution of ideas, as well as the in-depth analysis of content within scientific publications. The representation of heterogeneous graphs and the effective measurement, analysis, and mining of such graphs pose significant challenges. To address these challenges, we present AceMap, an academic system designed for knowledge discovery through academic graph. We present advanced database construction techniques to build the comprehensive AceMap database with large-scale academic entities that contain rich visual, textual, and numerical information. AceMap also employs innovative visualization, quantification, and analysis methods to explore associations and logical relationships among academic entities. AceMap introduces large-scale academic network visualization techniques centered on nebular graphs, providing a comprehensive view of academic networks from multiple perspectives. In addition, AceMap proposes a unified metric based on structural entropy to quantitatively measure the knowledge content of different academic entities. Moreover, AceMap provides advanced analysis capabilities, including tracing the evolution of academic ideas through citation relationships and concept co-occurrence, and generating concise summaries informed by this evolutionary process. In addition, AceMap uses machine reading methods to generate potential new ideas at the intersection of different fields. Exploring the integration of large language models and knowledge graphs is a promising direction for future research in idea evolution. Please visit \url{https://www.acemap.info} for further exploration.


Understanding the Role of Temperature in Diverse Question Generation by GPT-4

arXiv.org Artificial Intelligence

In this paper, we provide We conduct a preliminary study of the effect of GPT's temperature early results from our investigation into the effects of different parameter on the diversity of GPT4-generated questions. We find temperature settings in an MCQ generation pipeline. Existing literature that using higher temperature values leads to significantly higher suggests that content generated by GPT-4 can be homogenous; diversity, with different temperatures exposing different types of temperature provides a way to improve the diversity of the similarity between generated sets of questions. We also demonstrate generated content with minimal prompt engineering [2].


HyperCLOVA X Technical Report

arXiv.org Artificial Intelligence

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean, English, and code data, followed by instruction-tuning with high-quality human-annotated datasets while abiding by strict safety guidelines reflecting our commitment to responsible AI. The model is evaluated across various benchmarks, including comprehensive reasoning, knowledge, commonsense, factuality, coding, math, chatting, instruction-following, and harmlessness, in both Korean and English. HyperCLOVA X exhibits strong reasoning capabilities in Korean backed by a deep understanding of the language and cultural nuances. Further analysis of the inherent bilingual nature and its extension to multilingualism highlights the model's cross-lingual proficiency and strong generalization ability to untargeted languages, including machine translation between several language pairs and cross-lingual inference tasks. We believe that HyperCLOVA X can provide helpful guidance for regions or countries in developing their sovereign LLMs.


BERT-like Pre-training for Symbolic Piano Music Classification Tasks

arXiv.org Artificial Intelligence

This article presents a benchmark study of symbolic piano music classification using the masked language modelling approach of the Bidirectional Encoder Representations from Transformers (BERT). Specifically, we consider two types of MIDI data: MIDI scores, which are musical scores rendered directly into MIDI with no dynamics and precisely aligned with the metrical grid notated by its composer and MIDI performances, which are MIDI encodings of human performances of musical scoresheets. With five public-domain datasets of single-track piano MIDI files, we pre-train two 12-layer Transformer models using the BERT approach, one for MIDI scores and the other for MIDI performances, and fine-tune them for four downstream classification tasks. These include two note-level classification tasks (melody extraction and velocity prediction) and two sequence-level classification tasks (style classification and emotion classification). Our evaluation shows that the BERT approach leads to higher classification accuracy than recurrent neural network (RNN)-based baselines.


Navigating the Landscape of Large Language Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies

arXiv.org Artificial Intelligence

With the surge of ChatGPT,the use of large models has significantly increased,rapidly rising to prominence across the industry and sweeping across the internet. This article is a comprehensive review of fine-tuning methods for large models. This paper investigates the latest technological advancements and the application of advanced methods in aspects such as task-adaptive fine-tuning,domain-adaptive fine-tuning,few-shot learning,knowledge distillation,multi-task learning,parameter-efficient fine-tuning,and dynamic fine-tuning.


G-ACIL: Analytic Learning for Exemplar-Free Generalized Class Incremental Learning

arXiv.org Artificial Intelligence

Class incremental learning (CIL) trains a network on sequential tasks with separated categories but suffers from catastrophic forgetting, where models quickly lose previously learned knowledge when acquiring new tasks. The generalized CIL (GCIL) aims to address the CIL problem in a more real-world scenario, where incoming data have mixed data categories and unknown sample size distribution, leading to intensified forgetting. Existing attempts for the GCIL either have poor performance, or invade data privacy by saving historical exemplars. To address this, in this paper, we propose an exemplar-free generalized analytic class incremental learning (G-ACIL). The G-ACIL adopts analytic learning (a gradient-free training technique), and delivers an analytical solution (i.e., closed-form) to the GCIL scenario. This solution is derived via decomposing the incoming data into exposed and unexposed classes, allowing an equivalence between the incremental learning and its joint training, i.e., the weight-invariant property. Such an equivalence is theoretically validated through matrix analysis tools, and hence contributes interpretability in GCIL. It is also empirically evidenced by experiments on various datasets and settings of GCIL. The results show that the G-ACIL exhibits leading performance with high robustness compared with existing competitive GCIL methods. Codes will be ready at \url{https://github.com/ZHUANGHP/Analytic-continual-learning}.


Provable Interactive Learning with Hindsight Instruction Feedback

arXiv.org Machine Learning

We study interactive learning in a setting where the agent has to generate a response (e.g., an action or trajectory) given a context and an instruction. In contrast, to typical approaches that train the system using reward or expert supervision on response, we study learning with hindsight instruction where a teacher provides an instruction that is most suitable for the agent's generated response. This hindsight labeling of instruction is often easier to provide than providing expert supervision of the optimal response which may require expert knowledge or can be impractical to elicit. We initiate the theoretical analysis of interactive learning with hindsight labeling. We first provide a lower bound showing that in general, the regret of any algorithm must scale with the size of the agent's response space. We then study a specialized setting where the underlying instruction-response distribution can be decomposed as a low-rank matrix. We introduce an algorithm called LORIL for this setting and show that its regret scales as $\sqrt{T}$ where $T$ is the number of rounds and depends on the intrinsic rank but does not depend on the size of the agent's response space. We provide experiments in two domains showing that LORIL outperforms baselines even when the low-rank assumption is violated.


Artificial Intelligence in Everyday Life 2.0: Educating University Students from Different Majors

arXiv.org Artificial Intelligence

The integration With the surge in data-centric AI and its increasing capabilities, AI of AI into everyday life will only increase as new applications applications have become a part of our everyday lives. However, are developed for use in our homes, schools, governments, social misunderstandings regarding their capabilities, limitations, and lives, and workplaces. But despite the progress made, we have associated advantages and disadvantages are widespread. Consequently, also seen how serious the consequences of misunderstanding or in the university setting, there is a crucial need to educate failing to question AI decisions can be - leading to issues such as not only computer science majors but also students from various disciplines viral misinformation [8], biased systems that disproportionately about AI. In this experience report, we present an overview impact marginalized communities [1], and serious concerns about of an introductory course that we offered to students coming from data privacy. This situation highlights the need to bridge the gap different majors. Moreover, we discuss the assignments and quizzes between AI's everyday presence and people's lack of knowledge, so of the course, which provided students with a firsthand experience we can clear up misconceptions, reduce fears, and embrace a more of AI processes and insights into their learning patterns. Additionally, informed relationship with the AI that is shaping our future [4].


Deep Learning for Educational Data Science

arXiv.org Artificial Intelligence

As artificial intelligence (AI) continues to penetrate ever deeper into modern life, one particular family of machine learning algorithms--namely, deep neural networks--have come to be seen as the solution to many of the challenges that have stumped more classical algorithms in the past. Modeled loosely on the structure of biological neural networks, artificial neural networks consist of chains of simple mathematical transformations that can model complex non-linear decision boundaries in large problem spaces. In particular, deep neural networks--artificial neural networks that consist of multiple layers of transformations--allow for sufficient complexity to tackle tasks in a wide variety of fields. These models are collectively and more colloquially referred to as deep learning. A growing body of education researchers are now also turning their attention to leveraging the power of deep learning algorithms for the tasks of improving and understanding human learning. Researchers in educational data science, a field consisting of various interrelated research communities such as Educational Data Mining (EDM), Learning Analytics (LA), and AI in Education (AIED), have been involved in this endeavor.


Collaborative-Enhanced Prediction of Spending on Newly Downloaded Mobile Games under Consumption Uncertainty

arXiv.org Artificial Intelligence

With the surge in mobile gaming, accurately predicting user spending on newly downloaded games has become paramount for maximizing revenue. However, the inherently unpredictable nature of user behavior poses significant challenges in this endeavor. To address this, we propose a robust model training and evaluation framework aimed at standardizing spending data to mitigate label variance and extremes, ensuring stability in the modeling process. Within this framework, we introduce a collaborative-enhanced model designed to predict user game spending without relying on user IDs, thus ensuring user privacy and enabling seamless online training. Our model adopts a unique approach by separately representing user preferences and game features before merging them as input to the spending prediction module. Through rigorous experimentation, our approach demonstrates notable improvements over production models, achieving a remarkable \textbf{17.11}\% enhancement on offline data and an impressive \textbf{50.65}\% boost in an online A/B test. In summary, our contributions underscore the importance of stable model training frameworks and the efficacy of collaborative-enhanced models in predicting user spending behavior in mobile gaming.