Education
Automated Code Extraction from Discussion Board Text Dataset
Saravani, Sina Mahdipour, Ghaffari, Sadaf, Luther, Yanye, Folkestad, James, Moraes, Marcia
This study introduces and investigates the capabilities of three different text mining approaches, namely Latent Semantic Analysis, Latent Dirichlet Analysis, and Clustering Word Vectors, for automating code extraction from a relatively small discussion board dataset. We compare the outputs of each algorithm with a previous dataset that was manually coded by two human raters. The results show that even with a relatively small dataset, automated approaches can be an asset to course instructors by extracting some of the discussion codes, which can be used in Epistemic Network Analysis.
A Survey of Adversarial Defences and Robustness in NLP
Goyal, Shreya, Doddapaneni, Sumanth, Khapra, Mitesh M., Ravindran, Balaraman
In the past few years, it has become increasingly evident that deep neural networks are not resilient enough to withstand adversarial perturbations in input data, leaving them vulnerable to attack. Various authors have proposed strong adversarial attacks for computer vision and Natural Language Processing (NLP) tasks. As a response, many defense mechanisms have also been proposed to prevent these networks from failing. The significance of defending neural networks against adversarial attacks lies in ensuring that the model's predictions remain unchanged even if the input data is perturbed. Several methods for adversarial defense in NLP have been proposed, catering to different NLP tasks such as text classification, named entity recognition, and natural language inference. Some of these methods not only defend neural networks against adversarial attacks but also act as a regularization mechanism during training, saving the model from overfitting. This survey aims to review the various methods proposed for adversarial defenses in NLP over the past few years by introducing a novel taxonomy. The survey also highlights the fragility of advanced deep neural networks in NLP and the challenges involved in defending them.
Planning to Practice: Efficient Online Fine-Tuning by Composing Goals in Latent Space
Fang, Kuan, Yin, Patrick, Nair, Ashvin, Levine, Sergey
General-purpose robots require diverse repertoires of behaviors to complete challenging tasks in real-world unstructured environments. To address this issue, goal-conditioned reinforcement learning aims to acquire policies that can reach configurable goals for a wide range of tasks on command. However, such goal-conditioned policies are notoriously difficult and time-consuming to train from scratch. In this paper, we propose Planning to Practice (PTP), a method that makes it practical to train goal-conditioned policies for long-horizon tasks that require multiple distinct types of interactions to solve. Our approach is based on two key ideas. First, we decompose the goal-reaching problem hierarchically, with a high-level planner that sets intermediate subgoals using conditional subgoal generators in the latent space for a low-level model-free policy. Second, we propose a hybrid approach which first pre-trains both the conditional subgoal generator and the policy on previously collected data through offline reinforcement learning, and then fine-tunes the policy via online exploration. This fine-tuning process is itself facilitated by the planned subgoals, which breaks down the original target task into short-horizon goal-reaching tasks that are significantly easier to learn. We conduct experiments in both the simulation and real world, in which the policy is pre-trained on demonstrations of short primitive behaviors and fine-tuned for temporally extended tasks that are unseen in the offline data. Our experimental results show that PTP can generate feasible sequences of subgoals that enable the policy to efficiently solve the target tasks.
Sharp-SSL: Selective high-dimensional axis-aligned random projections for semi-supervised learning
Wang, Tengyao, Dobriban, Edgar, Gataric, Milana, Samworth, Richard J.
We propose a new method for high-dimensional semi-supervised learning problems based on the careful aggregation of the results of a low-dimensional procedure applied to many axis-aligned random projections of the data. Our primary goal is to identify important variables for distinguishing between the classes; existing low-dimensional methods can then be applied for final class assignment. Motivated by a generalized Rayleigh quotient, we score projections according to the traces of the estimated whitened between-class covariance matrices on the projected data. This enables us to assign an importance weight to each variable for a given projection, and to select our signal variables by aggregating these weights over high-scoring projections. Our theory shows that the resulting Sharp-SSL algorithm is able to recover the signal coordinates with high probability when we aggregate over sufficiently many random projections and when the base procedure estimates the whitened between-class covariance matrix sufficiently well. The Gaussian EM algorithm is a natural choice as a base procedure, and we provide a new analysis of its performance in semi-supervised settings that controls the parameter estimation error in terms of the proportion of labeled data in the sample. Numerical results on both simulated data and a real colon tumor dataset support the excellent empirical performance of the method.
Creating Large Language Model Resistant Exams: Guidelines and Strategies
The proliferation of Large Language Models (LLMs), such as ChatGPT, has raised concerns about their potential impact on academic integrity, prompting the need for LLM-resistant exam designs. This article investigates the performance of LLMs on exams and their implications for assessment, focusing on ChatGPT's abilities and limitations. We propose guidelines for creating LLM-resistant exams, including content moderation, deliberate inaccuracies, real-world scenarios beyond the model's knowledge base, effective distractor options, evaluating soft skills, and incorporating non-textual information. The article also highlights the significance of adapting assessments to modern tools and promoting essential skills development in students. By adopting these strategies, educators can maintain academic integrity while ensuring that assessments accurately reflect contemporary professional settings and address the challenges and opportunities posed by artificial intelligence in education.
METAM: Goal-Oriented Data Discovery
Galhotra, Sainyam, Gong, Yue, Fernandez, Raul Castro
Data is a central component of machine learning and causal inference tasks. The availability of large amounts of data from sources such as open data repositories, data lakes and data marketplaces creates an opportunity to augment data and boost those tasks' performance. However, augmentation techniques rely on a user manually discovering and shortlisting useful candidate augmentations. Existing solutions do not leverage the synergy between discovery and augmentation, thus under exploiting data. In this paper, we introduce METAM, a novel goal-oriented framework that queries the downstream task with a candidate dataset, forming a feedback loop that automatically steers the discovery and augmentation process. To select candidates efficiently, METAM leverages properties of the: i) data, ii) utility function, and iii) solution set size. We show METAM's theoretical guarantees and demonstrate those empirically on a broad set of tasks. All in all, we demonstrate the promise of goal-oriented data discovery to modern data science applications.
FedTP: Federated Learning by Transformer Personalization
Li, Hongxia, Cai, Zhongyi, Wang, Jingya, Tang, Jiangnan, Ding, Weiping, Lin, Chin-Teng, Shi, Ye
Federated learning is an emerging learning paradigm where multiple clients collaboratively train a machine learning model in a privacy-preserving manner. Personalized federated learning extends this paradigm to overcome heterogeneity across clients by learning personalized models. Recently, there have been some initial attempts to apply Transformers to federated learning. However, the impacts of federated learning algorithms on self-attention have not yet been studied. This paper investigates this relationship and reveals that federated averaging algorithms actually have a negative impact on self-attention where there is data heterogeneity. These impacts limit the capabilities of the Transformer model in federated learning settings. Based on this, we propose FedTP, a novel Transformer-based federated learning framework that learns personalized self-attention for each client while aggregating the other parameters among the clients. Instead of using a vanilla personalization mechanism that maintains personalized self-attention layers of each client locally, we develop a learn-to-personalize mechanism to further encourage the cooperation among clients and to increase the scablability and generalization of FedTP. Specifically, the learn-to-personalize is realized by learning a hypernetwork on the server that outputs the personalized projection matrices of self-attention layers to generate client-wise queries, keys and values. Furthermore, we present the generalization bound for FedTP with the learn-to-personalize mechanism. Notably, FedTP offers a convenient environment for performing a range of image and language tasks using the same federated network architecture - all of which benefit from Transformer personalization. Extensive experiments verify that FedTP with the learn-to-personalize mechanism yields state-of-the-art performance in non-IID scenarios. Our code is available online.
Future of Education: Application not Regurgitation of Knowledge – Part II - DataScienceCentral.com
AI technologies like ChatGPT are necessitating a fundamental overhaul of our educational systems and institutions. Getting the right answers to predetermined tests is no longer sufficient in an age where AI can access, integrate, and recite knowledge billions if not trillions of times faster than the human mind. So, what are the skills, capabilities, and experiences that our students and citizens will need to prosper in an age where personal and professional success will be based on the application, not the memorization and regurgitation, of knowledge? Let's continue that conversation here in Part II to define the requirements for humans to excel in creating organizational and societal value in a world dominated by AI and Big Data. Many organizations engage in a "wear'em down" decision-making process when dealing with wicked hard challenges with multiple opposing views.
Artificial Intelligence & Higher Ed: What Lies Ahead? - Higher Education Digest
Opportunities notwithstanding, there are many outstanding questions and potential downsides to the use of AI in teaching and learning that we must also attend to. These are related to the call for the pause on AI training by the likes of Elon Musk and Steve Wozniak. If a tool like ChatGPT formulates a poem, for example, who owns it--the machine or the human who published it? Further, because most of these AI tools are trained only on information openly available on the internet, the chance of spreading inaccurate information, even disinformation, is very high. Rampant bias in generated responses has been reported because these systems cannot access information located behind a login or paywall, such as peer-reviewed materials and textbooks written by vetted and more objective experts.
How Much Can Duolingo Teach Us?
In the fall of 2000, as the first dot-com bubble was bursting, the Guatemalan computer scientist Luis von Ahn attended a talk, at Carnegie Mellon, about ten problems that Yahoo couldn't solve. Von Ahn, who had just begun his Ph.D., liked solving problems. He had planned to study math until he realized that many mathematicians were still toiling away over questions that had proved unanswerable for centuries. "I talked to some computer-science professors and they would say, 'Oh, yeah, I solved an open problem last week,' " he told me recently. "That seemed just a lot more interesting."