Education
Attention Based Transformer for Student Answers Assessment
Khayi, Nisrine Ait (Institute for Intelligent Systems and the University of Memphis) | Rus, Vasile (Institute for Intelligent Systems and the University of Memphis )
Inspired by Vaswani’s transformer, we propose in this paper an attention-based transformer neural network with a multi-head attention mechanism for the task of student answer assessment. Results show the competitiveness of our proposed model. A highest accuracy of 71.5% was achieved when using ELMo embeddings, 10 heads of attention, and 2 layers. This is very competitive and rivals the highest accuracy achieved by a previously proposed BI-GRU-Capsnet deep network (72.5%) on the same dataset. The main advantages of using transformers over BI-GRU-Capsnet is reducing the training time and giving more space for parallelization.
Experiments with a Socratic Intelligent Tutoring System for Source Code Understanding
Alshaikh, Zeyad (University of Memphis ) | Tamang, Lasang (University of Memphis) | Rus, Vasile (University of Memphis)
Computer Science (CS) education is critical in todays world, and introductory programming courses are considered extremely difficult and frustrating, often considered a major stumbling block for students willing to pursue computer programming related careers. In this paper, we describe the design of Socratic Tutor, an Intelligent Tutoring System that can help novice programmers to better understand programming concepts. The system was inspired by the Socratic method of teaching in which the main goal is to ask a set of guiding questions about key concepts and major steps or segments of complete code examples. To evaluate the Socratic Tutor, we conducted a pilot study with 34 computer science students and the results are promising in terms of learning gains.
ByteDance Launches New App for AI English Learning Aimed at Beginners
ByteDdance's subsidiary Beijing Diandiankankan Technology announced yesterday to release a new AI English learning App named KaiYanJianDanXue(开言简单学, literally translated as Open Language Easy Learning), which is regarded as the beginner-friendly version of Open Language. According to the introduction of the product in the App Store, the functions of this new product mainly include providing scenario learning videos and online courses from North American teachers, improving pronunciation via AI technology, and offering learners individualized learning and reviewing plans. Zhang Yiming, the founder and CEO of ByteDance, regards that the combination with technology will be an inevitable trend in future education sector. From 2017 onwards, ByteDance started to launch educational products in succession such as Learning app Haohao Xuexi (means study well), online English learning platforms GoGoKid and aiKID, English learning app Tangyuan English, and AI English learning product for children from 2 to 8-year-old named GuaGuaLong.
UF Becomes First U.S. University to Acquire Cutting-Edge NVIDIA DGX A100 System, Advancing Its Artificial Intelligence Initiative
Moving forward on its sweeping vision to transform the future of education through artificial intelligence (AI), the University of Florida is collaborating with technology company NVIDIA to acquire the world's most advanced AI system to boost the performance of UF's powerful supercomputer. UF today announced it will be the country's first higher education institution to acquire the new NVIDIA DGX A100 -- the world's most advanced AI system. The new systems are scheduled to arrive at UF at the end of May and mark a significant step in the university's bold initiative to become a national leader in the application of AI, an expansive plan that will elevate UF in research, teaching, and economic development. The initiative includes a commitment from UF to hire 100 faculty members specifically focused on AI, in addition to the 500 new faculty hired across disciplines -- many of whom will integrate AI into their teaching and research. UF is also embarking on a unique plan to infuse AI across academic majors, creating a next generation AI-enabled workforce and democratizing a technology that has the potential to solve some of the globe's most formidable challenges.
Learning Computer Vision Technology and Applications from #EmergingTechnologies Leaders
Computer vision is an Artificial Intelligence technology that allows computers to understand and label images, is now used in convenience stores, driverless car testing, daily medical diagnostics and in monitoring the health of crops and livestock. In this session, we will have Dr. Nicholas Nicoloudis and Abinaya Seenivasan from @SAP along with Aruna Kolluru from @Dell Technologies to discuss various applications of Computer Vision. The content is as below: 1. What is Computer Vision? 2. Technologies involved, including Machine Learning techniques. Please join me and feel free to ask questions.
Quantum-Classical Machine learning by Hybrid Tensor Networks
Liu, Ding, Yao, Zekun, Zhang, Quan
Tensor networks (TN) have found a wide use in machine learning, and in particular, TN and deep learning bear striking similarities. In this work, we propose the quantum-classical hybrid tensor networks (HTN) which combine tensor networks with classical neural networks in a uniform deep learning framework to overcome the limitations of regular tensor networks in machine learning. We first analyze the limitations of regular tensor networks in the applications of machine learning involving the representation power and architecture scalability. We conclude that in fact the regular tensor networks are not competent to be the basic building blocks of deep learning. Then, we discuss the performance of HTN which overcome all the deficiency of regular tensor networks for machine learning. In this sense, we are able to train HTN in the deep learning way which is the standard combination of algorithms such as Back Propagation and Stochastic Gradient Descent. We finally provide two applicable cases to show the potential applications of HTN, including quantum states classification and quantum-classical autoencoder. These cases also demonstrate the great potentiality to design various HTN in deep learning way.
Stopping criterion for active learning based on deterministic generalization bounds
Ishibashi, Hideaki, Hino, Hideitsu
Active learning is a framework in which the learning machine can select the samples to be used for training. This technique is promising, particularly when the cost of data acquisition and labeling is high. In active learning, determining the timing at which learning should be stopped is a critical issue. In this study, we propose a criterion for automatically stopping active learning. The proposed stopping criterion is based on the difference in the expected generalization errors and hypothesis testing. We derive a novel upper bound for the difference in expected generalization errors before and after obtaining a new training datum based on PAC-Bayesian theory. Unlike ordinary PAC-Bayesian bounds, though, the proposed bound is deterministic; hence, there is no uncontrollable trade-off between the confidence and tightness of the inequality. We combine the upper bound with a statistical test to derive a stopping criterion for active learning. We demonstrate the effectiveness of the proposed method via experiments with both artificial and real datasets.
Joint Progressive Knowledge Distillation and Unsupervised Domain Adaptation
Nguyen-Meidine, Le Thanh, Granger, Eric, Kiran, Madhu, Dolz, Jose, Blais-Morin, Louis-Antoine
Currently, the divergence in distributions of design and operational data, and large computational complexity are limiting factors in the adoption of CNNs in real-world applications. For instance, person re-identification systems typically rely on a distributed set of cameras, where each camera has different capture conditions. This can translate to a considerable shift between source (e.g. lab setting) and target (e.g. operational camera) domains. Given the cost of annotating image data captured for fine-tuning in each target domain, unsupervised domain adaptation (UDA) has become a popular approach to adapt CNNs. Moreover, state-of-the-art deep learning models that provide a high level of accuracy often rely on architectures that are too complex for real-time applications. Although several compression and UDA approaches have recently been proposed to overcome these limitations, they do not allow optimizing a CNN to simultaneously address both. In this paper, we propose an unexplored direction -- the joint optimization of CNNs to provide a compressed model that is adapted to perform well for a given target domain. In particular, the proposed approach performs unsupervised knowledge distillation (KD) from a complex teacher model to a compact student model, by leveraging both source and target data. It also improves upon existing UDA techniques by progressively teaching the student about domain-invariant features, instead of directly adapting a compact model on target domain data. Our method is compared against state-of-the-art compression and UDA techniques, using two popular classification datasets for UDA -- Office31 and ImageClef-DA. In both datasets, results indicate that our method can achieve the highest level of accuracy while requiring a comparable or lower time complexity.
Learning Composable Energy Surrogates for PDE Order Reduction
Beatson, Alex, Ash, Jordan T., Roeder, Geoffrey, Xue, Tianju, Adams, Ryan P.
Meta-materials are an important emerging class of engineered materials in which complex macroscopic behaviour--whether electromagnetic, thermal, or mechanical--arises from modular substructure. Simulation and optimization of these materials are computationally challenging, as rich substructures necessitate high-fidelity finite element meshes to solve the governing PDEs. To address this, we leverage parametric modular structure to learn component-level surrogates, enabling cheaper high-fidelity simulation. We use a neural network to model the stored potential energy in a component given boundary conditions. This yields a structured prediction task: macroscopic behavior is determined by the minimizer of the system's total potential energy, which can be approximated by composing these surrogate models. Composable energy surrogates thus permit simulation in the reduced basis of component boundaries. Costly ground-truth simulation of the full structure is avoided, as training data are generated by performing finite element analysis with individual components. Using dataset aggregation to choose training boundary conditions allows us to learn energy surrogates which produce accurate macroscopic behavior when composed, accelerating simulation of parametric meta-materials.
Automatic Dialogic Instruction Detection for K-12 Online One-on-one Classes
Xu, Shiting, Ding, Wenbiao, Liu, Zitao
Online one-on-one class is created for highly interactive and immersive learning experience. It demands a large number of qualified online instructors. In this work, we develop six dialogic instructions and help teachers achieve the benefits of one-on-one learning paradigm. Moreover, we utilize neural language models, i.e., long short-term memory (LSTM), to detect above six instructions automatically. Experiments demonstrate that the LSTM approach achieves AUC scores from 0.840 to 0.979 among all six types of instructions on our real-world educational dataset.