Goto

Collaborating Authors

 teacher and student


Napster Is Back, and It Wants to Digitally Clone Teachers

WIRED

Once the music industry's biggest headache, Napster's next act is bringing AI to the classroom. The name that struck fear into the hearts of music executives in the early dotcom era may be coming to a classroom near you. In Dubai, United Arab Emirates, it has signed a strategic partnership with Gems Education to build AI agents and digital personas for education. In doing so, the once-notorious music-sharing company is bringing its technology into the classroom and testing how AI can augment the work of both teachers and students. Today's Napster, of course, bears no resemblance to its original incarnation as a peer-to-peer music-file-sharing service.


What We Lost When Big Tech Took Over the Classroom

The New Yorker

Natasha Singer's book "Coding Kids" tracks how public schools were bewitched by private corporations, setting the stage for the ed-tech backlash of today. Over Labor Day weekend, Springfield Public Schools, a district in western Massachusetts that serves twenty-three thousand students across some sixty schools, fell prey to a cyberattack that shuttered classes for four days. Staff members lost access to students' contact information, transcripts, and medical records; internal phone lines, security-camera systems, and even copy machines went down. Teachers and students couldn't use their school-issued laptops or log into their e-mail or other school applications from home devices. Springfield's schools use the Microsoft 365 Education and PowerSchool software platforms to house most of the stuff of everyday learning--schedules, assignments, slide decks, readings, handouts, quizzes--but everything stored there was now out of reach.


Supplementary Material MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel Maps

Neural Information Processing Systems

Specifically, robustness with only ACM loss is 48.38%, the addition of soft-labels improves it to 49.53%, the addition of mixup improves it to 52.29%, and the addition of both of these components make final robustness to 56.65%. Also, note that only soft labels are not enough to transfer robustness in this case, as shown by KDOnly column. This is in line with the observations of Goldblum et al. [4]. A.4.2 Role of Intermediate Features To understand the role of low, mid, and high-level features, we performed experiments on CIFAR-10 by progressively changing blocks used for distillation. For this ablation study, we kept all the standard settings reported in the Section A.1. Our correspondence of blocks and features is as follows: block 2: low-level features; block 3: mid-level features; block 4: high-level features. Please note that block 1 corresponds to the output of the first layer only. Therefore, we do not call it low-level features.



Structural Knowledge Distillation for Object Detection

Neural Information Processing Systems

Knowledge Distillation (KD) is a well-known training paradigm in deep neural networks where knowledge acquired by a large teacher model is transferred to a small student. KD has proven to be an effective technique to significantly improve the student's performance for various tasks including object detection. As such, KD techniques mostly rely on guidance at the intermediate feature level, which is typically implemented by minimizing an โ„“p-norm distance between teacher and student activations during training. In this paper, we propose a replacement for the pixel-wise independent โ„“p-norm based on the structural similarity (SSIM) [28]. By taking into account additional contrast and structural cues, feature importance, correlation and spatial dependence in the feature space are considered in the loss formulation. Extensive experiments on MSCOCO [16] demonstrate the effectiveness of our method across different training schemes and architectures. Our method adds only little computational overhead, is straightforward to implement and at the same time it significantly outperforms the standard โ„“p-norms. Moreover, more complex state-of-the-art KD methods [13, 33] using attention-based sampling mechanisms are outperformed, including a +3.5 AP gain using a Faster R-CNN R-50 [21] compared to a vanilla model.