madi
She was accused of faking an incriminating video of teenage cheerleaders. She was arrested, outcast and condemned. The problem? Nothing was fake after all
Madi Hime is taking a deep drag on a blue vape in the video, her eyes shut, her face flushed with pleasure. The 16-year-old exhales with her head thrown back, collapsing into laughter that causes smoke to billow out of her mouth. The clip is grainy and shaky – as if shot in low light by someone who had zoomed in on Madi's face – but it was damning. Madi was a cheerleader with the Victory Vipers, a highly competitive "all-star" squad based in Doylestown, Pennsylvania. The Vipers had a strict code of conduct; being caught partying and vaping could have got her thrown out of the team. And in July 2020, an anonymous person sent the incriminating video directly to Madi's coaches. Eight months later, that footage was the subject of a police news conference. "The police reviewed the video and other photographic images and found them to be what we now know to be called deepfakes," district attorney Matt Weintraub told the assembled journalists at the Bucks County courthouse on 15 March 2021. Someone was deploying cutting-edge technology to tarnish a teenage cheerleader's reputation. The vaping video was just one of many disturbing communications brought to the attention of Hilltown Township police department, Weintraub said. Madi had been receiving messages telling her she should kill herself. Her mother, Jennifer Hime, had told officers someone had been taking images from Madi's social media and manipulating them "to make her appear to be drinking".
MaDi: Learning to Mask Distractions for Generalization in Visual Deep Reinforcement Learning
Grooten, Bram, Tomilin, Tristan, Vasan, Gautham, Taylor, Matthew E., Mahmood, A. Rupam, Fang, Meng, Pechenizkiy, Mykola, Mocanu, Decebal Constantin
The visual world provides an abundance of information, but many input pixels received by agents often contain distracting stimuli. Autonomous agents need the ability to distinguish useful information from task-irrelevant perceptions, enabling them to generalize to unseen environments with new distractions. Existing works approach this problem using data augmentation or large auxiliary networks with additional loss functions. We introduce MaDi, a novel algorithm that learns to mask distractions by the reward signal only. In MaDi, the conventional actor-critic structure of deep reinforcement learning agents is complemented by a small third sibling, the Masker. This lightweight neural network generates a mask to determine what the actor and critic will receive, such that they can focus on learning the task. The masks are created dynamically, depending on the current input. We run experiments on the DeepMind Control Generalization Benchmark, the Distracting Control Suite, and a real UR5 Robotic Arm. Our algorithm improves the agent's focus with useful masks, while its efficient Masker network only adds 0.2% more parameters to the original structure, in contrast to previous work. MaDi consistently achieves generalization results better than or competitive to state-of-the-art methods.
MADI: Inter-domain Matching and Intra-domain Discrimination for Cross-domain Speech Recognition
Zhou, Jiaming, Zhao, Shiwan, Jiang, Ning, Zhao, Guoqing, Qin, Yong
End-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target domain. To improve transferability, existing UDA approaches mainly focus on matching the distributions of the source and target domains globally and/or locally, while ignoring the model discriminability. In this paper, we propose a novel UDA approach for ASR via inter-domain MAtching and intra-domain DIscrimination (MADI), which improves the model transferability by fine-grained inter-domain matching and discriminability by intra-domain contrastive discrimination simultaneously. Evaluations on the Libri-Adapt dataset demonstrate the effectiveness of our approach. MADI reduces the relative word error rate (WER) on cross-device and cross-environment ASR by 17.7% and 22.8%, respectively.