Media
Self-supervised Auxiliary Loss for Metric Learning in Music Similarity-based Retrieval and Auto-tagging
Akama, Taketo, Kitano, Hiroaki, Takematsu, Katsuhiro, Miyajima, Yasushi, Polouliakh, Natalia
In the realm of music information retrieval, similarity-based retrieval and auto-tagging serve as essential components. Given the limitations and non-scalability of human supervision signals, it becomes crucial for models to learn from alternative sources to enhance their performance. Self-supervised learning, which exclusively relies on learning signals derived from music audio data, has demonstrated its efficacy in the context of auto-tagging. In this study, we propose a model that builds on the self-supervised learning approach to address the similarity-based retrieval challenge by introducing our method of metric learning with a self-supervised auxiliary loss. Furthermore, diverging from conventional self-supervised learning methodologies, we discovered the advantages of concurrently training the model with both self-supervision and supervision signals, without freezing pre-trained models. We also found that refraining from employing augmentation during the fine-tuning phase yields better results. Our experimental results confirm that the proposed methodology enhances retrieval and tagging performance metrics in two distinct scenarios: one where human-annotated tags are consistently available for all music tracks, and another where such tags are accessible only for a subset of tracks.
Keeping the Questions Conversational: Using Structured Representations to Resolve Dependency in Conversational Question Answering
Zaib, Munazza, Sheng, Quan Z., Zhang, Wei Emma, Mahmood, Adnan
Having an intelligent dialogue agent that can engage in conversational question answering (ConvQA) is now no longer limited to Sci-Fi movies only and has, in fact, turned into a reality. These intelligent agents are required to understand and correctly interpret the sequential turns provided as the context of the given question. However, these sequential questions are sometimes left implicit and thus require the resolution of some natural language phenomena such as anaphora and ellipsis. The task of question rewriting has the potential to address the challenges of resolving dependencies amongst the contextual turns by transforming them into intent-explicit questions. Nonetheless, the solution of rewriting the implicit questions comes with some potential challenges such as resulting in verbose questions and taking conversational aspect out of the scenario by generating self-contained questions. In this paper, we propose a novel framework, CONVSR (CONVQA using Structured Representations) for capturing and generating intermediate representations as conversational cues to enhance the capability of the QA model to better interpret the incomplete questions. We also deliberate how the strengths of this task could be leveraged in a bid to design more engaging and eloquent conversational agents. We test our model on the QuAC and CANARD datasets and illustrate by experimental results that our proposed framework achieves a better F1 score than the standard question rewriting model.
TimelyFL: Heterogeneity-aware Asynchronous Federated Learning with Adaptive Partial Training
Zhang, Tuo, Gao, Lei, Lee, Sunwoo, Zhang, Mi, Avestimehr, Salman
In cross-device Federated Learning (FL) environments, scaling synchronous FL methods is challenging as stragglers hinder the training process. Moreover, the availability of each client to join the training is highly variable over time due to system heterogeneities and intermittent connectivity. Recent asynchronous FL methods (e.g., FedBuff) have been proposed to overcome these issues by allowing slower users to continue their work on local training based on stale models and to contribute to aggregation when ready. However, we show empirically that this method can lead to a substantial drop in training accuracy as well as a slower convergence rate. The primary reason is that fast-speed devices contribute to many more rounds of aggregation while others join more intermittently or not at all, and with stale model updates. To overcome this barrier, we propose TimelyFL, a heterogeneity-aware asynchronous FL framework with adaptive partial training. During the training, TimelyFL adjusts the local training workload based on the real-time resource capabilities of each client, aiming to allow more available clients to join in the global update without staleness. We demonstrate the performance benefits of TimelyFL by conducting extensive experiments on various datasets (e.g., CIFAR-10, Google Speech, and Reddit) and models (e.g., ResNet20, VGG11, and ALBERT). In comparison with the state-of-the-art (i.e., FedBuff), our evaluations reveal that TimelyFL improves participation rate by 21.13%, harvests 1.28x - 2.89x more efficiency on convergence rate, and provides a 6.25% increment on test accuracy.
Delta Denoising Score
Hertz, Amir, Aberman, Kfir, Cohen-Or, Daniel
We introduce Delta Denoising Score (DDS), a novel scoring function for text-based image editing that guides minimal modifications of an input image towards the content described in a target prompt. DDS leverages the rich generative prior of text-to-image diffusion models and can be used as a loss term in an optimization problem to steer an image towards a desired direction dictated by a text. DDS utilizes the Score Distillation Sampling (SDS) mechanism for the purpose of image editing. We show that using only SDS often produces non-detailed and blurry outputs due to noisy gradients. To address this issue, DDS uses a prompt that matches the input image to identify and remove undesired erroneous directions of SDS. Our key premise is that SDS should be zero when calculated on pairs of matched prompts and images, meaning that if the score is non-zero, its gradients can be attributed to the erroneous component of SDS. Our analysis demonstrates the competence of DDS for text based image-to-image translation. We further show that DDS can be used to train an effective zero-shot image translation model. Experimental results indicate that DDS outperforms existing methods in terms of stability and quality, highlighting its potential for real-world applications in text-based image editing.
A Framework for Fast Prototyping of Photo-realistic Environments with Multiple Pedestrians
Casao, Sara, Otero, Andrรฉs, Serra-Gรณmez, รlvaro, Murillo, Ana C., Alonso-Mora, Javier, Montijano, Eduardo
Robotic applications involving people often require advanced perception systems to better understand complex real-world scenarios. To address this challenge, photo-realistic and physics simulators are gaining popularity as a means of generating accurate data labeling and designing scenarios for evaluating generalization capabilities, e.g., lighting changes, camera movements or different weather conditions. We develop a photo-realistic framework built on Unreal Engine and AirSim to generate easily scenarios with pedestrians and mobile robots. The framework is capable to generate random and customized trajectories for each person and provides up to 50 ready-to-use people models along with an API for their metadata retrieval. We demonstrate the usefulness of the proposed framework with a use case of multi-target tracking, a popular problem in real pedestrian scenarios. The notable feature variability in the obtained perception data is presented and evaluated.
Dramatic video captures hammerhead going after group of sharks
Researchers captured drone footage of blacktip sharks evading a 12-foot-long hammerhead shark in Florida. Researchers have captured dramatic drone footage of blacktip sharks quickly evading a 12-foot-long hammerhead shark by swimming into shallow waters off Florida's coast. The drone footage captured by researchers with Florida's Atlantic University is the first evidence of large adult sharks using the shallows to flee predators. In the dramatic footage, the adult blacktop sharks are seen fleeing for shallow waters when faced with the hammerhead. The hammerhead shark was caught approaching the smaller blacktip sharks off the coast of Florida.
Can artificial intelligence predict the weather months out? This company says it can
FOX Business correspondent Lydia Hu has the latest on jobs at risk as AI further develops on "America's Newsroom." Artificial intelligence is being used and introduced across all sectors, aiding the research of oncologists and NASA scientists. Algorithms and machine-learning models, like the newly popular ChatGPT and Google's Bard, have helped students and professionals โ although the technology comes with a warning as governments around the world rush to devise regulations and standards. The potential of the industry and AI may appear to be boundless at this phase, with new research and tools publicly announced every week. Just days after the Biden administration called for public input on proposed artificial intelligence policies, tropical cyclones are already a topic of discussion.
DJI Encourage 3 Cinema Drone - Channel969
Immediately, drone and digicam know-how chief DJI introduced the DJI Encourage 3, a full-frame 8K cinema drone designed for top-level film productions. Its built-in design contains a 161 ultra-wide FOV night-vision FPV and the O3 Professional transmission and management system. DJI's first and solely cinema-grade drone, the Encourage 3 helps each RTK-powered Waypoint Professional and omnidirectional sensing to conduct safer and extra correct flight missions. "The Encourage 3 is the professional-level aerial platform all filmmakers have been ready for," mentioned DJI Artistic Director Ferdinand Wolf. "It empowers customers to totally maximize the potential of any shot as they'll report in cinematic-grade picture high quality beforehand solely out there with massive and clunky digicam programs. We're wanting ahead to seeing how the Encourage 3 will push aerial cinematography to a totally new stage."
I Cloned My Voice and My Mother Couldn't Tell the Difference
This article is from Understanding AI, a newsletter that explores how A.I. works and how it's changing our world. A couple of weeks ago, I used A.I. software to clone my voice. The resulting audio sounded pretty convincing to me, but I wanted to see what others thought. So I created a test audio file based on the first 12 paragraphs of this article that I wrote. Seven randomly chosen paragraphs were my real voice, while the other five were generated by A.I. I asked members of my family to see if they could tell the difference.