Media
NCCR Robotics: A documentary
This short film documents some of the most innovative projects that emerged from the work of NCCR Robotics, the Swiss-wide consortium coordinated from 2010 to 2022 by EPFL professor Dario Floreano and ETHZ professor Robert Riener, including other major research institutions across Switzerland. Shot over the course of six months in Lausanne, Geneva, Zurich, Wangen an der Aare, Leysin, Lugano, the documentary is a unique look at the state of the art of medical, educational and rescue robotics, and at the specific contributions that Swiss researchers have given to the field over the last decade. In addition to showing the robots in action, the film features extended interviews with top experts including Stéphanie Lacour, Silvestro Micera, Davide Scaramuzza, Robert Riener, Pierre Dillenbourg, Margarita Chli, Dario Floreano.
Qualitative Analysis of a Graph Transformer Approach to Addressing Hate Speech: Adapting to Dynamically Changing Content
Hebert, Liam, Chen, Hong Yi, Cohen, Robin, Golab, Lukasz
Our work advances an approach for predicting hate speech in social media, drawing out the critical need to consider the discussions that follow a post to successfully detect when hateful discourse may arise. Using graph transformer networks, coupled with modelling attention and BERT-level natural language processing, our approach can capture context and anticipate upcoming anti-social behaviour. In this paper, we offer a detailed qualitative analysis of this solution for hate speech detection in social networks, leading to insights into where the method has the most impressive outcomes in comparison with competitors and identifying scenarios where there are challenges to achieving ideal performance. Included is an exploration of the kinds of posts that permeate social media today, including the use of hateful images. This suggests avenues for extending our model to be more comprehensive. A key insight is that the focus on reasoning about the concept of context positions us well to be able to support multi-modal analysis of online posts. We conclude with a reflection on how the problem we are addressing relates especially well to the theme of dynamic change, a critical concern for all AI solutions for social impact. We also comment briefly on how mental health well-being can be advanced with our work, through curated content attuned to the extent of hate in posts.
GSim: A Graph Neural Network based Relevance Measure for Heterogeneous Graphs
Luo, Linhao, Fang, Yixiang, Lu, Moli, Cao, Xin, Zhang, Xiaofeng, Zhang, Wenjie
Heterogeneous graphs, which contain nodes and edges of multiple types, are prevalent in various domains, including bibliographic networks, social media, and knowledge graphs. As a fundamental task in analyzing heterogeneous graphs, relevance measure aims to calculate the relevance between two objects of different types, which has been used in many applications such as web search, recommendation, and community detection. Most of existing relevance measures focus on homogeneous networks where objects are of the same type, and a few measures are developed for heterogeneous graphs, but they often need the pre-defined meta-path. Defining meaningful meta-paths requires much domain knowledge, which largely limits their applications, especially on schema-rich heterogeneous graphs like knowledge graphs. Recently, the Graph Neural Network (GNN) has been widely applied in many graph mining tasks, but it has not been applied for measuring relevance yet. To address the aforementioned problems, we propose a novel GNN-based relevance measure, namely GSim. Specifically, we first theoretically analyze and show that GNN is effective for measuring the relevance of nodes in the graph. We then propose a context path-based graph neural network (CP-GNN) to automatically leverage the semantics in heterogeneous graphs. Moreover, we exploit CP-GNN to support relevance measures between two objects of any type. Extensive experiments demonstrate that GSim outperforms existing measures.
NewsPanda: Media Monitoring for Timely Conservation Action
Keh, Sedrick Scott, Shi, Zheyuan Ryan, Patterson, David J., Bhagabati, Nirmal, Dewan, Karun, Gopala, Areendran, Izquierdo, Pablo, Mallick, Debojyoti, Sharma, Ambika, Shrestha, Pooja, Fang, Fei
Non-governmental organizations for environmental conservation have a significant interest in monitoring conservation-related media and getting timely updates about infrastructure construction projects as they may cause massive impact to key conservation areas. Such monitoring, however, is difficult and time-consuming. We introduce NewsPanda, a toolkit which automatically detects and analyzes online articles related to environmental conservation and infrastructure construction. We fine-tune a BERT-based model using active learning methods and noise correction algorithms to identify articles that are relevant to conservation and infrastructure construction. For the identified articles, we perform further analysis, extracting keywords and finding potentially related sources. NewsPanda has been successfully deployed by the World Wide Fund for Nature teams in the UK, India, and Nepal since February 2022. It currently monitors over 80,000 websites and 1,074 conservation sites across India and Nepal, saving more than 30 hours of human efforts weekly. We have now scaled it up to cover 60,000 conservation sites globally.
Transfer of knowledge among instruments in automatic music transcription
Automatic music transcription (AMT) is one of the most challenging tasks in the music information retrieval domain. It is the process of converting an audio recording of music into a symbolic representation containing information about the notes, chords, and rhythm. Current research in this domain focuses on developing new models based on transformer architecture or using methods to perform semi-supervised training, which gives outstanding results, but the computational cost of training such models is enormous. This work shows how to employ easily generated synthesized audio data produced by software synthesizers to train a universal model. It is a good base for further transfer learning to quickly adapt transcription model for other instruments. Achieved results prove that using synthesized data for training may be a good base for pretraining general-purpose models, where the task of transcription is not focused on one instrument.
Blended Latent Diffusion
Avrahami, Omri, Fried, Ohad, Lischinski, Dani
The tremendous progress in neural image generation, coupled with the emergence of seemingly omnipotent vision-language models has finally enabled text-based interfaces for creating and editing images. Handling generic images requires a diverse underlying generative model, hence the latest works utilize diffusion models, which were shown to surpass GANs in terms of diversity. One major drawback of diffusion models, however, is their relatively slow inference time. In this paper, we present an accelerated solution to the task of local text-driven editing of generic images, where the desired edits are confined to a user-provided mask. Our solution leverages a recent text-to-image Latent Diffusion Model (LDM), which speeds up diffusion by operating in a lower-dimensional latent space. We first convert the LDM into a local image editor by incorporating Blended Diffusion into it. Next we propose an optimization-based solution for the inherent inability of this LDM to accurately reconstruct images. Finally, we address the scenario of performing local edits using thin masks. We evaluate our method against the available baselines both qualitatively and quantitatively and demonstrate that in addition to being faster, our method achieves better precision than the baselines while mitigating some of their artifacts.
Unions Representing Hollywood Writers and Actors Seek Limits on A.I. and Chatbots
In December, Apple introduced a service allowing book publishers to use human-sounding A.I. narrators, an innovation that could displace hundreds of voice actors who make a living performing audiobooks. The company's website says the service will benefit independent authors and small publishers. "I know someone always has to get there first, some company," said Chris Ciulla, who estimates that he has made $100,000 to $130,000 annually over the past five years narrating books under union contracts. "But for individuals not to understand how that can affect the pail-carrying narrator out there eventually is disappointing." Other actors fear that studios will use A.I. to replicate their voices while cutting them out of the process.
An Extensible Multimodal Multi-task Object Dataset with Materials
Standley, Trevor, Gao, Ruohan, Chen, Dawn, Wu, Jiajun, Savarese, Silvio
We present EMMa, an Extensible, Multimodal dataset of Amazon product listings that contains rich Material annotations. It contains more than 2.8 million objects, each with image(s), listing text, mass, price, product ratings, and position in Amazon's product-category taxonomy. Objects are annotated with one or more materials from this taxonomy. With the numerous attributes available for each object, we develop a Smart Labeling framework to quickly add new binary labels to all objects with very little manual labeling effort, making the dataset extensible. Each object attribute in our dataset can be included in either the model inputs or outputs, leading to combinatorial possibilities in task configurations. For example, we can train a model to predict the object category from the listing text, or the mass and price from the product listing image. EMMa offers a new benchmark for multi-task learning in computer vision and NLP, and allows practitioners to efficiently add new tasks and object attributes at scale. Perhaps the biggest problem faced by machine learning practitioners today is that of producing labeled datasets for their specific needs. Manually labeling large amounts of data is time-consuming and costly (Deng et al., 2009; Lin et al., 2014; Kuznetsova et al., 2020). Furthermore, it is often not possible to communicate how numerous ambiguous corner cases should be handled (e.g., is a hole puncher "sharp"?) to the human annotators we typically rely on to produce these labels. Could we solve this problem with the aid of machine learning? We hypothesized that we could accurately add new properties to every instance in a semi-automated fashion if given a rich dataset with substantial information about every instance. Consequently, we developed EMMa, a large, object-centric, multimodal, and multi-task dataset. We show that EMMa can be easily extended to contain any number of new object labels using a Smart Labeling technique we developed for large multi-task and multimodal datasets. Multi-task datasets contain labels for more than one attribute for each instance, whereas multimodal datasets contain data from more than one modality, such as images, text, audio, and tabular data. Derived from Amazon product listings, EMMa contains images, text, and a number of useful attributes, such as materials, mass, price, product category, and product ratings. Each attribute can be used as either a model input or a model output.
DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic Environments
Shen, Shihao, Cai, Yilin, Wang, Wenshan, Scherer, Sebastian
Learning-based visual odometry (VO) algorithms achieve remarkable performance on common static scenes, benefiting from high-capacity models and massive annotated data, but tend to fail in dynamic, populated environments. Semantic segmentation is largely used to discard dynamic associations before estimating camera motions but at the cost of discarding static features and is hard to scale up to unseen categories. In this paper, we leverage the mutual dependence between camera ego-motion and motion segmentation and show that both can be jointly refined in a single learning-based framework. In particular, we present DytanVO, the first supervised learning-based VO method that deals with dynamic environments. It takes two consecutive monocular frames in real-time and predicts camera ego-motion in an iterative fashion. Our method achieves an average improvement of 27.7% in ATE over state-of-the-art VO solutions in real-world dynamic environments, and even performs competitively among dynamic visual SLAM systems which optimize the trajectory on the backend. Experiments on plentiful unseen environments also demonstrate our method's generalizability.
VentureBeat is the latest publication to use AI in its articles
More media outlets are using AI to write articles, if not as aggressively as others. VentureBeat editorial director Michale Nuñez tells Bloomberg his publication is using Microsoft's Bing Chat to help edit and write stories. Reporters are encouraged to slip AI-written "sentences and fragments" into articles so long as they're accurate and independently verifiable. The OpenAI-powered tech is akin to having "another person on the team," Nuñez says. VentureBeat doesn't disclose the use of AI content provided it's limited and authentic, but also doesn't intend to create whole articles using the technology. Word surfaced in January that CNET had been using AI to produce entire financial explainer articles since November.