Media
UNSW researcher receives award recognising women in artificial intelligence
UNSW Engineering Professor Flora Salim has been honoured for her pioneering work in computing and machine learning by Women in AI, a global advocacy group for women in the artificial intelligence (AI) field. The 2022 Women in AI Awards Australia and New Zealand recognised women across various industries committed to excellence in AI. Finalists were judged on innovation, leadership and inspiring potential, global impact, and the ability of the AI solution to provide a social good for the community. Prof. Salim was recognised for her AI achievements in the Defence and Intelligence award category. The award acknowledged her research in the cross-cutting areas of ubiquitous computing and machine learning, with a focus on efficient, fair, and explainable machine learning for multi-dimensional sensor data, towards enabling situational and behaviour intelligence for multiple applications.
Leveraging Natural Supervision for Language Representation Learning and Generation
Recent breakthroughs in Natural Language Processing (NLP) have been driven by language models trained on a massive amount of plain text. While powerful, deriving supervision from textual resources is still an open question. For example, language model pretraining often neglects the rich, freely-available structures in textual data. In this thesis, we describe three lines of work that seek to improve the training and evaluation of neural models using naturally-occurring supervision. We first investigate self-supervised training losses to help enhance the performance of pretrained language models for various NLP tasks. Specifically, we alter the sentence prediction loss to make it better suited to other pretraining losses and more challenging to solve. We design an intermediate finetuning step that uses self-supervised training to promote models' ability in cross-task generalization. Then we describe methods to leverage the structures in Wikipedia and paraphrases. In particular, we propose training losses to exploit hyperlinks, article structures, and article category graphs for entity-, discourse-, entailment-related knowledge. We propose a framework that uses paraphrase pairs to disentangle semantics and syntax in sentence representations. We extend the framework for a novel generation task that controls the syntax of output text with a sentential exemplar. Lastly, we discuss our work on tailoring textual resources for establishing challenging evaluation tasks. We introduce three datasets by defining novel tasks using various fan-contributed websites, including a long-form data-to-text generation dataset, a screenplay summarization dataset, and a long-form story generation dataset. These datasets have unique characteristics offering challenges to future work in their respective task settings.
A Proposal for Foley Sound Synthesis Challenge
Choi, Keunwoo, Oh, Sangshin, Kang, Minsung, McFee, Brian
We during post-production to enhance its perceived acoustic properties, review recent machine learning challenges in audio, speech, and e.g., by simulating the sounds of footsteps, ambient environmental music research in Section 2 and existing works and datasets in Section sounds, or visible objects on the screen. While foley is traditionally 3. In Section 4, we provide a proposal for foley sound synthesis produced by foley artists, there is increasing interest in automatic challenge that includes problem definition, datasets, and evaluation or machine-assisted techniques building upon recent advances in metrics. We conclude the paper in Section 5. sound synthesis and generative models. To foster more participation in this growing research area, we propose a challenge for automatic 2. CASE STUDY: RESEARCH CHALLENGES foley synthesis. Through case studies on successful previous challenges in audio and machine learning, we set the goals of In this section, we review five existing research challenges: Blizzard the proposed challenge: rigorous, unified, and efficient evaluation Challenge, CHiME, DCASE, Music Demixing challenge, and of different foley synthesis systems, with an overarching goal of AI Song Contest. The former three are relatively mature while the drawing active participation from the research community. We outline latter two started after 2020. All of them started along with the increasing the details and design considerations of a foley sound synthesis popularity of the research problems and have contributed challenge, including task definition, dataset requirements, and evaluation to the continued growth by defining the tasks, providing common criteria.
EdiBERT, a generative model for image editing
Issenhuth, Thibaut, Tanielian, Ugo, Mary, Jérémie, Picard, David
Advances in computer vision are pushing the limits of im-age manipulation, with generative models sampling detailed images on various tasks. However, a specialized model is often developed and trained for each specific task, even though many image edition tasks share similarities. In denoising, inpainting, or image compositing, one always aims at generating a realistic image from a low-quality one. In this paper, we aim at making a step towards a unified approach for image editing. To do so, we propose EdiBERT, a bi-directional transformer trained in the discrete latent space built by a vector-quantized auto-encoder. We argue that such a bidirectional model is suited for image manipulation since any patch can be re-sampled conditionally to the whole image. Using this unique and straightforward training objective, we show that the resulting model matches state-of-the-art performances on a wide variety of tasks: image denoising, image completion, and image composition.
Learning Unsupervised Hierarchies of Audio Concepts
Afchar, Darius, Hennequin, Romain, Guigue, Vincent
Music signals are difficult to interpret from their low-level features, perhaps even more than images: e.g. highlighting part of a spectrogram or an image is often insufficient to convey high-level ideas that are genuinely relevant to humans. In computer vision, concept learning was therein proposed to adjust explanations to the right abstraction level (e.g. detect clinical concepts from radiographs). These methods have yet to be used for MIR. In this paper, we adapt concept learning to the realm of music, with its particularities. For instance, music concepts are typically non-independent and of mixed nature (e.g. genre, instruments, mood), unlike previous work that assumed disentangled concepts. We propose a method to learn numerous music concepts from audio and then automatically hierarchise them to expose their mutual relationships. We conduct experiments on datasets of playlists from a music streaming service, serving as a few annotated examples for diverse concepts. Evaluations show that the mined hierarchies are aligned with both ground-truth hierarchies of concepts -- when available -- and with proxy sources of concept similarity in the general case.
Causal Machine Learning: A Survey and Open Problems
Kaddour, Jean, Lynch, Aengus, Liu, Qi, Kusner, Matt J., Silva, Ricardo
Causal Machine Learning (CausalML) is an umbrella term for machine learning methods that formalize the data-generation process as a structural causal model (SCM). This perspective enables us to reason about the effects of changes to this process (interventions) and what would have happened in hindsight (counterfactuals). We categorize work in CausalML into five groups according to the problems they address: (1) causal supervised learning, (2) causal generative modeling, (3) causal explanations, (4) causal fairness, and (5) causal reinforcement learning. We systematically compare the methods in each category and point out open problems. Further, we review data-modality-specific applications in computer vision, natural language processing, and graph representation learning. Finally, we provide an overview of causal benchmarks and a critical discussion of the state of this nascent field, including recommendations for future work.
Surreal or too real? Breathtaking AI tool DALL-E takes its images to a bigger stage
DALL-E2, the AI image tool, generated these images of a giraffe shopping in a grocery store. When the Silicon Valley research lab OpenAI unveiled DALL-E earlier this year, it wowed the internet. The tool is seen as one of the most advanced artificial intelligence systems for creating images in the world. Type a description, and DALL-E instantly produces professional-looking art or hyperrealistic photographs. "It's incredibly powerful," said Hany Farid, a digital forensics expert at the University of California, Berkeley.
The 50 Greatest Fictional Deaths of All Time
"It is a far, far better thing that I do, than I have ever done," Sydney Carton thinks on his way to the guillotine. That far better thing is dying tragically, for many reasons: to save an innocent man, to fulfill his own redemption, and--of course--to make us cry at the end of A Tale of Two Cities. The death scene is one of the sharpest tools in a writer's toolbox, as likely to wound the writer themself as the reader--for if a well-written death scene can be thrilling, terrifying, or filled with despair, so can a poorly written one be bathetic, stupid, and eye-rolling. But let's not talk about those. Let's talk about the good ones, the deathless death scenes. We've assembled the 50 greatest fictional deaths of all time--the most moving, most funny, most shocking, most influential scenes from books, movies, TV, theater, video games, and more. Spoilers abound: It's a list that spans nearly 2,500 years of human culture, from Athens to A24, and is so competitive that even poor Sydney Carton and his famous last words couldn't make it. We've also talked to many of the creators behind the scenes on our list to ask them how they wrote them, why they killed off characters we loved, what makes a great death scene, and what final moments from fiction have stuck with them all their lives. We've made this list during a pandemic, as real-life death has stalked us all, more tangible than ever. After all, one of the many things art can do is to help us navigate the pitfalls of life, and there's no deeper pitfall than the final one. Here are the scenes that have shown us all what the big goodbye might actually be like, when it comes. Imagine Imagine the horror in Athens' Theatre of Dionysus at the premiere of Medea, as the audience heard the desperate cries of Medea's two sons while she ruthlessly stabbed them to death.
Aboriginal language could help solve complex AI problems
Jingulu – a language spoken by the Jingili people in the Northern Territory – has characteristics that allow it to be easily translated into AI commands. A new language inspired by Jingulu could be applied to any situation where communication between humans and a large number of AI agents is required. An Aboriginal language could hold the key to solving some of the most challenging communication problems between humans and artificial intelligence (AI) systems. A new paper, published by Frontiers in Physics and led by UNSW Canberra's Professor Hussein Abbass, explains how Jingulu – a language spoken by the Jingili people in the Northern Territory – has characteristics that allow it to be easily translated into AI commands. "The Aboriginal people have a long history of contributions to the defence of Australia," Professor Abbass said.