Media
Artist uses AI to 'resurrect' stars like Diana, John Lennon and Kurt Cobain who left us too soon
A photographer used artificial intelligence to bring stars who left us too soon back to life - creating eerie portraits of Princess Diana, Kurt Cobain, John Lennon, Janis Joplin, Freddie Mercury and others. The haunting and realistic images are the work of Alper Yesiltas, a photographer based in Turkey, created the portraits for a project titled'As If Nothing Happened.' He used artificial intelligence photo enhancer software and photo editing programs to create the pictures. 'With the development of AI technology, I've been excited for a while, thinking that "anything imaginable can be shown in reality,"' Yesiltas wrote about the project. The haunting and realistic images are the work of Alper Yesiltas, a photographer based in Turkey, created the portraits for a project titled'As If Nothing Happened.' 'With the development of AI technology, I've been excited for a while, thinking that "anything imaginable can be shown in reality,"' Yesiltas wrote about the project.
Technical Perspective: Traffic Classification in the Era of Deep Learning
Network traffic classification is a fundamental problem in networking. Given observations of network traffic, the goal is to infer properties of interest, such as what application generated the traffic. This enables network operators to monitor and optimize performance, detect anomalies or malware, block unwanted traffic, inform capacity planning, and so on. The problem has been extensively studied for more than 20 years, using a combination of heuristics, based on domain expertise, and automated methodologies. Some techniques rely on hard-coded rules, such as the use of well-known ports or servers. For example, a DNS request, the HTTP Host field, or the SNI field in TLS, may all reveal the name of the server contacted (for example, server.netflix.com),
The Joy and Dread of AI Image Generators Without Limits
For the past few months, Elle Simpson-Edin, a scientist by day, has been working with her wife on a novel, due out late this year, that she describes as a "grimdark queer science fantasy." As she prepared a website to promote the book, Simpson-Edin decided to experiment with illustrating its content using one of the powerful new artificial intelligence-powered art-making tools, which can create eye-catching and even photo-real images to match a text prompt. But most of these image generators are designed to restrict what users can depict, banning pornography, violence, and pictures showing the faces of real people. Every option she tried was too prudish. "The book is quite heavy on violence and sex, so art made in an environment where blood and sex is banned isn't really an option," Simpson-Edin says.
AI and the Future of Music Creation
My father and my brothers are amateur musicians who play various instruments; thus, music runs in my family. I was raised in a household where music listening is a tradition and where kids are encouraged to sing, play, and listen to music early in our days in Brazil. My earliest memories are of music; they are the things that moved me the most and that life taught me to value the most. I put together some rock bands when I was a teenager, and even now, there is never a shortage of musical instruments at my house, both analog and now digital. In addition to my musical education, adult life led me along other routes, including those in computing and artificial intelligence: a really rich experience that allowed me to learn new languages and develop technical skills.
DDGHM: Dual Dynamic Graph with Hybrid Metric Training for Cross-Domain Sequential Recommendation
Zheng, Xiaolin, Su, Jiajie, Liu, Weiming, Chen, Chaochao
Sequential Recommendation (SR) characterizes evolving patterns of user behaviors by modeling how users transit among items. However, the short interaction sequences limit the performance of existing SR. To solve this problem, we focus on Cross-Domain Sequential Recommendation (CDSR) in this paper, which aims to leverage information from other domains to improve the sequential recommendation performance of a single domain. Solving CDSR is challenging. On the one hand, how to retain single domain preferences as well as integrate cross-domain influence remains an essential problem. On the other hand, the data sparsity problem cannot be totally solved by simply utilizing knowledge from other domains, due to the limited length of the merged sequences. To address the challenges, we propose DDGHM, a novel framework for the CDSR problem, which includes two main modules, i.e., dual dynamic graph modeling and hybrid metric training. The former captures intra-domain and inter-domain sequential transitions through dynamically constructing two-level graphs, i.e., the local graphs and the global graph, and incorporating them with a fuse attentive gating mechanism. The latter enhances user and item representations by employing hybrid metric learning, including collaborative metric for achieving alignment and contrastive metric for preserving uniformity, to further alleviate data sparsity issue and improve prediction accuracy. We conduct experiments on two benchmark datasets and the results demonstrate the effectiveness of DDHMG.
CCR: Facial Image Editing with Continuity, Consistency and Reversibility
Yang, Nan, Luan, Xin, Jia, Huidi, Han, Zhi, Tang, Yandong
Three problems exist in sequential facial image editing: incontinuous editing, inconsistent editing, and irreversible editing. Incontinuous editing is that the current editing can not retain the previously edited attributes. Inconsistent editing is that swapping the attribute editing orders can not yield the same results. Irreversible editing means that operating on a facial image is irreversible, especially in sequential facial image editing. In this work, we put forward three concepts and corresponding definitions: editing continuity, consistency, and reversibility. Then, we propose a novel model to achieve the goal of editing continuity, consistency, and reversibility. A sufficient criterion is defined to determine whether a model is continuous, consistent, and reversible. Extensive qualitative and quantitative experimental results validate our proposed model and show that a continuous, consistent and reversible editing model has a more flexible editing function while preserving facial identity. Furthermore, we think that our proposed definitions and model will have wide and promising applications in multimedia processing. Code and data are available at https://github.com/mickoluan/CCR.
Current and Near-Term AI as a Potential Existential Risk Factor
Bucknall, Benjamin S., Dori-Hacohen, Shiri
There is a substantial and ever-growing corpus of evidence and literature exploring the impacts of Artificial intelligence (AI) technologies on society, politics, and humanity as a whole. A separate, parallel body of work has explored existential risks to humanity, including but not limited to that stemming from unaligned Artificial General Intelligence (AGI). In this paper, we problematise the notion that current and near-term artificial intelligence technologies have the potential to contribute to existential risk by acting as intermediate risk factors, and that this potential is not limited to the unaligned AGI scenario. We propose the hypothesis that certain already-documented effects of AI can act as existential risk factors, magnifying the likelihood of previously identified sources of existential risk. Moreover, future developments in the coming decade hold the potential to significantly exacerbate these risk factors, even in the absence of artificial general intelligence. Our main contribution is a (non-exhaustive) exposition of potential AI risk factors and the causal relationships between them, focusing on how AI can affect power dynamics and information security. This exposition demonstrates that there exist causal pathways from AI systems to existential risks that do not presuppose hypothetical future AI capabilities.
Learning Hierarchical Metrical Structure Beyond Measures
Jiang, Junyan, Chin, Daniel, Zhang, Yixiao, Xia, Gus
Music contains hierarchical structures beyond beats and measures. While hierarchical structure annotations are helpful for music information retrieval and computer musicology, such annotations are scarce in current digital music databases. In this paper, we explore a data-driven approach to automatically extract hierarchical metrical structures from scores. We propose a new model with a Temporal Convolutional Network-Conditional Random Field (TCN-CRF) architecture. Given a symbolic music score, our model takes in an arbitrary number of voices in a beat-quantized form, and predicts a 4-level hierarchical metrical structure from downbeat-level to section-level. We also annotate a dataset using RWC-POP MIDI files to facilitate training and evaluation. We show by experiments that the proposed method performs better than the rule-based approach under different orchestration settings. We also perform some simple musicological analysis on the model predictions. All demos, datasets and pre-trained models are publicly available on Github.
Understanding Aesthetics with Language: A Photo Critique Dataset for Aesthetic Assessment
Nieto, Daniel Vera, Celona, Luigi, Fernandez-Labrador, Clara
Computational inference of aesthetics is an ill-defined task due to its subjective nature. Many datasets have been proposed to tackle the problem by providing pairs of images and aesthetic scores based on human ratings. However, humans are better at expressing their opinion, taste, and emotions by means of language rather than summarizing them in a single number. In fact, photo critiques provide much richer information as they reveal how and why users rate the aesthetics of visual stimuli. In this regard, we propose the Reddit Photo Critique Dataset (RPCD), which contains tuples of image and photo critiques. RPCD consists of 74K images and 220K comments and is collected from a Reddit community used by hobbyists and professional photographers to improve their photography skills by leveraging constructive community feedback. The proposed dataset differs from previous aesthetics datasets mainly in three aspects, namely (i) the large scale of the dataset and the extension of the comments criticizing different aspects of the image, (ii) it contains mostly UltraHD images, and (iii) it can easily be extended to new data as it is collected through an automatic pipeline. To the best of our knowledge, in this work, we propose the first attempt to estimate the aesthetic quality of visual stimuli from the critiques. To this end, we exploit the polarity of the sentiment of criticism as an indicator of aesthetic judgment. We demonstrate how sentiment polarity correlates positively with the aesthetic judgment available for two aesthetic assessment benchmarks. Finally, we experiment with several models by using the sentiment scores as a target for ranking images. Dataset and baselines are available (https://github.com/mediatechnologycenter/aestheval).
Reconstructing Robot Operations via Radio-Frequency Side-Channel
Shah, Ryan, Ahmed, Mujeeb, Nagaraja, Shishir
While active attacks can be deadly to the Connected teleoperated robotic systems play a key role in ensuring operating environment and subject(s) involved, passive attacks can operational workflows are carried out with high levels of accuracy result in huge losses that stem from stealthy, unintentional information and low margins of error. In recent years, a variety of attacks have leakage. For example, if an attacker is able to identify what been proposed that actively target the robot itself from the cyber workflows a robot is carrying out, such as the movement of packages domain. However, little attention has been paid to the capabilities of in a warehouse between belts, they could use this information a passive attacker. In this work, we investigate whether an insider to sell on to competitors that can understand how competing warehousing adversary can accurately fingerprint robot movements and operational facilities operate and use this information to a malicious warehousing workflows via the radio frequency side channel advantage [21, 27]. in a stealthy manner. Using an SVM for classification, we found In this work we seek to explore other mechanisms to passively that an adversary can fingerprint individual robot movements with learn about robotic workflows. Side channels have previously been at least 96% accuracy, increasing to near perfect accuracy when used in different technological domains as a means to learn sensitive reconstructing entire warehousing workflows.