Media
A Perceptually Optimized and Self-Calibrated Tone Mapping Operator
Cao, Peibei, Le, Chenyang, Fang, Yuming, Ma, Kede
With the increasing popularity and accessibility of high dynamic range (HDR) photography, tone mapping operators (TMOs) for dynamic range compression are practically demanding. In this paper, we develop a two-stage neural network-based TMO that is self-calibrated and perceptually optimized. In Stage one, motivated by the physiology of the early stages of the human visual system, we first decompose an HDR image into a normalized Laplacian pyramid. We then use two lightweight deep neural networks (DNNs), taking the normalized representation as input and estimating the Laplacian pyramid of the corresponding LDR image. We optimize the tone mapping network by minimizing the normalized Laplacian pyramid distance (NLPD), a perceptual metric aligning with human judgments of tone-mapped image quality. In Stage two, the input HDR image is self-calibrated to compute the final LDR image. We feed the same HDR image but rescaled with different maximum luminances to the learned tone mapping network, and generate a pseudo-multi-exposure image stack with different detail visibility and color saturation. We then train another lightweight DNN to fuse the LDR image stack into a desired LDR image by maximizing a variant of the structural similarity index for multi-exposure image fusion (MEF-SSIM), which has been proven perceptually relevant to fused image quality. The proposed self-calibration mechanism through MEF enables our TMO to accept uncalibrated HDR images, while being physiology-driven. Extensive experiments show that our method produces images with consistently better visual quality. Additionally, since our method builds upon three lightweight DNNs, it is among the fastest local TMOs.
Three ways AI is transforming music
Each fall, I begin my course on the intersection of music and artificial intelligence by asking my students if they're concerned about AI's role in composing or producing music. So far, the question has always elicited a resounding "yes." Their fears can be summed up in a sentence: AI will create a world where music is plentiful, but musicians get cast aside. In the upcoming semester, I'm anticipating a discussion about Paul McCartney, who in June 2023 announced that he and a team of audio engineers had used machine learning to uncover a "lost" vocal track of John Lennon by separating the instruments from a demo recording. But resurrecting the voices of long-dead artists is just the tip of the iceberg in terms of what's possible โ and what's already being done. In an interview, McCartney admitted that AI represents a "scary" but "exciting" future for music.
You Are Not Responsible for Your Own Online Privacy
In 2010, Mark Zuckerberg told the audience at a TechCrunch awards ceremony that young people--especially social media users--no longer cared about privacy. "People have really gotten comfortable not only sharing more information and different kinds, but more openly and with more people," he said. "That social norm is just something that has evolved over time." While this statement obviously hasn't aged well, it reflects a common belief that privacy violations happen when individuals reveal their own information. In other words, when something posted to Reddit or TikTok goes viral, or a nude photo sent to an admirer leaks, it's first and foremost the fault of the person who posted it. This model of individualized accountability is very persistent.
Domain-specific ChatBots for Science using Embeddings
Artificial intelligence and machine-learning (AI/ML) methods are growing in sophistication and capability. The application of these methods to the physical sciences is correspondingly seeing enormous growth.[1] Recent years have seen the convergence of several new trends. Generative AI seeks to create novel outputs that conform to the structure of training data,[2, 3] for instance enabling image synthesis[4-6] or text generation. Large language models (LLMs) are generative neural networks trained on text completion, but which can be used for a variety of tasks, including sentiment analysis, code completion, document generation, or for interactive chatbots that respond to users in natural language.[7] The most successful implementations of this concept--such as the generative pre-trained transformer (GPT)[8]-- exploit the transformer architecture,[9] which has a self-attention mechanism, allowing the model to weigh the relevance of each input in a sequence and capture the contextual dependencies between words regardless of their distance from each other in the text sequence. LLMs are part of a general trend in ML towards foundation models--extensive training of large deep neural networks on enormous datasets in a task-agnostic manner.[7,
Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
Mozes, Maximilian, He, Xuanli, Kleinberg, Bennett, Griffin, Lewis D.
Spurred by the recent rapid increase in the development and distribution of large language models (LLMs) across industry and academia, much recent work has drawn attention to safety- and security-related threats and vulnerabilities of LLMs, including in the context of potentially criminal activities. Specifically, it has been shown that LLMs can be misused for fraud, impersonation, and the generation of malware; while other authors have considered the more general problem of AI alignment. It is important that developers and practitioners alike are aware of security-related problems with such models. In this paper, we provide an overview of existing - predominantly scientific - efforts on identifying and mitigating threats and vulnerabilities arising from LLMs. We present a taxonomy describing the relationship between threats caused by the generative capabilities of LLMs, prevention measures intended to address such threats, and vulnerabilities arising from imperfect prevention measures. With our work, we hope to raise awareness of the limitations of LLMs in light of such security concerns, among both experienced developers and novel users of such technologies.
Exploiting Time-Frequency Conformers for Music Audio Enhancement
Chae, Yunkee, Koo, Junghyun, Lee, Sungho, Lee, Kyogu
With the proliferation of video platforms on the internet, recording musical performances by mobile devices has become commonplace. However, these recordings often suffer from degradation such as noise and reverberation, which negatively impact the listening experience. Consequently, the necessity for music audio enhancement (referred to as music enhancement from this point onward), involving the transformation of degraded audio recordings into pristine high-quality music, has surged to augment the auditory experience. To address this issue, we propose a music enhancement system based on the Conformer architecture that has demonstrated outstanding performance in speech enhancement tasks. Our approach explores the attention mechanisms of the Conformer and examines their performance to discover the best approach for the music enhancement task. Our experimental results show that our proposed model achieves state-of-the-art performance on single-stem music enhancement. Furthermore, our system can perform general music enhancement with multi-track mixtures, which has not been examined in previous work.
Exploring the Integration Strategies of Retriever and Large Language Models
Liu, Ye, Yavuz, Semih, Meng, Rui, Moorthy, Meghana, Joty, Shafiq, Xiong, Caiming, Zhou, Yingbo
The integration of retrieved passages and large language models (LLMs), such as ChatGPTs, has significantly contributed to improving open-domain question answering. However, there is still a lack of exploration regarding the optimal approach for incorporating retrieved passages into the answer generation process. This paper aims to fill this gap by investigating different methods of combining retrieved passages with LLMs to enhance answer generation. We begin by examining the limitations of a commonly-used concatenation approach. Surprisingly, this approach often results in generating "unknown" outputs, even when the correct document is among the top-k retrieved passages. To address this issue, we explore four alternative strategies for integrating the retrieved passages with the LLMs. These strategies include two single-round methods that utilize chain-of-thought reasoning and two multi-round strategies that incorporate feedback loops. Through comprehensive analyses and experiments, we provide insightful observations on how to effectively leverage retrieved passages to enhance the answer generation capability of LLMs.
Controversial new AI app allows you to text with Jesus โ and Satan
CyberGuy shows you how to save money with these apps. Welcome to the world of "Text With Jesus," where you're just a tap away from a conversation with the holy โ and, for a price, the not-so-holy. CLICK TO GET KURT'S FREE CYBERGUY NEWSLETTER WITH SECURITY ALERTS, QUICK TIPS, TECH REVIEWS AND EASY HOW-TO'S TO MAKE YOU SMARTER For those longing for a more personal connection to their faith, this app might be the digital salvation they're seeking. Designed with devoted Christians in mind, "Text With Jesus" promises interaction with figures like Jesus, Mary, Joseph, Peter and Matthew. This app wears its spirituality on its screen, guiding you through its queries with responses mined from the depths of the Bible's rich text.
Fox News AI Newsletter: Teachers go back to school with AI amid cheating concerns
ChatGPT has proven it can help students with their homework, but now it is helping teachers create those very courses, a computer science professor told Fox News. LEARNING CURVE: Teachers claim ChatGPT is cheating, but then use the tech for their grading . BACK TO SCHOOL: How parents and educators can ensure AI's ethical use in the classroom. School districts across the country have been faced with whether the use of ChatGPT in the classroom should be allowed. IN DEMAND: Businesses are on the hunt for workers with these AI skills.