Media
Amazon Freevee adds terrifying AI-generated men to 12 Angry Men poster
The classic movie 12 Angry Men is titled as such because, well, it's about a jury comprised of 12 men. But viewers have noticed recently that the image Amazon uses for the movie has more than 12 characters in it. Further, their melting, inhuman faces look like they could be somebody's sleep paralysis monsters. The terrifying quality to the characters' faces is just one of the elements indicating the use of AI to generate the image. Their deformed and claw-like hands are another, along with the other obvious AI artifacts in the photo.
Engadget Podcast: How AI will shape Apple's WWDC 2024
We're gearing up to cover Apple's Worldwide Developers Conference (WWDC) next week! In this episode, Cherlynn and Devindra dive into everything they expect at WWDC: Tons of AI announcements; more on iOS 18, iPadOS 18, and macOS 15; and hopefully some improvements for Vision Pro and visionOS. In addition, we chat about what we expect to see at Summer Game Fest and demonstrate how we used an AI editing tool to clear up some awful podcast audio. Devindra also talks with Justin Samuels, the founder of Render ATL, about why he started a massive tech conference in Atlanta. Listen below or subscribe on your podcast app of choice. If you've got suggestions or topics you'd like covered on the show, be sure to email us or drop a note in the comments! And be sure to check out our other podcast, Engadget News! Humane AI warns users its battery case "may pose a fire risk" – 34:36 Welcome back to the Engadget podcast. This week we are getting ready for WWDC 2024 happening in a couple of days.
No, Drake's Cover of 'Hey There Delilah' Isn't AI
As if he didn't have enough to deal with amid his beef with Kendrick Lamar (or perhaps to distract from it), Drake showed up on a remix of parody rapper Snowd4y's cover of Plain White T's "Hey There Delilah," called "Wah Gwan Delilah," that has everyone … perplexed? Let's walk through this together, it's a mess. It had what appeared to be Drake joining the comedian in a series of quips about women and name-checks of Toronto landmarks like the Yonge-Dundas Square mall. As the track spread, it made its way to the Plain White T's themselves, who posted a video on X and TikTok with the caption "too stunned to speak." Frontman Tom Higgenson also says "it's crazy that everybody thinks that it's real," seemingly referencing early rumors that Drake's lyrics were generated using artificial intelligence.
'Top Gun' producer says he doesn't believe claims AI will replace key jobs
"Top Gun" producer Jerry Bruckheimer sees the overall benefit of artificial intelligence. "Anything that makes our lives easier that doesn't take jobs away from people that we work with every day is good for everybody. It gives them a better movie experience. We can make things look more real and things like that," he told Fox News Digital. However, he didn't see the technology eliminating important jobs in the industry.
Generative AI Models: Opportunities and Risks for Industry and Authorities
Alt, Tobias, Ibisch, Andrea, Meiser, Clemens, Wilhelm, Anna, Zimmer, Raphael, Berghoff, Christian, Droste, Christoph, Karschau, Jens, Laus, Friederike, Plaga, Rainer, Plesch, Carola, Sennewald, Britta, Thaeren, Thomas, Unverricht, Kristina, Waurick, Steffen
Generative AI models are capable of performing a wide range of tasks that traditionally require creativity and human understanding. They learn patterns from existing data during training and can subsequently generate new content such as texts, images, and music that follow these patterns. Due to their versatility and generally high-quality results, they, on the one hand, represent an opportunity for digitalization. On the other hand, the use of generative AI models introduces novel IT security risks that need to be considered for a comprehensive analysis of the threat landscape in relation to IT security. In response to this risk potential, companies or authorities using them should conduct an individual risk analysis before integrating generative AI into their workflows. The same applies to developers and operators, as many risks in the context of generative AI have to be taken into account at the time of development or can only be influenced by the operating company. Based on this, existing security measures can be adjusted, and additional measures can be taken.
Sales Whisperer: A Human-Inconspicuous Attack on LLM Brand Recommendations
Lin, Weiran, Gerchanovsky, Anna, Akgul, Omer, Bauer, Lujo, Fredrikson, Matt, Wang, Zifan
Large language model (LLM) users might rely on others (e.g., prompting services), to write prompts. However, the risks of trusting prompts written by others remain unstudied. In this paper, we assess the risk of using such prompts on brand recommendation tasks when shopping. First, we found that paraphrasing prompts can result in LLMs mentioning given brands with drastically different probabilities, including a pair of prompts where the probability changes by 100%. Next, we developed an approach that can be used to perturb an original base prompt to increase the likelihood that an LLM mentions a given brand. We designed a human-inconspicuous algorithm that perturbs prompts, which empirically forces LLMs to mention strings related to a brand more often, by absolute improvements up to 78.3%. Our results suggest that our perturbed prompts, 1) are inconspicuous to humans, 2) force LLMs to recommend a target brand more often, and 3) increase the perceived chances of picking targeted brands.
MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
Chowdhury, Sanjoy, Nag, Sayan, Joseph, K J, Srinivasan, Balaji Vasan, Manocha, Dinesh
Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are predominantly conditioned on textual descriptions of it. Inspired by how musicians compose music not just from a movie script, but also through visualizations, we propose MeLFusion, a model that can effectively use cues from a textual description and the corresponding image to synthesize music. MeLFusion is a text-to-music diffusion model with a novel "visual synapse", which effectively infuses the semantics from the visual modality into the generated music. To facilitate research in this area, we introduce a new dataset MeLBench, and propose a new evaluation metric IMSM. Our exhaustive experimental evaluation suggests that adding visual information to the music synthesis pipeline significantly improves the quality of generated music, measured both objectively and subjectively, with a relative gain of up to 67.98% on the FAD score. We hope that our work will gather attention to this pragmatic, yet relatively under-explored research area.
How to Strategize Human Content Creation in the Era of GenAI?
Esmaeili, Seyed A., Bhawalkar, Kshipra, Feng, Zhe, Wang, Di, Xu, Haifeng
Generative AI (GenAI) will have significant impact on content creation platforms. In this paper, we study the dynamic competition between a GenAI and a human contributor. Unlike the human, the GenAI's content only improves when more contents are created by human over the time; however, GenAI has the advantage of generating content at a lower cost. We study the algorithmic problem in this dynamic competition model about how the human contributor can maximize her utility when competing against the GenAI for content generation over a set of topics. In time-sensitive content domains (e.g., news or pop music creation) where contents' value diminishes over time, we show that there is no polynomial time algorithm for finding the human's optimal (dynamic) strategy, unless the randomized exponential time hypothesis is false. Fortunately, we are able to design a polynomial time algorithm that naturally cycles between myopically optimizing over a short time window and pausing and provably guarantees an approximation ratio of $\frac{1}{2}$. We then turn to time-insensitive content domains where contents do not lose their value (e.g., contents on history facts). Interestingly, we show that this setting permits a polynomial time algorithm that maximizes the human's utility in the long run.
Semantic-Enhanced Relational Metric Learning for Recommender Systems
Li, Mingming, Zhu, Fuqing, Yuan, Feng, Hu, Songlin
Recently, relational metric learning methods have been received great attention in recommendation community, which is inspired by the translation mechanism in knowledge graph. Different from the knowledge graph where the entity-to-entity relations are given in advance, historical interactions lack explicit relations between users and items in recommender systems. Currently, many researchers have succeeded in constructing the implicit relations to remit this issue. However, in previous work, the learning process of the induction function only depends on a single source of data (i.e., user-item interaction) in a supervised manner, resulting in the co-occurrence relation that is free of any semantic information. In this paper, to tackle the above problem in recommender systems, we propose a joint Semantic-Enhanced Relational Metric Learning (SERML) framework that incorporates the semantic information. Specifically, the semantic signal is first extracted from the target reviews containing abundant item features and personalized user preferences. A novel regression model is then designed via leveraging the extracted semantic signal to improve the discriminative ability of original relation-based training process. On four widely-used public datasets, experimental results demonstrate that SERML produces a competitive performance compared with several state-of-the-art methods in recommender systems.
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
Davoodi, Arash Gholami, Davoudi, Seyed Pouyan Mousavi, Pezeshkpour, Pouya
Large language models (LLMs) demonstrate impressive capabilities in mathematical reasoning. However, despite these achievements, current evaluations are mostly limited to specific mathematical topics, and it remains unclear whether LLMs are genuinely engaging in reasoning. To address these gaps, we present the Mathematical Topics Tree (MaTT) benchmark, a challenging and structured benchmark that offers 1,958 questions across a wide array of mathematical subjects, each paired with a detailed hierarchical chain of topics. Upon assessing different LLMs using the MaTT benchmark, we find that the most advanced model, GPT-4, achieved a mere 54\% accuracy in a multiple-choice scenario. Interestingly, even when employing Chain-of-Thought prompting, we observe mostly no notable improvement. Moreover, LLMs accuracy dramatically reduced by up to 24.2 percentage point when the questions were presented without providing choices. Further detailed analysis of the LLMs' performance across a range of topics showed significant discrepancy even for closely related subtopics within the same general mathematical area. In an effort to pinpoint the reasons behind LLMs performances, we conducted a manual evaluation of the completeness and correctness of the explanations generated by GPT-4 when choices were available. Surprisingly, we find that in only 53.3\% of the instances where the model provided a correct answer, the accompanying explanations were deemed complete and accurate, i.e., the model engaged in genuine reasoning.