Media
'Minecraft' movie mayhem raises alarms for America's youth, 'bad for society': expert
"A Minecraft Movie," the big-screen adaptation of the popular video game "Minecraft," has been packing theaters with rowdy kids and teens since its release this month, spurring a social media phenomenon and sparking concern for America's youth. Videos on social media show young theatergoers huge reactions to one key scene, where one of the film's stars, Jack Black, yells out the phrase "Chicken Jockey!" as a small, Frankenstein-looking creature lands on top of a chicken in a boxing ring to face off with co-star Jason Momoa. The scene has prompted excited fans to scream, shout, throw popcorn around, jump up out of their seats, and in one instance in Provo, Utah, toss a live chicken in the air during a screening, according to the Salt Lake Tribune. Springs Cinema & Taphouse in Sandy Springs, Georgia, told FOX 5 Atlanta that its staff has had to clean up popcorn, ICEEs, ketchup and shattered glass. The scene featuring the "Chicken Jockey" in "A Minecraft Movie" has spawned some chaotic movie theater behavior from young audiences. "The movie-going experience has changed a lot since I was younger," Josh Gunderson, director of marketing and events at Oviedo Mall in Florida, told FOX Business.
Google Pixel 9a review: Engaging AI features and mighty battery life give Apple's 'budget' iPhone a run for its money
Apple released its latest'budget' phone, the 599 iPhone 16e, back in February after months of feverish anticipation. But not to be outdone, rival tech giant Google has released its own handset at an'unbeatable' price โ the Pixel 9a. The device โ which at 499 is 100 cheaper than Apple's equivalent โ has a 6.3-inch display, two rear cameras and more than 30 hours of battery life on a single charge. It's packed with'helpful' AI tools such as Gemini โ Google's chatbot which was built to rival OpenAI's ChatGPT, now on Apple phones. MailOnline tests the new Google handset, described as a more accessible alternative to the firm's flagship Pixel 9 ( 799).
LG's Integrated TV Ad Tech Analyzes Your Emotions
LG TVs will soon leverage an artificial intelligence model built for showing advertisements that more closely align with viewers' personal beliefs and emotions. This story originally appeared on Ars Technica, a trusted source for technology news, tech policy analysis, reviews, and more. Ars is owned by WIRED's parent company, Condรฉ Nast. The company plans to incorporate a partner company's AI tech into its TV software in order to interpret psychological factors impacting a viewer, such as personal interests, personality traits, and lifestyle choices. The aim is to show LG webOS users ads that will emotionally impact them.
Can this 70,000 robot transform AI research?
Reachy 2 is touted as a "lab partner for the AI era." The folks at Hugging Face, the open-source artificial intelligence gurus, just jumped into the world of robotics by acquiring Pollen Robotics. And right out of the gate, they are offering the Reachy 2, a super-interesting humanoid robot designed as a "lab partner for the AI era." Ready to dive in and see what all the buzz is about? GET SECURITY ALERTS & EXPERT TECH TIPS โ SIGN UP FOR KURT'S'THE CYBERGUY REPORT' NOW So, what makes Reachy 2 stand out?
'Terminator' director James Cameron flip-flops on AI, says Hollywood is 'looking at it all wrong'
Fox News Flash top entertainment and celebrity headlines are here. James Cameron's stance on artificial intelligence has evolved over the past few years, and he feels Hollywood needs to embrace it in a few different ways. Cameron joined the board of directors for Stability AI last year, explaining his decision on the "Boz to the Future" podcast last week. "The goal was to understand the space, to understand what's on the minds of the developers," he said. How much resources you have to throw at it to create a new model that does a purpose-built thing, and my goal was to try to integrate it into a VFX workflow." He continued by saying the shift to AI is a necessary one. James Cameron wants Hollywood to implement AI more for big-budget films. WHAT IS ARTIFICIAL INTELLIGENCE (AI)? If we want to continue to see the kinds of movies that I've always loved and that I like to make and that I will go to see โ 'Dune,' 'Dune: Part Two' or one of my films or big effects-heavy, CG-heavy films โ we've got to figure out how to cut the cost of that in half. That's about doubling their speed to completion on a given shot, so your cadence is faster and your throughput cycle is faster, and artists get to move on and do other cool things and then other cool things, right? Cameron doesn't think films are ultimately "a big target" for companies like OpenAI. "Their goal is not to make GenAI movies.
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
Yang, Siwei, Hui, Mude, Zhao, Bingchen, Zhou, Yuyin, Ruiz, Nataniel, Xie, Cihang
We introduce $\texttt{Complex-Edit}$, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to automatically collect a diverse set of editing instructions at scale. Our approach follows a well-structured ``Chain-of-Edit'' pipeline: we first generate individual atomic editing tasks independently and then integrate them to form cohesive, complex instructions. Additionally, we introduce a suite of metrics to assess various aspects of editing performance, along with a VLM-based auto-evaluation pipeline that supports large-scale assessments. Our benchmark yields several notable insights: 1) Open-source models significantly underperform relative to proprietary, closed-source models, with the performance gap widening as instruction complexity increases; 2) Increased instructional complexity primarily impairs the models' ability to retain key elements from the input images and to preserve the overall aesthetic quality; 3) Decomposing a complex instruction into a sequence of atomic steps, executed in a step-by-step manner, substantially degrades performance across multiple metrics; 4) A straightforward Best-of-N selection strategy improves results for both direct editing and the step-by-step sequential approach; and 5) We observe a ``curse of synthetic data'': when synthetic data is involved in model training, the edited images from such models tend to appear increasingly synthetic as the complexity of the editing instructions rises -- a phenomenon that intriguingly also manifests in the latest GPT-4o outputs.
InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task Planning
Wang, Zheng, Teo, Shu Xian, Chew, Jun Jie, Shi, Wei
Recent advancements in large language models (LLMs) have enabled their use as agents for planning complex tasks. Existing methods typically rely on a thought-action-observation (TAO) process to enhance LLM performance, but these approaches are often constrained by the LLMs' limited knowledge of complex tasks. Retrieval-augmented generation (RAG) offers new opportunities by leveraging external databases to ground generation in retrieved information. In this paper, we identify two key challenges (enlargability and transferability) in applying RAG to task planning. We propose InstructRAG, a novel solution within a multi-agent meta-reinforcement learning framework, to address these challenges. InstructRAG includes a graph to organize past instruction paths (sequences of correct actions), an RL-Agent with Reinforcement Learning to expand graph coverage for enlargability, and an ML-Agent with Meta-Learning to improve task generalization for transferability. The two agents are trained end-to-end to optimize overall planning performance. Our experiments on four widely used task planning datasets demonstrate that InstructRAG significantly enhances performance and adapts efficiently to new tasks, achieving up to a 19.2% improvement over the best existing approach.
Out of Sight Out of Mind, Out of Sight Out of Mind: Measuring Bias in Language Models Against Overlooked Marginalized Groups in Regional Contexts
Elsafoury, Fatma, Hartmann, David
We know that language models (LMs) form biases and stereotypes of minorities, leading to unfair treatments of members of these groups, thanks to research mainly in the US and the broader English-speaking world. As the negative behavior of these models has severe consequences for society and individuals, industry and academia are actively developing methods to reduce the bias in LMs. However, there are many under-represented groups and languages that have been overlooked so far. This includes marginalized groups that are specific to individual countries and regions in the English speaking and Western world, but crucially also almost all marginalized groups in the rest of the world. The UN estimates, that between 600 million to 1.2 billion people worldwide are members of marginalized groups and in need for special protection. If we want to develop inclusive LMs that work for everyone, we have to broaden our understanding to include overlooked marginalized groups and low-resource languages and dialects. In this work, we contribute to this effort with the first study investigating offensive stereotyping bias in 23 LMs for 270 marginalized groups from Egypt, the remaining 21 Arab countries, Germany, the UK, and the US. Additionally, we investigate the impact of low-resource languages and dialects on the study of bias in LMs, demonstrating the limitations of current bias metrics, as we measure significantly higher bias when using the Egyptian Arabic dialect versus Modern Standard Arabic. Our results show, LMs indeed show higher bias against many marginalized groups in comparison to dominant groups. However, this is not the case for Arabic LMs, where the bias is high against both marginalized and dominant groups in relation to religion and ethnicity. Our results also show higher intersectional bias against Non-binary, LGBTQIA+ and Black women.
SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation
Bougie, Nicolas, Watanabe, Narimasa
Recommender systems play a central role in numerous real-life applications, yet evaluating their performance remains a significant challenge due to the gap between offline metrics and online behaviors. Given the scarcity and limits (e.g., privacy issues) of real user data, we introduce SimUSER, an agent framework that serves as believable and cost-effective human proxies. SimUSER first identifies self-consistent personas from historical data, enriching user profiles with unique backgrounds and personalities. Then, central to this evaluation are users equipped with persona, memory, perception, and brain modules, engaging in interactions with the recommender system. SimUSER exhibits closer alignment with genuine humans than prior work, both at micro and macro levels. Additionally, we conduct insightful experiments to explore the effects of thumbnails on click rates, the exposure effect, and the impact of reviews on user engagement. Finally, we refine recommender system parameters based on offline A/B test results, resulting in improved user engagement in the real world.
Memorization vs. Reasoning: Updating LLMs with New Knowledge
Li, Aochong Oliver, Goyal, Tanya
Large language models (LLMs) encode vast amounts of pre-trained knowledge in their parameters, but updating them as real-world information evolves remains a challenge. Existing methodologies and benchmarks primarily target entity substitutions, failing to capture the full breadth of complex real-world dynamics. In this paper, we introduce Knowledge Update Playground (KUP), an automatic pipeline for simulating realistic knowledge updates reflected in an evidence corpora. KUP's evaluation framework includes direct and indirect probes to both test memorization of updated facts and reasoning over them, for any update learning methods. Next, we present a lightweight method called memory conditioned training (MCT), which conditions tokens in the update corpus on self-generated "memory" tokens during training. Our strategy encourages LLMs to surface and reason over newly memorized knowledge at inference. Our results on two strong LLMs show that (1) KUP benchmark is highly challenging, with the best CPT models achieving $<2\%$ in indirect probing setting (reasoning) and (2) MCT training significantly outperforms prior continued pre-training (CPT) baselines, improving direct probing (memorization) results by up to $25.4\%$.