Goto

Collaborating Authors

 hong


South Korea's new manager Robert Moreno dismisses 'AI coaching' allegations

Al Jazeera

South Korea's new manager Robert Moreno dismisses'AI coaching' allegations Share'Fake news': South Korea's manager Moreno dismisses AI coaching allegations on social media South Korea's new interim football coach Robert Moreno has denied allegations that he relied on artificial intelligence during his last coaching role but admitted to using technology in his work. Addressing his first news conference as the Taegeuk Warriors' temporary coach on Monday, Moreno promised to make South Korean football fans "happy". The Spaniard was appointed as the Asian giants' temporary coach following a disastrous few months for the team and the country's football federation, during which on-field results and off-field controversies left the fans bitterly disappointed. Moreno, who was forced to deny earlier this year that he used ChatGPT to prepare for games, will be in charge for friendly matches in the coming months. He addressed allegations that he used artificial intelligence during his time in charge of Russian club Sochi, saying "It was a piece of news that is completely false and went viral."


David Crowley Beats Former Frontrunner Francesca Hong in Democratic Primary for Wisconsin Governor

TIME - Tech

Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW?


Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion Transformers

Neural Information Processing Systems

Text-to-image diffusion models excel at translating language prompts into photorealistic images by implicitly grounding textual concepts through their cross-modal attention mechanisms. Recent multi-modal diffusion transformers extend this by introducing joint self-attention over concatenated image and text tokens, enabling richer and more scalable cross-modal alignment. However, a detailed understanding of how and where these attention maps contribute to image generation remains limited. In this paper, we introduce Seg4Diff (Segmentation for Diffusion), a systematic framework for analyzing the attention structures of MM-DiT, with a focus on how specific layers propagate semantic information from text to image. Through comprehensive analysis, we identify a semantic grounding expert layer, a specific MM-DiT block that consistently aligns text tokens with spatially coherent image regions, naturally producing high-quality semantic segmentation masks. We further demonstrate that applying a lightweight fine-tuning scheme with mask-annotated image data enhances the semantic grouping capabilities of these layers and thereby improves both segmentation performance and generated image fidelity. Our findings demonstrate that semantic grouping is an emergent property of diffusion transformers and can be selectively amplified to advance both segmentation and generation performance, paving the way for unified models that bridge visual perception and generation.


The People vs. AI

TIME - Tech

One icy morning in February, nearly 200 people gathered in a church in downtown Richmond, Va. Most had awakened before dawn and driven in from across the state. There were Republicans and Democrats from rural farms and D.C. exurbs. They shared one goal: to fight back against AI development in a region with the largest concentration of data centers in the world. "Aren't you tired of being ignored by both parties, and having your quality of life and your environment absolutely destroyed by corporate greed?" state senator Danica Roem said, to a standing ovation. The activists--wearing homemade shirts with slogans like Boondoggle: Data Center in Botetourt County--marched to the state capitol and spent the day testifying to lawmakers about their fears over data centers' impacts on electricity, water, noise pollution, and more. Some lawmakers pledged to help: "You're getting a sh-t deal," state delegate John McAuliff told activists. The phrase captured many people's feelings toward the AI industry as a whole. Not much unites Americans these days.




A New AI Math Startup Just Cracked 4 Previously Unsolved Problems

WIRED

Axiom says its AI found solutions to several long-standing math problems, a sign of the technology's steadily advancing reasoning capabilities. Five years ago, mathematicians Dawei Chen and Quentin Gendron were trying to untangle a difficult area of algebraic geometry involving differentials, elements of calculus used to measure distance along curved surfaces . While working on one theorem, they ran into an unexpected roadblock: Their argument depended on a strange formula from number theory, but they were unable to solve or justify it. In the end, Chen and Gendron wrote a paper presenting their idea as a conjecture, rather than a theorem. Chen recently spent hours prompting ChatGPT in the hopes of getting the AI to come up with a solution to the still unsolved problem, but it wasn't working.


AI Models Get Brain Rot, Too

WIRED

A new study shows that feeding large language models low-quality, high-engagement content from social media lowers their cognitive abilities. AI models may be a bit like humans, after all. A new study from the University of Texas at Austin, Texas A&M, and Purdue University shows that large language models fed a diet of popular but low-quality social media content experience a kind of "brain rot" that may be familiar to anyone who has spent too long doomscrolling on X or TikTok. We live in an age where information grows faster than attention spans--and much of it is engineered to capture clicks, not convey truth or depth," says Junyuan Hong, an incoming assistant professor at the National University of Singapore who worked on the study as a graduate student at UT Austin. "We wondered: What happens when AIs are trained on the same stuff?"


MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations

arXiv.org Artificial Intelligence

Electrocardiogram (ECG) plays a foundational role in modern cardiovascular care, enabling non-invasive diagnosis of arrhythmias, myocardial ischemia, and conduction disorders. While machine learning has achieved expert-level performance in ECG interpretation, the development of clinically deployable multimodal AI systems remains constrained, primarily due to the lack of publicly available datasets that simultaneously incorporate raw signals, diagnostic images, and interpretation text. Most existing ECG datasets provide only single-modality data or, at most, dual modalities, making it difficult to build models that can understand and integrate diverse ECG information in real-world settings. To address this gap, we introduce MEETI (MIMIC-IV-Ext ECG-Text-Image), the first large-scale ECG dataset that synchronizes raw waveform data, high-resolution plotted images, and detailed textual interpretations generated by large language models. In addition, MEETI includes beat-level quantitative ECG parameters extracted from each lead, offering structured parameters that support fine-grained analysis and model interpretability. Each MEETI record is aligned across four components: (1) the raw ECG waveform, (2) the corresponding plotted image, (3) extracted feature parameters, and (4) detailed interpretation text. This alignment is achieved using consistent, unique identifiers. This unified structure supports transformer-based multimodal learning and supports fine-grained, interpretable reasoning about cardiac health. By bridging the gap between traditional signal analysis, image-based interpretation, and language-driven understanding, MEETI established a robust foundation for the next generation of explainable, multimodal cardiovascular AI. It offers the research community a comprehensive benchmark for developing and evaluating ECG-based AI systems.


ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

arXiv.org Artificial Intelligence

In this paper, we present details of the 1st W-CODA workshop, held in conjunction with the ECCV 2024. W-CODA aims to explore next-generation solutions for autonomous driving corner cases, empowered by state-of-the-art multimodal perception and comprehension techniques. 5 Speakers from both academia and industry are invited to share their latest progress and opinions. We collect research papers and hold a dual-track challenge, including both corner case scene understanding and generation. As the pioneering effort, we will continuously bridge the gap between frontier autonomous driving techniques and fully intelligent, reliable self-driving agents robust towards corner cases.