Asia
TaiwanVQA: Benchmarking and Enhancing Cultural Understanding in Vision-Language Models
Vision-language models (VLMs) often struggle with culturally specific content -- a challenge largely overlooked by existing benchmarks that focus on dominant languages and globalized datasets. We introduce TAIWANVQA, a VQA benchmark designed for Taiwanese culture to evaluate recognition and reasoning in regional contexts. TAIWANVQA contains 2,736 images and 5,472 manually curated questions covering topics such as traditional foods, public signs, festivals, and landmarks. The official benchmark set includes 1,000 images and 2,000 questions for systematic assessment, with the remainder of the data used as training material. Evaluations on state-of-the-art VLMs reveal strong visual recognition but notable weaknesses in cultural reasoning.
SteerConf: Steering LLMs for Confidence Elicitation
Large Language Models (LLMs) exhibit impressive performance across diverse domains but often suffer from overconfidence, limiting their reliability in critical applications. We propose SteerConf, a novel framework that systematically steers LLMs' confidence scores to improve their calibration and reliability. SteerConf introduces three key components: (1) a steering prompt strategy that guides LLMs to produce confidence scores in specified directions (e.g., conservative or optimistic) by leveraging prompts with varying steering levels; (2) a steered confidence consistency measure that quantifies alignment across multiple steered confidences to enhance calibration; and (3) a steered confidence calibration method that aggregates confidence scores using consistency measures and applies linear quantization for answer selection. SteerConf operates without additional training or fine-tuning, making it broadly applicable to existing LLMs. Experiments on seven benchmarks spanning professional knowledge, common sense, ethics, and reasoning tasks, using advanced LLM models (GPT-3.5, LLaMA 3, GPT-4), demonstrate that SteerConf significantly outperforms existing methods, often by a significant margin. Our findings highlight the potential of steering the confidence of LLMs to enhance their reliability for safer deployment in real-world applications.
Mamba Modulation On the Length Generalization of Mamba
The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mambas performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behavior of its state-space dynamics, particularly within the parameterization of the state transition matrix A. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, exp( PN t=1 t), we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix A, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of Amatrices in each layer. We show that this can significantly improve performance in settings where simply modulating t fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices.
1bf3dbbd6346f50627e2ab1795f90435-Paper-Conference.pdf
Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization (ยตP) was proposed for vanilla Transformers, which enables stable HP transfer from small to large language models, and dramatically reduces tuning costs. However, it remains unclear whether ยตP of vanilla Transformers extends to diffusion Transformers, which differ architecturally and objectively. In this work, we generalize standard ยตP to diffusion Transformers and validate its effectiveness through large-scale experiments. First, we rigorously prove that ยตP of mainstream diffusion Transformers, including DiT, U-ViT, PixArt-ฮฑ, and MMDiT, aligns with that of the vanilla Transformer, enabling the direct application of existing ยตP methodologies. Leveraging this result, we systematically demonstrate that DiT-ยตP enjoys robust HP transferability. Notably, DiT-XL-2-ยตP with transferred learning rate achieves 2.9 faster convergence than the original DiT-XL-2.
Jรผrgen Habermas Defended Reason in a Darkening Age
The great German philosopher, who died in March, understood how much depended on a principled public sphere. Habermas emerged from the uncompromising Frankfurt School, but his work was considerably less fatalistic. You wake up and brace yourself for the barrage of toxic gibberish that constitutes the modern public sphere. Your e-mail is overrun with spam, scams, and smut. There are voice mails from no one about nothing. A glance at the news reveals that the President is continuing to spew lies and obscenities; that a trillionaire is peddling white-supremacist propaganda on a social-media platform he owns; that a chart-topping musical artist is praising Hitler, or apologizing for praising Hitler, or praising Hitler once again. Publications from the on down employ clickbait headlines that treat you like a starving rat in a Pavlovian experiment. A.I. systems simulate the experience of talking to an arrogant ten-year-old boy who knows far less than he thinks he does. When pressed, the chatbots admit that they cannot "naturally understand human morality, dignity, culture, or meaning." It all adds up to a continuous discursive tinnitus--a buzz of random, fake, stupid, sinister chatter that nobody wants and nobody can stop. The person who should have been best able to explain how we got here was the great German philosopher Jรผrgen Habermas, who illuminated how a feisty, principled public sphere is integral to democracy. But Habermas died in March, at the age of ninety-six, and, although he remained active until his final months, commenting on Ukraine, Gaza, and Eurobonds, he struggled to understand the turn history had taken. As a teen-ager in 1945, he had witnessed American soldiers enter his home town of Gummersbach, near Cologne, carrying messages of freedom and openness. Eight decades later, he watched American voters choose a leader who had advertised his fascistic bent in blood-and-soil rhetoric, fantasies of punitive violence, and a taste for bombastic architectural kitsch.
ReinAD: Towards Real-world Industrial Anomaly Detection with a Comprehensive Contrastive Dataset
Recent years have witnessed significant advancements in industrial anomaly detection (IAD) thanks to existing anomaly detection datasets. However, the large performance gap between these benchmarks and real industrial practice reveals critical limitations in existing datasets. We argue that the mismatch between current datasets and real industrial scenarios becomes the primary barrier to practical IAD deployment. To this end, we propose ReinAD dataset, a comprehensive contrastive dataset towards Real-world industrial Anomaly Detection. Our dataset prioritizes three critical real-world requirements: 1) Contrast-based anomaly definition that is essential for industrial practice, 2) Fine-grained unaligned image pairs reflecting real inspections, and 3) Large-scale data from active production lines spanning multiple industrial categories. Based on our dataset, we introduce the ReinADNet. It takes both normal reference and test images as inputs, achieving anomaly detection through normal-anomaly comparison. To address the fine-grained and unaligned properties of real industrial scenes, our method integrates pyramidal similarity aggregation for comprehensive anomaly characterization and globallocal feature fusion for spatial misalignment tolerance. Our method outperforms all baselines on the ReinAD dataset (e.g., 64.5% v.s.
Last-Iterate Convergence of Smooth Regret Matching + Variants in Learning Nash Equilibria
Regret Matching+ (RM+) variants are widely used to build superhuman Poker AIs, yet few studies investigate their last-iterate convergence in learning a Nash equilibrium (NE). Although their last-iterate convergence is established for games satisfying the Minty Variational Inequality (MVI), no studies have demonstrated that these algorithms achieve such convergence in the broader class of games satisfying the weak MVI. A key challenge in proving last-iterate convergence for RM+ variants in games satisfying the weak MVI is that even if the game's loss gradient satisfies the weak MVI, RM+ variants operate on a transformed loss feedback which does not satisfy the weak MVI. To provide last-iterate convergence for RM+ variants, we introduce a concise yet novel proof paradigm that involves: (i) transforming an RM+ variant into an Online Mirror Descent (OMD) instance that updates within the original strategy space of the game to recover the weak MVI, and (ii) showing last-iterate convergence by proving the distance between accumulated regrets converges to zero via the recovered weak MVI of the feedback. Inspired by our proof paradigm, we propose Smooth Optimistic Gradient Based RM+ (SOGRM+) and show that it achieves last-iterate and finite-time best-iterate convergence in learning an NE of games satisfying the weak MVI, the weakest condition among all known RM+ variants. Experiments show that SOGRM+ significantly outperforms other algorithms. Our code is available at https://github.
NASA's 'Son of Concorde' breaks the sound barrier: 247 million supersonic jet hits 713mph during test flight - paving the way for flights from London to New York in under 4 hours
Furious Trump EXPLODES over war talks as he threatens to'hit Iran very hard again' and tells rival leader he'better watch his mouth'... while also slamming Israel for continuing to drop bombs Ilhan Omar cries poor as she claims her millionaire husband only made TWO HUNDRED dollars last year... despite his empire being worth $30million Angelina Jolie's son Pax, 22, surfaces in LA after bombshell revelation about his relationship to Brad Pitt'Media-obsessed' Anna Paulina Luna reveals secret to her rising power as she turns into Republicans' 'favorite headache' Call me cynical, but the real reason Gruesome Twosome Harry and Meghan are returning to the UK is just so obvious... and highly humiliating: MAUREEN CALLAHAN No one can see the real reason Jelly Roll divorced Bunnie XO. Royals wish Prince William happy birthday and Father's Day with sweet photo of him and Charlotte after King's Trooping the Colour - as Charles pays tribute to Philip Family-man facade of award-winning children's ...