Law
bd20ff18345f0ded89242bf9ef58e46c-Paper-Position_Paper_Track.pdf
This position paper argues that human pose estimation (HPE) cannot be considered privacy-preserving or human-centric unless privacy is measured and evaluated. Although privacy concerns have become more visible in recent years, HPE systems are still assessed almost exclusively using accuracy metrics. Privacy is neither defined in measurable terms nor linked to regulatory requirements, and common deployment architectures introduce additional risks due to data transmission and storage. We highlight the limitations of current practices, including the continued reliance on RGB inputs and the lack of benchmarks that reflect legal and ethical constraints. We call for a shift in evaluation practices: privacy must become part of how HPE systems are designed, tested, and compared.
Analyzing Vulnerabilities of MoE Based LLMs via Stable Safety critical Expert Identification
Large language models with Mixture-of-Experts (MoE) architectures achieve efficiency and scalability, yet their routing mechanisms introduce safety alignment challenges insufficiently addressed by techniques developed for dense models. In this work, the MoE-specific safety risk of positional vulnerability--that safetyaligned behaviors rely on specific expert modules--is formalized and systematically analyzed. An analytical framework, SAFEX, is presented to robustly identify, characterize, and validate safety-critical experts via a stability-based expert selection procedure, and to decompose them into two functional groups: the Harmful Content Detection Group (HCDG), which specializes in identifying and recognizing harmful content within user inputs, and the Harmful Response Control Group (HRCG), which specializes in controlling and enforcing model behaviors to generate appropriate safety responses. Expert-level interventions are conducted to probe causality and to test mitigation. Targeted masking of SAFEX-selected experts reveals that safety behavior is highly concentrated. On Qwen3-30B-A3B, configured with 48 MoE-FFN layers and 128 experts per layer under top-8 routing (48 128 = 6,144 experts in total), disabling 12 selected experts reduces the refusal rate by 22%. In addition, lightweight adaptation is performed using LoRA under three configurations--the HRCG, the union of HCDG and HRCG, and all experts--and the resulting updates are composed through negative weight merging targeted at the HRCG, leading to improved refusal under adversarial prompts without full-model retraining. These results establish positional vulnerability as a distinct MoE-specific safety challenge and provide a practical, computeefficient pathway for expert-level safety interventions within routed architectures (https://github.com/Bearisbug/SAFEx).
Pro3D-Editor: AProgressive-Views Perspective for Consistent and Precise 3DEditing
T gions, ext-guided which 3D has editing significant aims potential to precisely for edit various semantically practical applications relevant local ranging 3D refrom 3D games to film production. Existing methods typically follow a viewindiscriminate paradigm: editing 2D views indiscriminately and projecting them back dencies, into resulting 3D space. in Ho inconsistent wever, the multi-vie y overlook w editing.
People training new AI models admit they just get chatbots to do it
The next generation of AI models are meant to be trained by people paid to have conversations with them, but several of these workers have admitted to that they simply get chatbots to do it instead. People who are paid to train new AI models by supplying them with high-quality conversation and tests are cheating and using chatbots like ChatGPT to do the job instead, multiple whistleblowers have told . The seemingly widespread practice risks undermining the future of AI, as it could lead to the "collapse" of more advanced models. Most AI models operating today were trained on text and data scraped from the internet . But as models have scaled up, requiring yet more training data, AI firms have begun using workers who carry out conversations and tests with AI, in the hope that the resulting high-quality data can improve the power and usefulness of future large language models (LLMs). These workers are normally employed by third parties, rather than AI companies directly, and are often working without full-time contracts and for low pay.
Preference Learning with Lie Detectors can Induce Honesty or Evasion
As AI systems become more capable, deceptive behaviors can undermine evaluation and mislead users at deployment. Recent work has shown that lie detectors can accurately classify deceptive behavior, but they are not typically used in the training pipeline due to concerns around contamination and objective hacking. We examine these concerns by incorporating a lie detector into the labelling step of LLM post-training and evaluating whether the learned policy is genuinely more honest, or instead learns to fool the lie detector while remaining deceptive. Using DolusChat, a novel 65k-example dataset with paired truthful/deceptive responses, we identify three key factors that determine the honesty of learned policies: amount of exploration during preference learning, lie detector accuracy, and KL regularization strength. We find that preference learning with lie detectors and GRPO can lead to policies which evade lie detectors, with deception rates of over 85%. However, if the lie detector true positive rate (TPR) or KL regularization is sufficiently high, GRPO learns honest policies. In contrast, off-policy algorithms (DPO) consistently lead to deception rates under 25% for realistic TPRs. Our results illustrate a more complex picture than previously assumed: depending on the context, lie-detector-enhanced training can be a powerful tool for scalable oversight, or a counterproductive method encouraging undetectable misalignment.
See Stonehenge's construction like NEVER before: Incredible visual reveals the vast manpower needed to haul the 25-tonne stones into position 5,000 years ago
Keir Starmer cries as he quits No 10 claiming a deluded list of'achievements' - now Britain awaits its seventh PM in ten years Putin'prepares mass call-up' for the Ukraine meat-grinder - as video shows Russian veteran with no legs threatening recruiter with a knife in sign of growing resistance facing desperate Kremlin'Al Roker is an absolute ****': KENNEDY's Today show insider gives brutal behind-the-scenes verdict on beloved weatherman and names other two-faced NBC hosts No one can see the real reason Jelly Roll divorced Bunnie XO. Boston's Scotland-loving residents claim England fans are'ruining the vibe' compared to the Tartan Army Secret life of John Travolta's daughter Ella Bleu: New details about'unusual' relationship with her dad revealed by insiders amid fears that aspiring actress is'stuck' Johnny Depp's ex Amber Heard gives rare glimpse of daughter Oonagh, five, after finishing 10k race in Spain Colorado siblings VANISH from home in middle of the night... and police ...
The NY-12 Primary Is Awash with Money but Short on Belief
The race--whose candidates include Micah Lasher, Alex Bores, George Conway, and Jack Schlossberg--is at once glitzy, confusing, and uninspiring. Alex Bores is one of many candidates in the hotly contested race for New York's Twelfth Congressional district. A good seat in Congress can be hard to find, and difficult to get up from. The average district--and there are four hundred and thirty-five of them--is roughly the size of Wales, or New Jersey. New York's Twelfth District, which spans the Upper East Side, the Upper West Side, midtown, and Chelsea, is one of the richest, smallest, and most solidly Democratic districts in the country. It has the most people with college degrees and is in the ninety-fifth percentile for members of the Silent Generation. After its incumbent, Jerry Nadler, who has been in Congress since 1992, announced his retirement last year, the race to fill his seat has also become one of the most contested.
Detecting High-Stakes Interactions with Activation Probes
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting "high-stakes" interactions--where the text indicates that the interaction might lead to significant harm--as a critical, yet underexplored, target for such monitoring. We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data. Probes' performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude. These savings are enabled by reusing activations of the model that is being monitored. Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis.
PHANTOM: ABenchmark for Hallucination Detection in Financial Long-Context QA
While Large Language Models (LLMs) show great promise, their tendencies to hallucinate pose significant risks in high-stakes domains like finance, especially when used for regulatory reporting and decision-making. Existing hallucination detection benchmarks fail to capture the complexities of financial benchmarks, which require high numerical precision, nuanced understanding of the language of finance, and ability to handle long-context documents. To address this, we introduce PHANTOM, a novel benchmark dataset for evaluating hallucination detection in long-context financial QA. Our approach first generates a seed dataset of high-quality "query-answer-document (chunk)" triplets, with either hallucinated or correct answers - that are validated by human annotators and subsequently expanded to capture various context lengths and information placements. We demonstrate how PHANTOM allows fair comparison of hallucination detection models and provides insights into LLM performance, offering a valuable resource for improving hallucination detection in financial applications. Further, our benchmarking results highlight the severe challenges out-of-the-box models face in detecting real-world hallucinations on long context data, and establish some promising directions towards alleviating these challenges, by fine-tuning open-source LLMs using PHANTOM.1
WolBanking77: Wolof Banking Speech Intent Classification Dataset
Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90% of the population, while the national illiteracy rate remains at of 42%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification.