Technology
Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences Training
We introduce Mini-Sequence Transformer (MsT), a simple and effective methodology for highly efficient and accurate LLM training with extremely long sequences. MsT partitions input sequences and iteratively processes mini-sequences to reduce intermediate memory usage. Integrated with activation recomputation, it enables significant memory savings in both forward and backward passes.
Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd Wisdom
The generation of label noise is often modeled as a process involving a probability transition matrix (also interpreted as the) imposed onto the label distribution. Under this model, learning the ``ground-truth classifier''---i.e., the classifier that can be learned if no noise was present---and the confusion matrix boils down to a model identification problem. Prior works along this line demonstrated appealing empirical performance, yet identifiability of the model was mostly established by assuming an instance-invariant confusion matrix. Having an (occasionally) instance-dependent confusion matrix across data samples is apparently more realistic, but inevitably introduces outliers to the model. Our interest lies in confusion matrix-based noisy label learning with such outliers taken into consideration. We begin with pointing out that under the model of interest, using labels produced by only one annotator is fundamentally insufficient to detect the outliers or identify the ground-truth classifier. Then, we prove that by employing a crowdsourcing strategy involving multiple annotators, a carefully designed loss function can establish the desired model identifiability under reasonable conditions. Our development builds upon a link between the noisy label model and a column-corrupted matrix factorization mode---based on which we show that crowdsourced annotations distinguish nominal data and instance-dependent outliers using a low-dimensional subspace. Experiments show that our learning scheme substantially improves outlier detection and the classifier's testing accuracy.
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
Training offline RL models using visual inputs poses two significant challenges,, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. This paper, in contrast, tries to build more flexible constraints for value estimation without impeding the exploration of potential advantages. The key idea is to leverage off-the-shelf RL simulators, which can be easily interacted with in an online manner, as the " " for offline policies. To enable effective online-to-offline knowledge transfer, we introduce CoWorld, a model-based RL approach that mitigates cross-domain discrepancies in state and reward spaces. Experimental results demonstrate the effectiveness of CoWorld, outperforming existing RL approaches by large margins.
SfPUEL: Shape from Polarization under Unknown Environment Light
Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment light. To handle this challenging light condition, we design a transformer-based framework for enhancing the perception of global context features. We further propose to integrate photometric stereo (PS) priors from pretrained models to enrich extracted features for high-quality normal predictions. As metallic and dielectric materials exhibit different BRDFs, SfPUEL additionally predicts dielectric and metallic material segmentation to further boost performance. Experimental results on synthetic and our collected real-world dataset demonstrate that SfPUEL significantly outperforms existing SfP and single-shot normal estimation methods.
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
We propose ID-to-3D, a method to generate identity-and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured'in-the-wild' image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometric and texture generation, via fine-tuning only 0.2% of its available training parameters.
MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models that addresses 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications.
A Simple yet Scalable Granger Causal Structural Learning Approach for Topological Event Sequences
Network operators need an efficient method to identify the root causes of these alarms to mitigate potential losses. This task is challenging due to the increasing scale of telecommunication networks and the interconnected nature of devices, where one fault can trigger a cascade of alarms across multiple devices within a topological network. Recent years have seen a growing focus on causal approaches to addressing this problem, emphasizing the importance of learning a Granger causal graph from topological event sequences. Such causal graphs delineate the relations among alarms and can significantly aid engineers in identifying and rectifying faults. However, existing methods either ignore the topological relationships among devices or suffer from relatively low scalability and efficiency, failing to deliver high-quality responses in a timely manner. To this end, this paper proposes $S^2GCSL$, a simple yet scalable Granger causal structural learning approach for topological event sequences.
What happens after the bombs drop: Scientists reveal the terrifying global aftermath of nuclear war
Furious Trump issues chilling threat to Iran demanding Strait of Hormuz is'FULLY OPENED' in hours or America will'obliterate their power plants'... and there's already a key target in sight Chappell Roan accused of'leaving Jude Law's 11-year-old daughter in tears and using security guard to threaten her' I was the only one JFK Jr and Carolyn Bessette trusted when they burdened me with an extraordinarily intimate secret. How Iran's ruthless enforcers use rape to crush dissent: Brutal sex attacks on victims as young as 12 used to strike fear into protesters, rights groups reveal amid fury over sickening nurse gang rape Shia LaBeouf suffers public meltdown in Rome as he's caught screaming'f*** off' at woman... after battery arrests'He just didn't protect him': Insiders reveal REAL reason Justin Bieber and Usher's secret feud hit'boiling point' at Oscars Mom-to-be finds out cop who got her pregnant has HIV after baby mama's text... as he is charged with felony I thought I was losing my mind... then doctors told me I had'exploding head syndrome'. America is about to be torn apart by a financial tsunami - and it's not just an oil crisis to fear. Denise Richards's plastic surgeon reveals stunning before-and-after photos of her facelift'Get the f*** out of my life,' JFK Jr screamed at Carolyn Bessette... what she cruelly told friends about his manhood... the cuckolding, cocaine - and moment that sent her truly psychotic: MAUREEN CALLAHAN has the untold REAL story Florida's Olivier Rioux, tallest player in college basketball history, dwarfs 6ft8 March Madness rival as defending champs roll to win YouTuber who exposed Somali'fraudsters' in bombshell investigation reveals terrifying threats from left-wing activists... as he begs for cash to help pay for security Charlie's Angels bombshell Jaclyn Smith looks nowhere near her 80 years in Beverly Hills... see her now Fury over plan for 110 homes near Yosemite Park that will tower up to 24ft and'cause road chaos' Gisele Pelicot tells how she thought she was dying from a brain tumor... then she discovered the horrific truth of her husband's abuse Iran ballistic missile hits Israeli city in terrifying strike near top-secret facility that is key to country's atomic weapons program Couple murdered outside Walgreens near golf's Players Championship were killed by jealous ex, says sheriff As the threat of a nuclear war intensifies, the terrifying reality of what could happen after the bombs explode may cause more fear than the initial cataclysm. For decades, worst-case scenarios have projected that tens of millions could perish within minutes as nuclear warheads struck major metropolitan areas such as New York, Washington, Chicago and Los Angeles .
French prosecutors suspect Musk encouraged deepfakes row to inflate X value
Elon Musk-owned X's Grok AI chatbot stirred outrage earlier this year over it generating images of naked women and girls without their consent. Paris - French prosecutors said Saturday they had alerted U.S. authorities to a suspicion that tech tycoon Elon Musk had encouraged controversy over sexualized deepfakes on X to artificially increase the value of his company. The social media network's Grok AI chatbot stirred outrage earlier this year over it generating images of naked women and girls without their consent. The controversy sparked by sexually explicit deepfakes generated by Grok (X's AI) may have been deliberately generated in order to artificially boost the value of companies X and xAI, the Paris prosecutor's office said, confirming a report in Le Monde newspaper on Friday. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right.
Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with Humans
Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes seven LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4, Mixtral-8x22B, and LLaMA-3.1. We find that, despite several misalignments, LLMs can generally respond appropriately to certain situations. Nevertheless, they fall short in alignment with the emotional behaviors of human beings and cannot establish connections between similar situations.