Goto

Collaborating Authors

 Technology



Leveraging Drift to Improve Sample Complexity of Variance Exploding Diffusion Models

Neural Information Processing Systems

Variance exploding (VE) based diffusion models, an important class of diffusion models, have shown state-of-the-art (SOTA) performance. However, only a few theoretical works analyze VE-based models, and those works suffer from a worse forward convergence rate $1/\text{poly}(T)$ than the $\exp{(-T)}$ of variance preserving (VP) based models, where $T$ is the forward diffusion time and the rate measures the distance between forward marginal distribution $q_T$ and pure Gaussian noise. The slow rate is due to the Brownian Motion without a drift term. In this work, we design a new drifted VESDE forward process, which allows a faster $\exp{(-T)}$ forward convergence rate. With this process, we achieve the first efficient polynomial sample complexity for a series of VE-based models with reverse SDE under the manifold hypothesis. Furthermore, unlike previous works, we allow the diffusion coefficient to be unbounded instead of a constant, which is closer to the SOTA models. Besides the reverse SDE, the other common reverse process is the probability flow ODE (PFODE) process, which is deterministic and enjoys faster sample speed. To deepen the understanding of VE-based models, we consider a more general setting considering reverse SDE and PFODE simultaneously, propose a unified tangent-based analysis framework, and prove the first quantitative convergence guarantee for SOTA VE-based models with reverse PFODE.We also show that the drifted VESDE can balance different error terms and improve generated samples without training through synthetic and real-world experiments.


Towards Flexible Visual Relationship Segmentation

Neural Information Processing Systems

Visual relationship understanding has been studied separately in human-object interaction(HOI) detection, scene graph generation(SGG), and referring relationships(RR) tasks. Given the complexity and interconnectedness of these tasks, it is crucial to have a flexible framework that can effectively address these tasks in a cohesive manner.In this work, we propose FleVRS, a single model that seamlessly integrates the above three aspects in standard and promptable visual relationship segmentation, and further possesses the capability for open-vocabulary segmentation to adapt to novel scenarios. FleVRS leverages the synergy between text and image modalities, to ground various types of relationships from images and use textual features from vision-language models to visual conceptual understanding.Empirical validation across various datasets demonstrates that our framework outperforms existing models in standard, promptable, and open-vocabulary tasks, e.g., +1.9 $mAP$ on HICO-DET, +11.4 $Acc$ on VRD, +4.7 $mAP$ on unseen HICO-DET.Our FleVRS represents a significant step towards a more intuitive, comprehensive, and scalable understanding of visual relationships.


Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model

Neural Information Processing Systems

Spatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that are excessively customized and devoid of causal connections further undermine the generalizability and interpretability. To this end, we establish a causal framework for ST predictions, termed CaPaint, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Going beyond this process, we utilize the back-door adjustment to specifically address the sub-regions identified as non-causal in the upstream phase.


xLSTM: Extended Long Short-Term Memory

Neural Information Processing Systems

In the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM). Since then, LSTMs have stood the test of time and contributed to numerous deep learning success stories, in particular they constituted the first Large Language Models (LLMs). However, the advent of the Transformer technology with parallelizable self-attention at its core marked the dawn of a new era, outpacing LSTMs at scale. We now raise a simple question: How far do we get in language modeling when scaling LSTMs to billions of parameters, leveraging the latest techniques from modern LLMs, but mitigating known limitations of LSTMs? Firstly, we introduce exponential gating with appropriate normalization and stabilization techniques. Secondly, we modify the LSTM memory structure, obtaining: (i) sLSTM with a scalar memory, a scalar update, and new memory mixing, (ii) mLSTM that is fully parallelizable with a matrix memory and a covariance update rule. Integrating these LSTM extensions into residual block backbones yields xLSTM blocks that are then residually stacked into xLSTM architectures. Exponential gating and modified memory structures boost xLSTM capabilities to perform favorably when compared to state-of-the-art Transformers and State Space Models, both in performance and scaling.


Energy-Guided Continuous Entropic Barycenter Estimation for General Costs

Neural Information Processing Systems

Optimal transport (OT) barycenters are a mathematically grounded way of averaging probability distributions while capturing their geometric properties. In short, the barycenter task is to take the average of a collection of probability distributions w.r.t.


Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising

Neural Information Processing Systems

Transformer-based diffusion models have achieved significant advancements across a variety of generative tasks. However, producing high-quality outputs typically necessitates large transformer models, which result in substantial training and inference overhead. In this work, we investigate an alternative approach involving multiple experts for denoising, and introduce RemixDiT, a novel method designed to enhance output quality at a low cost. The goal of RemixDiT is to craft N diffusion experts for different denoising timesteps, yet without the need for expensive training of N independent models. To achieve this, RemixDiT employs K basis models (where K < N) and utilizes learnable mixing coefficients to adaptively craft expert models. This design offers two significant advantages: first, although the total model size is increased, the model produced by the mixing operation shares the same architecture as a plain model, making the overall model as efficient as a standard diffusion transformer.


You're doing your laundry wrong! Experts reveal why you should NEVER close the washing machine door after a wash

Daily Mail - Science & tech

Furious Trump issues threat to Iran demanding Strait of Hormuz is'FULLY OPENED' in hours or America will'obliterate their power plants'... and there's already a key target in sight Nancy Guthrie's desperate family releases emotional new statement as they plead for'renewed attention to our mom's case' 50 days after she vanished Chappell Roan accused of'leaving Jude Law's 11-year-old daughter in tears and using security guard to threaten her' I was the only one JFK Jr and Carolyn Bessette trusted when they burdened me with an extraordinarily intimate secret. How Iran's ruthless enforcers use rape to crush dissent: Brutal sex attacks on victims as young as 12 used to strike fear into protesters, rights groups reveal amid fury over sickening nurse gang rape Shia LaBeouf suffers public meltdown in Rome as he's caught screaming'f*** off' at woman... after battery arrests'He just didn't protect him': Insiders reveal REAL reason Justin Bieber and Usher's secret feud hit'boiling point' at Oscars I thought I was losing my mind... then doctors told me I had'exploding head syndrome'. America is about to be torn apart by a financial tsunami - and it's not just an oil crisis to fear. Denise Richards's plastic surgeon reveals stunning before-and-after photos of her facelift'Get the f*** out of my life,' JFK Jr screamed at Carolyn Bessette... what she cruelly told friends about his manhood... the cuckolding, cocaine - and moment that sent her truly psychotic: MAUREEN CALLAHAN has the untold REAL story'Meteorite' CRASHES into woman's home as residents are left terrified by massive sonic boom YouTuber who exposed Somali'fraudsters' in bombshell investigation reveals terrifying threats from left-wing activists... as he begs for cash to help pay for security Sabine Getty's gown gets STUCK in escalator at Oscars during dramatic moment Charlie's Angels bombshell Jaclyn Smith looks nowhere near her 80 years in Beverly Hills... see her now Florida's Olivier Rioux, tallest player in college basketball history, dwarfs 6ft8 March Madness rival as defending champs roll to win RFK Jr reveals his strange daily routine... including not eating until noon and meditating with'dead people' Iran ballistic missile hits Israeli city in terrifying strike near top-secret facility that is key to country's atomic weapons program Ted Cruz proposes dramatic change to ICE funding in repsonse to'extreme and unreasonable' Democrats in effort to end bitter standoff causing airport chaos And now it turns out you've probably been doing your laundry wrong this entire time. Experts at AO.com have revealed why you should never close the washing machine door after a wash.


UniMTS: Unified Pre-training for Motion Time Series

Neural Information Processing Systems

Motion time series collected from low-power, always-on mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR. However, given security and privacy concerns, building large-scale motion time series datasets remains difficult, hindering the development of pre-trained models for human activity analysis. Typically, existing models are trained and tested on the same dataset, leading to poor generalizability across variations in device location, device mounting orientation, and human activity type. In this paper, we introduce UniMTS, the first unified pre-training procedure for motion time series that generalizes across diverse device latent factors and activities. Specifically, we employ a contrastive learning framework that aligns motion time series with text descriptions enriched by large language models.


RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier

Neural Information Processing Systems

State-of-the-art (SOTA) semi-supervised learning techniques, such as FixMatch and it's variants, have demonstrated impressive performance in classification tasks. However, these methods are not directly applicable to regression tasks. In this paper, we present RankUp, a simple yet effective approach that adapts existing semi-supervised classification techniques to enhance the performance of regression tasks. RankUp achieves this by converting the original regression task into a ranking problem and training it concurrently with the original regression objective.