Technology
xLSTM: Extended Long Short-Term Memory
In the 1990s, the constant error carousel and gating were introduced as the central ideas of the Long Short-Term Memory (LSTM). Since then, LSTMs have stood the test of time and contributed to numerous deep learning success stories, in particular they constituted the first Large Language Models (LLMs). However, the advent of the Transformer technology with parallelizable self-attention at its core marked the dawn of a new era, outpacing LSTMs at scale. We now raise a simple question: How far do we get in language modeling when scaling LSTMs to billions of parameters, leveraging the latest techniques from modern LLMs, but mitigating known limitations of LSTMs? Firstly, we introduce exponential gating with appropriate normalization and stabilization techniques. Secondly, we modify the LSTM memory structure, obtaining: (i) sLSTM with a scalar memory, a scalar update, and new memory mixing, (ii) mLSTM that is fully parallelizable with a matrix memory and a covariance update rule. Integrating these LSTM extensions into residual block backbones yields xLSTM blocks that are then residually stacked into xLSTM architectures. Exponential gating and modified memory structures boost xLSTM capabilities to perform favorably when compared to state-of-the-art Transformers and State Space Models, both in performance and scaling.
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
Transformer-based diffusion models have achieved significant advancements across a variety of generative tasks. However, producing high-quality outputs typically necessitates large transformer models, which result in substantial training and inference overhead. In this work, we investigate an alternative approach involving multiple experts for denoising, and introduce RemixDiT, a novel method designed to enhance output quality at a low cost. The goal of RemixDiT is to craft N diffusion experts for different denoising timesteps, yet without the need for expensive training of N independent models. To achieve this, RemixDiT employs K basis models (where K < N) and utilizes learnable mixing coefficients to adaptively craft expert models. This design offers two significant advantages: first, although the total model size is increased, the model produced by the mixing operation shares the same architecture as a plain model, making the overall model as efficient as a standard diffusion transformer.
You're doing your laundry wrong! Experts reveal why you should NEVER close the washing machine door after a wash
Furious Trump issues threat to Iran demanding Strait of Hormuz is'FULLY OPENED' in hours or America will'obliterate their power plants'... and there's already a key target in sight Nancy Guthrie's desperate family releases emotional new statement as they plead for'renewed attention to our mom's case' 50 days after she vanished Chappell Roan accused of'leaving Jude Law's 11-year-old daughter in tears and using security guard to threaten her' I was the only one JFK Jr and Carolyn Bessette trusted when they burdened me with an extraordinarily intimate secret. How Iran's ruthless enforcers use rape to crush dissent: Brutal sex attacks on victims as young as 12 used to strike fear into protesters, rights groups reveal amid fury over sickening nurse gang rape Shia LaBeouf suffers public meltdown in Rome as he's caught screaming'f*** off' at woman... after battery arrests'He just didn't protect him': Insiders reveal REAL reason Justin Bieber and Usher's secret feud hit'boiling point' at Oscars I thought I was losing my mind... then doctors told me I had'exploding head syndrome'. America is about to be torn apart by a financial tsunami - and it's not just an oil crisis to fear. Denise Richards's plastic surgeon reveals stunning before-and-after photos of her facelift'Get the f*** out of my life,' JFK Jr screamed at Carolyn Bessette... what she cruelly told friends about his manhood... the cuckolding, cocaine - and moment that sent her truly psychotic: MAUREEN CALLAHAN has the untold REAL story'Meteorite' CRASHES into woman's home as residents are left terrified by massive sonic boom YouTuber who exposed Somali'fraudsters' in bombshell investigation reveals terrifying threats from left-wing activists... as he begs for cash to help pay for security Sabine Getty's gown gets STUCK in escalator at Oscars during dramatic moment Charlie's Angels bombshell Jaclyn Smith looks nowhere near her 80 years in Beverly Hills... see her now Florida's Olivier Rioux, tallest player in college basketball history, dwarfs 6ft8 March Madness rival as defending champs roll to win RFK Jr reveals his strange daily routine... including not eating until noon and meditating with'dead people' Iran ballistic missile hits Israeli city in terrifying strike near top-secret facility that is key to country's atomic weapons program Ted Cruz proposes dramatic change to ICE funding in repsonse to'extreme and unreasonable' Democrats in effort to end bitter standoff causing airport chaos And now it turns out you've probably been doing your laundry wrong this entire time. Experts at AO.com have revealed why you should never close the washing machine door after a wash.
UniMTS: Unified Pre-training for Motion Time Series
Motion time series collected from low-power, always-on mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR. However, given security and privacy concerns, building large-scale motion time series datasets remains difficult, hindering the development of pre-trained models for human activity analysis. Typically, existing models are trained and tested on the same dataset, leading to poor generalizability across variations in device location, device mounting orientation, and human activity type. In this paper, we introduce UniMTS, the first unified pre-training procedure for motion time series that generalizes across diverse device latent factors and activities. Specifically, we employ a contrastive learning framework that aligns motion time series with text descriptions enriched by large language models.
RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier
State-of-the-art (SOTA) semi-supervised learning techniques, such as FixMatch and it's variants, have demonstrated impressive performance in classification tasks. However, these methods are not directly applicable to regression tasks. In this paper, we present RankUp, a simple yet effective approach that adapts existing semi-supervised classification techniques to enhance the performance of regression tasks. RankUp achieves this by converting the original regression task into a ranking problem and training it concurrently with the original regression objective.
GFT: Graph Foundation Model with Transferable Tree Vocabulary
Inspired by the success of foundation models in applications such as ChatGPT, as graph data has been ubiquitous, one can envision the far-reaching impacts that can be brought by Graph Foundation Models (GFMs) with broader applications in the areas such as scientific research, social network analysis, drug discovery, and e-commerce. Despite the significant progress of pre-trained graph neural networks, there haven't been GFMs that can achieve desired performance on various graph-learning-related tasks. Building GFMs may rely on a vocabulary that encodes transferable patterns shared among different tasks and domains. Unlike image and text, defining such transferable patterns for graphs remains an open question. In this paper, we aim to bridge this gap by rethinking the transferable patterns on graphs as computation trees -- i.e., tree structures derived from the message-passing process. Based on this insight, we propose a cross-task, cross-domain graph foundation model named GFT, short for Graph Foundation model with transferable Tree vocabulary. By treating computation trees as tokens within the transferable vocabulary, GFT improves model generalization and reduces the risk of negative transfer. The theoretical analyses and extensive experimental studies have demonstrated the transferability of computation trees and shown the effectiveness of GFT across diverse tasks and domains in graph learning.
ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose Estimation
We propose ManiPose, a manifold-constrained multi-hypothesis model for human-pose 2D-to-3D lifting. We provide theoretical and empirical evidence that, due to the depth ambiguity inherent to monocular 3D human pose estimation, traditional regression models suffer from pose-topology consistency issues, which standard evaluation metrics (MPJPE, P-MPJPE and PCK) fail to assess. ManiPose addresses depth ambiguity by proposing multiple candidate 3D poses for each 2D input, each with its estimated plausibility.
FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor Scenes
Empowering 3D Gaussian Splatting with generalization ability is appealing. However, existing generalizable 3D Gaussian Splatting methods are largely confined to narrow-range interpolation between stereo images due to their heavy backbones, thus lacking the ability to accurately localize 3D Gaussian and support free-view synthesis across wide view range. In this paper, we present a novel framework FreeSplat that is capable of reconstructing geometrically consistent 3D scenes from long sequence input towards free-view synthesis.Specifically, we firstly introduce Low-cost Cross-View Aggregation achieved by constructing adaptive cost volumes among nearby views and aggregating features using a multi-scale structure. Subsequently, we present the Pixel-wise Triplet Fusion to eliminate redundancy of 3D Gaussians in overlapping view regions and to aggregate features observed across multiple views. Additionally, we propose a simple but effective free-view training strategy that ensures robust view synthesis across broader view range regardless of the number of views. Our empirical results demonstrate state-of-the-art novel view synthesis peformances in both novel view rendered color maps quality and depth maps accuracy across different numbers of input views. We also show that FreeSplat performs inference more efficiently and can effectively reduce redundant Gaussians, offering the possibility of feed-forward large scene reconstruction without depth priors. Our code will be made open-source upon paper acceptance.