FlowMotion: Target-Predictive Conditional Flow Matching for Jitter-Reduced Text-Driven Human Motion Generation

Cuba, Manolo Canales, Melício, Vinícius do Carmo, Gois, João Paulo

arXiv.org Artificial Intelligence 

Achieving high-fidelity and temporally smooth 3D human motion generation remains a challenge, particularly within resource-constrained environments. We introduce FlowMotion, a novel method leveraging Conditional Flow Matching (CFM). FlowMotion incorporates a training objective within CFM that focuses on more accurately predicting target motion in 3D human motion generation, resulting in enhanced generation fidelity and temporal smoothness while maintaining the fast synthesis times characteristic of flow-matching-based methods. FlowMotion achieves state-of-the-art jitter performance, achieving the best jitter in the KIT dataset and the second-best jitter in the HumanML3D dataset, and a competitive FID value in both datasets. This combination provides robust and natural motion sequences, off ering a promising equilibrium between generation quality and temporal naturalness. Introduction The synthesis of 3D human body motion has diverse applications across fields such as robotics [1, 2, 3], VR/ AR [4, 5], entertainment [6, 7, 8, 9], and social interaction in virtual 3D spaces [10, 11, 12, 13]. While recent advances in 3D human motion generation are significant, several challenges remain. The inherent complexities of motion generation are exacerbated by the need to incorporate diverse constraints, such as spatial trajectories [14, 15], interactions with surrounding objects [16, 17], or temporal specifications defined by keyframes [18], all aimed at producing lifelike movements. To achieve realistic and context-aware motion synthesis, motion generation techniques frequently leverage data-driven motion capture data. Recent studies in generative models have driven the development of new techniques for synthesizing 3D human body motion. These methods primarily focus on generating realistic movements based on user inputs, especially descriptive text that specifies the intended action. A key advantage of these generative approaches is their ability to produce a diverse range of plausible motion sequences from a single prompt. This allows users to explore multiple interpretations of a desired movement and select the sequence that best aligns with their creative vision.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found