Most currently deployed LLMs undergo continuous training or additional finetun-ing. By contrast, most research into LLMs' internal mechanisms focuses on models
Based on GAIA, we evaluate a suite of popular text-to-video (T2V) models on their ability to generate visually rational actions, revealing their pros and cons on different categories of actions.