Interview with William Yijiang Li: vision language models and the physical world
At the International Conference on Machine Learning (ICML 2026),, and presented their work Vision Language Models Cannot Reason About Physical Transformation . In this interview, William Yijiang Li tells us more about the research, the method the team used, and the controlled experiments they carried out. What is the topic of the research in your paper and why is it an interesting area for study? Our research asks a fairly fundamental question: do vision-language models actually understand how the physical world changes over time? Modern vision-language models are increasingly being used in areas such as robotics, embodied agents, and video understanding.
Oct-5-2026, 08:00:57 GMT
- Genre:
- Personal > Interview (0.67)
- Research Report > New Finding (0.47)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Natural Language (1.00)
- Representation & Reasoning > Agents (0.34)
- Information Technology > Artificial Intelligence