TinyClick: Single-Turn Agent for Empowering GUI Automation
Pawlowski, Pawel, Zawistowski, Krystian, Lapacz, Wojciech, Skorupa, Marcin, Wiacek, Adam, Postansque, Sebastien, Hoscilowicz, Jakub
–arXiv.org Artificial Intelligence
We present a single-turn agent for graphical user interface (GUI) interaction tasks, using Vision-Language Model Florence-2-Base. The agent's primary task is identifying the screen coordinates of the UI element corresponding to the user's command. It demonstrates strong performance on Screenspot and OmniAct, while maintaining a compact size of 0.27B parameters and minimal latency. Relevant improvement comes from multi-task training and MLLM-based data augmentation. Manually annotated corpora are scarce, but we show that MLLM augmentation might produce better results. On Screenspot and OmniAct, our model outperforms both GUI-specific models (e.g., SeeClick) and MLLMs (e.g., GPT-4V).
arXiv.org Artificial Intelligence
Oct-17-2024
- Country:
- North America
- United States (0.05)
- Canada > Ontario
- Toronto (0.04)
- Europe > Poland
- Masovia Province > Warsaw (0.04)
- North America
- Genre:
- Research Report (0.64)
- Technology:
- Information Technology
- Graphics (1.00)
- Communications (1.00)
- Human Computer Interaction > Interfaces (0.91)
- Artificial Intelligence
- Vision (1.00)
- Natural Language (1.00)
- Machine Learning (1.00)
- Information Technology