FaceLLM: A Multimodal Large Language Model for Face Understanding
Shahreza, Hatef Otroshi, Marcel, Sébastien
–arXiv.org Artificial Intelligence
Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as those in facial images. In particular, tasks that require detailed understanding of facial structure, expression, emotion, and demographic features remain un-derexplored by MLLMs due to the lack of large-scale annotated face image-text datasets. In this work, we introduce F aceLLM, a multimodal large language model trained specifically for facial image understanding. T o construct the training data, we propose a novel weakly supervised pipeline that uses ChatGPT with attribute-aware prompts to generate high-quality question-answer pairs based on images from the FairFace dataset. The resulting corpus, called F airF aceGPT, covers a diverse set of attributes including expression, pose, skin texture, and forensic information. Our experiments demonstrate that FaceLLM improves the performance of MLLMs on various face-centric tasks and achieves state-of-the-art performance. This work highlights the potential of synthetic supervision via language models for building domain-specialized MLLMs, and sets a precedent for trustworthy, human-centric multimodal AI systems. FairFaceGPT dataset and pretrained FaceLLM models are publicly available in the project page.
arXiv.org Artificial Intelligence
Jul-15-2025
- Country:
- Europe > Switzerland (0.28)
- Genre:
- Research Report (0.64)
- Industry:
- Information Technology > Security & Privacy (0.69)
- Technology: