SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
–arXiv.org Artificial Intelligence
These limitations lead to reduced robustness and generalization, making it difficult to meet the stringent demands of intelligent manufacturing for high accuracy, adaptability, and real-time responsiveness [8]. For these challenges, Multimodal Large Language Models (MLLMs) have recently attracted significant attention as a promising paradigm for industrial vision inspection [9]. By integrating heterogeneous modalities such as images and text, MLLMs exhibit enhanced semantic understanding and cross-modal reasoning capabilities [10], which are crucial for accurately interpreting intricate industrial scenes. This capability enables them to identify subtle defect patterns that would otherwise be overlooked by unimodal vision systems [11]. However, the adoption of MLLMs in real-world industrial environments is hindered by their massive parameter scale, high computational cost, and substantial memory footprint [12]. Deploying these models directly on resource-constrained edge devices often proves infeasible [13], while cloud-only processing introduces latency that disrupts the stringent real-time requirements of industrial pipelines [14]. To address these issues, we propose SAEC, a scene-aware enhanced edge-cloud collaborative industrial vision inspection framework with MLLM. The central idea of SAEC is to integrate scene-aware mechanisms with an edge-cloud collaborative architecture in order to balance accuracy and efficiency. By harmonizing multimodal reasoning with adaptive task scheduling across edge and cloud resources, SAEC achieves robust defect detection under complex scenarios, while maintaining scalability, resource efficiency, and practicality for deployment in real-world manufacturing systems.
arXiv.org Artificial Intelligence
Sep-23-2025