Enhancing Large Vision Language Models with Self-Training on Image Comprehension 1,3, Fan Yin