internalized self-correction
Internalized Self-Correction for Large Language Models
Upadhyaya, Nishanth, Sridharamurthy, Raghavendra
In this article, we introduce'Internalized Self-Correction' (InSeC) for large language models (LLMs). While many approaches exist for self-reflect ion at inference time, we propose a novel method that combines ideas from neg ative sampling, self-reflection during training, and inference time. InSeC allow s LLMs to correct themselves by introducing mistakes and their correspond ing corrections during training, thereby converting the learning process into a tru e supervised learning task with both positive and negative examples. This approac h can be extended to improve instruction following and correct hallucinations or incorrect sentences generated by LLMs.