Goto

Collaborating Authors

 Large Language Model






ReST-MCTS: LLM Self-Training via Process Reward Guided Tree Search Dan Zhang

Neural Information Processing Systems

Recent methodologies in LLM self-training mostly rely on LLM generating responses and filtering those with correct output answers as training data. This approach often yields a low-quality fine-tuning training set (e.g., incorrect plans or intermediate reasoning).


Supplementary Material for CableInspect AD An Expert Annotated Anomaly Detection

Neural Information Processing Systems

For more information, please refer to the Distribution and Maintenance subsections of the datasheet provided in J. The annotations are in the COCO format. We provide detailed explanations on how the dataset can be read in the code repository.