Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization

Open in new window