Learning to Search Better Than Your Teacher

Chang, Kai-Wei, Krishnamurthy, Akshay, Agarwal, Alekh, Daumé, Hal III, Langford, John

May-20-2015–arXiv.org Machine Learning

Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Can learning to search work even when the reference is poor? We provide a new learning to search algorithm, LOLS, which does well relative to the reference policy, but additionally guarantees low regret compared to deviations from the learned policy: a local-optimality guarantee. Consequently, LOLS can improve upon the reference policy, unlike previous algorithms. This enables us to develop structured contextual bandits, a partial information structured prediction setting with many potential applications.

artificial intelligence, machine learning, natural language, (18 more...)

arXiv.org Machine Learning

May-20-2015

arXiv.org PDF

Add feedback

Country:
- North America > United States > Maryland (0.28)

Genre:
- Research Report (0.64)

Technology:
- Information Technology > Artificial Intelligence
  - Representation & Reasoning > Search (1.00)
  - Machine Learning (1.00)
  - Natural Language > Grammars & Parsing (0.94)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found