Finally, we want our algorithm to have low computational overhead, so that itcan be applied asawrapper on top ofarbitrary prediction methods, for both regression and classification.
This paper is about the problem of learning a stochastic policy for generating an object (like a molecular graph) from a sequence of actions, such that the probability of generating an object isproportional to agiven positivereward for that object.