Goto

Collaborating Authors

 Country


UnderstandingEnd-to-EndModel-Based ReinforcementLearningMethodsasImplicit Parameterization

Neural Information Processing Systems

While knowntobesample efficient, these methods havefailed tofully leverage recent advances indeep learning, forcing the use of less efficient but more scalable model-free methods which try to learn the values directly.




07f560092a0edceabf55af32a40eaee3-Paper-Datasets_and_Benchmarks.pdf

Neural Information Processing Systems

First,theirsemantic feature extractions are outdated while state-of-the-art large-scale pre-trained language models like BERT cannot be utilized due to the lack of original text.


As stated in Section A, we apply the softmax function such thatRAPsoftmax outputs a synthetic datasetdrawnfromsomeprobabilisticfamilyofdistributionsD = n σ(M)| M Rn

Neural Information Processing Systems

Pt i=1eqi(x)(eai eqi(Di 1)) which is the exactly the distribution computed byMWEM. D(x)log(D(x)) (6) The optimization problem becomesDt = argminD (X)Lmwem(D, eQt, eAt). We show the exact details ofGEM in Algorithms 2 and 3. Note that given a vector of queries Qt = hq1,...,qti,wedefinefQt() = hfq1(),...,fqt()i. B.1 Lossfunction(fork-waymarginals)anddistributionalfamily For anyz R,G(z)outputs a distribution over each attribute, which we can use to calculate the answer toaquery viafq. Empirically,wefindthatour model tends to better capture the distribution of the overall private dataset in this way (Figure 3).