Country
2172fde49301047270b2897085e4319d-Supplemental.pdf
We define a Gibbs distribution associated with the problem and evaluate its normalization constant -thepartition functionZ. The partition function, and in particular its logarithm divided by the temperature - the free energy -, contains all the information we aim to understand in the problem. By taking the proper derivatives, and possibly add an external field, we can compute relevant macroscopic properties such as: average overlapwiththegroundtruth,averagelossachieved. Indisordered systemwehavetoconsider the additional complication given by the randomness.
216f44e2d28d4e175a194492bde9148f-Paper.pdf
We assume the environment modeled as discrete-time factored-action MDP (FA-MDP)M = hS,A,P,R,γi where S is the set of states s, A is the set of vector-represented actionsa = (a1,...,am),P(s0|s,a) = Pr(st+1 = s0|st = s,at = a)isthe transition probability,R(s,a) R is the immediate reward for taking actiona in state s, and γ [0,1) is the discount factor.