Country
4a1d69d1f64c6b6df105b15984ca527a-Supplemental-Conference.pdf
The cluster sizes are unbalanced with|G 1|/|G 2| = 2, where we randomly choose200 number "0" and100 number of "5" for each repetition. Here we choose three clusters:G1 containing the "T-shirt/top",G2 containing the "Trouser" andG3 containing the "Dress", so that the number of clusters isK = 3 in the algorithms.
5631e6ee59a4175cd06c305840562ff3-Paper.pdf
Ateachtimestepoftheepisode,thelearnerobserves the current state of the environment, chooses one of theK available actions, and earns a reward. Consequently, the state of the environment changes according to the transition function of the underlying MDP, as a function of the previous state and the action taken by the learner.