End-to-End Learning of Communications Systems Without a Channel Model
Aoudia, Fayçal Ait, Hoydis, Jakob
–arXiv.org Artificial Intelligence
REINFORCEMENT LEARNING RL aims to optimize the behavior of agents that interact with an environment by taking actions in order to minimize a loss. An agent in a state s P S, takes an action a P A according to some policy π. After taking an action, the agent receives a perexample lossl. The expected per-example loss given a state and an action is denoted by Lps, aq, i.e., Lps, aq " E rl s, as. L is assumed to be unknown, and the aim of the agent is to find a policy which minimizes the per-example loss. During the agent's training, the policy π is usually chosen to be stochastic, i.e., πp sq is a probability distribution over the action space A conditional on a state s. Using a stochastic policy enables exploration of the agent's environment, which is fundamental in RL. Indeed, training in RL is similar to a tryand-fail process:the agent takes an action chosen according to its state, and afterwards improves its policy according to the loss received from the environment. Using a stochastic policy, the agent aims to minimize the loss Jps, πq, defined as ż Jps, πq " πpa sqLps, aq da.
arXiv.org Artificial Intelligence
Dec-5-2018