Vanishing gradient in Deep Neural Network
Let us consider the succession of neurons in Figure 1. We update the weight w (1) by calculating the derivative of Loss function respect to it. For simplicity, we assume that the bias term b is zero for each neuron. Let's ask, what can stop or slow down the gradient? We analyze the product a / z.
Sep-25-2021, 22:30:15 GMT
- Technology: