Free, sign in with Google. A ready quiz does not use your hourly limit.
What the lecture covers
The lecture derives backpropagation using a network with one neuron per layer. For a single training example, the cost depends on the output activation, which is computed from a weighted sum and then passed through an activation function. The chain rule breaks the cost’s sensitivity to a weight into three parts: how the weight changes the weighted sum, how the sum changes the activation, and how the activation changes the cost. For squared error, these derivatives depend on the output error, the activation function’s derivative, and the previous neuron’s activation. The weight’s gradient is then averaged across training examples to contribute to the gradient of the full cost.
The same reasoning applies to biases and earlier layers. The derivative of the weighted sum with respect to a bias is one; its derivative with respect to the previous activation is the connecting weight. Tracking that sensitivity makes it possible to repeat the chain-rule calculation backward through the network. With multiple neurons per layer, the equations gain indices, and a neuron’s effect on cost must include contributions from every path through which it influences the output. These derivatives form the gradient used to update weights and biases so as to reduce the cost.
Key ideas
A simple network provides a setting for measuring how its weights and biases affect the cost.
The output cost depends on an activation computed by applying a nonlinear function to a weighted sum.
The chain rule expresses a weight’s effect on cost as a product of sensitivities along the computational path.
The output-error derivative, activation-function derivative, and previous activation together determine the weight sensitivity.
The single-example derivative is averaged across examples, and the bias derivative follows the same pattern.
Tracking sensitivity to earlier activations allows the chain-rule calculation to continue backward through the network.
For layers with multiple neurons, the same method applies with additional indices and summed contributions from multiple paths.
The resulting derivatives make up the gradient used to step downhill and reduce network cost.
Sample questions
For a neuron with incoming activations a₁, …, aₙ, corresponding weights w₁, …, wₙ, and bias b, let z = Σᵢ wᵢaᵢ + b and let the neuron’s activation be g(z), where g is a nonlinear function such as sigmoid or ReLU. Which statement correctly distinguishes the weighted input from the activation?
Az is the nonlinear function’s output; g(z) is the weighted sum before the bias is added.
Bz and g(z) are interchangeable because applying a nonlinear function does not change the neuron’s value.
Cz is the weighted sum plus the bias; g(z) is the value passed on after applying the nonlinear function.
Dz is the bias alone; g(z) is the weighted sum of the incoming activations.
Show answer
Correct answer: C. The weighted input combines incoming activations with their weights and adds the bias. Applying the nonlinear function to that result gives the neuron’s activation.
In a neural network, suppose the cost C depends on an activation A_L, which depends on a weighted input Z_L, which depends on a weight W_L. Which expression applies the chain rule to compute how the cost changes with that weight?
A∂C/∂W_L = ∂C/∂A_L, because intermediate quantities do not affect the gradient
B∂C/∂W_L = (∂C/∂A_L) + (∂A_L/∂Z_L) + (∂Z_L/∂W_L)
C∂C/∂W_L = (∂C/∂A_L)(∂Z_L/∂A_L)(∂W_L/∂Z_L)
D∂C/∂W_L = (∂C/∂A_L)(∂A_L/∂Z_L)(∂Z_L/∂W_L)
Show answer
Correct answer: D. The cost changes through a sequence of dependencies, so its sensitivity to the weight is the product of the sensitivities at each link in that sequence.