Home › Lectures › 3Blue1Brown — Neural networks

Backpropagation calculus

Explains how the chain rule computes neural-network cost sensitivities and how those derivatives are propagated backward through layers.

3Blue1Brown⏱ 10 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

The lecture derives backpropagation using a network with one neuron per layer. For a single training example, the cost depends on the output activation, which is computed from a weighted sum and then passed through an activation function. The chain rule breaks the cost’s sensitivity to a weight into three parts: how the weight changes the weighted sum, how the sum changes the activation, and how the activation changes the cost. For squared error, these derivatives depend on the output error, the activation function’s derivative, and the previous neuron’s activation. The weight’s gradient is then averaged across training examples to contribute to the gradient of the full cost.

The same reasoning applies to biases and earlier layers. The derivative of the weighted sum with respect to a bias is one; its derivative with respect to the previous activation is the connecting weight. Tracking that sensitivity makes it possible to repeat the chain-rule calculation backward through the network. With multiple neurons per layer, the equations gain indices, and a neuron’s effect on cost must include contributions from every path through which it influences the output. These derivatives form the gradient used to update weights and biases so as to reduce the cost.

Key ideas

Sample questions

For a neuron with incoming activations a₁, …, aₙ, corresponding weights w₁, …, wₙ, and bias b, let z = Σᵢ wᵢaᵢ + b and let the neuron’s activation be g(z), where g is a nonlinear function such as sigmoid or ReLU. Which statement correctly distinguishes the weighted input from the activation?

  1. Az is the nonlinear function’s output; g(z) is the weighted sum before the bias is added.
  2. Bz and g(z) are interchangeable because applying a nonlinear function does not change the neuron’s value.
  3. Cz is the weighted sum plus the bias; g(z) is the value passed on after applying the nonlinear function.
  4. Dz is the bias alone; g(z) is the weighted sum of the incoming activations.
Show answer

Correct answer: C. The weighted input combines incoming activations with their weights and adds the bias. Applying the nonlinear function to that result gives the neuron’s activation.

In a neural network, suppose the cost C depends on an activation A_L, which depends on a weighted input Z_L, which depends on a weight W_L. Which expression applies the chain rule to compute how the cost changes with that weight?

  1. A∂C/∂W_L = ∂C/∂A_L, because intermediate quantities do not affect the gradient
  2. B∂C/∂W_L = (∂C/∂A_L) + (∂A_L/∂Z_L) + (∂Z_L/∂W_L)
  3. C∂C/∂W_L = (∂C/∂A_L)(∂Z_L/∂A_L)(∂W_L/∂Z_L)
  4. D∂C/∂W_L = (∂C/∂A_L)(∂A_L/∂Z_L)(∂Z_L/∂W_L)
Show answer

Correct answer: D. The cost changes through a sequence of dependencies, so its sensitivity to the weight is the product of the sensitivities at each link in that sequence.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →