Home › Lectures › 3Blue1Brown — Neural networks

Backpropagation, intuitively

An intuitive explanation of how backpropagation computes weight and bias updates, and how mini-batches make gradient descent faster.

3Blue1Brown⏱ 13 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

Backpropagation calculates how a training example should influence a network’s weights and biases to reduce its cost. For a handwritten-digit example, the output neuron for the correct digit should become more active, while the others should become less active. The size of each desired change depends on how far the output is from its target. To affect an output neuron, the network can adjust its bias and incoming weights. Connections from more active neurons have greater influence, so changing their weights can have a larger effect. The algorithm also tracks how earlier-layer activations would need to change, with the influence of each connection depending on its weight.

Each neuron’s desired changes are combined to determine the nudges for the preceding layer. Repeating this process backward through the network gives the effects on earlier weights and biases. A single example’s nudges are not enough: the network combines contributions from many examples to estimate the negative gradient, the direction that reduces cost. Because processing the full dataset for every update is slow, training commonly uses randomly shuffled mini-batches. Their updates approximate the full-data gradient while requiring less computation; repeating them moves the network toward a lower-cost solution. The method relies on sufficient labeled training data, such as the MNIST handwritten-digit examples.

Key ideas

Sample questions

In gradient-based neural network learning, the weights and biases are adjusted using examples with known desired outputs. What objective should these adjustments serve?

  1. ATo increase the cost whenever the network’s output differs from the desired output
  2. BTo ensure the network produces the same output for every training example
  3. CTo make every weight and bias as large as possible
  4. DTo make the network’s cost across the training examples as small as possible
Show answer

Correct answer: D. Learning aims to find parameter values that reduce the cost associated with the network’s outputs on the training examples.

In a neural network, the cost depends on its weights and biases. What does backpropagation provide to support choosing parameter updates that reduce this cost?

  1. AA guarantee that every update reaches the global minimum
  2. BThe gradient of the cost with respect to the weights and biases
  3. CThe minimum possible value of the cost function
  4. DA new set of weights and biases without using the cost
Show answer

Correct answer: B. The gradient describes how the cost changes with the network’s parameters, so it can guide updates toward lower cost.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →