Home › Lectures › 3Blue1Brown — Neural networks

Gradient Descent, How Neural Networks Learn

Learn how gradient descent adjusts a neural network’s weights by minimizing a cost function, and what its successes and limitations reveal.

3Blue1Brown⏱ 21 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

A neural network for digit recognition maps 784 pixel values to ten output activations. Its weights and biases determine how it classifies an image. To train it, compare its output with the correct label using a cost function: the squared differences for one example, averaged across the training set. Learning means changing the weights and biases to reduce this cost and improve performance on examples the network has not seen.

Gradient descent provides a way to make those changes. The gradient indicates how the cost changes with each parameter; stepping in the negative-gradient direction reduces it most quickly locally. Repeating small steps moves the parameters toward a local minimum, though the result depends on where training starts and need not be the global minimum. The gradient’s components also indicate which parameter adjustments have the greatest effect. Efficiently computing it for a neural network is the job of backpropagation.

A basic network can classify most test digits correctly, but that does not mean its hidden layers learn intuitive features such as edges and loops. In the example, their weights look nearly random, and the network confidently labels random noise. The lecture also notes that larger networks can memorize randomly shuffled labels, while structured data can make good solutions easier to find. Minimizing cost alone does not guarantee human-like understanding.

Key ideas

Sample questions

A neuron receives inputs 2 and −1, with corresponding weights 0.5 and 3, and has a bias of 1. If it applies ReLU, which activation does it output?

  1. A2
  2. B0
  3. C1
  4. D−1
Show answer

Correct answer: B. First combine each input with its weight and add the bias: 2(0.5) + (−1)(3) + 1 = −1. ReLU outputs 0 for a negative value.

After training an image-classification network, you want to check whether it generalizes beyond its training examples. Which test best serves this purpose?

  1. ARemove the labels from its training images and inspect its predictions.
  2. BTrain it longer on its existing training images and inspect its training performance.
  3. CEvaluate it again on the same images used to train it.
  4. DEvaluate it on labeled images that were not used during training.
Show answer

Correct answer: D. Previously unseen labeled images show whether the network performs well beyond the examples it learned from.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →