Free, sign in with Google. A ready quiz does not use your hourly limit.
What the lecture covers
Backpropagation calculates how a training example should influence a network’s weights and biases to reduce its cost. For a handwritten-digit example, the output neuron for the correct digit should become more active, while the others should become less active. The size of each desired change depends on how far the output is from its target. To affect an output neuron, the network can adjust its bias and incoming weights. Connections from more active neurons have greater influence, so changing their weights can have a larger effect. The algorithm also tracks how earlier-layer activations would need to change, with the influence of each connection depending on its weight.
Each neuron’s desired changes are combined to determine the nudges for the preceding layer. Repeating this process backward through the network gives the effects on earlier weights and biases. A single example’s nudges are not enough: the network combines contributions from many examples to estimate the negative gradient, the direction that reduces cost. Because processing the full dataset for every update is slow, training commonly uses randomly shuffled mini-batches. Their updates approximate the full-data gradient while requiring less computation; repeating them moves the network toward a lower-cost solution. The method relies on sufficient labeled training data, such as the MNIST handwritten-digit examples.
Key ideas
Backpropagation computes the gradient information needed to adjust weights and biases so the network’s cost decreases.
For a digit-recognition example, the correct output should rise and the incorrect outputs should fall, with larger adjustments for outputs farther from their targets.
An output activation can be influenced by its bias, its incoming weights, and the activations in the preceding layer.
Weights connected to more active neurons have greater influence on the output and can warrant larger adjustments.
Backpropagation combines the desired effects of output neurons to determine how the preceding layer should change, then repeats the process backward.
The network averages the desired weight and bias changes across training examples to approximate the negative gradient.
Stochastic gradient descent uses shuffled mini-batches to make faster, approximate updates instead of processing the entire dataset each time.
The method depends on having enough labeled training data to guide learning.
Sample questions
In gradient-based neural network learning, the weights and biases are adjusted using examples with known desired outputs. What objective should these adjustments serve?
ATo increase the cost whenever the network’s output differs from the desired output
BTo ensure the network produces the same output for every training example
CTo make every weight and bias as large as possible
DTo make the network’s cost across the training examples as small as possible
Show answer
Correct answer: D. Learning aims to find parameter values that reduce the cost associated with the network’s outputs on the training examples.
In a neural network, the cost depends on its weights and biases. What does backpropagation provide to support choosing parameter updates that reduce this cost?
AA guarantee that every update reaches the global minimum
BThe gradient of the cost with respect to the weights and biases
CThe minimum possible value of the cost function
DA new set of weights and biases without using the cost
Show answer
Correct answer: B. The gradient describes how the cost changes with the network’s parameters, so it can guide updates toward lower cost.