Home › Lectures › 3Blue1Brown — Neural networks

Attention in transformers, step-by-step

Explains how attention uses queries, keys, and values to update token embeddings with context, and how multiple heads operate in a transformer.

3Blue1Brown⏱ 26 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

A transformer begins by assigning each token an embedding that represents the token and its position. Attention lets these initially context-free vectors exchange information, so a word’s representation can reflect its use in a sentence. For example, context can distinguish different meanings of “mole” or make “tower” refer more specifically to the Eiffel Tower. This matters for next-token prediction: the final token’s vector must eventually carry information from the wider context that helps predict what comes next.

In one attention head, learned query and key maps produce vectors whose dot products measure how relevant tokens are to one another. A column-wise softmax turns these scores into weights; causal masking sets the influence of future tokens to zero while keeping the weights normalized. A learned value map transforms each token into information that can be added to other tokens’ embeddings, weighted by the attention pattern. Multiple heads perform different kinds of contextual updates in parallel, and their changes are combined. Repeated attention blocks and other transformer operations can build increasingly nuanced representations. The lecture also notes that attention’s parallelizable computations support efficient scaling, while its pattern grows quadratically with context length.

Key ideas

Sample questions

In a transformer’s embedding space, what can a consistent direction between token representations indicate?

  1. AThe order in which tokens appear in a sentence
  2. BA guarantee that the tokens have identical meanings
  3. CThe normalized attention weight assigned to each token
  4. DA semantic feature or relationship shared by changes along that direction
Show answer

Correct answer: D. Embedding-space directions can correspond to semantic meaning, so movement along a direction may reflect a semantic feature or relationship.

In a Transformer, the same token appears in several different sentences. Before surrounding tokens are used to update its representation, what determines its initial embedding?

  1. AIts position in the sentence, which gives it a different initial embedding each time.
  2. BThe meanings of nearby tokens, which change its embedding immediately.
  3. CThe attention weights computed from the full sentence.
  4. DThe token identity alone, so its initial embedding is the same across those sentences.
Show answer

Correct answer: D. The initial embedding is retrieved from a token-based lookup and does not yet incorporate sentence context.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →