Explains how attention uses queries, keys, and values to update token embeddings with context, and how multiple heads operate in a transformer.
Free, sign in with Google. A ready quiz does not use your hourly limit.
A transformer begins by assigning each token an embedding that represents the token and its position. Attention lets these initially context-free vectors exchange information, so a word’s representation can reflect its use in a sentence. For example, context can distinguish different meanings of “mole” or make “tower” refer more specifically to the Eiffel Tower. This matters for next-token prediction: the final token’s vector must eventually carry information from the wider context that helps predict what comes next.
In one attention head, learned query and key maps produce vectors whose dot products measure how relevant tokens are to one another. A column-wise softmax turns these scores into weights; causal masking sets the influence of future tokens to zero while keeping the weights normalized. A learned value map transforms each token into information that can be added to other tokens’ embeddings, weighted by the attention pattern. Multiple heads perform different kinds of contextual updates in parallel, and their changes are combined. Repeated attention blocks and other transformer operations can build increasingly nuanced representations. The lecture also notes that attention’s parallelizable computations support efficient scaling, while its pattern grows quadratically with context length.
Correct answer: D. Embedding-space directions can correspond to semantic meaning, so movement along a direction may reflect a semantic feature or relationship.
Correct answer: D. The initial embedding is retrieved from a token-based lookup and does not yet incorporate sentence context.
Paste a video link — LearnReplay builds a comprehension quiz.
Create a quiz →