Home › Lectures › 3Blue1Brown — Neural networks

How Might LLMs Store Facts?

The lecture explains how an MLP could encode a fact such as “Michael Jordan plays basketball” and why real model features may be distributed across neurons.

3Blue1Brown⏱ 23 minOpen on YouTube ↗
Take the full quiz — 6 questions →

Free, sign in with Google. A ready quiz does not use your hourly limit.

What the lecture covers

The lecture explores how facts might be represented inside a transformer, focusing on its multilayer perceptron (MLP) blocks. In a simplified example, vectors encode the first name Michael, the last name Jordan, and the sport basketball as directions in a high-dimensional space. An MLP first uses a matrix to detect features of an input vector, then applies a nonlinear function such as ReLU. This can make a neuron respond only when both name features are present. A second matrix maps that activation back into the embedding space, where it can add the basketball direction to the vector. The MLP applies the same operations independently to each token position, while attention is what allows information to pass between positions.

The example illustrates a possible mechanism, not a complete account of how real models store facts. In GPT-3, MLPs contain about 116 billion parameters, roughly two thirds of the model’s total. Research suggests that individual neurons rarely correspond neatly to single features. The lecture discusses superposition: in high-dimensional spaces, many more than one feature per dimension can be represented by directions that are nearly, rather than perfectly, perpendicular. As a result, a feature may be encoded by a combination of neurons instead of one neuron activating on its own. This may help explain both the difficulty of interpreting models and their capacity to represent many ideas.

Key ideas

Sample questions

In a simplified account of a transformer that must retain the association “Michael Jordan plays basketball,” what role is attributed to its MLP component?

  1. AIt stores facts only by changing the model’s vocabulary.
  2. BIt replaces the need for other components of the transformer.
  3. CIt provides additional capacity in which factual associations can be stored.
  4. DIt guarantees that every factual association will be recalled correctly.
Show answer

Correct answer: C. The MLP is presented as adding capacity for storing facts; this does not imply perfect recall or that it replaces the rest of the model.

If a transformer represents the full name “Michael Jordan” with one vector and uses separate directions for “Michael” and “Jordan,” what must be true for that vector to score one against both directions, given that the name is split across two tokens?

  1. AAn earlier attention block must have made information from the first token available at the second token.
  2. BA later MLP must merge the two directions after the name has already been represented.
  3. CThe two token directions must be identical, so the vector only needs to encode one name part.
  4. DThe full-name vector must be stored independently at both token positions, without information passing between them.
Show answer

Correct answer: A. A single vector associated with the complete name can reflect both token-specific directions only if information about the first token reaches the second token’s position before that representation is formed.

Take the full quiz — 6 questions →

A quiz for any lecture

Paste a video link — LearnReplay builds a comprehension quiz.

Create a quiz →