Ted’s Public Workbench Interactive Learning
Inside AI: Explore a Transformer
Follow a sentence through a transformer.
A language model builds an answer one token at a time. Look inside the calculation: how text becomes numbers, how attention connects the pieces, and how the next token is chosen. Change a setting and see what moves.
Three Things to Try
- Inspect tokenization
Compare a familiar word with an unusual name. Notice where the pieces split.
- Compare attention heads
Keep the same sentence and token. Look at which connections change.
- Change sampling settings
Adjust one setting at a time and watch the candidate probabilities.
The Interactive Model
Follow the Next Token
Begin with a prepared example. Select a token or an attention connection to inspect the calculation. Live mode downloads the model after you choose to load it; the explorer shows its progress and lets you cancel.
The interactive explorer could not load. Reload the page to try again, or keep reading the guide below.
Your text and model computations stay in this browser. Live mode downloads model files; this page does not send your prompts or record your session.
A Guided Explanation
From Text to the Next Token
Split the Text Into Tokens
A tokenizer turns your text into pieces from a fixed vocabulary. A token can be a whole word, part of a word, punctuation, or text that includes a leading space. The same visible word can therefore have different tokenizations in different contexts.
Each token receives a numeric ID. An ID is a vocabulary lookup key, not a score: a larger number does not mean a more important word.
Try it: Compare an ordinary word with an unusual name. Inspect the token pieces before looking at the prediction.
Give Each Token a Vector and a Position
The model looks up a learned list of numbers, called an embedding, for each token ID. GPT-2 adds a learned position embedding so the calculation can distinguish the first token from the fifth.
These vectors begin the model’s working representation of the text. Individual coordinates are learned values; they do not come with reliable labels such as “kindness” or “certainty.” A colored cell in the diagram is a number, not a named idea.
Look for: The separate token and position contributions. Follow their combined vector into the first transformer block.
Use Attention to Combine Context
Each attention head makes three projections of the input vectors: queries, keys, and values. Comparing a token’s query with the available keys produces scores. Softmax turns the scores into weights, and those weights determine how much of each value vector to combine.
GPT-2 uses causal attention: a token can use itself and earlier tokens, but it cannot look ahead. Several heads do this in parallel with different learned projections. Their outputs are combined and projected back into the model’s working representation.
Read the attention equation
Attention(Q, K, V) = softmax(QKᵀ / √dₖ + M)V
QKᵀ compares queries with keys. Dividing by the square root of the key dimension keeps the scores on a useful scale. The causal mask M rules out future positions. Softmax makes the permitted weights add to one, and multiplying by V forms their weighted combination.
Try it: Select a token and compare its attention weights across heads. A large weight shows a strong connection in that calculation; it does not by itself explain the entire prediction.
Refine the Representation in Each Block
Attention is one part of a transformer block. Layer normalization adjusts the scale of the working values. A feed-forward network applies learned transformations and a nonlinear activation to each token position.
Residual connections add a block’s updates to the existing representation. Repeating the blocks lets later calculations work with the changes made by earlier ones. The model’s weights remain fixed during this demonstration: generating text is inference, not training.
Look for: The path around each transformation. That is the residual connection carrying the existing representation forward.
Score the Possible Next Tokens
The final representation at the last position is mapped to a score for every token in the vocabulary. These scores are called logits. Softmax converts them into a probability distribution whose values add to one.
A high probability means the token is favored as a continuation under this model and context. It is not a probability that a statement is true. The visible candidate list shows only part of the vocabulary unless the explorer explicitly says otherwise.
Try it: Compare two prompts that differ in one token. Watch how the highest-ranked continuations change.
Choose a Token, Then Repeat
A generation rule selects the next token from the distribution. Greedy selection chooses the highest-scoring token. Sampling can choose another candidate according to its probability.
- Temperature
- Rescales the logits before softmax. A lower positive temperature concentrates probability on leading candidates; a higher value spreads it more broadly.
- Top-k
- Keeps only the k highest-ranked candidates before sampling, then renormalizes their probabilities.
- Top-p
- Keeps a set of leading candidates whose cumulative probability reaches the selected threshold, then renormalizes. The number kept depends on the distribution.
The chosen token is appended to the context. The model then calculates another next-token distribution. An answer that appears all at once in a chat window is built through repeated steps like these.
Try it: Change one sampling control at a time. Notice whether it changes the ranking, the eligible candidates, or their relative probabilities.
Know What the Picture Shows
This explorer uses GPT-2 to make transformer mechanics visible. Modern language models can use different architectures, tokenizers, training methods, context lengths, and generation rules. The diagram is a teaching view of one model, not a complete picture of every AI system.
Prepared examples show recorded model results. Live mode runs the downloaded model in your browser. Illustrative vectors and simplified diagrams are labeled where they are used; they should not be read as measurements of a live prompt.
The explorer limits context to 32 tokens to keep the visualization manageable. Text generation can be inaccurate, biased, or offensive. Use the results to inspect a calculation rather than as a source of factual answers.
Keep asking: Which numbers did the model produce, which choices did the generation rule make, and which parts did the diagram simplify?
Built on Transformer Explainer
Inside AI adapts Transformer Explainer by the Polo Club of Data Science. The original work is available under the MIT license. Ted Tschopp’s adaptation adds this site’s design, guided explanations, and browser-loading controls.