AI Academy / Simulation

See what attention actually calculates

Edit three tiny vectors and inspect dot products, normalized weights and the weighted output. No model downloads or API calls.

You will learn to

  • Compute scaled dot-product attention
  • Explain causal masking
  • Observe temperature and normalization

Before you start

Vectors, dot products and weighted averages

A deliberately small model

Three positions each have a two-dimensional vector. This experiment uses the same vectors as queries, keys and values, equivalent to identity projections. Real transformers learn different projections and combine multiple heads. The position labels are just A, B and C; they have no learned language meaning.

From similarity to weights

For each query, multiply matching coordinates with a key and sum them. Divide by the square root of the dimension, here √2. We additionally divide by the temperature control to explore concentration. Softmax exponentiates these scores and divides each result by the sum. Each row of weights sums to one. The implementation subtracts the largest available score before exponentiation for numerical stability.

Keep future positions out

With the causal mask enabled, row A can attend only to A, row B to A and B, and row C to all three. Masked scores receive zero weight after normalization. Without the mask, the first position can use later positions. A model trained to predict the next token must not peek at the future answer in its context.

Test an invariant

Set all vectors to 0,0 and disable the mask. Every score is zero, so each row has three weights of one third. Enable the mask: the rows become [1,0,0], [0.5,0.5,0], and [1/3,1/3,1/3]. The output is still zero because every value is zero. This is a real computation of one operation, not a language model, training run or generated response.

Try it yourself

Enable JavaScript for this interactive activity. You can read all lesson explanations above without it.

Continue exploring