In a Transformer, attention lets each position in a sequence select and combine information from other positions. A token’s query is compared with other tokens’ keys; those comparisons become weights that determine how much of each corresponding value to use. That simple operation helped make it possible to build sequence models that process positions in parallel during training.
A visual mental model: ask, match, retrieve
Imagine a library information desk. A visitor asks a question, the catalog has labels for possible sources, and the selected books contain the information to retrieve. In attention, these are useful analogies for a query, keys, and values—but they are learned numerical vectors, not literal questions, labels, or books.
- Query (Q): what the current position is looking for.
- Keys (K): representations compared against the query to determine relevance.
- Values (V): the information each position makes available to combine.
The computation can be read from left to right:
Q × Kᵀ → divide by √dₖ → optional mask → softmax weights → weighted sum with V
For each query, the model scores the keys, converts the scores into weights, and uses those weights to blend the values. The result is a context-aware representation for that position.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
How scaled dot-product attention works
The original Transformer paper defines the operation as:
Attention(Q, K, V) = softmax(QKᵀ / √dₖ)V
- Compare query and keys. The dot product in
QKᵀgives a compatibility score for each query-key pair. - Scale the scores. Divide by the square root of the key dimension,
√dₖ. The paper explains that this helps keep large dot products from pushing softmax into regions with very small gradients. - Apply a mask when needed. A mask can prevent attention to disallowed positions, such as future target tokens during autoregressive generation.
- Turn scores into weights. Softmax converts the scores into normalized weights.
- Blend the values. Multiply the weights by
Vand sum, producing a weighted combination.
Dot-product attention was attractive in the paper partly because it can use optimized matrix multiplication. In the authors’ comparison, it was faster and more space-efficient in practice than additive attention; that historical comparison should not be read as a claim about every modern implementation.
Rank #2
What attention is doing inside a Transformer
Self-attention connects positions in one sequence
In self-attention, queries, keys, and values are all derived from the same sequence representation. Each position can therefore gather information from other positions in that sequence. A word representation, for example, can be updated using information from the surrounding tokens rather than being processed only through its immediate neighbors.
Encoder-decoder attention connects two sequences
In encoder-decoder attention, decoder queries are compared with keys from the encoder output, and the resulting weights combine encoder values. This gives the decoder a way to draw on the encoded input while producing an output sequence.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
Masking keeps autoregressive predictions from looking ahead
When a decoder predicts a sequence one position at a time, it masks future target positions. A prediction at position i can then use earlier target outputs, but not later ones. Without that restriction, training could expose the model to information unavailable when generating the next token.
Position information is added separately
Attention by itself does not encode token order. In the original Transformer, the authors added positional encodings to token embeddings, using sine and cosine functions at different frequencies. This is the design in that paper, not a requirement that every later Transformer use the same positional scheme.
Attention is also not the entire Transformer layer. The original encoder and decoder layers include feed-forward sublayers, residual connections, and normalization alongside attention.
Why use multiple attention heads?
Multi-head attention runs several attention computations in parallel, each with its own learned query, key, and value projections. The outputs from the heads are concatenated and projected again. This lets the model combine information from different representation subspaces and positions.
Heads are not guaranteed to have neat, human-readable jobs. A particular head may show a pattern that is interesting to inspect, but assigning it a simple linguistic role requires evidence beyond its existence. In the original paper’s base configuration, the authors used eight heads, with 64-dimensional keys and values per head.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the original Transformer changed—and what its results mean
Ashish Vaswani and coauthors described the architectural shift directly: “We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.” Their 2017 paper emphasized parallelizability and training time alongside translation quality. Its reported results included:
| Original reported result | Context |
|---|---|
| 28.4 BLEU | WMT 2014 English-to-German, reported by Vaswani et al. in 2017. |
| 41.0 BLEU | WMT 2014 English-to-French, reported by Vaswani et al. in 2017. |
| 3.5 days on eight GPUs | Training time for the paper’s English-to-French model, reported by Vaswani et al. in 2017. |
These are the paper’s experimental results, not current records or a comparison of present-day training costs.
How attention heatmaps should—and should not—be read
A heatmap can show how strongly positions attend to one another for a selected input, head, and layer. Jesse Vig’s 2019 work describes head-level, whole-model, and neuron-level visualization views and demonstrates them with BERT and GPT-2. Such views can reveal patterns worth investigating, including positional or lexical patterns.
Recommended Free Tools
A visualization displays scores and patterns; by itself, it does not prove why a model produced a particular answer or provide a causal explanation of its behavior. Vig’s paper identifies empirical evaluation of attention’s impact on predictions as future work, a reminder that a vivid heatmap is not a transparent window into all model reasoning.
Quick Recap
Further reading
- Vaswani et al., Attention Is All You Need (NeurIPS 2017) — the original paper and its equations, architecture, and experiments.
- Google Research paper record — the publication abstract and reported results.
- Jesse Vig, Visualizing Attention in Transformer-Based Language Representation Models (2019) — approaches to viewing attention patterns.
- Harvard NLP, The Annotated Transformer — a line-by-line educational implementation.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




