October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

Encoder-Only vs. Decoder-Only vs. Encoder-Decoder: How Transformer Attention Works

The attention equation is shared across Transformer families; masks and information flow distinguish bidirectional encoders, causal decoders, and encoder-decoder models.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The difference between encoder-only, decoder-only, and encoder-decoder Transformers is mainly how their attention blocks are arranged and which tokens each position may use—not three different attention equations. Encoders commonly read a full input bidirectionally; causal decoders limit each position to the prefix; encoder-decoder models combine both patterns and add cross-attention from output positions to encoded input.

The shared attention equation

All three designs build on scaled dot-product attention:

As an Amazon Associate I earn from qualifying purchases.

Attention(Q, K, V) = softmax(QKᵀ / √dₖ)V

Q contains queries, K keys, and V values. The product QKᵀ scores how strongly each query matches each key. Dividing by the square root of the key dimension, √dₖ, controls the score scale before softmax. Softmax turns each row of scores into weights, and multiplying those weights by V produces a weighted sum of value vectors. The original Transformer paper describes the mechanism in Attention Is All You Need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Self-attention and cross-attention

In self-attention, queries, keys, and values are learned projections of the same sequence representation. In cross-attention, queries come from decoder states, while keys and values come from encoder states. That difference in the source of the vectors is what lets a decoder consult a separately encoded input.

#1 Best Overall
BLOKEES - Transformers - Action Edition 06 - Optimus Prime - Transformers: Prime - Model Kit - Assembly Required - 14+
  • MAGNETIC LED SYSTEM WITH BREATHING LIGHT: Touch-activated magnetic LEDs illuminate the chest and eyes with a 10-second breathing light effect, letting you trigger the Leader Module's awakening moment on demand.
  • ENHANCED DYNAMIC ARTICULATION: Unassembled 328-piece kit with two-stage elbow joints bending up to 160 degrees, highly flexible two-stage knees, and a Human System mechanical skeleton for stable, action-packed posing.
  • FULL WEAPON & ACCESSORY SET: Includes arm-cannons, arm-swords, the Star Saber, the Matrix, alternative faces, alternative shoulder armor, a battle mode mask, alternative vehicle windows, and interchangeable hands for recreating legendary battle scenes.
  • EASY SNAP-FIT ASSEMBLY: No glue or tools required — all parts and accessories snap securely into place for tool-free customization, complete with a display stand and instruction booklet.

What a mask changes

A mask is applied to attention scores before softmax. Connections a position cannot use receive a prohibitive score—conventionally negative infinity—and therefore zero weight after softmax. The equation stays the same; the mask changes which query-key connections are available.

Why use multiple heads?

Multi-head attention uses multiple learned query, key, and value projections. Each head computes attention separately; the outputs are concatenated and projected. Heads can learn different relationships among positions, but that does not mean every head has a reliably human-readable linguistic role.

How the three architectures differ

Architecture Typical attention pattern What a position can use Common task pattern Examples
Encoder-only Bidirectional self-attention Input tokens on either side Representing or classifying a complete input BERT-like encoders
Decoder-only Causal self-attention Current and earlier tokens; later target tokens are masked Next-token prediction and autoregressive generation GPT-like causal language models
Encoder-decoder Bidirectional encoder self-attention, causal decoder self-attention, and decoder cross-attention The decoder uses earlier target tokens and can attend to encoded source positions Conditional sequence-to-sequence tasks, such as translation The original Transformer; T5 and BART are common examples

These are typical patterns, not rigid rules for every implementation. Hugging Face documents that a causal decoder model can be run with bidirectional attention for a particular use, while noting that changing its attention mode does not make it an encoder architecture. The Attention Interface documentation distinguishes the selected attention mode from the model’s block architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Flame Toys Transformers Megatron Furai Model Kit (G1 Version)
  • Good articulation with over 40 movable joints, any pose can be set easily.
  • The design reveals a modernized and shape optimized Megatron (G1 version).
  • With different injection color of runner parts and simple assembly design, it is suitable for model kit beginner.
  • No glue required.

How each architecture uses attention

Encoder-only: use the whole input as context

An encoder processes the supplied sequence into contextualized representations. With bidirectional attention, a token can use input tokens to its left and right. This suits tasks where the complete input is available and the goal is to understand or represent it, including classification and embeddings. Google’s Transformer course describes these common encoder uses.

Decoder-only: predict from a causal prefix

A causal decoder generates left to right. When predicting a token, its mask blocks later target positions, so it cannot see the answer token in advance. The sequence probability is factorized into next-token probabilities conditioned on the preceding prefix. During generation, the model predicts a token, appends it to the prefix, and repeats the process.

Encoder-decoder: generate from a separate source

The encoder reads the source sequence and produces contextualized states. The decoder applies causal self-attention to the target prefix and cross-attention to the encoder output. At each output position, decoder queries compare with encoder keys and use encoder values, allowing the output to draw on relevant source positions. The generated sequence is conditioned on both the encoded input and the target tokens generated so far. Hugging Face’s encoder-decoder explanation describes this arrangement.

Rank #3
Sale
Transformers, Classic Class, CC24, Transformers Dark of the Moon, Sentinel Prime, Model Kit, Assembly Required, 14+
  • OFFICIALLY LICENSED TRANSFORMERS: DARK OF THE MOON COLLECTIBLE WITH FAITHFUL MECHANICAL DETAIL – Crafted under full official Transformers authorization, this 90-piece Classic Class Sentinel Prime model kit faithfully recreates his iconic Dark of the Moon design standing approximately 5.12 inches tall with sharp mechanical detailing, true-to-character proportions, and a refined head sculpt that captures every commanding, battle-hardened aspect of his legendary Transformers presence.
  • SIGNATURE LIGHT-UP EYES FOR MAXIMUM DISPLAY IMPACT – CC24 Sentinel Prime features a striking light-up eyes design that enhances his expression and brings powerful visual impact and commanding presence to every display configuration, making him one of the most visually dramatic and display-worthy figures in the entire Transformers Classic Class lineup and an instant centerpiece for any serious Transformers collection.
  • 20+ MOVABLE JOINTS WITH UPGRADED FRAME FOR DYNAMIC BATTLE POSES – Featuring an upgraded frame design with 20+ articulated joints throughout the body, Sentinel Prime delivers improved articulation and enhanced stability for a wide range of powerful battle stances and commanding action poses that faithfully recreate his most iconic and treacherous moments from Transformers: Dark of the Moon.
  • EXCLUSIVE WEAPON CONFIGURATION FOR BATTLE-READY DISPLAY – Sentinel Prime arrives fully armed with an exclusive weapon configuration including dedicated firearm weapon accessories and a character-specific display stand, delivering everything needed to recreate his most powerful and commanding battle moments from Transformers: Dark of the Moon straight out of the box.
  • TOOL-FREE SNAP-FIT ASSEMBLY FOR TRANSFORMERS COLLECTORS AGES 14+ – Simple snap-fit construction requires no tools, glue, or paint, making CC24 Sentinel Prime quick and satisfying to assemble and delivering a professional-quality, display-ready finish worthy of any dedicated Transformers fan, Dark of the Moon enthusiast, model kit builder, or Classic Class collector's shelf, desk, or display case.

Which architecture fits the task?

Choose by the task’s information flow rather than assuming one family is universally better.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Understanding a complete input: An encoder-only pattern lets each position use context on both sides.
  • Continuing a prefix or generating text autoregressively: A causal decoder matches the left-to-right prediction process.
  • Turning one sequence into another: An encoder-decoder design keeps source representations separate and exposes them to the decoder through cross-attention.

Four questions make the distinction practical:

  1. What can each position see? Does it need right-side context, or must it obey a causal prefix?
  2. What is the input-output structure? Is the task to represent a complete input, continue a prefix, or map a source sequence to a target sequence?
  3. Where does conditioning live? Is input context part of the same causal sequence, or held in a separate encoder representation?
  4. What are the sequence and implementation constraints? Consider sequence lengths, attention kernels, caching, batch shape, and hardware.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Attention cost: what the quadratic term does and does not tell you

Google’s course gives a simplified self-attention scaling expression of O(N² · S · D), where N is context length, S is the number of self-attention layers, and D is the number of heads per layer. The key implication is that the sequence-length term is quadratic in this account: increasing sequence length can substantially increase attention computation.

This expression is not a universal wall-clock or memory estimate, and it does not by itself establish that one architecture family is faster. Practical costs also depend on model dimensions, implementation, hardware, batch shape, and optimization. Compare actual systems only with those factors controlled.

Rank #4
Sale
BLOKEES - Transformers Classic Class Megatronus Prime Model Kit
  • OFFICIALLY LICENSED TRANSFORMERS ONE COLLECTIBLE WITH SCREEN-ACCURATE MOVIE DETAILING – Crafted under full official Transformers One authorization, this 107-piece Classic Class Megatronus stands approximately 12.5 cm tall, faithfully recreating the legendary guardian of Cybertron and one of the Thirteen Original Primes with meticulously sculpted armor texturing, authentic color schemes, and screen-accurate proportions that capture every detail of his iconic miner-turned-warrior appearance from the Transformers One film.
  • DUAL LED LIGHTING SYSTEM — GLOWING EYES & ILLUMINATED CHEST – CC20 Megatronus features built-in LED modules in both his eyes and chest that bring authentic Cybertronian energy signatures to life with dramatic glowing illumination, making him one of the most visually striking and display-worthy figures in the entire Transformers Classic Class lineup and an instant commanding centerpiece for any Transformers One or Thirteen Original Primes collection.
  • 20-POINT SUPER ARTICULATION WITH ENHANCED FULL-BODY MOBILITY – Featuring 20 highly adjustable articulated joints throughout the body with enhanced mobility upgrades including enhanced knee bending for powerful forward kick angles, lateral shoulder movement, double-jointed elbows, and hip extension, Megatronus delivers complete freedom of movement and total control over head, limbs, and torso for explosive, dynamic combat poses worthy of Cybertron's most powerful and rebellious Prime.
  • PREMIUM COMBAT-READY ACCESSORY SET WITH BLAST EFFECTS – Megatronus arrives fully equipped for battle with a complete premium accessories package including signature character-specific weapons, multiple interchangeable hand sets featuring fist, gripping, and commanding gesture options, dynamic blast effects parts, and a dedicated display stand — delivering everything needed to recreate the most powerful and legendary combat moments from Transformers One straight out of the box.
  • 107-PIECE TOOL-FREE SNAP-FIT ASSEMBLY FOR TRANSFORMERS COLLECTORS AGES 14+ – Built using a revolutionary panel and component dual-structure design from 107 pre-colored snap-fit parts requiring no glue, brushes, or cutting tools, CC20 Megatronus delivers a low barrier-to-entry assembly experience with professional-grade results for builders of all skill levels — the perfect addition for dedicated Transformers fans, Transformers One enthusiasts, model kit builders, and Classic Class collectors ready to add the legendary first Megatron to their display.

What the original Transformer’s translation results show

The original Transformer paper reported 28.4 BLEU on the WMT 2014 English-to-German task. For English-to-French, its abstract reported 41.8 BLEU for a single model after 3.5 days of training on eight GPUs. Google Research’s publication page displays 41.0 for the English-to-French result, a discrepancy between that page and the arXiv abstract; the two figures should not be silently reconciled. These are historical results from the 2017 paper, not a current head-to-head comparison of modern language-model architectures. See the paper abstract and Google Research publication page.

Further reading

Natural Language Processing with Transformers, Revised Edition by Lewis Tunstall, Leandro von Werra, and Thomas Wolf is an intermediate-to-advanced practical guide to NLP and Transformers. Its 408 pages cover attention mechanisms, the encoder-decoder framework, Transformer anatomy, self-attention, and the encoder, decoder, and encoder-decoder branches. See the O’Reilly publisher listing.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 2
Flame Toys Transformers Megatron Furai Model Kit (G1 Version)
Flame Toys Transformers Megatron Furai Model Kit (G1 Version)
Good articulation with over 40 movable joints, any pose can be set easily.; The design reveals a modernized and shape optimized Megatron (G1 version).
$67.39

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.