Decoder vs Encoder in Transformer Models
1. Core Components of Transformers
Core Components of Transformers
Encoder and Decoder: Architectural Distinctions
The transformer architecture, introduced by Vaswani et al. (2017), consists of two primary components: the encoder and the decoder. While both share similarities in their use of self-attention mechanisms, their roles and internal structures differ significantly.
The encoder processes input sequences (e.g., source language tokens in machine translation) and generates a continuous representation that captures contextual relationships. It consists of multiple identical layers, each containing:
- A multi-head self-attention mechanism
- A position-wise feed-forward network
- Residual connections and layer normalization
The decoder, in contrast, generates output sequences (e.g., target language tokens) autoregressively. Its layers include:
- Masked multi-head self-attention to prevent future token visibility
- Encoder-decoder attention that incorporates encoder outputs
- Identical feed-forward and normalization components as the encoder
Mathematical Formulation
The self-attention mechanism common to both components computes scaled dot-product attention:
where Q, K, and V represent queries, keys, and values matrices respectively, and dk is the dimension of the keys. The encoder computes this over its input sequence, while the decoder applies masking to the attention weights:
where M is a lower triangular matrix with values of -∞ for positions i > j to enforce causality.
Information Flow and Practical Implications
In sequence-to-sequence tasks, the encoder processes the entire input in parallel, creating a memory-efficient representation. The decoder then attends to both:
- Its own previous outputs (via masked self-attention)
- The encoder's final representation (via cross-attention)
This separation enables transformers to handle variable-length sequences while maintaining parallel processing capabilities during training. During inference, the decoder operates sequentially, generating one token at a time while referencing the static encoder output.
Variants and Specialized Architectures
Recent architectures have modified this paradigm:
- Encoder-only models (e.g., BERT) discard the decoder for classification tasks
- Decoder-only models (e.g., GPT) use masked self-attention throughout
- Sparse attention variants optimize computational complexity in both components
The choice between encoder-decoder versus single-component architectures depends on task requirements—bidirectional context understanding favors encoders, while generative tasks typically require decoders.

Self-Attention Mechanism Overview
Core Mathematical Formulation
The self-attention mechanism computes a weighted sum of input representations, where the weights are dynamically derived from pairwise interactions between elements. Given an input sequence X ∈ ℝn×d (n tokens, d dimensions), three learnable matrices project X into query (Q), key (K), and value (V) spaces:
where WQ, WK, WV ∈ ℝd×dk are projection matrices. The attention scores A are computed via scaled dot-product:
The scaling factor 1/√dk prevents gradient vanishing in high-dimensional spaces. The final output is a convex combination of value vectors:
Multi-Head Extension
Multi-head attention (MHA) applies h independent attention mechanisms in parallel, enabling the model to jointly attend to information from different representation subspaces. Each head i computes:
The outputs are concatenated and linearly projected:
where WO ∈ ℝhdv×d is the output projection matrix. This architecture allows differential focus on position-based versus content-based information.
Computational Complexity Analysis
The self-attention mechanism exhibits O(n2d) time and space complexity due to the pairwise attention matrix. For sequence length n and hidden dimension d, this creates quadratic memory demands that motivate sparse attention variants in long-context applications.
Positional Encoding
Since self-attention is permutation-equivariant, sinusoidal positional encodings P ∈ ℝn×d are added to input embeddings:
This provides explicit position information while maintaining the model's ability to attend to arbitrary sequence positions.
Practical Implementation Considerations
- Memory optimization: FlashAttention reduces memory overhead through tiling and recomputation techniques.
- Numerical stability: Attention logits are shifted by their maximum value before softmax application.
- Causal masking: In decoder layers, upper-triangular masking prevents attending to future tokens.

2. Role and Function of the Encoder
Role and Function of the Encoder
The encoder in a transformer model processes input sequences into a rich, context-aware representation that captures both local and global dependencies. It consists of multiple identical layers, each containing two primary sub-components: a multi-head self-attention mechanism and a position-wise feed-forward network. Residual connections and layer normalization stabilize training by mitigating vanishing gradients.
Self-Attention Mechanism
The self-attention mechanism computes dynamic weightings between all pairs of tokens in the input sequence. Given an input matrix X ∈ ℝn×d where n is the sequence length and d is the embedding dimension, the mechanism first projects X into query (Q), key (K), and value (V) matrices:
where WQ, WK, WV ∈ ℝd×dk are learnable projection matrices. The attention scores are computed as:
The scaling factor √dk prevents gradient saturation in the softmax. Multi-head attention extends this by performing h parallel attention operations, concatenating the results:
where each headi = Attention(QWiQ, KWiK, VWiV) and WO ∈ ℝhdv×d.
Position-Wise Feed-Forward Network
Following self-attention, each encoder layer applies a position-wise feed-forward network (FFN) to each token independently:
where W1 ∈ ℝd×dff, W2 ∈ ℝdff×d, and dff is typically 4×d. This expands the model's capacity to learn non-linear transformations.
Residual Connections and Normalization
Each sub-layer (attention and FFN) employs residual connections followed by layer normalization:
This architecture enables stable training of deep networks by preserving gradient flow. The encoder's output is a sequence of contextualized embeddings where each token's representation incorporates information from all other tokens in the input.
Practical Applications
Encoder-only architectures like BERT leverage this design for tasks requiring bidirectional context, such as text classification and named entity recognition. In encoder-decoder models (e.g., T5), the encoder's output serves as cross-attention keys/values for the decoder, enabling sequence-to-sequence tasks like machine translation.

Encoder Layer Structure
The encoder in a transformer model is composed of a stack of identical layers, each containing two primary sub-components: a multi-head self-attention mechanism and a position-wise feed-forward neural network. Residual connections and layer normalization are applied around each of these sub-layers, forming the core of the encoder's processing pipeline.
Multi-Head Self-Attention
The multi-head attention mechanism allows the model to jointly attend to information from different representation subspaces at different positions. For each attention head i, the input embeddings X are linearly projected into queries (Qi), keys (Ki), and values (Vi) through learned weight matrices:
where WiQ, WiK, and WiV are learned projection matrices for head i. The attention scores are computed as:
The outputs of all attention heads are concatenated and projected through a final linear layer:
Position-wise Feed-Forward Network
Following the attention mechanism, each encoder layer contains a fully connected feed-forward network (FFN) applied independently to each position. The FFN consists of two linear transformations with a ReLU activation in between:
This architecture allows the model to process each position's representation while maintaining the ability to incorporate information from the attention mechanism's output.
Residual Connections and Layer Normalization
Each sub-layer (self-attention and FFN) employs residual connections followed by layer normalization. The output of each sub-layer is computed as:
where Sublayer(x) represents either the multi-head attention or feed-forward network. This architecture helps mitigate the vanishing gradient problem and enables training of deeper networks.
Positional Encoding
Since the transformer lacks recurrent or convolutional operations, positional information is injected through sinusoidal positional encodings:
These encodings are added to the input embeddings before being processed by the encoder layers, providing the model with information about the relative or absolute position of tokens in the sequence.
Applications of Encoder-Only Models (e.g., BERT)
Encoder-only models like BERT (Bidirectional Encoder Representations from Transformers) leverage the transformer encoder architecture to process input sequences bidirectionally, capturing contextual relationships between tokens. Unlike autoregressive decoder models, encoder-only architectures do not generate sequences but instead produce dense representations useful for downstream tasks.
Natural Language Understanding (NLU) Tasks
BERT's bidirectional attention mechanism enables state-of-the-art performance on NLU benchmarks. The model's pre-trained representations are fine-tuned for:
- Text Classification: Sentiment analysis, topic labeling, and intent detection benefit from BERT's contextual embeddings. For a document D with tokens {x1, ..., xn}, the [CLS] token's hidden state h[CLS] serves as the aggregate representation for classification:
- Named Entity Recognition (NER): BERT computes token-level embeddings hi that capture surrounding context, improving entity boundary detection compared to word-level models like Word2Vec.
- Question Answering: In extractive QA (e.g., SQuAD), BERT predicts answer spans by modeling the joint probability of start and end positions given a question-context pair.
Information Retrieval and Semantic Search
Dense retrieval systems use BERT to encode queries and documents into a shared embedding space. The similarity score between query q and document d is computed via:
where E is the encoder output. Advanced variants like ANCE (Approximate Nearest Neighbor Negative Contrastive Learning) optimize the embedding space for retrieval efficiency.
Transfer Learning and Parameter Efficiency
Encoder-only models enable parameter-efficient adaptation through:
- Adapter Layers: Task-specific modules inserted between transformer layers, freezing the base model while training only the adapters.
- Prompt Tuning: Reformulating downstream tasks as masked language modeling by adding learned soft prompts to the input.
The adapter architecture modifies layer outputs as:
where fθ is a small feedforward network and LN denotes layer normalization.
Multimodal Extensions
Vision-language models like VL-BERT extend the encoder to process both text and image regions. For an image-text pair, region features vi and word embeddings wj are concatenated as input:
Cross-modal attention in the encoder learns alignments between visual and textual elements, enabling tasks like visual question answering.
3. Role and Function of the Decoder
Role and Function of the Decoder
Architecture and Components
The decoder in a transformer model is responsible for generating output sequences autoregressively, conditioned on the encoded representation from the encoder. Unlike the encoder, which processes input tokens in parallel, the decoder generates output tokens sequentially, masking future positions to prevent information leakage. The decoder consists of multiple layers, each containing three key sub-components:
- Masked Multi-Head Self-Attention: Ensures each output token attends only to previous tokens in the sequence.
- Multi-Head Cross-Attention: Allows the decoder to attend to the encoder's output representation.
- Position-wise Feed-Forward Network: Applies non-linear transformations to each token independently.
Autoregressive Generation Mechanism
The decoder generates sequences token-by-token using teacher forcing during training and autoregressive sampling during inference. At each step t, the decoder produces a probability distribution over the vocabulary conditioned on the previously generated tokens y<t and the encoder's output X:
where ht(L) is the hidden state at layer L, and Wo is the output projection matrix. The masking in self-attention ensures the model cannot attend to future positions during training:
where M is a lower triangular matrix with -∞ in positions corresponding to future tokens.
Cross-Attention and Encoder-Decoder Interaction
The decoder's cross-attention mechanism creates dynamic, context-aware connections between the target sequence and encoded source representation. For each decoder layer, the queries come from the decoder's previous layer, while keys and values are derived from the encoder's output:
This allows the model to learn which parts of the input sequence are most relevant for generating each output token, enabling tasks like machine translation where alignment between source and target is crucial.
Practical Considerations
Modern implementations often employ techniques to improve decoder efficiency and performance:
- Beam Search: Maintains multiple candidate sequences during generation to reduce the impact of local optima.
- Length Normalization: Adjusts for sequence length bias in beam search scoring.
- Top-k/Top-p Sampling: Stochastic decoding methods that improve output diversity while maintaining coherence.
In large language models, the decoder's autoregressive nature becomes computationally expensive for long sequences due to the quadratic complexity of self-attention and the serial dependency in token generation. Recent architectures like sparse attention and memory mechanisms attempt to mitigate these limitations while preserving the decoder's essential functionality.

3.2 Decoder Layer Structure
The decoder in transformer models is responsible for autoregressive sequence generation, leveraging masked self-attention and encoder-decoder attention to produce outputs conditioned on previous tokens and encoded inputs. Unlike the encoder, which processes all input tokens in parallel, the decoder operates sequentially to preserve causality.
Core Components of a Decoder Layer
Each decoder layer consists of three primary sub-layers:
- Masked Multi-Head Self-Attention – Ensures each output token attends only to previous positions via a causal mask.
- Encoder-Decoder Attention – Allows the decoder to focus on relevant parts of the encoded input sequence.
- Position-wise Feed-Forward Network – Applies non-linear transformations to each token independently.
Residual connections and layer normalization are applied around each sub-layer, following the formula:
Masked Self-Attention Mechanism
The masked self-attention prevents information leakage from future tokens by applying a lower-triangular mask to the attention scores. For a sequence length n, the mask M is defined as:
This ensures that during training, the attention weights for position i depend only on positions 1 to i.
Encoder-Decoder Attention
This sub-layer uses queries from the decoder and keys/values from the encoder output, enabling cross-attention between the source and target sequences. The attention scores are computed as:
where Q is derived from the decoder’s previous layer, and K, V come from the encoder’s final output.
Position-wise Feed-Forward Network
Identical in structure to the encoder’s FFN, this sub-layer applies two linear transformations with a ReLU activation:
The hidden dimension is typically expanded by a factor of 4 (e.g., 512 → 2048) before projecting back to the model dimension.
Practical Implementation Considerations
Modern implementations optimize decoder layers through:
- Key-Value Caching – Reusing computed key-value pairs in autoregressive generation to avoid redundant computation.
- Memory-Efficient Attention – Techniques like memory-efficient flash attention reduce memory overhead for long sequences.
- Parallel Decoding – Methods like speculative decoding enable partial parallelization while maintaining strict causality.
Masked Self-Attention Mechanism
The masked self-attention mechanism is a critical component in transformer decoders, enabling autoregressive generation by preventing positions from attending to future tokens. Unlike standard self-attention in encoders, which allows full bidirectional context, masked self-attention enforces a causal constraint through a carefully constructed attention mask.
Mathematical Formulation
Given an input sequence X of length n, the attention scores A are computed as:
where Q, K are the query and key matrices, dk is the key dimension, and M is the mask matrix defined as:
This mask ensures that when computing attention weights for position i, only positions 1 through i contribute to the weighted sum. The -∞ values become zero after softmax normalization, effectively blocking attention to future tokens.
Implementation Considerations
In practice, the mask is typically implemented as:
- Additive masking: Adding -1e9 to prohibited positions before softmax
- Multiplicative masking: Element-wise multiplication with a binary mask
- Memory optimization: Using triangular matrices to avoid storing full n×n masks
Modern implementations often combine these approaches. For example, PyTorch's nn.TransformerDecoder uses additive masking with a cached upper-triangular boolean mask.
Autoregressive Properties
The masking creates several important properties:
- Causality: Ensures predictions depend only on past inputs
- Parallelizability: All positions can be computed simultaneously during training
- Teacher forcing: Enables efficient training via shifted outputs
During inference, this manifests as sequential generation where each new token's computation depends on all previously generated tokens.
Variants and Extensions
Several modified masking approaches have been developed:
- Local window masking: Restricts attention to a fixed neighborhood around each position
- Strided patterns: Allows sparse attention while maintaining some global connectivity
- Dynamic masking: Learns mask patterns adaptively during training
These variants trade off between computational efficiency and modeling capacity, with applications in long-sequence processing and specialized generation tasks.
Practical Implications
The choice of masking strategy affects:
- Memory usage: Full attention requires O(n²) memory
- Training stability: Improper masking can lead to gradient issues
- Generation quality: Influences the model's ability to maintain long-range coherence
In transformer-based language models like GPT, the masked self-attention mechanism enables the model to learn powerful autoregressive distributions while maintaining training efficiency through parallel computation of all positions.

Applications of Decoder-Only Models (e.g., GPT)
Decoder-only transformer models, such as OpenAI's GPT family, have revolutionized natural language processing (NLP) by leveraging autoregressive generation to produce coherent and contextually relevant text. Unlike encoder-decoder architectures (e.g., BERT or T5), these models exclusively use the decoder stack with masked self-attention, enabling them to predict the next token in a sequence while preventing information leakage from future positions.
Autoregressive Language Modeling
The core mechanism of decoder-only models is autoregressive language modeling, where the probability of a sequence is factorized as:
Here, each token \(x_t\) is generated conditioned on all preceding tokens \(x_{ where \(M\) is a lower-triangular mask matrix with \(M_{ij} = -\infty\) for \(j > i\) and \(0\) otherwise. Decoder-only models excel at open-ended text generation tasks, including: GPT-3 demonstrated that large decoder-only models can perform tasks with minimal examples via prompt engineering. For instance, providing a prompt like: enables the model to infer the translation task without explicit fine-tuning. While traditionally dominated by encoder-decoder models, decoder-only architectures achieve summarization by conditioning on the input text followed by a;
"TL;DR:"). The model then generates a condensed version autoregressively. Modern decoder-only models employ several optimizations to enhance scalability and performance: Despite their versatility, decoder-only models face critical limitations: The encoder and decoder in transformer models share foundational building blocks—multi-head self-attention, feed-forward networks, and layer normalization—but differ in their connectivity and masking mechanisms. The encoder processes input sequences bidirectionally, allowing each token to attend to all other tokens in the input. In contrast, the decoder employs masked self-attention to prevent leftward information flow, ensuring autoregressive generation. Encoders exclusively use self-attention, computing attention scores between tokens within the same sequence. Decoders implement two attention layers: While both components use positional embeddings, decoders require stricter position handling. The encoder's bidirectional attention allows global positional context, whereas the decoder must preserve temporal ordering through: Encoder layers exhibit O(n²) complexity for sequence length n due to full self-attention. Decoders have O(n² + nm) complexity when processing encoder outputs of length m, with the cross-attention term becoming dominant in many-to-many sequence tasks like machine translation. The encoder and decoder in transformer models are optimized using distinct but complementary training objectives. While the encoder focuses on creating rich contextual representations, the decoder is trained to generate coherent sequences conditioned on these representations. The choice of objective function directly impacts the model's ability to handle tasks like machine translation, text summarization, or dialogue generation. The encoder is typically trained using Masked Language Modeling (MLM), where a fraction of input tokens are randomly masked, and the model must predict the original tokens based on bidirectional context. Given an input sequence X = [x1, ..., xn], a random subset of tokens is replaced with a [MASK] token, and the objective is to minimize the negative log-likelihood of the original tokens: where ℳ is the set of masked positions. This forces the encoder to develop a deep bidirectional understanding of language structure, as each prediction depends on both left and right contexts. The decoder employs autoregressive language modeling, where it predicts each token conditioned on previous tokens in a left-to-right manner. For a target sequence Y = [y1, ..., ym], the objective is: Here, X represents the encoder's output, and y<t denotes all tokens before position t. This unidirectional constraint is crucial for tasks like text generation, where the model must produce plausible continuations without access to future tokens. In full transformer architectures (e.g., BART, T5), the encoder and decoder are trained jointly using a combination of denoising and autoregressive objectives. For instance, BART corrupts the input text with multiple noising strategies (token deletion, masking, rotation) and trains the decoder to reconstruct the original sequence: where X̃ is the corrupted input. This approach bridges the gap between bidirectional understanding and autoregressive generation, making the model versatile for both comprehension and production tasks. Recent variants like ELECTRA replace MLM with replaced token detection, where a generator network produces plausible alternatives for masked tokens, and the discriminator (encoder) learns to distinguish original tokens from replacements. The objective becomes: This method improves sample efficiency, as every token contributes to the loss rather than just the masked subset. Encoder-only architectures, such as BERT (Bidirectional Encoder Representations from Transformers), excel in tasks requiring deep bidirectional context understanding. These models process input sequences in their entirety, leveraging self-attention mechanisms to capture relationships between all tokens simultaneously. The encoder's bidirectional nature allows it to generate contextualized embeddings, making it ideal for: Decoder-only models, such as GPT (Generative Pre-trained Transformer), specialize in autoregressive generation. Unlike encoders, decoders use masked self-attention to ensure each token only attends to previous tokens, making them inherently unidirectional. This architecture is optimized for: The full Transformer architecture, combining both encoder and decoder, is designed for sequence-to-sequence (seq2seq) tasks. The encoder processes the input sequence, while the decoder generates the output sequence autoregressively, attending to the encoder's outputs via cross-attention. Key applications include: In encoder-decoder models, cross-attention enables the decoder to attend to the encoder's hidden states. Given encoder outputs Henc and decoder hidden states Hdec, the cross-attention mechanism computes: where Q = HdecWQ, K = HencWK, and V = HencWV. Here, WQ, WK, and WV are learned projection matrices, and dk is the dimension of the key vectors. Recent advancements have introduced hybrid architectures tailored for specific use cases: Choosing between encoder-only, decoder-only, or encoder-decoder models depends on the task: Encoder-decoder models in the transformer architecture consist of two primary components: the encoder, which processes input sequences, and the decoder, which generates output sequences. The encoder maps an input sequence to a continuous representation, while the decoder autoregressively produces an output sequence conditioned on this representation. Models like T5 (Text-to-Text Transfer Transformer) and BART (Bidirectional and Auto-Regressive Transformers) exemplify this architecture with distinct design choices. The encoder comprises multiple identical layers, each containing two sub-layers: a multi-head self-attention mechanism and a position-wise feed-forward network. Residual connections and layer normalization stabilize training. For an input sequence X = (x1, ..., xn), the encoder computes: where Q, K, and V are learned linear projections of the input. The encoder's bidirectional self-attention allows each token to attend to all positions in the input, enabling rich contextual representations. The decoder also consists of stacked layers but includes three sub-layers: masked self-attention, encoder-decoder attention, and a feed-forward network. The masked self-attention ensures autoregressive properties by preventing positions from attending to subsequent tokens. The encoder-decoder attention layer allows the decoder to focus on relevant parts of the encoder's output. Here, M is a mask matrix with Mij = −∞ if i < j to enforce causality. T5 treats all NLP tasks as text-to-text problems, using the same encoder-decoder architecture for tasks like translation, summarization, and classification. The encoder processes the input text, and the decoder generates a target sequence. T5 employs a relative position bias instead of sinusoidal position embeddings, improving generalization to longer sequences. BART combines bidirectional encoder representations (like BERT) with autoregressive decoding. It is pretrained by corrupting text with noise (e.g., masking, deletion, permutation) and learning to reconstruct the original sequence. The encoder maps corrupted input to latent representations, while the decoder reconstructs the clean sequence autoregressively. Encoder-decoder models excel in tasks requiring sequence generation conditioned on input, such as machine translation (T5), summarization (BART), and dialogue systems. Their modular architecture allows fine-tuning for domain-specific applications, leveraging transfer learning from large-scale pretraining. Sequence-to-sequence (seq2seq) tasks, such as machine translation, text summarization, and dialogue generation, rely on the interplay between the encoder and decoder in transformer models. The encoder processes the input sequence into a context-rich representation, while the decoder generates the output sequence autoregressively, conditioned on both the encoder's output and its own previous predictions. The encoder maps an input sequence X = (x1, ..., xn) to a continuous representation Z = (z1, ..., zn) through stacked self-attention and feed-forward layers. The decoder then generates the output sequence Y = (y1, ..., ym) one token at a time, using masked self-attention to prevent information leakage from future tokens and cross-attention to incorporate Z. The decoder operates autoregressively, meaning each predicted token yt becomes part of the input for predicting yt+1. This requires teacher forcing during training and beam search or sampling during inference. The probability of the output sequence is factorized as: The decoder's cross-attention layer aligns each decoding step with relevant parts of the encoded input. For a query qt (current decoder state), keys K, and values V (encoder outputs), the attention weights are computed as: This allows the decoder to dynamically focus on different parts of the input sequence at each step, enabling precise context-aware generation. Recent advancements address these challenges: The cross-attention mechanism is a critical component in Transformer-based architectures, enabling dynamic interaction between sequences from different modalities or contextual sources. Unlike self-attention, where queries, keys, and values originate from the same sequence, cross-attention computes attention scores between two distinct sequences—typically the decoder’s hidden states (queries) and the encoder’s outputs (keys and values). Given an encoder output E ∈ ℝn×d and decoder hidden state D ∈ ℝm×d, cross-attention computes: where: The scaling factor √dk stabilizes gradients by normalizing the dot product magnitudes. The softmax operation generates a probability distribution over encoder tokens for each decoder query. Cross-attention allows the decoder to focus adaptively on relevant encoder positions during autoregressive generation. For instance, in machine translation, the decoder might attend to different source-language words when predicting each target token. The mechanism’s efficacy stems from: In standard Transformer architectures, cross-attention layers are interleaved between self-attention and feed-forward layers in the decoder stack. The computational flow for a decoder layer is: This design ensures the decoder first consolidates its internal state before integrating cross-sequence information. Key implementation challenges include: In multimodal models (e.g., vision-language Transformers), cross-attention bridges heterogeneous representations—such as attending image regions when generating descriptive text. The computational efficiency of encoder and decoder components in transformer models is governed by their distinct architectural roles and the resulting algorithmic complexity. The encoder processes input sequences in a fully parallel manner, leveraging self-attention to compute relationships between all tokens simultaneously. For an input sequence of length N, the self-attention mechanism in the encoder scales quadratically with O(N²) due to the pairwise token interactions. However, this computation is embarrassingly parallelizable across attention heads and layers, making it highly efficient on modern hardware accelerators like GPUs and TPUs. In contrast, the decoder introduces an additional computational constraint: autoregressive generation requires sequential processing of outputs with a masked self-attention mechanism. During training with teacher forcing, the decoder can parallelize computations across the output sequence length M, but inference must process tokens one at a time due to the causal masking constraint. This results in a step-wise complexity of O(M²) per generated token, accumulating to O(M³) for full sequence generation. The decoder's incremental processing pattern creates unique memory bandwidth challenges. While the encoder's key-value (K, V) pairs for all input tokens can be precomputed and reused across decoding steps, the decoder must dynamically update its own K, V caches for previously generated tokens. This leads to a memory access complexity of O(M·d_{\text{model}}) per token in autoregressive decoding, often becoming the bottleneck in large-language model inference. Recent architectural innovations like cross-attention pruning and memory-compressed attention reduce the decoder's computational overhead. Sparse attention patterns in models like Longformer and BigBird approximate full attention with O(N log N) complexity, while retrieval-augmented decoders (e.g., RETRO) offload context processing to external memory networks. The table below contrasts the computational profiles: Practical implementations often employ techniques like kv-cache sharing across beam search candidates and mixed-precision quantization to mitigate the decoder's computational demands. The encoder-decoder attention layer adds another O(N·M·d_{\text{model}}) term, but its impact is typically dwarfed by the decoder's self-attention costs during long-sequence generation. The relationship between model size and training data requirements in transformer architectures is governed by scaling laws, which empirically describe how performance improves with increased compute, parameters, and data. For autoregressive decoder-only models (e.g., GPT-3), the scaling behavior follows a power-law relationship: where N is the number of parameters, D is training tokens, Nc and Dc are critical thresholds, and αN, αD ≈ 0.07 are scaling exponents derived empirically. The loss floor L∞ represents irreducible error. Bidirectional encoder models (e.g., BERT) exhibit different scaling characteristics compared to autoregressive decoders due to their masked language modeling objective. The compute-optimal training regime for encoder models requires: whereas decoder models achieve optimal performance with Dopt ≈ 300N. This 15× difference stems from the decoder's need to learn stronger sequential prediction capabilities. Recent architectures like Mixture-of-Experts (MoE) decoders challenge traditional scaling laws by enabling sparse activation patterns. For a MoE model with E experts: while maintaining the compute cost of a model with N/E active parameters. This allows decoders to scale beyond 1T parameters while keeping training costs manageable. The Chinchilla scaling laws suggest that current large language models are significantly undertrained, with optimal training compute allocation favoring more data over larger models. For a fixed compute budget C (in FLOPs): This implies that doubling model size should be accompanied by quadrupling the training data to maintain compute-optimal performance. The choice between encoder-only, decoder-only, or encoder-decoder architectures depends fundamentally on the task requirements and data characteristics. Encoder-only models like BERT excel at bidirectional context understanding, making them ideal for tasks requiring deep analysis of input data such as text classification, named entity recognition, or sentiment analysis. The bidirectional attention mechanism allows each token to attend to all other tokens in both directions, capturing rich contextual relationships. Decoder-only models such as GPT specialize in autoregressive generation, where outputs are produced sequentially with each step conditioned on previous outputs. This architecture is particularly effective for open-ended generation tasks like story writing, code completion, or conversational responses. The causal attention mask prevents the model from "seeing" future tokens, enforcing the autoregressive property essential for coherent generation. Encoder-decoder models like T5 or BART combine both components, making them suitable for sequence-to-sequence tasks requiring comprehensive input understanding and sophisticated output generation. Machine translation, text summarization, and question answering often benefit from this architecture. The encoder processes the input to create a dense representation, which the decoder then uses to generate outputs token-by-token while attending to both the encoded representation and its own previous outputs. Transformer-XL and other memory-augmented variants introduce recurrence mechanisms to handle longer sequences than standard transformers. These models maintain a memory of previous segments, allowing them to capture dependencies beyond the fixed context window of vanilla transformers. This is particularly valuable for tasks involving long documents or continuous streams of data. When selecting an architecture, consider these key factors: Sparse attention mechanisms and mixture-of-experts approaches are pushing the boundaries of traditional architectures. Models like Switch Transformer demonstrate how conditional computation can make massive models more practical. For multimodal tasks, architectures like Perceiver IO show how shared encoder-decoder structures can handle diverse input and output modalities while maintaining efficiency. Recent work in retrieval-augmented generation combines the benefits of encoder and decoder models with external knowledge retrieval. This hybrid approach, exemplified by models like RAG and REALM, achieves state-of-the-art performance on knowledge-intensive tasks while maintaining the fluency of decoder-based generation.Key Applications
1. Text Generation
2. Few-Shot Learning
Translate English to French:
"Hello" → "Bonjour"
"Goodbye" → "Au revoir"
"Thank you" →3. Text Summarization
Architectural Optimizations
Limitations and Challenges
4. Architectural Differences
Architectural Differences
Core Structural Components
Attention Mechanism Variants
Positional Processing
Memory and Computational Complexity

4.2 Training Objectives
Encoder Training: Masked Language Modeling (MLM)
Decoder Training: Autoregressive Language Modeling
Joint Training in Sequence-to-Sequence Models
Contrastive Learning Objectives
Practical Considerations
4.3 Use Cases and Model Types
Encoder-Only Models
Decoder-Only Models
Encoder-Decoder Models
Mathematical Formulation of Cross-Attention
Specialized Variants
Practical Considerations

5. Architecture of Encoder-Decoder Models (e.g., T5, BART)
Architecture of Encoder-Decoder Models (e.g., T5, BART)
Encoder Structure
Decoder Structure
T5: Unified Text-to-Text Framework
BART: Denoising Autoencoder
Key Architectural Differences
Practical Applications

Sequence-to-Sequence Tasks
Encoder-Decoder Architecture in Seq2Seq
Autoregressive Generation
Cross-Attention Mechanism
Practical Challenges
Advanced Techniques

5.3 Cross-Attention Mechanism
Mathematical Formulation
Mechanism Dynamics
Architectural Integration
Practical Considerations

6. Computational Efficiency
6.1 Computational Efficiency
Memory Bandwidth Considerations
Hybrid Architectures and Optimizations
Component
Training Complexity
Inference Complexity
Parallelizability
Encoder
O(N²)
O(N²)
Full
Decoder
O(M²)
O(M³)
Partial (causal mask)
6.2 Model Size and Training Data Requirements
Encoder-Decoder Asymmetry
Parameter Efficiency Frontiers
Practical Implications

6.3 Choosing Between Encoder, Decoder, or Combined Models
Architectural Trade-offs in Model Selection
Hybrid Architectures for Complex Tasks
Practical Considerations
Emerging Trends and Specialized Variants

7. Key Research Papers
7.1 Key Research Papers
7.2 Recommended Books and Articles
7.3 Online Resources and Tutorials








