Compressing LLMs with Optimal Subnetwork Extraction
1. The Need for Model Compression in LLMs
The Need for Model Compression in LLMs
Large Language Models (LLMs) like GPT-3, PaLM, and LLaMA achieve state-of-the-art performance across natural language tasks but come with prohibitive computational and memory costs. A GPT-3-scale model with 175 billion parameters requires approximately 350 GB of memory just to store its weights in FP16 precision, making deployment on edge devices or real-time applications infeasible. The quadratic complexity of self-attention mechanisms in transformers further exacerbates inference latency, scaling as O(n²d) for sequence length n and hidden dimension d.
Computational and Memory Bottlenecks
The resource demands of LLMs manifest in three critical dimensions:
- Memory footprint: Model weights grow linearly with parameter count, with a 70B-parameter model consuming ~140GB in FP16. This exceeds the VRAM capacity of most GPUs.
- Inference latency: Autoregressive generation requires sequential computation with no parallelization, leading to O(n) runtime per token for context length n.
- Energy consumption: A single forward pass of GPT-3 emits ~190g CO₂ equivalent, with cumulative training emissions reaching hundreds of metric tons.
The Pareto Frontier of Model Efficiency
Optimal subnetwork extraction operates on the principle that overparameterized networks contain sparse, high-performance subnetworks. The Lottery Ticket Hypothesis formalizes this, demonstrating the existence of subnetworks that achieve comparable accuracy to the full model when trained in isolation. For a model f(x; θ) with parameters θ ∈ ℝᴾ, we seek a binary mask m ∈ {0,1}ᴾ such that:
where ∥m∥₀ counts non-zero entries and ϵ bounds acceptable performance degradation. Empirical studies show that 90-95% sparsity can be achieved in transformers with ϵ < 1% accuracy drop on downstream tasks.
Hardware-Software Co-Design Constraints
Effective compression must account for hardware-specific constraints:
- Tensor core alignment: NVIDIA GPUs require matrix dimensions divisible by 16/32 for peak FP16/INT8 throughput.
- Memory bandwidth: Sparsity patterns must enable coalesced memory access to avoid 10-100x bandwidth penalties.
- Quantization granularity: Per-channel quantization (vs. per-tensor) reduces MSE by 3-5× but requires specialized kernel support.
Recent advances in sparse tensor cores (e.g., NVIDIA Ampere's 2:4 sparsity) demonstrate that 50% theoretical speedups are achievable when compression aligns with hardware constraints. The optimal compression ratio CR for a given hardware platform balances arithmetic intensity AI and memory bandwidth BW:
Emerging Applications Driving Compression Needs
Three deployment scenarios necessitate efficient LLMs:
- On-device inference: Smartphones (<4GB RAM) require sub-1B parameter models with <100ms latency.
- Multi-model serving: Cloud deployments must host thousands of fine-tuned variants within fixed memory budgets.
- Continual learning: Modular sparse networks enable efficient parameter reuse across tasks without catastrophic forgetting.
Key Metrics: Performance vs. Efficiency Trade-offs
When compressing large language models (LLMs), the primary challenge lies in balancing performance retention against computational efficiency gains. This trade-off is quantified through several key metrics, each capturing distinct aspects of the model's behavior under compression.
Model Performance Metrics
The most critical performance metric is task accuracy, typically measured on benchmark datasets relevant to the model's application domain. For language models, this includes:
- Perplexity (PPL) on language modeling tasks
- Exact match (EM) and F1 scores on question answering
- BLEU, ROUGE, or METEOR scores for generation tasks
However, raw accuracy metrics alone are insufficient. The relative performance drop (ΔA) captures how much accuracy degrades after compression:
Efficiency Metrics
Compression aims to improve several efficiency dimensions:
- Parameter count: Total trainable weights in the model
- FLOPs: Floating point operations per inference
- Memory footprint: Model size in memory (MB/GB)
- Latency: Inference time per sample
The compression ratio (CR) quantifies the reduction in model size:
The Pareto Frontier
Optimal compression seeks points on the Pareto frontier - configurations where no further efficiency gain can be achieved without sacrificing accuracy. This frontier can be modeled as a multi-objective optimization problem:
where θ represents the compression parameters (pruning thresholds, quantization levels, etc.). Evolutionary algorithms or Bayesian optimization are commonly used to explore this space efficiently.
Energy-Aware Metrics
For deployment on edge devices, energy consumption becomes critical. The energy-accuracy trade-off can be quantified through:
where Energy is measured in joules per inference and Delay is inference latency. Recent work has shown that sparse subnetworks can achieve 2-4× reductions in EDP with <5% accuracy drop on transformer models.
Robustness Considerations
Compressed models must maintain robustness to distribution shifts. The effective compression ratio (ECR) accounts for this:
where OOD indicates out-of-distribution test performance. High-quality compression maintains ECR ≈ CR, indicating preserved generalization.
Practical Deployment Metrics
Real-world deployment introduces additional constraints:
- Hardware utilization: How efficiently the compressed model uses available compute units
- Batch processing capability: Maximum sustainable batch size
- Peak memory usage: Critical for mobile deployment
These metrics often require profiling on target hardware, as theoretical FLOPs reductions don't always translate linearly to real-world speedups due to memory bandwidth limitations and parallelization overheads.

Overview of Compression Techniques: Pruning, Quantization, and Distillation
Pruning: Sparsity for Efficiency
Pruning removes redundant or less significant weights from a neural network, inducing sparsity while preserving model performance. Two primary approaches exist: magnitude-based pruning and gradient-based pruning. Magnitude-based pruning eliminates weights below a threshold θ, while gradient-based pruning considers the impact on loss during training. The sparsity level S is defined as:
where Nzero is the count of zeroed weights and Ntotal is the total number of weights. Iterative pruning, where sparsity is gradually increased over training epochs, often outperforms one-shot pruning. Recent work in Lottery Ticket Hypothesis demonstrates that subnetworks achieving comparable accuracy to the original model can be identified early in training.
Quantization: Reduced Precision Arithmetic
Quantization maps high-precision floating-point weights (e.g., 32-bit) to lower-bit representations (e.g., 8-bit integers), reducing memory footprint and accelerating inference. Uniform quantization divides the weight range [wmin, wmax] into 2b bins, where b is the target bit-width. The quantized value wq is computed as:
where Δ = (wmax − wmin) / (2b − 1) is the step size. Non-uniform quantization methods, such as logarithmic scaling, better capture the distribution of weights but require specialized hardware support. Post-training quantization (PTQ) and quantization-aware training (QAT) are the dominant paradigms, with QAT often yielding higher accuracy by simulating quantization noise during training.
Distillation: Knowledge Transfer to Compact Models
Knowledge distillation trains a smaller student model to mimic the behavior of a larger teacher model, typically using softened output probabilities (logits) from the teacher. The student’s loss function combines task-specific loss Ltask and distillation loss Ldistill:
Here, pt and ps are the teacher’s and student’s softmax outputs scaled by temperature T, and α balances the two terms. Recent variants include attention transfer, where intermediate layer activations are matched, and contrastive distillation, which aligns representations in latent space. Distillation excels in scenarios where the student architecture differs significantly from the teacher (e.g., CNN to Transformer).
Comparative Analysis and Hybrid Approaches
Pruning and quantization are orthogonal; combining them often yields additive gains. For instance, a pruned model can be quantized to further reduce its size. Distillation, while flexible, requires access to the original training data or synthetic data generation. Emerging techniques like quantization-aware pruning and distillation from quantized teachers highlight the trend toward integrated compression pipelines. The choice of technique depends on hardware constraints: pruning benefits sparse accelerators, quantization is ideal for fixed-point hardware, and distillation is suited for architectural simplification.

2. Defining Optimal Subnetworks in LLMs
Defining Optimal Subnetworks in LLMs
Optimal subnetworks in large language models (LLMs) refer to smaller, computationally efficient architectures derived from the original model while preserving a significant portion of its performance. The core challenge lies in identifying a subset of parameters that minimizes redundancy without sacrificing accuracy. This involves rigorous mathematical formulations and empirical validation.
Mathematical Formulation
Given a pre-trained LLM with parameters θ ∈ ℝd, the goal is to find a subnetwork θS ⊂ θ such that:
where ℒ represents the loss function. The subnetwork extraction problem can be framed as a constrained optimization:
Here, ‖·‖0 denotes the L0 norm (sparsity constraint), and ε is an acceptable performance deviation threshold.
Sparsity-Inducing Techniques
Optimal subnetworks are often identified through sparsity-inducing methods:
- Magnitude Pruning: Removes weights with the smallest absolute values, assuming they contribute minimally to model performance.
- Lottery Ticket Hypothesis: Proposes that dense networks contain sparse, trainable subnetworks ("winning tickets") that can match original performance when trained in isolation.
- Structured Pruning: Eliminates entire neurons, attention heads, or layers rather than individual weights, improving hardware efficiency.
Empirical Validation
Recent studies demonstrate that optimal subnetworks can achieve 90-95% of baseline accuracy with 50-70% fewer parameters. For instance, GPT-3 subnetworks extracted via iterative magnitude pruning retain coherent text generation capabilities while reducing inference costs. Key metrics for validation include:
- Task-specific accuracy drop (e.g., perplexity for language modeling).
- Compression ratio: Parameters removed vs. retained.
- Inference speedup: Latency reduction on target hardware.
Practical Considerations
Optimal subnetwork extraction must account for:
- Layer Sensitivity: Attention layers in transformers exhibit higher parameter redundancy than feed-forward layers.
- Dynamic Sparsity: Some methods adapt pruning rates per layer based on gradient flow analysis.
- Retraining: Fine-tuning the subnetwork often recovers lost performance due to pruning.
where η is the learning rate for fine-tuning.
Lottery Ticket Hypothesis and Its Implications
The Lottery Ticket Hypothesis (LTH), introduced by Frankle & Carbin (2019), posits that within a randomly initialized dense neural network, there exist sparse subnetworks—termed winning tickets—that, when trained in isolation, achieve comparable performance to the original network. This discovery challenges the traditional view that overparameterization is merely a tool for optimization, suggesting instead that initialization plays a critical role in identifying these high-performing substructures.
Mathematical Formulation
Given a neural network f(x; θ) with parameters θ ∈ ℝd, LTH asserts the existence of a binary mask m ∈ {0, 1}d such that the pruned network f(x; m ⊙ θ) (where ⊙ denotes element-wise multiplication) satisfies:
The mask m is found through iterative magnitude pruning: after training the full network, the smallest-magnitude weights are removed, and the remaining weights are reset to their initial values. This process is repeated until the desired sparsity is achieved.
Implications for LLM Compression
For large language models (LLMs), LTH offers a framework for extreme compression without significant performance loss. Key implications include:
- Weight Reinitialization: Winning tickets often require resetting to initial values rather than retaining trained weights, suggesting that optimization trajectories are less important than initialization.
- Sparsity Patterns: The hypothesis implies that early pruning (e.g., during pre-training) can identify structurally important connections, reducing computational overhead.
- Dynamic Sparsity: Recent extensions like RigL (Evci et al., 2020) show that dynamically updating the mask during training can improve performance at high sparsity levels.
Practical Considerations
Applying LTH to LLMs introduces unique challenges:
where si is the layer-wise sparsity and ni is the hidden dimension. While unstructured pruning reduces FLOPs theoretically, hardware efficiency depends on support for sparse operations. Structured pruning (e.g., removing entire attention heads) often yields better practical speedups.
Extensions and Limitations
Recent work has generalized LTH to:
- Structured Lottery Tickets: Pruning entire blocks or layers while preserving performance (You et al., 2020).
- Superposition Tickets: Allocating multiple subnetworks within a single model (Wortsman et al., 2022).
However, the hypothesis assumes i.i.d. data and may not hold for out-of-distribution tasks, requiring careful evaluation in real-world LLM deployments.

Iterative Magnitude Pruning for Subnetwork Discovery
Iterative Magnitude Pruning (IMP) is a structured approach to discovering sparse, high-performing subnetworks within large neural networks. The method leverages the empirical observation that many overparameterized models contain smaller subnetworks that achieve comparable performance to the dense original network when trained in isolation. IMP operates by progressively removing low-magnitude weights while retaining the structural integrity of the network.
Algorithmic Framework
The IMP procedure consists of three primary phases: training, pruning, and rewinding. Given a neural network f(x; θ) with parameters θ ∈ ℝd, the algorithm proceeds as follows:
where p represents the pruning ratio (fraction of weights removed) and ⊙ denotes element-wise multiplication. The mask m is constructed by zeroing out the smallest p% of weights by magnitude.
Learning Dynamics and Rewinding
Critical to IMP's success is the rewinding step, which resets remaining weights to their values from an earlier training iteration k < T while maintaining the sparsity pattern. This addresses the optimization challenges caused by pruning:
Theoretical work suggests this rewinding approximates training the subnetwork from initialization while benefiting from the original network's optimization trajectory. The optimal rewinding point k is typically early in training (10-20% of total iterations).
Convergence Properties
Under mild assumptions about the loss landscape, IMP converges to a sparse subnetwork with performance comparable to the original network. Let L(θ) be the loss function and θ* the optimal parameters. For a pruning schedule removing pt parameters at iteration t, the subnetwork error bound satisfies:
where C is a constant, n is the dataset size, and Rt represents the approximation error at pruning step t.
Practical Implementation
Effective application of IMP requires careful tuning of several hyperparameters:
- Pruning schedule: Linear, exponential, or one-shot removal of weights
- Rewinding epoch: Typically 10-20% of total training budget
- Final sparsity level: Often 90-99% for modern LLMs
- Warmup period: Initial training before first pruning iteration
The method shows particular effectiveness when combined with dynamic sparse training techniques, allowing the network to recover from overly aggressive pruning steps by temporarily reactivating promising connections.
Extensions and Variants
Recent advancements have produced several IMP derivatives:
- Global IMP: Prunes across all layers simultaneously rather than layer-wise
- Structured IMP: Removes entire neurons or attention heads
- Lottery Ticket IMP: Identifies subnetworks trainable from scratch
- Progressive IMP: Gradually increases sparsity during training
These variants trade off between computational efficiency, final model performance, and hardware compatibility, with structured pruning often yielding more practical speedups on conventional hardware.

Gradient-Based Methods for Subnetwork Identification
Gradient-based methods leverage the information contained in the gradients of the loss function with respect to the model parameters to identify critical subnetworks within large language models (LLMs). These approaches are grounded in the hypothesis that parameters with higher gradient magnitudes contribute more significantly to the model's performance, making them prime candidates for retention during compression.
Theoretical Foundation
The core idea stems from the first-order Taylor expansion of the loss function L around a parameter configuration θ. For a small perturbation Δθ, the change in loss can be approximated as:
This implies that parameters with larger gradient components |∂L/∂θᵢ| will induce more significant changes in the loss when modified. By preserving these high-gradient parameters and pruning others, we can maintain model performance while reducing size.
Implementation Strategies
Several gradient-based techniques have emerged for subnetwork identification:
- Gradient Magnitude Pruning: Ranks parameters by the L2-norm of their gradients over multiple training batches and retains the top-k.
- Gradient Flow Analysis: Tracks how gradients propagate through the network to identify critical pathways.
- Hessian-Gradient Product: Uses second-order information to account for parameter interactions.
Gradient Magnitude Scoring
The scoring function for parameter importance typically takes the form:
where N is the number of samples in the scoring batch. This empirical expectation smooths out stochastic variations in individual gradient estimates.
Practical Considerations
Effective implementation requires addressing several challenges:
- Gradient Instability: Gradients can vary significantly across batches, necessitating large enough sample sizes for reliable estimates.
- Memory Overhead: Storing gradients for all parameters during the scoring phase requires substantial memory resources.
- Dynamic Importance: Parameter importance can shift during training, suggesting iterative re-evaluation may be beneficial.
Recent work has shown that combining gradient information with activation patterns can yield more robust subnetworks. The gradient-activation product metric:
where a_i represents the activation at a given layer, captures both the parameter sensitivity and its actual usage during inference.
Advanced Variants
More sophisticated approaches incorporate gradient information into learnable masks:
where σ is the sigmoid function, and α, β are learnable parameters. This allows for soft, differentiable pruning during training while still converging to a hard subnetwork for inference.
Empirical studies have demonstrated that gradient-based methods can identify subnetworks comprising as little as 10-20% of original parameters while maintaining 90-95% of the full model's performance on benchmark tasks. The quality of identified subnetworks strongly correlates with the diversity and representativeness of the data used during the gradient computation phase.

3. Data Preparation and Model Initialization
Data Preparation and Model Initialization
Dataset Curation for Subnetwork Discovery
The selection and preprocessing of training data directly impacts the quality of discovered subnetworks. For language model compression, we require:
- Domain-representative samples covering the full linguistic distribution the model was trained on
- Balanced task distribution ensuring all capabilities (translation, QA, generation) are equally represented
- Computationally tractable size - typically 1-5% of original pretraining data
The data sampling process follows:
where τ controls the sharpness of sampling distribution, favoring more challenging examples that better expose model capabilities.
Model Initialization Strategies
Three initialization approaches prove effective for subnetwork extraction:
1. Warm-start from Pretrained Weights
Starting from the full pretrained model enables gradient-based mask learning. The initialization preserves:
- Embedding layer distributions
- Attention head specialization patterns
- Feedforward network feature hierarchies
2. Random Rewinding
For more aggressive compression, we rewind to early training checkpoints while preserving:
where η is the original learning rate and k is the rewind step.
3. Lottery Ticket Initialization
Iterative magnitude pruning identifies winning tickets - subnetworks that achieve comparable performance when trained in isolation. The mask m is initialized as:
where θ is the layer-specific percentile threshold.
Gradient Mask Initialization
The subnetwork mask gradients require careful initialization to avoid premature convergence:
- Bernoulli sampling with p=0.5 for unbiased exploration
- Gumbel-softmax relaxation for differentiable sampling during backpropagation
- Layer-wise sparsity constraints based on target compression ratios
The mask update rule incorporates both gradient signals and structural constraints:
where τ controls the softmax temperature and η is the mask learning rate.
Computational Considerations
Memory-efficient implementations leverage:
- Gradient checkpointing for transformer layers
- Sparse matrix operations for mask updates
- Mixed-precision training (FP16/FP32)
The initialization overhead remains manageable, typically adding <15% to baseline training time while enabling 5-10x compression ratios in subsequent steps.
3.2 Step-by-Step Extraction Pipeline
The extraction of optimal subnetworks from large language models (LLMs) involves a systematic pipeline that balances computational efficiency with minimal performance degradation. Below is a detailed breakdown of the process, including mathematical formulations and practical considerations.
1. Initialization and Pruning Criteria
The pipeline begins by defining a pruning criterion to identify less critical weights. A common approach is to use magnitude-based pruning, where weights below a threshold τ are removed. The threshold is often determined dynamically based on the desired sparsity level s:
Here, W represents the weight matrix, and s is the target sparsity (e.g., 0.5 for 50% sparsity). Alternatively, gradient-based criteria can be used to assess the importance of weights during fine-tuning.
2. Iterative Pruning and Fine-Tuning
Pruning is performed iteratively to avoid abrupt performance drops. At each step t, a fraction of weights is pruned, followed by fine-tuning to recover lost accuracy. The sparsity at step t is given by:
where si is the initial sparsity, sf is the final sparsity, and T is the total number of iterations. The cubic decay ensures gradual pruning, allowing the model to adapt.
3. Subnetwork Extraction via Lottery Ticket Hypothesis
The Lottery Ticket Hypothesis suggests that dense networks contain smaller subnetworks ("winning tickets") capable of matching the original performance. To extract such a subnetwork:
- Step 1: Train the full model and record the final weights Wf.
- Step 2: Reinitialize the model with early training weights W0.
- Step 3: Apply a binary mask M to retain only weights where |Wf| > τ.
The mask M is defined element-wise as:
4. Dynamic Sparsity Adaptation
To optimize the subnetwork further, dynamic sparsity adaptation adjusts the pruning threshold during training. A common method uses the movement pruning criterion, where weights are pruned based on their gradient movement rather than magnitude:
Weights with the smallest |ΔWij| are pruned first, as they contribute least to learning.
5. Validation and Performance Benchmarking
After extraction, the subnetwork is evaluated on a validation set to ensure performance parity with the original model. Key metrics include:
- Perplexity for language models.
- Task-specific accuracy (e.g., GLUE score for NLP tasks).
- Inference latency and memory footprint.
If performance drops significantly, the pruning threshold or fine-tuning duration is adjusted iteratively.
6. Deployment and Scalability Considerations
For deployment, the subnetwork is converted to a sparse format (e.g., CSR or CSC) to leverage hardware acceleration. Practical considerations include:
- Kernel support for sparse operations (e.g., NVIDIA’s Sparse Tensor Cores).
- Quantization to further reduce memory usage without significant accuracy loss.
- Distributed inference for large-scale applications.

Evaluating Subnetwork Performance: Benchmarks and Metrics
Performance Metrics for Compressed LLMs
The evaluation of compressed subnetworks requires multiple complementary metrics that capture different aspects of model quality. The primary metrics fall into three categories:
- Task Performance: Accuracy, perplexity, F1-score, BLEU (for translation tasks)
- Computational Efficiency: FLOPs, latency, memory footprint
- Compression Rate: Parameter count reduction, pruning ratio
For language models, perplexity remains the most fundamental metric, calculated as:
where N is the sequence length and p(w_i|w_{<i}) is the model's predicted probability for token w_i given previous tokens.
Benchmarking Protocols
Standardized evaluation requires carefully designed benchmarks that isolate compression effects from other variables. The most rigorous approach combines:
- Zero-shot evaluation on held-out test sets
- Fine-tuned evaluation on downstream tasks
- Stress testing with long-context or out-of-distribution examples
For transformer-based models, the compression-performance tradeoff curve provides critical insights. This plots model quality (y-axis) against compression ratio (x-axis), revealing the Pareto frontier of optimal subnetworks.
Latency and Throughput Measurement
Real-world deployment requires measuring inference speed under realistic conditions:
Key considerations include:
- End-to-end latency including tokenization and detokenization
- Memory bandwidth constraints
- Hardware-specific optimizations (e.g., tensor cores on GPUs)
Cross-Architecture Comparisons
When comparing subnetworks across different base architectures, normalized metrics become essential. The compression efficiency ratio accounts for baseline model differences:
This metric must be interpreted alongside absolute performance numbers, as it can mask quality degradation in highly compressed models.
Robustness Evaluation
Compression can affect model robustness in subtle ways. Comprehensive evaluation should include:
- Adversarial attack resistance (textual perturbations, synonym substitutions)
- Calibration metrics (expected calibration error, reliability diagrams)
- Failure mode analysis (error clustering by input type)
The effective robustness metric quantifies this relationship:
3.4 Case Study: Extracting a Subnetwork from GPT-3
Optimal subnetwork extraction from large language models like GPT-3 involves identifying a sparse, high-performance subset of weights that retains most of the original model's capabilities. The process begins with a pretrained GPT-3 model, typically with 175 billion parameters, and applies structured pruning techniques to isolate a computationally efficient subnetwork.
Mathematical Framework for Subnetwork Extraction
The core objective is to solve the constrained optimization problem:
where θ represents the full parameter set, θs is the subnetwork, ℒ is the loss function, and k is the target parameter count. The L0 norm enforces sparsity by limiting the number of non-zero parameters.
Iterative Magnitude Pruning with Rewinding
The extraction process follows an iterative procedure:
- Train the full GPT-3 model to convergence on the target task.
- Compute weight importance scores using magnitude-based criteria:
$$ I_{ij} = |W_{ij}| $$where Wij are the model weights.
- Prune the lowest-magnitude weights, retaining only the top-k by importance.
- Rewind the remaining subnetwork to its initialization state early in training.
- Retrain the pruned subnetwork to recover performance.
Architectural Considerations for GPT-3
When applied to GPT-3's transformer architecture, special attention must be paid to:
- Attention head pruning: Removing entire attention heads rather than individual weights preserves matrix operation efficiency.
- Layer-wise sparsity distribution: Allocating more parameters to middle layers, which typically exhibit higher sensitivity.
- Feed-forward network sparsity: Applying higher compression rates to the FFN layers compared to attention weights.
Performance Metrics and Tradeoffs
Experimental results on GPT-3 show that:
can be achieved with subnetworks containing only 10-15% of the original parameters, where Perf measures task-specific accuracy. The compression ratio depends heavily on the target task complexity, with simpler tasks allowing more aggressive pruning.
Practical Implementation Challenges
Key implementation hurdles include:
- Memory constraints during weight importance computation for 175B parameters
- Non-uniform GPU utilization during sparse matrix operations
- Maintaining stable training dynamics when rewinding large subnetworks
Recent advances in distributed pruning algorithms and block-sparse tensor operations have made subnetwork extraction feasible at GPT-3's scale. The resulting compressed models demonstrate comparable few-shot learning capabilities while reducing inference costs by 5-10x.

4. Scalability Issues in Very Large Models
4.1 Scalability Issues in Very Large Models
The rapid growth of large language models (LLMs) has exposed fundamental scalability challenges that emerge when model size exceeds a critical threshold. These issues manifest across computational, memory, and energy dimensions, often following non-linear scaling laws that defy naive expectations.
Computational Complexity Breakdown
The self-attention mechanism in transformers scales quadratically with sequence length N due to the pairwise token interaction computation:
where d represents the embedding dimension. For models like GPT-3 with N=2048 and d=12288, this results in approximately 2.4 × 1011 FLOPs per layer per forward pass. The total computational cost becomes:
where L is the number of layers, and Cffn accounts for the feed-forward network operations.
Memory Bottlenecks
Model parameters and activations create severe memory constraints during both training and inference. The parameter memory for a transformer with L layers scales as:
For a 175B parameter model using 16-bit precision, this requires 350GB just for parameters. Activation memory grows linearly with batch size B and sequence length N:
creating prohibitive memory demands for large B and N values.
Energy Consumption
The energy cost of training scales superlinearly with model size. Recent studies show the relationship follows:
where P is the parameter count. Training a 1B parameter model consumes approximately 27 MWh, while a 175B model requires over 1,000 MWh - comparable to the annual energy usage of 100 US households.
Communication Overhead
Distributed training introduces additional scaling constraints. The communication-to-computation ratio for data-parallel training is:
where β is the inverse network bandwidth, P is the parameter count, and C is the computational throughput. This ratio grows linearly with model size, creating fundamental scaling limits for synchronous training approaches.
Practical Implications
These scaling laws have forced several architectural adaptations:
- Model parallelism becomes necessary when single devices cannot hold model parameters
- Gradient checkpointing trades computation for memory by recomputing activations
- Sparse attention patterns reduce the O(N2) scaling to more manageable forms
- Mixed precision training reduces memory bandwidth pressure
Recent work on mixture-of-experts architectures demonstrates one promising direction, where the computational cost scales with the number of active parameters rather than total parameters:
where k is the number of experts per token and e is the expert hidden dimension.

Retaining Generalization Capabilities
When extracting subnetworks from large language models, a critical challenge is maintaining the model's ability to generalize beyond its training distribution. The lottery ticket hypothesis suggests that dense networks contain sparse, trainable subnetworks that can match the original model's performance when trained in isolation. However, naively pruning weights often degrades out-of-distribution generalization, even when in-distribution task performance remains high.
Generalization Metrics for Subnetwork Evaluation
To quantify generalization, we measure both:
- In-distribution (ID) accuracy: Performance on the training task distribution
- Out-of-distribution (OOD) robustness: Performance on shifted test distributions
The generalization gap G between ID and OOD performance can be formalized as:
where fθ represents the subnetwork's predictions and the expectations are taken over training and test distributions respectively.
Stabilizing OOD Performance Through Gradient Alignment
Recent work demonstrates that subnetworks maintaining similar gradient directions to the original model tend to preserve better generalization. We can measure this alignment via:
where α ∈ [-1,1] indicates the cosine similarity between original and subnetwork gradients. Subnetworks with α > 0.8 empirically show < 5% OOD performance degradation.
Practical Implementation via Gradient Preservation
To enforce gradient alignment during subnetwork extraction:
- Compute the full model's gradients on a diverse calibration set
- During pruning, preserve weights whose removal most impacts gradient direction
- Optimize the subnetwork mask m to minimize:
This formulation selectively keeps weights that contribute most to the original model's learning dynamics. The resulting subnetworks maintain 92-97% of the original model's OOD performance across common NLP benchmarks while reducing parameter counts by 60-80%.
Architectural Considerations
Attention heads in transformer layers show particularly strong gradient alignment properties. Preserving:
- At least 30% of attention heads per layer
- Full dimensionality for remaining heads
maintains >90% of the original model's few-shot learning capabilities. In contrast, uniformly pruning attention dimensions across all heads degrades few-shot performance by 15-20% even at identical parameter counts.

4.3 Computational Costs of Extraction Methods
The computational overhead of subnetwork extraction scales non-linearly with model size and depends critically on the search algorithm's complexity class. For a transformer with L layers, d attention heads, and hidden dimension h, the brute-force search space grows as:
Three dominant computational bottlenecks emerge during extraction:
1. Gradient Computation Overhead
First-order methods like Magnitude Pruning require only a single backward pass (O(n)), while second-order methods like Optimal Brain Surgeon must compute the Hessian inverse:
For a weight matrix W ∈ ℝm×n, this requires O(m3n3) operations - prohibitive for modern LLMs.
2. Subnetwork Evaluation Cost
Each candidate subnetwork requires validation on a holdout set. The Lottery Ticket Hypothesis approach evaluates k subnetworks through iterative magnitude pruning, requiring:
Where Tfwd and Tbwd are the forward/backward pass times for the full model.
3. Memory Bandwidth Constraints
Weight shuffling during Dynamic Sparse Training creates irregular memory access patterns. The Amdahl's Law-limited speedup is:
Where p is the parallelizable fraction and s is the sparsity level. For 90% sparsity, theoretical speedup plateaus at 10× even with infinite compute.
Practical Tradeoffs in Extraction Methods
- One-shot pruning: Lowest compute (single pass) but poor subnetwork quality
- Iterative pruning: O(log n) passes with progressive sparsification
- Reinforcement learning search: O(n2) policy updates but discovers better topologies
Recent work on sublinear extraction (Chen et al., 2023) approximates the Hessian-vector product using finite differences, reducing the complexity from O(n3) to O(n log n). The key insight is that most eigenvalues of the Hessian in LLMs cluster near zero, allowing low-rank approximation:
Empirical measurements on GPT-3 show that 99% of the Hessian's spectral energy is captured in the top 0.1% of eigenvectors, enabling practical computation.

5. Key Research Papers on Subnetwork Extraction
5.1 Key Research Papers on Subnetwork Extraction
- Data extraction methods for systematic review (semi)automation: Update ... — Version Changes Updated. Changes from Version 2. This version of the LSR includes 41 new papers. The article text was updated to reflect changes and new research trends such as large language models (LLMs) being used to extract data, as well as continuing trends in increased availability of datasets, source code, relation extraction and summarisation.
- Compressing Large Language Models with Automated Sub-Network Search — Large Language Models (LLMs) demonstrate exceptional reasoning abilities, enabling strong generalization across diverse tasks such as commonsense reasoning and instruction following. However, as LLMs scale, inference costs become increasingly prohibitive, accumulating significantly over their life cycle. In this paper we consider model compression for LLMs to reduce model size while improving ...
- [PDF] Compressing Large Language Models with ... - Semantic Scholar — This paper considers model compression for LLMs to reduce model size while improving downstream task performance, phrase this as a neural architecture search problem that automatically prunes structural components by searching for the Pareto-optimal set of sub-networks balancing between performance and on-device latency. Large Language Models (LLMs) demonstrate exceptional reasoning abilities ...
- Efficient Compressing and Tuning Methods for Large Language Models: A ... — With the advent of large language models (LLMs), a significant shift has occurred in the research community, with many scholars focusing on the intricate mechanisms that underpin language models at scale within the realm of natural language processing (NLP).Meanwhile, a diverse group of researchers, multinational corporations, and organizations have turned their efforts toward developing ...
- Compressing LLMs: The Truth is Rarely Pure and Never Simple — Despite their remarkable achievements, modern Large Language Models (LLMs) face exorbitant computational and memory footprints. Recently, several works have shown significant success in training-free and data-free compression (pruning and quantization) of LLMs that achieve 50 - 60% sparsity and reduce the bit width to 3 or 4 bits per weight, with negligible degradation of perplexity over the ...
- Compressing LLMs: The Truth is Rarely Pure and Never Simple — Despite their remarkable achievements, modern Large Language Models (LLMs) encounter exorbitant computational and memory footprints. Recently, several works have shown significant success in *training-free* and *data-free* compression (pruning and quantization) of LLMs achieving 50-60\% sparsity and reducing the bit-width down to 3 or 4 bits per weight, with negligible perplexity degradation ...
- A Survey on Model Compression for Large Language Models — Abstract. Large Language Models (LLMs) have transformed natural language processing tasks successfully. Yet, their large size and high computational needs pose challenges for practical use, especially in resource-limited settings. Model compression has emerged as a key research area to address these challenges. This paper presents a survey of model compression techniques for LLMs. We cover ...
- GitHub - vllm-project/llm-compressor: Transformers-compatible library ... — Big updates have landed in LLM Compressor! Check out these exciting new features: Axolotl Sparse Finetuning Integration: Easily finetune sparse LLMs through our seamless integration with Axolotl.Learn more here.; AutoAWQ Integration: Perform low-bit weight-only quantization efficiently using AutoAWQ, now part of LLM Compressor.Note: This integration should be considered experimental for now.
- Harnessing Small Models and LLMs: A Key Points Extraction Approach for ... — Wireless network tuning logs, pivotal in maintaining and improving network health, offer invaluable insights about network issues. However, the variance in their structure and content can make extracting essential details time-consuming and challenging. This thesis presents a novel method combining the power of small models with Large Language Models (LLMs) to efficiently generate high-quality ...
- arXiv:2308.07633v4 [cs.CL] 30 Jul 2024 — A Survey on Model Compression for Large Language Models Xunyu Zhu 1,2, Jian Li ∗, Yong Liu 3, Can Ma 1,2, Weiping Wang 1Institute of Information Engineering, Chinese Academy of Sciences 2School of Cyber Security, University of Chinese Academy of Sciences 3Gaoling School of Artificial Intelligence, Renmin University of China {zhuxunyu, lijian9026, macan, wangweiping}@iie.ac.cn, liuyonggsai ...
5.2 Tools and Libraries for Model Compression
- Model Compression - an overview | ScienceDirect Topics — Model compression tries to reduce the expenses associated with large model sizes by representing the model more efficiently with little performance effect. Model compression may be classified in numerous ways and is especially utilized for deployments in embedded devices, accelerators, and mobile platforms [264-266]. Pruning, quantization ...
- A Survey on Model Compression for Large Language Models — Abstract. Large Language Models (LLMs) have transformed natural language processing tasks successfully. Yet, their large size and high computational needs pose challenges for practical use, especially in resource-limited settings. Model compression has emerged as a key research area to address these challenges. This paper presents a survey of model compression techniques for LLMs. We cover ...
- PDF A Survey on Model Compression for Large Language Models - ACL Anthology — as model compression (Han et al., 2016) offers ∗Corresponding author. a solution. Model compression involves trans-forming a large, resource-intensive model into a compact version suitable for deployment on resource-constraineddevices.Additionally,model compression can enhance LLM inference speed and optimizes resource efficiency.
- Compressing Large Language Models with Automated Sub-Network Search — model compression for LLMs to reduce model size while improving downstream task perfor-mance. We phrase this as a neural architec-ture search problem that automatically prunes structural components, such as attention heads, neurons, and layers by searching for the Pareto-optimal set of sub-networks balancing between performance and on-device ...
- [2402.09748] Model Compression and Efficient Inference for Large ... — 1) High compression ratio: quantizing the weights in LLMs from 32-bit float to 4-bit integer could drastically compress the model size to approximately 1 / 8 1 8 1/8, essential for memory-bound 1 1 1 "memory-bound" means that the transfer between the device and global memory nearly reaches the limitation or fetching data from the memory is ...
- Efficient Compressing and Tuning Methods for Large Language Models: A ... — A distilled model trained on the output of a larger teacher model can maintain its performance of the larger model despite aggressive compression. By aligning the output distributions of the student and teacher models, knowledge distillation can help the student model retain rich, generalized knowledge, even when quantization is applied [ 68 ].
- PDF Combining Improvements in the Compression of Large Language Models — to the Kronecker factors instead of the original matrix. For compression of LLMs, Kronecker decomposition uses a small matrix size (i.e. 2×1, 2×2) to compress matrices 2-4 times and impose explicit block structure in the weight matrices. 4 Methods We have identified two techniques which all, in their own way, decrease the number of model
- LLMLingua: Innovating LLM efficiency with prompt compression — During our test, we used LLaMA-7B as the small language model and GPT-3.5-Turbo-0301, one of OpenAI's LLMs, as the closed LLM. The results show that LLMLingua maintains the original reasoning, summarization, and dialogue capabilities of the prompt, even at a maximum compression ratio of 20x, as reflected in the evaluation metric (EM) columns ...
- Introduction of LLM Compression.md - GitHub — Quantization involves reducing the precision of the model's weights and activations from higher bit-widths (e.g., 32-bit floating-point) to lower bit-widths (e.g., 8-bit integers). This process decreases the model size and accelerates inference by enabling faster arithmetic operations. Types of Quantization: Post-Training Quantization (PTQ): Applied after the model is trained.
- (PDF) Efficient Compression of Large Language Models: A Case Study on ... — Efficient compression of large language models is important for enhancing computational efficiency and reducing the necessary virtual memory requirements for deployment in resource-limited ...
5.3 Advanced Topics and Ongoing Research
- Compressing LLMs: The Truth is Rarely Pure and Never Simple — Despite their remarkable achievements, modern Large Language Models (LLMs) encounter exorbitant computational and memory footprints. Recently, several works have shown significant success in training-free and data-free compression (pruning and quantization) of LLMs achieving 50-60% sparsity and reducing the bit-width down to 3 or 4 bits per weight, with negligible perplexity degradation over ...
- Efficient Compressing and Tuning Methods for Large Language Models: A ... — With the advent of large language models (LLMs), a significant shift has occurred in the research community, with many scholars focusing on the intricate mechanisms that underpin language models at scale within the realm of natural language processing (NLP).Meanwhile, a diverse group of researchers, multinational corporations, and organizations have turned their efforts toward developing ...
- A Survey of Research in Large Language Models for Electronic Design ... — In recent years, Large Language Models (LLMs) have risen prominently in the field of machine learning. These models are typically characterized by their extensive training on web-scale datasets and exceptional ability in Natural Language Processing (NLP).In NLP, models such as GPT-3 [] and its successors [] have significantly advanced the capabilities of natural language generation, enabling ...
- Compressing Large Language Models with Automated Sub-Network Search — model compression for LLMs to reduce model size while improving downstream task perfor-mance. We phrase this as a neural architec-ture search problem that automatically prunes structural components, such as attention heads, neurons, and layers by searching for the Pareto-optimal set of sub-networks balancing between performance and on-device ...
- Optimizing LLMs for Resource-Constrained Environments: A Survey of ... — View a PDF of the paper titled Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques, by Sanjay Surendranath Girija and 5 other authors View PDF Abstract: Large Language Models (LLMs) have revolutionized many areas of artificial intelligence (AI), but their substantial resource requirements limit their ...
- PDF DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs — fine-tuned LLMs concurrently is challenging due to the spo-radic, bursty, and varying request patterns of different LLMs. To bridge this gap, we present DeltaZip, an LLM serving sys- ... tem that efficiently serves multiple full-parameter fine-tuned models concurrently by aggressively compressing model deltas by up to 10×while maintaining high ...
- Semantic Compression with Large Language Models - IEEE Xplore — The rise of large language models (LLMs) is revolutionizing information retrieval, question answering, summarization, and code generation tasks. However, in addition to confidently presenting factually inaccurate information at times (known as "hallucinations"), LLMs are also inherently limited by the number of input and output tokens that can be processed at once, making them potentially ...
- Title: A Survey on Model Compression for Large Language Models - arXiv.org — Large Language Models (LLMs) have transformed natural language processing tasks successfully. Yet, their large size and high computational needs pose challenges for practical use, especially in resource-limited settings. Model compression has emerged as a key research area to address these challenges. This paper presents a survey of model compression techniques for LLMs. We cover methods like ...
- PDF Retrieval-based Knowledge Transfer: An Effective Approach for Extreme ... — LLMs. This indicates the effectiveness and prac-ticality of the retrieval-based knowledge transfer paradigm for extreme model compression. Our contributions can be summarized as follows: We propose a new compression paradigm called Retrieval-based Knowledge Transfer, which aims to transfer knowledge from LLMs to extremely small-scale models. This
- D C Trust Scrutinizing the Trustworthiness of Efficient Compression — Compressing high-capability Large Language Models (LLMs) has emerged as a favored strategy for resource-efficient inferences. While state-of-the-art (SoTA) compression methods boast impressive advancements in preserving benign task performance, the potential risks of compression in terms of safety and trustworthi-ness have been largely neglected.






