Secure Aggregation Techniques
1. Definition and Core Principles
Secure Aggregation Techniques: Definition and Core Principles
Secure aggregation is a cryptographic protocol that enables multiple parties to compute the sum of their private inputs without revealing individual values. It is a foundational technique in privacy-preserving machine learning, particularly in federated learning settings where clients collaboratively train a model without exposing their raw data. The core principle hinges on additive homomorphic encryption or masking schemes that allow aggregation while preserving data confidentiality.
Mathematical Foundations
The protocol relies on the additive property of certain cryptographic schemes. Given n clients each holding a private vector xi, secure aggregation computes:
while ensuring no party learns any xi beyond what can be inferred from the sum. A common approach uses secret sharing: each client splits their input into shares distributed among other clients such that only the aggregate can be reconstructed. For two clients, this can be expressed as:
where sij is client i's share sent to client j. The server then computes:
Key Properties
- Input Privacy: Individual inputs remain hidden even if up to t parties collude, where t is a threshold parameter.
- Correctness: The output is guaranteed to equal the sum of all valid inputs.
- Dropout Resilience: The protocol tolerates a fraction of clients disconnecting mid-computation.
Practical Implementation
Modern implementations often use pairwise Diffie-Hellman key agreements to generate correlated random masks that cancel out upon aggregation. Each client i generates a shared secret kij with every other client j, then masks their input as:
where R is a large integer range. When all yi are summed, the pairwise masks cancel out, leaving only the sum of xi.
Security Considerations
The protocol must withstand both passive and active adversaries. Passive attackers observe communication but follow the protocol, while active attackers may deviate arbitrarily. Robust secure aggregation requires:
- Zero-knowledge proofs to verify correct share generation
- Digital signatures to authenticate messages
- Secure channels to prevent man-in-the-middle attacks
Recent advances incorporate lattice-based cryptography for post-quantum security and functional encryption for more complex aggregation functions beyond simple sums.

Threat Models and Security Requirements
Adversarial Capabilities in Secure Aggregation
Secure aggregation protocols must account for adversaries with varying capabilities. A semi-honest (passive) adversary follows the protocol but attempts to infer private data from observed messages. In contrast, a malicious (active) adversary may deviate arbitrarily—injecting false inputs, dropping messages, or manipulating computations. Federated learning systems often assume semi-honest participants but must guard against malicious clients or a compromised central server.
where View is the adversary's observation during real execution and Sim is a simulated view. The protocol is secure if this advantage is negligible.
Security Requirements
Four core properties must be guaranteed:
- Input Privacy: Individual user inputs remain hidden from other participants and the server. Achieved via cryptographic techniques like additive secret sharing or homomorphic encryption.
- Correctness: The aggregated result must be verifiably accurate despite adversarial interference. This requires mechanisms like zero-knowledge proofs or redundant cross-checks.
- Robustness: The system should tolerate up to t malicious participants without compromising results. Threshold cryptosystems are commonly employed.
- Forward Secrecy: Compromised keys from past rounds must not reveal historical data. Key rotation and ephemeral keys address this.
Real-World Attack Vectors
Practical threats include:
- Model Inversion: Reconstructing training data from gradient updates (e.g., using GAN-based attacks).
- Membership Inference: Determining if a specific sample was in the training set.
- Sybil Attacks: An adversary controls multiple fake clients to skew aggregation.
where Δ is noise scale in differential privacy. Higher sensitivity increases vulnerability.
Case Study: Federated Learning with Secure Aggregation
In Google's 2017 implementation, clients encrypt local updates using pairwise Diffie-Hellman keys. The server only sees the sum of updates, not individual contributions. The protocol withstands client dropouts and maintains privacy against a honest-but-curious server.
1.3 Key Cryptographic Primitives Used
Secure aggregation relies on cryptographic primitives that enable privacy-preserving computation over distributed data. The most critical primitives include homomorphic encryption, secret sharing, and secure multi-party computation (MPC) protocols. Each serves a distinct role in ensuring data confidentiality while permitting meaningful computation.
Homomorphic Encryption
Partially homomorphic encryption (PHE) schemes allow specific algebraic operations on ciphertexts without decryption. For secure aggregation, additive homomorphism is particularly useful. Given two ciphertexts E(a) and E(b), the property ensures:
where ⊕ denotes a homomorphic addition operation. The Paillier cryptosystem is widely adopted for this purpose due to its efficiency and provable security under the decisional composite residuosity assumption. Its encryption function for a message m and random r is:
where n is an RSA modulus and g is a generator. The decryption function exploits the Carmichael function to recover m.
Secret Sharing
Shamir's secret sharing (SSS) enables distributed storage of sensitive data by splitting a secret s into n shares, where any t shares can reconstruct s. The scheme operates over a finite field using polynomial interpolation:
Each share is a point (x_i, f(x_i)). Secure aggregation protocols often use verifiable secret sharing (VSS) to detect malicious share distribution, employing Pedersen commitments or Feldman verifiability.
Secure Multi-Party Computation
MPC protocols like GMW or SPDZ extend secret sharing to active computation. For secure aggregation, garbled circuits and oblivious transfer (OT) are frequently combined. A 1-out-of-2 OT protocol allows a receiver to learn one of two sender values without revealing which was chosen. The Naor-Pinkas OT scheme achieves this under the DDH assumption:
where c is the receiver's choice bit. These primitives compose to enable privacy-preserving federated learning and other distributed analytics.
Zero-Knowledge Proofs
Non-interactive zero-knowledge (NIZK) proofs verify correctness of encrypted computations without revealing inputs. Groth16 zk-SNARKs are particularly efficient for arithmetic circuit satisfiability, with proofs consisting of only three group elements:
The verification equation involves pairing operations e(A, B) = e(g^α, h^β) · e(g^f, h) · e(C, h^γ), where α, β, γ are toxic waste discarded after setup. This enables validation of aggregated model updates in federated learning while preserving user privacy.

2. Homomorphic Encryption for Aggregation
2.1 Homomorphic Encryption for Aggregation
Homomorphic encryption (HE) enables computations on encrypted data without decryption, making it a cornerstone of secure aggregation in federated learning and privacy-preserving analytics. Unlike traditional encryption, which requires decryption before processing, HE allows arithmetic operations directly on ciphertexts, producing encrypted results that decrypt to the correct output.
Mathematical Foundations
Partially Homomorphic Encryption (PHE) schemes support either addition or multiplication, while Fully Homomorphic Encryption (FHE) permits both. The most widely used additive HE scheme is Paillier encryption, defined as follows:
where m is the plaintext, r is a random integer, n is the product of two large primes, and g is a generator. The homomorphic property ensures:
This additive property allows secure aggregation of encrypted gradients or model updates in federated learning. For example, if clients submit encrypted updates E(Δw₁), E(Δw₂), the server computes their product to obtain E(Δw₁ + Δw₂) without accessing individual Δwᵢ.
Practical Implementation Challenges
Despite its theoretical promise, HE introduces computational overhead. Paillier encryption expands ciphertext size to O(n²) bits, and FHE operations are orders of magnitude slower than plaintext computations. Recent optimizations include:
- Lattice-based schemes (e.g., CKKS, BFV) that support approximate arithmetic for machine learning.
- Batching: Packing multiple values into a single ciphertext to amortize costs.
- Hybrid approaches: Combining HE with secure multi-party computation (SMPC) for efficiency.
Case Study: Federated Learning with HE
In a federated averaging (FedAvg) scenario, clients encrypt local gradients using Paillier before transmission. The server aggregates ciphertexts multiplicatively, then decrypts the sum once. This prevents the server from inferring individual data points while preserving the global model's accuracy. However, the scheme must address:
- Client dropout: Incomplete updates can bias the aggregated result.
- Quantization errors: Fixed-point encoding for HE may degrade model performance.
Limitations and Mitigations
HE alone cannot prevent all privacy leaks. For instance, the number of non-zero updates may reveal client activity patterns. To address this, differential privacy noise is often added before encryption. Additionally, newer schemes like threshold HE distribute decryption keys among multiple parties to prevent single-point failures.
Recent advances in GPU-accelerated HE libraries (e.g., Microsoft SEAL, PALISADE) have reduced latency, but real-world deployment still requires careful trade-offs between security, accuracy, and performance.

2.2 Secure Multi-party Computation (SMPC) Approaches
Secure Multi-party Computation (SMPC) enables multiple parties to jointly compute a function over their private inputs without revealing those inputs to each other. This cryptographic primitive is foundational for privacy-preserving federated learning, where model updates must be aggregated without exposing individual contributions. SMPC protocols achieve this by decomposing computations into shares distributed across participants, ensuring no single party can reconstruct another's private data.
Garbled Circuits
Yao's Garbled Circuits protocol allows two parties to evaluate arbitrary Boolean circuits without revealing their inputs. The circuit generator (typically the server) encrypts each gate's truth table using symmetric keys, while the evaluator (client) obliviously decrypts only the gates relevant to their input via oblivious transfer. For a simple AND gate with inputs x and y:
Where kxb denotes the key for bit value b of input x. The evaluator decrypts only one ciphertext per gate using keys obtained via oblivious transfer, learning nothing about other input combinations.
Secret Sharing
Shamir's Secret Sharing splits a value s into n shares using a random polynomial of degree t:
Each party receives a point (i, f(i)). The original secret can be reconstructed via Lagrange interpolation when at least t+1 shares are combined. For additive secret sharing, a simpler variant where s = s1 + s2 + \cdots + sn mod p is often used in federated learning aggregation.
Homomorphic Encryption
Partially homomorphic schemes like Paillier encryption enable secure aggregation of model updates. For encrypted weights E(w1) and E(w2):
Where N is the Paillier modulus. This allows the server to compute the sum of encrypted gradients without decrypting individual contributions. Fully homomorphic encryption (FHE) extends this to arbitrary computations but incurs prohibitive computational overhead for large neural networks.
Hybrid Approaches
Practical implementations often combine techniques. The Prio system uses additive secret sharing for scalability with lightweight verification, while Secure Aggregation for federated learning employs masked model updates with pairwise cryptographic seeds. For a network with n participants:
- Each client generates random seeds with every other client via Diffie-Hellman key exchange
- Model updates are masked by pairwise pseudorandom numbers derived from these seeds
- The server sums all masked updates, where pairwise masks cancel out:
$$ \sum_{i=1}^n (w_i + \sum_{j \neq i} \text{PRG}(s_{ij}) - \text{PRG}(s_{ji})) = \sum_{i=1}^n w_i $$
This approach provides information-theoretic security against colluding parties while maintaining practical communication costs linear in the number of participants.

2.3 Differential Privacy Integration
Integrating differential privacy (DP) with secure aggregation ensures that even if an adversary gains access to aggregated data, individual contributions remain statistically obfuscated. The core mechanism involves adding calibrated noise to each client's update before aggregation, adhering to formal privacy guarantees such as (ε, δ)-DP. This noise is typically drawn from distributions like the Gaussian or Laplace, scaled to the sensitivity of the function being computed.
Mathematical Formulation
Given a function f with L2-sensitivity Δf, the Gaussian mechanism ensures (ε, δ)-DP by adding noise sampled from N(0, σ2), where:
For federated learning, the sensitivity Δf is often bounded by gradient clipping, ensuring updates satisfy ∥g∥2 ≤ C. The noise scale σ then depends on the clipping norm C and the desired privacy budget.
Implementation in Secure Aggregation
In a federated setting, DP integration occurs at two levels:
- Local DP: Clients add noise to their updates before submission, ensuring privacy even if the server is compromised. This requires careful coordination to avoid excessive noise accumulation.
- Global DP: The server adds noise post-aggregation, leveraging the secure aggregation protocol to hide individual contributions while maintaining utility.
A hybrid approach combines both: clients clip gradients and apply local noise, while the server injects additional noise during aggregation. The total privacy cost composes via advanced composition theorems or the moments accountant method.
Privacy-Accuracy Trade-offs
The choice of ε and δ directly impacts model performance. Smaller ε values provide stronger privacy but degrade accuracy due to higher noise. Empirical studies show that for ε ∈ [0.1, 1.0] and δ = 10−5, the accuracy drop in image classification tasks is typically under 5%.
where d is the model dimension and n is the number of clients. This highlights the tension between high-dimensional models and privacy preservation.
Case Study: Federated Learning with DP
In a 2022 implementation of DP-SGD for federated learning, researchers used the following parameters:
- Gaussian noise with σ = 1.1
- Clipping norm C = 1.0
- Privacy budget ε = 0.5, δ = 10−6
The resulting model achieved 92% of the non-private baseline accuracy on CIFAR-10, demonstrating practical viability for privacy-sensitive applications.

3. Federated Learning with Secure Aggregation
Federated Learning with Secure Aggregation
Federated learning (FL) enables decentralized model training across multiple clients while preserving data privacy by keeping raw data localized. Secure aggregation (SecAgg) enhances this framework by ensuring that individual client updates remain confidential during the aggregation phase, even from the central server. This is achieved through cryptographic techniques that allow the server to compute only the sum of updates without accessing individual contributions.
Cryptographic Foundations
Secure aggregation relies on two primary cryptographic primitives: secret sharing and homomorphic encryption. Secret sharing distributes a client's model update into shares, which are then distributed among other clients or servers. Homomorphic encryption allows computations to be performed on encrypted data without decryption, enabling the server to aggregate updates while preserving privacy.
Here, ΔWi represents the model update from client i, and Enc(·) denotes homomorphic encryption. The server decrypts only the aggregated result, ensuring individual updates remain hidden.
Protocol Design
The secure aggregation protocol in federated learning typically follows these steps:
- Setup Phase: Clients generate cryptographic keys and establish secure communication channels.
- Masked Submission: Each client masks their update with a random value (secret-shared among other clients) before transmission.
- Aggregation Phase: The server combines the masked updates, and the random masks cancel out, revealing only the sum.
This approach ensures robustness against dropout attacks, where malicious clients attempt to disrupt aggregation by withholding shares.
Practical Considerations
Implementing secure aggregation introduces computational and communication overhead. The complexity scales with the number of clients and model parameters, making optimizations such as gradient quantization and sparse aggregation essential for scalability. Recent advancements, like the FastSecAgg protocol, reduce latency by leveraging efficient key exchange mechanisms.
Case Study: Cross-Silo Federated Learning
In healthcare, multiple hospitals collaboratively train a model without sharing patient data. Secure aggregation ensures compliance with regulations like HIPAA by preventing the central server from accessing individual hospital updates. A real-world deployment by Google demonstrated a 30% reduction in communication rounds using SecAgg while maintaining model accuracy.
Security Guarantees
Secure aggregation provides formal privacy guarantees under the honest-but-curious adversary model, where the server follows the protocol but may attempt to infer client data. For stronger security against active adversaries, techniques like differential privacy can be combined with SecAgg.
Here, ε quantifies privacy loss, Δf is the sensitivity of the aggregation function, and λ controls noise magnitude. This ensures that even if an adversary observes the aggregated output, individual contributions remain indistinguishable.

3.2 Gossip-based Protocols for Decentralized Aggregation
Gossip-based protocols, inspired by epidemic spreading models, provide a robust mechanism for decentralized aggregation in distributed systems. These protocols operate through pairwise, asynchronous communication between nodes, where each node periodically exchanges state information with a randomly selected neighbor. The stochastic nature of this communication ensures eventual consistency while maintaining resilience against node failures and network partitions.
Mathematical Foundations
The convergence properties of gossip protocols can be analyzed using Markov chains or Lyapunov stability theory. Consider a network of N nodes where each node i maintains a local value xi. During each gossip round:
where j is a randomly selected neighbor. This averaging process leads to global consensus at the network average value:
The convergence rate depends on the spectral gap of the network's Laplacian matrix, with complete graphs achieving fastest convergence.
Secure Aggregation Variants
For privacy-preserving aggregation, cryptographic techniques can be integrated with gossip protocols:
- Differential Privacy: Nodes add carefully calibrated noise to their values before gossiping, ensuring (ε,δ)-differential privacy guarantees.
- Homomorphic Encryption: Computations occur on encrypted values, with pairwise interactions preserving the homomorphic property.
- Secret Sharing: Each node's input is split into shares distributed among neighbors, with aggregation performed on the shares.
Practical Considerations
Real-world implementations must address several challenges:
- Network Dynamics: Gossip protocols naturally adapt to changing network topologies, making them suitable for mobile or ad-hoc environments.
- Resource Constraints: The O(log N) convergence time makes gossip efficient for large networks, though message complexity grows with network diameter.
- Byzantine Resilience: Variants like verified averaging can detect and exclude malicious nodes attempting to corrupt the aggregation.
Case Study: Federated Learning
In federated learning systems, gossip protocols enable decentralized model aggregation without a central coordinator. Each device:
- Computes a local model update
- Exchanges updates with randomly selected peers
- Performs weighted averaging
This approach reduces communication bottlenecks while preserving data locality, with privacy benefits over centralized aggregation.
Performance Optimization
The gossip period τ must balance convergence speed against network load:
where α represents computation costs and β communication costs. Adaptive algorithms can dynamically adjust τ based on network conditions.

Handling Dropouts and Byzantine Faults
Secure aggregation protocols must account for two critical failure modes: dropouts (participants leaving the computation unexpectedly) and Byzantine faults (participants submitting maliciously crafted inputs). Both scenarios threaten the correctness and privacy guarantees of federated learning systems.
Dropout Resilience
Dropouts occur when clients disconnect during aggregation due to network instability or device failures. A robust protocol must ensure the aggregated result remains computable even if a subset of participants vanish. The key challenge lies in reconstructing partial contributions without violating privacy.
Here, S denotes surviving clients, D represents dropouts, and 𝔼[g_j] estimates missing gradients. Practical implementations often use:
- Secret sharing thresholds: Shamir's scheme with redundancy factor t+1 allows recovery from t dropouts
- Commitment schemes: Clients submit cryptographic commitments before participating
- Lagrange interpolation: Reconstructs missing shares from available participants
Byzantine Robustness
Byzantine participants may submit arbitrary values to corrupt the aggregation. Defenses typically combine cryptographic verification with statistical methods:
Where g_(k) is the k-th ordered gradient and τ a robustness threshold. Advanced techniques include:
- Krum aggregation: Selects the vector closest to its nearest neighbors
- Bulyan defense: Iteratively filters outliers using coordinate-wise median
- Zeno++: Uses historical trends to detect anomalous updates
Hybrid Approaches
State-of-the-art frameworks like Elastic combine dropout tolerance with Byzantine resilience through:
- Redundant secret sharing with verifiable reconstruction
- Adaptive clipping bounds based on participant history
- Differential privacy noise calibrated to failure rates
Where c_t dynamically adjusts based on observed dropout rates and Byzantine detection statistics.

4. Computational Overhead Analysis
4.1 Computational Overhead Analysis
Secure aggregation protocols introduce computational overhead due to cryptographic operations, masking mechanisms, and distributed computation. The primary contributors to this overhead include:
- Homomorphic encryption — Modular exponentiation and polynomial interpolation dominate runtime.
- Secret sharing — Polynomial generation and reconstruction scale with the number of participants.
- Multi-party computation (MPC) — Communication rounds and verifiable computations add latency.
Mathematical Modeling of Overhead
The computational cost for a federated learning system with N clients using additive secret sharing can be modeled as:
where:
- Tenc is the time for encrypting a local model update (e.g., using Paillier encryption).
- Tshare is the time to generate secret shares (Shamir's scheme).
- Tagg is the aggregation time, which depends on the reconstruction complexity.
Case Study: Federated Averaging with Secure Aggregation
For a federated averaging (FedAvg) scenario, the overhead grows quadratically with the model dimension d due to masking operations. The per-client computation time is:
where p is the prime modulus size, and k is the number of shares required for reconstruction. For large models (e.g., d > 106), this becomes prohibitive without optimization.
Optimization Techniques
To mitigate overhead, modern systems employ:
- Batching — Aggregating multiple updates in a single cryptographic operation.
- Lattice-based cryptography — Replacing RSA/Paillier with faster homomorphic schemes like CKKS.
- Parallelization — Distributing share generation across GPUs or TPUs.
Benchmark Example
A ResNet-18 model (d ≈ 11M parameters) secured with 256-bit Paillier encryption requires:
whereas switching to CKKS reduces this to ~0.3 seconds/client at the cost of approximate arithmetic.
4.2 Communication Efficiency Trade-offs
Secure aggregation protocols must balance cryptographic security with communication overhead, a challenge exacerbated in distributed settings with resource-constrained devices. The trade-offs arise from three primary factors: encryption overhead, coordination complexity, and bandwidth constraints. For instance, homomorphic encryption enables privacy-preserving aggregation but introduces multiplicative communication costs due to ciphertext expansion. A typical additive homomorphic scheme like Paillier expands a 32-bit integer to 2048 bits, increasing bandwidth by 64×.
Quantifying Communication Costs
The total communication cost C for a secure aggregation protocol can be modeled as:
where N is the number of participants, Splain and Senc are the sizes of plaintext and ciphertext messages, and R(d) is the redundancy factor due to network diameter d. For a star topology with pairwise secure channels, R(d) = 2, while tree-based aggregation reduces it to R(d) = O(\log N) at the cost of increased latency.
Optimization Strategies
Three approaches mitigate communication overhead:
- Compression: Applying quantization or sparsification to model updates before encryption. For example, 1-bit stochastic gradient descent reduces Splain by 32× while maintaining convergence.
- Topology-aware aggregation: Leveraging network structure to minimize R(d). Hierarchical aggregation trees can reduce cross-device transmissions by 40% compared to flat topologies.
- Hybrid schemes: Combining lightweight symmetric encryption for intra-cluster communication with asymmetric crypto for global aggregation. The SecAgg+ protocol demonstrates 22% lower bandwidth than pure homomorphic approaches.
Case Study: Federated Learning with Secure Aggregation
In a 1000-device federated learning scenario, vanilla secure aggregation requires 2.3 TB of total communication per round for ResNet-18 updates (45 MB/device). Employing the above optimizations reduces this to 98 GB:
The 1.4 MB plaintext size results from 8-bit quantization and 90% gradient sparsification, while 2.8 MB ciphertexts use elliptic curve cryptography (ECC) with point compression. The redundancy factor of 1.5 accounts for a two-level aggregation hierarchy.
Latency-Bandwidth Trade-offs
Communication efficiency gains often come at the expense of increased computation or latency. For example:
- Adding one hierarchy level reduces bandwidth by 30% but increases round-trip time by 40 ms due to intermediate aggregation steps.
- Using lattice-based cryptography instead of ECC decreases ciphertext size by 35% but raises encryption time by 8× on mobile CPUs.
The optimal operating point depends on the specific constraints. Wireless sensor networks prioritize energy efficiency, favoring larger compression ratios, while data center deployments may minimize latency through flatter topologies.

4.3 Optimizations for Large-scale Deployments
Efficient Communication Protocols
Large-scale deployments of secure aggregation require minimizing communication overhead while preserving privacy. Traditional secure multi-party computation (MPC) protocols suffer from quadratic communication complexity, making them impractical for federated learning with thousands of participants. Recent advances leverage ring-based topology or tree-structured aggregation to reduce this to O(n log n) or even O(n) in some cases.
Where d is the model dimension, p is the modulus size, k is the number of neighbors in the topology, and λ is the security parameter. For a 1M-parameter model with 1000 participants, this reduces communication from 100GB to under 10GB per round.
Quantization and Sparsification
Model updates in federated learning often contain redundant information. Two key optimizations:
- Stochastic quantization: Represent gradients using 8-bit fixed-point instead of 32-bit floats, with error compensation to maintain convergence.
- Top-k sparsification: Only transmit the largest k% of gradient values, reconstructing the remainder via zero-filling or local estimation.
Where β is a compensation factor and e tracks quantization error. When combined with secure aggregation, these techniques must preserve the additive homomorphic properties of the encryption scheme.
Hierarchical Aggregation
For geo-distributed deployments, a two-tier hierarchy improves scalability:
Edge devices first aggregate within local clusters (green), then regional servers (blue) perform cross-cluster secure aggregation before forwarding to the global server (orange). This reduces WAN traffic by 60-80% in empirical studies.
Hardware Acceleration
Modern cryptographic operations for secure aggregation (Paillier, CKKS, or lattice-based schemes) benefit from GPU/TPU acceleration. Key optimizations include:
- Batched modular exponentiation using NVIDIA CUDA cores
- Parallelized number-theoretic transform (NTT) for RLWE operations
- Custom ASICs for polynomial multiplication in post-quantum schemes
# Example: GPU-accelerated Paillier in PyTorch
import torch
import tenseal as ts
ctx = ts.context(ts.SCHEME_TYPE.PAILLIER, n_threads=4)
ctx.generate_galois_keys()
ctx.generate_relin_keys()
ctx = ctx.to(device='cuda') # Offload to GPU
Dynamic Participant Scheduling
In real-world deployments, device availability follows power-law distributions. An adaptive scheduling algorithm maximizes throughput:
Where πt(i) is the participation probability for device i at round t, ri is its historical reliability score, and η controls exploration-exploitation tradeoff. This achieves 92% cohort completion rates compared to 67% with random selection.
5. Privacy-preserving Healthcare Analytics
Privacy-preserving Healthcare Analytics
Secure aggregation techniques in healthcare analytics must balance data utility with strict privacy guarantees, particularly when dealing with sensitive patient records. Federated learning (FL) has emerged as a leading paradigm, enabling distributed model training without raw data exchange. However, standard FL frameworks like FedAvg still expose gradient updates to inference attacks, necessitating cryptographic enhancements.
Differential Privacy in Federated Healthcare
Differential privacy (DP) provides mathematically provable guarantees against membership inference attacks. In healthcare FL, each client adds calibrated noise to gradients before aggregation. For a query function f over dataset D, (ε, δ)-DP ensures:
where D and D' differ by one record. The Gaussian mechanism achieves this by sampling noise proportional to the L2-sensitivity Δ2f:
Secure Multi-party Computation (SMPC) Protocols
When DP alone cannot meet regulatory requirements (e.g., HIPAA), SMPC enables cryptographic aggregation. The Shamir's Secret Sharing scheme allows n hospitals to split gradients into t-out-of-n shares. Reconstruction requires at least t parties to collaborate, preventing any single entity from accessing raw data. For additive sharing across k clients:
where rij are pairwise random masks. The aggregated result emerges only after summing all shares:
Hybrid Approaches for Clinical Data
Real-world deployments often combine DP and SMPC. The Prio system exemplifies this by:
- Applying local DP to patient-level features
- Using verifiable SMPC for cross-institution aggregation
- Employing zero-knowledge proofs to detect malicious clients
This hybrid approach demonstrated a 92% AUC in predicting ICU mortality across 23 hospitals while reducing re-identification risk below 0.1% in the iDASH 2019 genomic challenge.
Optimization Challenges
Non-IID data distribution across healthcare providers creates convergence bottlenecks. The FedProx algorithm mitigates this by introducing a proximal term to the local objective:
where μ controls the penalty for deviation from the global model wt. Clinical trials show FedProx reduces communication rounds by 38% compared to FedAvg when training COVID-19 prognosis models.

5.2 Secure Aggregation in Smart Grids
Secure aggregation in smart grids addresses the challenge of preserving privacy while enabling efficient data collection from distributed energy resources (DERs), such as solar panels, wind turbines, and smart meters. Unlike traditional federated learning settings, smart grids impose strict latency constraints, require real-time decision-making, and must comply with regulatory frameworks like NISTIR 7628 and IEC 62351.
Threat Model and Security Requirements
Smart grids face adversarial threats including:
- Data inference attacks – Adversaries may attempt to reconstruct individual consumption patterns from aggregated data.
- False data injection – Malicious actors can manipulate meter readings to destabilize grid operations.
- Collusion attacks – Multiple compromised nodes may collaborate to expose private data.
Secure aggregation must satisfy:
- Differential privacy (DP) – Ensures individual contributions cannot be distinguished.
- Byzantine robustness – Tolerates a fraction of malicious nodes.
- Forward secrecy – Compromised keys do not reveal past aggregated data.
Cryptographic Techniques for Smart Grid Aggregation
Homomorphic encryption (HE) and secure multi-party computation (SMPC) are commonly employed. The Paillier cryptosystem, a partially homomorphic scheme, allows additive aggregation of encrypted meter readings:
where \( E \) denotes encryption, and \( m_i \) are individual measurements. For large-scale deployments, lattice-based schemes like CKKS enable efficient approximate arithmetic on encrypted data.
Distributed Key Generation (DKG)
To prevent single-point failures, threshold cryptosystems distribute key shares among \( n \) nodes, requiring \( t \) participants to decrypt. The Feldman verifiable secret sharing (VSS) protocol ensures correctness:
where \( a_0 \) is the secret, and commitments \( g^{a_i} \) allow verification of shares.
Case Study: Privacy-Preserving Demand Response
In a real-world implementation by Pacific Northwest National Laboratory, secure aggregation enabled privacy-preserving load forecasting. Each smart meter adds Gaussian noise \( \mathcal{N}(0, \sigma^2) \) to satisfy \( (\epsilon, \delta) \)-DP before encryption. The control center decrypts only the aggregated noisy sum:
where \( x_i \) is the true reading and \( \eta_i \sim \mathcal{N}(0, \sigma^2) \). The variance \( \sigma^2 \) is calibrated to the sensitivity \( \Delta \) of the aggregation query:
Performance Optimizations
To meet sub-second latency requirements, hybrid approaches combine:
- Batching – Aggregate readings in fixed time windows.
- Hierarchical aggregation – Local aggregators (e.g., substations) pre-process data before final consolidation.
- Lightweight primitives – Elliptic curve cryptography (ECC) reduces computational overhead versus RSA.
Recent work by Zhang et al. (IEEE TPWRS 2023) demonstrates that lattice-based post-quantum schemes can achieve 1.2 ms encryption latency per smart meter at 128-bit security.

5.3 Cross-silo Federated Learning Use Cases
Cross-silo federated learning (FL) enables multiple organizations to collaboratively train machine learning models without sharing raw data, making it particularly valuable in industries with stringent data privacy requirements. Unlike cross-device FL, where training occurs across numerous edge devices, cross-silo FL involves a smaller number of large, trusted entities—such as hospitals, financial institutions, or research labs—each contributing substantial datasets.
Healthcare: Multi-Institutional Medical Imaging
In healthcare, cross-silo FL allows hospitals to train diagnostic models on distributed patient data while complying with regulations like HIPAA or GDPR. For instance, a federated model for tumor detection in MRI scans can be trained across multiple hospitals, where each institution maintains control over its data. The global model aggregates updates via secure aggregation protocols, ensuring no single participant's data is exposed. A typical workflow involves:
- Local training: Each hospital trains on its dataset using a shared architecture.
- Secure aggregation: Model updates are encrypted using techniques like homomorphic encryption or differential privacy before aggregation.
- Global update: The aggregated model is redistributed to all participants.
Here, ΔWg represents the global model update, N is the number of participants, and Enc(·) denotes an encryption function applied to local updates ΔWi.
Finance: Fraud Detection Across Banks
Banks leverage cross-silo FL to improve fraud detection models without sharing sensitive transaction data. By federating learning across financial institutions, the model benefits from diverse transaction patterns while preserving customer confidentiality. Secure multi-party computation (SMPC) is often employed to compute aggregated gradients without revealing individual contributions. For example:
- Feature alignment: Banks preprocess transaction data to ensure consistent feature spaces.
- Partial gradient masking: Each bank computes gradients masked with secret shares before transmission.
- Secure aggregation: A central coordinator reconstructs the global gradient only after receiving all masked contributions.
Pharmaceutical Research: Drug Discovery Collaborations
Pharmaceutical companies use cross-silo FL to accelerate drug discovery by pooling molecular activity data while protecting intellectual property. Federated learning frameworks like FATE or OpenFL enable secure collaboration by:
- Vertical partitioning: Participants contribute different features (e.g., one provides chemical structures, another provides biological assay results).
- Homomorphic encryption: Ensures computations on encrypted data yield valid encrypted results.
- Federated averaging: Model parameters are averaged after each round of local training.
Where θt denotes the model parameters at step t, η is the learning rate, ni is the sample size of the i-th silo, and gi is the gradient computed locally.
Challenges and Mitigations
Cross-silo FL introduces unique challenges, including:
- Non-IID data: Data distributions vary across silos, leading to biased models. Techniques like federated meta-learning or weighted aggregation can mitigate this.
- Communication overhead: Large model updates between silos require efficient compression (e.g., gradient quantization).
- Trust boundaries: Even among trusted entities, secure aggregation must prevent collusion attacks.
Advanced protocols like HybridAlpha combine SMPC and differential privacy to balance privacy and utility, while blockchain-based FL frameworks provide auditability for regulatory compliance.

6. Foundational Research Papers
6.1 Foundational Research Papers
- Practical Secure Aggregation by Combining Cryptography and Trusted ... — Several research work have explored secure aggregation protocols based on purely cryptographic mechanisms [7, 36]. Others have examined outsourcing computations to TEEs to alleviate the computational burden of purely cryptographic approaches [18, 51]. Our work conducts a thorough investiga-tion of previously unexplored secure aggregation ...
- Two Secure Privacy-Preserving Data Aggregation Schemes for IoT — In this paper, two secure and efficient data aggregation schemes are proposed for IoT devices. Both of them support "plug and play" and preserve IoT devices' private data by blending their data before reported. ... This research was supported partially by the Fundamental Research Funds for the Central Universities (No. 2019CDQYRJ006 ...
- A Flexible and Scalable Malicious Secure Aggregation Protocol for ... — Secure aggregation becomes a major solution to providing privacy for federated learning. Secure aggregation for mobile devices typically relies on Shamir secret sharing (SSS) to achieve dropout robustness, but limits the system's corruption and dropout tolerance. Although Prio+, a state-of-the-art method utilizing two non-colluding servers, avoids such limitations, its effectiveness is only ...
- A Data Attack Detection Framework for Cryptography-Based Secure ... - MDPI — To solve these issues, this paper proposes a data attack detection framework based on a cryptographic secure aggregation method, which aims to prevent data-aggregation errors, data poisoning, and illegal data sources through encrypted-data auditing techniques, thereby enhancing the security and correctness of the cryptographic secure ...
- An Efficient Multi-Party Secure Aggregation Method Based on Multi ... — The federated learning on large-scale mobile terminals and Internet of Things (IoT) devices faces the issues of privacy leakage, resource limitation, and frequent user dropouts. This paper proposes an efficient secure aggregation method based on multi-homomorphic attributes to realize the privacy-preserving aggregation of local models while ensuring low overhead and tolerating user dropouts.
- Secure and efficient multi-key aggregation for federated learning — This paper introduces a secure multi-key aggregation protocol called MKAgg, that utilizes homomorphic encryption. MKAgg enables clients to leave the process at any point, and it achieves aggregation on ciphertexts by transforming encrypted data using different public keys into ciphertexts under the same public key.
- An effective and verifiable secure aggregation scheme with privacy ... — Federated learning has gained significant attention for enabling collaborative model training on distributed devices while maintaining data privacy. However, sharing gradients poses risks to local data privacy. This paper presents a secure aggregation scheme that addresses privacy protection and verifiability in federated learning.
- Secure Data Aggregation Based on End-to-End Homomorphic Encryption in ... — An alternative solution to secure the aggregation process is to provide an end-to-end security protocol, wherein intermediary nodes combine the data without decoding the acquired data. As a consequence, the intermediary aggregating nodes do not have to maintain confidential key values, enabling end-to-end security across sensor devices and base ...
6.2 Open-source Implementations and Libraries
- ACCESS-FL: Agile Communication and Computation for Efficient Secure ... — To address these privacy threats, Google proposed the Secure Aggregation (SecAgg) protocol (bonawitz2017practical, ) as a secure multi-party computation (MPC) (cramer2015secure, ) method based on Diffie-Hellman (DH) key agreement (diffie1976new, ) and Shamir's secret sharing (dawson1994breadth, ; pang2005new, ).In SecAgg, each client generates a shared secret for every other client ...
- Segam: Secure and Efficient Group-by-Aggregation Queries across ... — SMCQL is the only interactive-based open-source system supporting secure group-by-aggregation queries. Thus, we used the reported performance of other works for comparison. 6.2 Overall Evaluation. We evaluate query time and communication consumption in Segam using the lineitem table from the TPC-H benchmark. It is essential to note that the ...
- Practical Secure Aggregation by Combining Cryptography and Trusted ... — Secure aggregation enables a group of mutually distrustful parties, each holding private inputs, to collaboratively compute an aggregate value while preserving the privacy of their individual inputs. ... Secure aggregation variant implementations. ... Open-Source Fully Homomorphic Encryption Library. In Proceedings of the 10th Workshop on ...
- An effective and verifiable secure aggregation scheme with privacy ... — Secure multiparty computation employs techniques such as garbled circuits and secret sharing to achieve secure computation without disclosing private input information. For instance, Bonawitz et al. proposed a classic secure aggregation protocol based on the principles of secure multiparty computation PPML [15]. They adopted a dual masking ...
- Privacy-enhancing Aggregation Techniques For Smart Grid ... - Library — Offering a comprehensive exploration of various privacy preserving data aggregation techniques, this book is an exceptional resource for the academics, researchers, and graduate students seeking to exploit secure data aggregation techniques in smart grid communications and Internet of Things (IoT) scenarios.
- Fast Authentication from Aggregate Signatures with Improved Security — We ran the open-source implementations of our state-of-the-art counterparts in our experimental setup, to draw a fair comparison. We benchmarked the ECDSA in MIRACL library and RSA in GMP library . We benchmarked Ed25519 and SPHINCS using their Supercop implementations. Lastly, we used the open-source implementation of pqNTRUsign .
- Security and Communication Networks - Wiley Online Library — Spatial ciphertext aggregation computing scheme and system architecture are discussed in Section 3. The secure computing protocol and a verifiable aggregating signature algorithm are proposed in Sections 4 and 5. Security analysis and the simulation results are shown in Section 6. The conclusion is drawn in Section 7. 2. Related Work
- EPPDA: An Efficient and Privacy‐Preserving Data Aggregation Scheme with ... — In , Othman et al. present an end-to-end secure data aggregation scheme, namely, Robust and Efficient Secure Data Aggregation Scheme in Healthcare Using IoT (RESDA). The main objective of the proposed scheme is the security of the data aggregation to be achieved without introducing significant overheads on the sensors limited by the battery.
- Secure multiparty learning from the aggregation of locally trained ... — In this paper, we propose a new framework for secure multi-party learning and construct a concrete scheme by incorporating aggregate signature and proxy re-encryption techniques. Unlike the previous solutions for multi-party privacy-preserving machine learning, we don't use encryption algorithm to encrypt the whole dataset or the intermediate ...
- PDF Practical Secure Aggregation for Privacy-Preserving Machine Learning - IACR — withdistinctfieldelementsinF.Giventheseparameters,the scheme consists of two algorithms. The sharing algorithm SS.share(s,t,U) →{(u,s u)} u∈U takes as input a secret s, a set Uof nfield elements representing user IDs, and
6.3 Advanced Topics and Ongoing Research
- Practical Secure Aggregation by Combining Cryptography and Trusted ... — secure aggregation. (2) An in-depth evaluation of said combinations with varying numbers of parties and input data sizes. (3) A detailed analysis of the performance and security trade-offs across the different approaches. 2Background 2.1Secure aggregation Secure aggregation (SA) is a type of secure multi-party com-
- An effective and verifiable secure aggregation scheme with privacy ... — So, many of these techniques require further optimization to improve their efficiency and practicality for real-world applications. To achieve an efficient lightweight secure aggregation scheme and verifiability, we propose a verifiable and seamless secure aggregation method for privacy-preserving in FL.
- SAFED: secure and adaptive framework for edge-based data aggregation in ... — Section 2 reviews the current research on privacy-preserving data aggregation schemes in IoT and ... many are too complex to implement easily. Therefore, there is a need for privacy-preserving data aggregation techniques that are both secure and efficient, specifically designed to meet the constraints of IoT devices with limited resources ...
- SAMFL: Secure Aggregation Mechanism for Federated Learning with ... — AggDec: Finally, this aggregation decryption step utilizes the m s k, the G i, and the s k to perform secure gradient aggregation. The process involves calculating C ≔ (∏ i ∈ [n] g w i ⊤ c i d i t i) / g z, where w, (d 1, d 2, …, d n), and z are components of s k. Subsequently, the secure gradient aggregation result is obtained by ...
- Contract-based hierarchical security aggregation scheme for enhancing ... — Here are some of the trending research topics: Abadi et al. [16] ... proposed a double-masking scheme that uses Shamir secret sharing and a series of cryptographic techniques to achieve secure aggregation against semi-honest and malicious attackers, as well as embedding protection against user loss through threshold Shamir secret sharing. The ...
- An Efficient Multi-Party Secure Aggregation Method Based on Multi ... — The federated learning on large-scale mobile terminals and Internet of Things (IoT) devices faces the issues of privacy leakage, resource limitation, and frequent user dropouts. This paper proposes an efficient secure aggregation method based on multi-homomorphic attributes to realize the privacy-preserving aggregation of local models while ensuring low overhead and tolerating user dropouts.
- Multiple-Round Aggregation of Abstract Semantics for Secure ... — Current secure aggregation (SA) methods [7,8,9] struggle with efficiency, particularly in environments that necessitate multiple rounds of aggregation. The inefficiencies are primarily due to the overhead associated with managing extensive communication rounds and the computational burden of handling complex data structures.
- Secure Data Aggregation Based on End-to-End Homomorphic Encryption in ... — With cloud-assisted healthcare WSNs, a secure and reliable certificateless publicly monitored approach that provides adaptive data exchange and confidentiality control, as well as effective grouped user revoking, was introduced in . Table 1 presents a comparative review of the numerous methodologies suggested for secure data aggregation. The ...
- Secure Data Aggregation Based on End-to-End Homomorphic Encryption in ... — Thus, according to Table 1, there have been diverse secure data aggregation methods that resolve the challenges of authentication, confidentiality, integrity, and threats such as sensor node compromised attacks, misleading data injectors, replay attacks, spoofing, defined simple text threats, encrypted assessment, and unauthorized access ...
- Secure Data Aggregation Aided by Privacy Preserving in Internet of ... — 1. Introduction. The Internet of Things (IoT) is an intelligent system which can realize the sensing, the communication, and the decision-making through its underlying technology [].With the emerging of fifth generation (5G) communication, the reasonable implement of distributed devices has become an important research topic in IoT [2, 3].The rapid development of sensor techniques, including ...








