Reference implementation for the IEEE Access paper:
Federated Learning for Privacy-Preserving Employee Performance Analytics Jay Barach (Senior Member, IEEE). IEEE Access, vol. 13, 2025, pp. 132724–132738. DOI: 10.1109/ACCESS.2025.3591360
HFAN-Priv is a Hierarchical Federated Attention Network with Privacy-aware optimization and Interpretability. It predicts employee resignation risk across decentralized organizational data silos without sharing raw data, combining:
- Dual-attention modeling — feature-level and instance-level attention (HFAN).
- Differential privacy — Privacy-Aware Gradient Masking (gradient clipping + Gaussian mechanism) for (ε, δ)-DP.
- Local explainability — SHAP + LIME computed on-device.
- Adaptive federated aggregation — validation-loss / update-variance weighting.
License. Code is released under CC BY 4.0. © 2025 The Authors. If you use it, please cite the IEEE Access paper above and the Kaggle dataset.
Every formal component of the paper:
- HFAN (
hfan_priv.model) — feature encoder (Eq. 2), feature attention (Eqs. 3–5), instance attention (Eqs. 6–8), sigmoid classifier (Eq. 9). - PAGM / DP (
hfan_priv.privacy) — L2 gradient clipping (Eq. 13), Gaussian mechanism (Eq. 10), analytic noise scale (Eq. 19), and an ε-composition tracker. - Federated loop (
hfan_priv.federated) — the full Algorithm 1: local training, DP-masked updates, and adaptive aggregationw_i = 1/(L_i + γσ_i)(Eqs. 14–15). - Explainability (
hfan_priv.explain) — SHAP + LIME with a top-k agreement metric (Table 8). - Metrics (
hfan_priv.metrics) — accuracy, precision, recall, F1, AUC, MCC, ECE (Tables 3, 4, 6), with per-client and unseen-client (zero-shot) evaluation.
The paper reports near-perfect results (~99% accuracy/F1/AUC) on the author's evaluation of the employee dataset. This repository is built to be fully runnable and honest:
- It ships a synthetic dataset generator matching Table 2's schema, so the whole pipeline runs offline with no download.
- On the synthetic data, the code reproduces the paper's methodology and trends faithfully: the dual-attention model trains and generalizes across federated clients, differential privacy is correctly applied, and the privacy/utility trade-off (more noise → lower utility) is visible as in Table 6.
- Absolute metrics on the synthetic set are lower than the paper's (the synthetic resignation signal is deliberately not tuned to inflate results). Running on the real Kaggle dataset (
scripts/download_data.py) with the same code exercises the identical pipeline on the paper's data.
A note on the DP noise scale. The analytic Gaussian bound (Eq. 19) gives σ ≈ 4.85 for ε=1, δ=1e-5, C=1. For a compact attention network that absolute scale overwhelms the (small) aggregated update, so the default uses a practical noise multiplier (0.1) applied to the server-aggregated update, and configs/privacy_sweep.yaml sweeps it to reproduce the Table 6 trade-off trend. compute_noise_multiplier() still exposes the exact Eq. 19 value.
- Python 3.9+, PyTorch 2.0+ (CPU is fine)
git clone https://github.com/<your-username>/hfan-priv.git
cd hfan-priv
python -m venv .venv && source .venv/bin/activate
pip install --upgrade pip
pip install -e . # core
pip install -e ".[explain]" # SHAP + LIME (else a permutation fallback is used)
pip install -e ".[dev]" # pytest, ruffConsole command: hfan-priv-run.
The paper uses the Kaggle Employee Performance and Productivity Analysis dataset (100k records, 20 attributes — Table 2):
pip install kaggle # put your token at ~/.kaggle/kaggle.json
python scripts/download_data.py --source kaggle --out dataOr place the CSV at data/employee_performance.csv manually. If no CSV is present, the pipeline auto-generates a synthetic dataset with the same schema — nothing else is required to run.
bash scripts/run_all.sh # quick offline run (synthetic data, 1 seed)
bash scripts/run_all.sh full # full protocol (50 rounds, 3 seeds)Or:
python -m hfan_priv.run --config configs/default.yamlResults (per-client and unseen-client metrics, privacy accounting, SHAP/LIME agreement) are written to outputs/<name>/summary.json, averaged over seeds.
hfan-priv/
├── src/hfan_priv/
│ ├── config.py # typed configs (paper hyperparameters as defaults)
│ ├── data.py # Kaggle loading, preprocessing, client partitioning, synthetic gen
│ ├── metrics.py # accuracy, F1, AUC, MCC, ECE + aggregation
│ ├── run.py # experiment runner (federated + CV + baselines + explainability)
│ ├── model/hfan.py # HFAN dual-attention network
│ ├── privacy/pagm.py # gradient clipping + Gaussian mechanism + ε accounting
│ ├── federated/ # client (local DP training) + server (adaptive aggregation)
│ └── explain/ # SHAP + LIME + agreement
├── configs/ # default, smoke, privacy_sweep
├── scripts/ # download_data, make_synthetic_data, run_all
├── tests/ # unit + end-to-end tests
├── .github/workflows/ # CI
├── LICENSE NOTICE CITATION.cff pyproject.toml requirements.txt
| Paper element | Where |
|---|---|
| Global objective (Eq. 1) | federated.server aggregation |
| Feature encoder (Eq. 2) | model.hfan.FeatureEncoder |
| Feature attention (Eqs. 3–5) | model.hfan.FeatureAttention |
| Instance attention (Eqs. 6–8) | model.hfan.InstanceAttention |
| Sigmoid classifier (Eq. 9) | model.hfan.HFAN.head |
| Gaussian mechanism (Eq. 10) | privacy.pagm.GaussianMechanism |
| (ε, δ)-DP guarantee (Eq. 11) | privacy.pagm (accounting) |
| Noise scale (Eqs. 12, 19) | privacy.pagm.compute_noise_multiplier |
| Gradient clipping (Eq. 13) | privacy.pagm.clip_gradients |
| Global update (Eq. 14) | federated.server._aggregate |
| Adaptive weights (Eq. 15) | federated.server._aggregate |
| SHAP (Eqs. 16–17) | explain.explainers.explain_shap |
| LIME (Eq. 18) | explain.explainers.explain_lime |
| Federated algorithm (Algorithm 1) | federated.server.fit + federated.client |
pip install -e ".[dev]"
pytest -q # includes an end-to-end federated learning test
pytest -q -m "not slow" # fast unit tests onlyCI runs the unit tests plus an end-to-end smoke run on every push.
Paper (IEEE Access 2025):
@article{barach2025hfanpriv,
author = {Barach, Jay},
title = {Federated Learning for Privacy-Preserving Employee Performance Analytics},
journal = {IEEE Access},
year = {2025},
volume = {13},
pages = {132724--132738},
doi = {10.1109/ACCESS.2025.3591360}
}Software:
@software{barach2025hfanpriv_code,
author = {Barach, Jay},
title = {{HFAN-Priv}: Federated Learning for Privacy-Preserving Employee Performance Analytics},
year = {2025},
version = {1.0.0},
url = {https://github.com/<your-username>/hfan-priv},
note = {Reference implementation for the IEEE Access paper, DOI: 10.1109/ACCESS.2025.3591360}
}Code: CC BY 4.0 — see LICENSE. © 2025 The Authors.