Skip to content

About

No description, website, or topics provided.

Resources

Stars

82 stars

Watchers

0 watching

Forks

Repository files navigation

HFAN-Priv: Federated Learning for Privacy-Preserving Employee Performance Analytics

Reference implementation for the IEEE Access paper:

Federated Learning for Privacy-Preserving Employee Performance Analytics Jay Barach (Senior Member, IEEE). IEEE Access, vol. 13, 2025, pp. 132724–132738. DOI: 10.1109/ACCESS.2025.3591360

HFAN-Priv is a Hierarchical Federated Attention Network with Privacy-aware optimization and Interpretability. It predicts employee resignation risk across decentralized organizational data silos without sharing raw data, combining:

  • Dual-attention modeling — feature-level and instance-level attention (HFAN).
  • Differential privacy — Privacy-Aware Gradient Masking (gradient clipping + Gaussian mechanism) for (ε, δ)-DP.
  • Local explainability — SHAP + LIME computed on-device.
  • Adaptive federated aggregation — validation-loss / update-variance weighting.

License. Code is released under CC BY 4.0. © 2025 The Authors. If you use it, please cite the IEEE Access paper above and the Kaggle dataset.


What's implemented

Every formal component of the paper:

  • HFAN (hfan_priv.model) — feature encoder (Eq. 2), feature attention (Eqs. 3–5), instance attention (Eqs. 6–8), sigmoid classifier (Eq. 9).
  • PAGM / DP (hfan_priv.privacy) — L2 gradient clipping (Eq. 13), Gaussian mechanism (Eq. 10), analytic noise scale (Eq. 19), and an ε-composition tracker.
  • Federated loop (hfan_priv.federated) — the full Algorithm 1: local training, DP-masked updates, and adaptive aggregation w_i = 1/(L_i + γσ_i) (Eqs. 14–15).
  • Explainability (hfan_priv.explain) — SHAP + LIME with a top-k agreement metric (Table 8).
  • Metrics (hfan_priv.metrics) — accuracy, precision, recall, F1, AUC, MCC, ECE (Tables 3, 4, 6), with per-client and unseen-client (zero-shot) evaluation.

Reproduction note (please read)

The paper reports near-perfect results (~99% accuracy/F1/AUC) on the author's evaluation of the employee dataset. This repository is built to be fully runnable and honest:

  • It ships a synthetic dataset generator matching Table 2's schema, so the whole pipeline runs offline with no download.
  • On the synthetic data, the code reproduces the paper's methodology and trends faithfully: the dual-attention model trains and generalizes across federated clients, differential privacy is correctly applied, and the privacy/utility trade-off (more noise → lower utility) is visible as in Table 6.
  • Absolute metrics on the synthetic set are lower than the paper's (the synthetic resignation signal is deliberately not tuned to inflate results). Running on the real Kaggle dataset (scripts/download_data.py) with the same code exercises the identical pipeline on the paper's data.

A note on the DP noise scale. The analytic Gaussian bound (Eq. 19) gives σ ≈ 4.85 for ε=1, δ=1e-5, C=1. For a compact attention network that absolute scale overwhelms the (small) aggregated update, so the default uses a practical noise multiplier (0.1) applied to the server-aggregated update, and configs/privacy_sweep.yaml sweeps it to reproduce the Table 6 trade-off trend. compute_noise_multiplier() still exposes the exact Eq. 19 value.


Requirements & installation

  • Python 3.9+, PyTorch 2.0+ (CPU is fine)
git clone https://github.com/<your-username>/hfan-priv.git
cd hfan-priv
python -m venv .venv && source .venv/bin/activate
pip install --upgrade pip
pip install -e .              # core
pip install -e ".[explain]"  # SHAP + LIME (else a permutation fallback is used)
pip install -e ".[dev]"      # pytest, ruff

Console command: hfan-priv-run.


Dataset

The paper uses the Kaggle Employee Performance and Productivity Analysis dataset (100k records, 20 attributes — Table 2):

pip install kaggle          # put your token at ~/.kaggle/kaggle.json
python scripts/download_data.py --source kaggle --out data

Or place the CSV at data/employee_performance.csv manually. If no CSV is present, the pipeline auto-generates a synthetic dataset with the same schema — nothing else is required to run.


Quickstart

bash scripts/run_all.sh          # quick offline run (synthetic data, 1 seed)
bash scripts/run_all.sh full     # full protocol (50 rounds, 3 seeds)

Or:

python -m hfan_priv.run --config configs/default.yaml

Results (per-client and unseen-client metrics, privacy accounting, SHAP/LIME agreement) are written to outputs/<name>/summary.json, averaged over seeds.


Repository layout

hfan-priv/
├── src/hfan_priv/
│   ├── config.py            # typed configs (paper hyperparameters as defaults)
│   ├── data.py              # Kaggle loading, preprocessing, client partitioning, synthetic gen
│   ├── metrics.py           # accuracy, F1, AUC, MCC, ECE + aggregation
│   ├── run.py               # experiment runner (federated + CV + baselines + explainability)
│   ├── model/hfan.py        # HFAN dual-attention network
│   ├── privacy/pagm.py      # gradient clipping + Gaussian mechanism + ε accounting
│   ├── federated/           # client (local DP training) + server (adaptive aggregation)
│   └── explain/             # SHAP + LIME + agreement
├── configs/                 # default, smoke, privacy_sweep
├── scripts/                 # download_data, make_synthetic_data, run_all
├── tests/                   # unit + end-to-end tests
├── .github/workflows/       # CI
├── LICENSE  NOTICE  CITATION.cff  pyproject.toml  requirements.txt

Paper → code mapping

Paper element Where
Global objective (Eq. 1) federated.server aggregation
Feature encoder (Eq. 2) model.hfan.FeatureEncoder
Feature attention (Eqs. 3–5) model.hfan.FeatureAttention
Instance attention (Eqs. 6–8) model.hfan.InstanceAttention
Sigmoid classifier (Eq. 9) model.hfan.HFAN.head
Gaussian mechanism (Eq. 10) privacy.pagm.GaussianMechanism
(ε, δ)-DP guarantee (Eq. 11) privacy.pagm (accounting)
Noise scale (Eqs. 12, 19) privacy.pagm.compute_noise_multiplier
Gradient clipping (Eq. 13) privacy.pagm.clip_gradients
Global update (Eq. 14) federated.server._aggregate
Adaptive weights (Eq. 15) federated.server._aggregate
SHAP (Eqs. 16–17) explain.explainers.explain_shap
LIME (Eq. 18) explain.explainers.explain_lime
Federated algorithm (Algorithm 1) federated.server.fit + federated.client

Testing

pip install -e ".[dev]"
pytest -q                 # includes an end-to-end federated learning test
pytest -q -m "not slow"   # fast unit tests only

CI runs the unit tests plus an end-to-end smoke run on every push.


Citation

Paper (IEEE Access 2025):

@article{barach2025hfanpriv,
  author  = {Barach, Jay},
  title   = {Federated Learning for Privacy-Preserving Employee Performance Analytics},
  journal = {IEEE Access},
  year    = {2025},
  volume  = {13},
  pages   = {132724--132738},
  doi     = {10.1109/ACCESS.2025.3591360}
}

Software:

@software{barach2025hfanpriv_code,
  author  = {Barach, Jay},
  title   = {{HFAN-Priv}: Federated Learning for Privacy-Preserving Employee Performance Analytics},
  year    = {2025},
  version = {1.0.0},
  url     = {https://github.com/<your-username>/hfan-priv},
  note    = {Reference implementation for the IEEE Access paper, DOI: 10.1109/ACCESS.2025.3591360}
}

License

Code: CC BY 4.0 — see LICENSE. © 2025 The Authors.

About

No description, website, or topics provided.

Resources

Stars

82 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages