Skip to main navigation Skip to search Skip to main content

Synthetic Data-Driven Explainability for Federated Learning-Based Intrusion Detection System

Research output: Contribution to journalArticlepeer-review

Abstract

An intrusion detection system (IDS) is vital for monitoring network traffic and alerting users to threats. Unlike traditional IDS, which relies on centralized data processing and raises privacy concerns, federated learning (FL)-based IDSs enable collaborative model training among multiple clients while keeping user data private. However, explaining model behavior in FL using explainable artificial intelligence (XAI) is challenging due to its distributed nature and lack of access to client data. Traditional XAI methods such as local interpretable model-agnostic explanations (LIME) and SHapley Additive Explanations (SHAP) require input data, which conflicts with FL’s privacy constraints. In this work, we develop a deep neural network (DNN)-based IDS in an FL setup in nonindependent and identically distributed (non-i.i.d.) settings. Our FL-DNN model achieves high performance in binary classification for detecting malicious network traffic. In this work, we propose a novel privacy preserving, explainable FL framework that uses high-quality synthetic data to enable explainability of the global DNN model without exposing client data to the server. To generate synthetic data, we train multiple federated generative models in non-i.i.d. settings. Among them, the Federated learning-based Wasserstein conditional Generative adversarial network with gradient penalty (FL-WCGAN-GP) produces synthetic samples with high data quality at the server. These synthetic samples on the server side are then used as reference inputs for post hoc XAI methods for explaining the global DNN model. We assess the sufficiency of synthetic data-based explanations for the global DNN model using SHAP showing that synthetic data-based explanations closely approximate the explanations derived from real client data. Furthermore, we quantitatively evaluate postlocal explanations of LIME and SHAP based on faithfulness and robustness. Results show that SHAP provides more faithful and robust explanations than LIME for client-side models using real data and server-side models using synthetic data, supporting privacy-preserving explainability in trustworthy FL-based IDS.

Original languageEnglish (US)
Pages (from-to)5918-5944
Number of pages27
JournalIEEE Internet of Things Journal
Volume13
Issue number4
DOIs
StatePublished - 2026

Keywords

  • Explainable artificial intelligence (XAI)
  • federated learning
  • generative adversarial networks (GANs)
  • intrusion detection system (IDS)
  • synthetic data
  • trustworthy artificial intelligence (AI)

ASJC Scopus subject areas

  • Signal Processing
  • Information Systems
  • Hardware and Architecture
  • Computer Science Applications
  • Computer Networks and Communications

Fingerprint

Dive into the research topics of 'Synthetic Data-Driven Explainability for Federated Learning-Based Intrusion Detection System'. Together they form a unique fingerprint.

Cite this