AlignAD-VAE: A Variational Autoencoder with MMD-Based Dataset Alignment for Network Anomaly Detection



Saka, Samed, Selis, Valerio ORCID: 0000-0002-1856-4707 and Marshall, Alan ORCID: 0000-0002-8058-5242
(2025) AlignAD-VAE: A Variational Autoencoder with MMD-Based Dataset Alignment for Network Anomaly Detection In: 2025 IEEE 24th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), 2025-11-14 - 2025-11-17.

[thumbnail of 2025 - AlignAD-VAE - A Variational Autoencoder with MMD-Based Dataset Alignment for Network Anomaly Detection.pdf] Text
2025 - AlignAD-VAE - A Variational Autoencoder with MMD-Based Dataset Alignment for Network Anomaly Detection.pdf - Author Accepted Manuscript
Available under License Creative Commons Attribution.

Download (592kB) | Preview

Abstract

This study addresses the persistent challenge of cross-dataset generalisability in intrusion detection systems by both assessing whether concatenating datasets improves generalisability and proposing AlignAD-VAE, a new unsupervised variational autoencoder model augmented with maximum mean discrepancy (MMD)-based alignment. The model aims to reduce the distribution shift between datasets by aligning their latent representations in a common feature space. We systematically evaluate AlignAD-VAE against modern architectures such as autoencoder and variational autoencoder baselines across multiple cross-dataset configurations using the CIC-IDS2017, CSE-CIC-IDS2018, and CIC-DDoS2019 datasets. Our experiments cover both single-dataset training and concatenated multidataset training, assessing model performance on completely unseen datasets. Concatenating training datasets improves generalisability by up to 10%, as it exposes models to a broader range of normal patterns and traffic variations, thereby reducing overfitting to dataset-specific artefacts. While all models benefit from the richer training data, AlignAD-VAE outperforms the VAE baseline by up to 2%, indicating that the integration of MMD-based domain alignment provides additional, although modest, improvements in cross-domain adaptation, as reflected in AUC-ROC, F1-score, and accuracy metrics. These findings highlight that combining diverse datasets with domain alignment can make IDS more robust to unseen network environments, a critical requirement for real-world deployment.

Item Type: Conference Item (Unspecified)
Uncontrolled Keywords: Anomaly Detection, Cross-Dataset Generalisability, Domain Adaptation, Network Security
Divisions: Faculty of Science & Engineering
Faculty of Science & Engineering > School of Computer Science & Informatics
Faculty of Science & Engineering > School of Computer Science & Informatics > Trustworthy Computing
Depositing User: Symplectic Admin
Date Deposited: 09 Feb 2026 09:42
Last Modified: 23 May 2026 11:15
DOI: 10.1109/Trustcom66490.2025.00100
Related Websites:
URI: https://livrepository.liverpool.ac.uk/id/eprint/3196920
Disclaimer: The University of Liverpool is not responsible for content contained on other websites from links within repository metadata. Please contact us if you notice anything that appears incorrect or inappropriate.