TY - GEN
T1 - Packing Induced Bias in Deep Learning Malware Classifiers
T2 - 5th IEEE International Conference on AI in Cybersecurity, ICAIC 2026
AU - Kalapala, Jeevana Swaroop
AU - Zhang, Lan
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Packing is a highly used malware evasion technique that compresses, encrypts, or obfuscates executable content, significantly altering the structural characteristics that static malware detectors rely on. While prior work has studied the impact of packing on traditional feature based classifiers, limited attention has been given to understanding how modern deep learning based static detectors behave under realistic packing conditions. This paper investigates the robustness and generalization capabilities of a CNN based malware detector trained on image representations of Windows Portable Executable (PE) binaries by conducting two controlled experiments. The first evaluates how increasing exposure to packed benign ware affects a model's ability to correctly distinguish packed malware from packed benign binaries. Results show initial improvement in true negative rates, followed by instability as packing artifacts dominate learned representations. The second experiment analyzes cross packer generalization and demonstrates strong packer dependency, where models perform well only on packers seen during training and collapse when confronted with unseen packers. Overall, our findings demonstrate that packing significantly undermines the reliability of static deep learning based malware detectors, highlighting the need for packing aware training strategies and more resilient detection models.
AB - Packing is a highly used malware evasion technique that compresses, encrypts, or obfuscates executable content, significantly altering the structural characteristics that static malware detectors rely on. While prior work has studied the impact of packing on traditional feature based classifiers, limited attention has been given to understanding how modern deep learning based static detectors behave under realistic packing conditions. This paper investigates the robustness and generalization capabilities of a CNN based malware detector trained on image representations of Windows Portable Executable (PE) binaries by conducting two controlled experiments. The first evaluates how increasing exposure to packed benign ware affects a model's ability to correctly distinguish packed malware from packed benign binaries. Results show initial improvement in true negative rates, followed by instability as packing artifacts dominate learned representations. The second experiment analyzes cross packer generalization and demonstrates strong packer dependency, where models perform well only on packers seen during training and collapse when confronted with unseen packers. Overall, our findings demonstrate that packing significantly undermines the reliability of static deep learning based malware detectors, highlighting the need for packing aware training strategies and more resilient detection models.
KW - Deep Learning
KW - Malware Detection
KW - Software Packing
UR - https://www.scopus.com/pages/publications/105041764226
UR - https://www.scopus.com/pages/publications/105041764226#tab=citedBy
U2 - 10.1109/ICAIC67076.2026.11395794
DO - 10.1109/ICAIC67076.2026.11395794
M3 - Conference contribution
AN - SCOPUS:105041764226
T3 - 2026 IEEE 5th International Conference on AI in Cybersecurity, ICAIC 2026
BT - 2026 IEEE 5th International Conference on AI in Cybersecurity, ICAIC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 February 2026 through 20 February 2026
ER -