Generative Data Augmentation for Low-Resource Smishing Detection: A Controlled Evaluation of Robustness and Calibration

Authors

  • Mariam Nadeem Department of Cybersecurity and Artificial Intelligence, Alexandria University Author
  • Mira Noori Department of Cybersecurity and Artificial Intelligence, Sorbonne University Author

Abstract

Generative augmentation may help low-resource smishing detection, but artificial examples can introduce unrealistic language patterns and distort model calibration. We compared three augmentation strategies—paraphrase generation, class-conditional message generation, and targeted code-mixed generation—using a Bangla smishing dataset with controlled training-set sizes. Augmented models were evaluated on untouched temporal and external test sets. At the smallest training size, augmentation improved macro-F1 by up to 0.08, with the greatest gain from targeted code-mixed examples. Benefits narrowed as the amount of real training data increased. Unfiltered class-conditional generation produced lexical artefacts that improved in-sample performance but reduced external precision. A quality-filtering step based on duplication, language balance, and classifier uncertainty removed 17% of generated messages and improved calibration error from 0.091 to 0.052. Generative augmentation can strengthen low-resource detection when it targets underrepresented linguistic patterns and is evaluated on real, temporally separated data. Quantity alone was not beneficial; filtering and external validation were necessary to prevent artificial regularities from becoming shortcuts.

References

1. Azam MA, et al. TransFuse: an interpretable transformer ensemble framework for Bangla smishing detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546480/

2. Khan MFI, Begum MH, Rahman MA, Limon GQ, Azam MA, Masum AKM. A comprehensive review of advances in transformer, GAN, and attention mechanisms: their role in multimodal learning and applications across NLP. International Journal of Science and Research Archive. 2025;15(1):454-459. doi:10.30574/ijsra.2025.15.1.0980.

3. Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative adversarial nets. Adv Neural Inf Process Syst. 2014;27. Available from: https://arxiv.org/abs/1406.2661

4. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1706.03762

5. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. 2019. p. 4171-4186. doi:10.18653/v1/N19-1423.

6. Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, et al. Unsupervised cross-lingual representation learning at scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 8440-8451. doi:10.18653/v1/2020.acl-main.747.

7. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1705.07874

8. Ribeiro MT, Singh S, Guestrin C. "Why should I trust you?": explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 1135-1144. doi:10.1145/2939672.2939778.

9. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

10. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.

11. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.

12. Ovadia Y, Fertig E, Ren J, Nado Z, Sculley D, Nowozin S, et al. Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift. Adv Neural Inf Process Syst. 2019;32. Available from: https://arxiv.org/abs/1906.02530

13. Jain S, Wallace BC. Attention is not explanation. In: Proceedings of NAACL-HLT. 2019. p. 3543-3556. doi:10.18653/v1/N19-1357.

14. Jacovi A, Goldberg Y. Towards faithfully interpretable NLP systems: how should we define and evaluate faithfulness?. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 4198-4205. doi:10.18653/v1/2020.acl-main.386.

15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

Published

2026-06-01