Generative Data Augmentation for Low-Resource Smishing Detection: A Controlled Evaluation of Robustness and Calibration
Abstract
Generative augmentation may help low-resource smishing detection, but artificial examples can introduce unrealistic language patterns and distort model calibration. We compared three augmentation strategies—paraphrase generation, class-conditional message generation, and targeted code-mixed generation—using a Bangla smishing dataset with controlled training-set sizes. Augmented models were evaluated on untouched temporal and external test sets. At the smallest training size, augmentation improved macro-F1 by up to 0.08, with the greatest gain from targeted code-mixed examples. Benefits narrowed as the amount of real training data increased. Unfiltered class-conditional generation produced lexical artefacts that improved in-sample performance but reduced external precision. A quality-filtering step based on duplication, language balance, and classifier uncertainty removed 17% of generated messages and improved calibration error from 0.091 to 0.052. Generative augmentation can strengthen low-resource detection when it targets underrepresented linguistic patterns and is evaluated on real, temporally separated data. Quantity alone was not beneficial; filtering and external validation were necessary to prevent artificial regularities from becoming shortcuts.
References
1. Azam MA, et al. TransFuse: an interpretable transformer ensemble framework for Bangla smishing detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546480/
2. Khan MFI, Begum MH, Rahman MA, Limon GQ, Azam MA, Masum AKM. A comprehensive review of advances in transformer, GAN, and attention mechanisms: their role in multimodal learning and applications across NLP. International Journal of Science and Research Archive. 2025;15(1):454-459. doi:10.30574/ijsra.2025.15.1.0980.
3. Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative adversarial nets. Adv Neural Inf Process Syst. 2014;27. Available from: https://arxiv.org/abs/1406.2661
4. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1706.03762
5. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. 2019. p. 4171-4186. doi:10.18653/v1/N19-1423.
6. Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, et al. Unsupervised cross-lingual representation learning at scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 8440-8451. doi:10.18653/v1/2020.acl-main.747.
7. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1705.07874
8. Ribeiro MT, Singh S, Guestrin C. "Why should I trust you?": explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 1135-1144. doi:10.1145/2939672.2939778.
9. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html
10. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.
11. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.
12. Ovadia Y, Fertig E, Ren J, Nado Z, Sculley D, Nowozin S, et al. Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift. Adv Neural Inf Process Syst. 2019;32. Available from: https://arxiv.org/abs/1906.02530
13. Jain S, Wallace BC. Attention is not explanation. In: Proceedings of NAACL-HLT. 2019. p. 3543-3556. doi:10.18653/v1/N19-1357.
14. Jacovi A, Goldberg Y. Towards faithfully interpretable NLP systems: how should we define and evaluate faithfulness?. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 4198-4205. doi:10.18653/v1/2020.acl-main.386.
15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.
16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.
17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.
18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.
Published
Issue
Section
License
Authors retain copyright. Articles published under the Creative Commons Attribution 4.0 International licence (CC BY 4.0) may be shared and adapted for any purpose, including commercially, provided appropriate credit is given, a link to the licence is supplied, and changes are indicated. No additional legal or technological restrictions may be imposed. Third-party material is included only where its credit line permits. Licence: https://creativecommons.org/licenses/by/4.0/. Earlier publications remain subject to their stated licence and author agreements unless the rights holder authorizes a change.