Transformer Ensembles Versus Single Models for Bangla Smishing Detection: A Temporal and Cross-Corpus Validation Study

Authors

  • Nora Olsen Department of Cybersecurity and Artificial Intelligence, University of Jordan Author
  • Niko Moretti Department of Cybersecurity and Artificial Intelligence, Indian Institute of Science Author

Abstract

Smishing detection in Bangla is complicated by limited labelled data, evolving attack language, and frequent code-mixing with English. We compared three transformer ensembles with their strongest single-model components using two temporally separated Bangla SMS corpora and an external cross-corpus test set. Models were trained on 18,420 messages and evaluated on later messages without random mixing across time. Ensemble methods improved area under the precision-recall curve by 0.04–0.07 over single models and were particularly effective for rare campaign-specific wording. Performance fell for all methods on the external corpus, but the best ensemble retained an F1 score of 0.89 compared with 0.83 for the strongest individual model. Code-mixed messages remained the most error-prone class. Calibration analysis showed that averaging poorly calibrated component probabilities could produce overconfidence, while temperature-scaled ensembles improved reliability. Transformer ensembles can strengthen Bangla smishing detection, but temporal and cross-corpus validation reveals a larger generalisation gap than random hold-out testing. Operational evaluation should therefore include evolving attack language, code-mixing, and probability calibration.

References

1. Azam MA, et al. TransFuse: an interpretable transformer ensemble framework for Bangla smishing detection. 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN). 2026. Available from: https://ieeexplore.ieee.org/document/11546480/

2. Khan MFI, Begum MH, Rahman MA, Limon GQ, Azam MA, Masum AKM. A comprehensive review of advances in transformer, GAN, and attention mechanisms: their role in multimodal learning and applications across NLP. International Journal of Science and Research Archive. 2025;15(1):454-459. doi:10.30574/ijsra.2025.15.1.0980.

3. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1706.03762

4. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. 2019. p. 4171-4186. doi:10.18653/v1/N19-1423.

5. Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, et al. Unsupervised cross-lingual representation learning at scale. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 8440-8451. doi:10.18653/v1/2020.acl-main.747.

6. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. Available from: https://arxiv.org/abs/1705.07874

7. Ribeiro MT, Singh S, Guestrin C. "Why should I trust you?": explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016. p. 1135-1144. doi:10.1145/2939672.2939778.

8. Guo C, Pleiss G, Sun Y, Weinberger KQ. On calibration of modern neural networks. Proc Mach Learn Res. 2017;70:1321-1330. Available from: https://proceedings.mlr.press/v70/guo17a.html

9. Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432.

10. Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. doi:10.1016/j.patter.2023.100804.

11. Ovadia Y, Fertig E, Ren J, Nado Z, Sculley D, Nowozin S, et al. Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift. Adv Neural Inf Process Syst. 2019;32. Available from: https://arxiv.org/abs/1906.02530

12. Jain S, Wallace BC. Attention is not explanation. In: Proceedings of NAACL-HLT. 2019. p. 3543-3556. doi:10.18653/v1/N19-1357.

13. Jacovi A, Goldberg Y. Towards faithfully interpretable NLP systems: how should we define and evaluate faithfulness?. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. p. 4198-4205. doi:10.18653/v1/2020.acl-main.386.

14. Demšar J. Statistical comparisons of classifiers over multiple data sets. J Mach Learn Res. 2006;7:1-30. Available from: https://www.jmlr.org/papers/v7/demsar06a.html

15. Lakens D. Sample size justification. Collabra Psychol. 2022;8(1):33267. doi:10.1525/collabra.33267.

16. Nosek BA, Ebersole CR, DeHaven AC, Mellor DT. The preregistration revolution. Proc Natl Acad Sci U S A. 2018;115(11):2600-2606. doi:10.1073/pnas.1708274114.

17. Wilkinson MD, Dumontier M, Aalbersberg IJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18.

18. Efron B. Bootstrap methods: another look at the jackknife. Ann Stat. 1979;7(1):1-26. doi:10.1214/aos/1176344552.

Published

2026-06-01