Transformer Architectures for Long-Horizon Time-Series Forecasting: Sparse Attention, Decomposition, and Frequency-Domain Modeling

Authors

  • Elena Marković Author

Keywords:

attention, deep learning, temporal modeling, long-term forecasting, FEDformer

Abstract

Accurate long-horizon time-series forecasting is important in energy systems, transportation, finance, meteorology, industrial monitoring, healthcare, and capacity planning. Traditional statistical approaches remain effective in many structured settings, while recurrent neural networks and convolutional models introduced nonlinear representation learning for complex temporal sequences. Following the success of Transformers in natural-language processing, self-attention architectures were increasingly adapted for long-sequence forecasting. This review examines major developments in Transformer-based time-series forecasting through 2022. Conventional self-attention can represent long-range dependencies but exhibits quadratic computational complexity with respect to sequence length. LogSparse Transformer and Informer therefore introduced sparse attention strategies intended to reduce computational and memory requirements. Autoformer incorporated progressive seasonal-trend decomposition and an Auto-Correlation mechanism to model periodic dependencies at the subsequence level. FEDformer extended this direction by combining decomposition with frequency-enhanced representations based on Fourier and wavelet transformations. Pyraformer introduced pyramidal attention for multiresolution temporal relationships. These approaches are compared with recurrent forecasting, DeepAR, N-BEATS, Temporal Fusion Transformers, deep state-space models, and probabilistic multivariate forecasting. The review examines deterministic and probabilistic objectives, covariate integration, long-range dependency modeling, decomposition, computational complexity, and evaluation protocols. The 2022 literature demonstrated substantial interest in specialized temporal architectures but also highlighted concerns regarding benchmark consistency, sensitivity to forecasting horizon, model complexity, and whether increasingly elaborate attention mechanisms always outperform simpler forecasting baselines.

References

Box GEP, Jenkins GM, Reinsel GC, Ljung GM. Time Series Analysis: Forecasting and Control. 5th ed. Hoboken: Wiley; 2015.

Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735-1780.

Lai G, Chang WC, Yang Y, Liu H. Modeling long- and short-term temporal patterns with deep neural networks. In: Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval. 2018. p. 95-104.

Salinas D, Flunkert V, Gasthaus J, Januschowski T. DeepAR: probabilistic forecasting with autoregressive recurrent networks. Int J Forecast. 2020;36(3):1181-1191.

Oreshkin BN, Carpov D, Chapados N, Bengio Y. N-BEATS: neural basis expansion analysis for interpretable time series forecasting. In: International Conference on Learning Representations. 2020.

Lim B, Arık SÖ, Loeff N, Pfister T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. Int J Forecast. 2021;37(4):1748-1764.

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.

Li S, Jin X, Xuan Y, Zhou X, Chen W, Wang YX, Yan X. Enhancing the locality and breaking the memory bottleneck of Transformer on time series forecasting. Adv Neural Inf Process Syst. 2019;32:5243-5253.

Zhou H, Zhang S, Peng J, Zhang S, Li J, Xiong H, Zhang W. Informer: beyond efficient Transformer for long sequence time-series forecasting. Proc AAAI Conf Artif Intell. 2021;35(12):11106-11115.

Wu H, Xu J, Wang J, Long M. Autoformer: decomposition Transformers with Auto-Correlation for long-term series forecasting. Adv Neural Inf Process Syst. 2021;34:22419-22430.

Zhou T, Ma Z, Wen Q, Wang X, Sun L, Jin R. FEDformer: frequency enhanced decomposed Transformer for long-term series forecasting. Proc Mach Learn Res. 2022;162:27268-27286.

Liu S, Yu H, Liao C, Li J, Lin W, Liu AP, Dustdar S. Pyraformer: low-complexity pyramidal attention for long-range time series modeling and forecasting. In: International Conference on Learning Representations. 2022.

Woo G, Liu C, Sahoo D, Kumar A, Hoi S. ETSformer: exponential smoothing Transformers for time-series forecasting. arXiv. 2022;2202.01381.

Qin Y, Song D, Chen H, Cheng W, Jiang G, Cottrell G. A dual-stage attention-based recurrent neural network for time series prediction. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. 2017. p. 2627-2633.

Rangapuram SS, Seeger MW, Gasthaus J, Stella L, Wang Y, Januschowski T. Deep state space models for time series forecasting. Adv Neural Inf Process Syst. 2018;31:7785-7794.

Sen R, Yu HF, Dhillon IS. Think globally, act locally: a deep neural network approach to high-dimensional time series forecasting. Adv Neural Inf Process Syst. 2019;32:4837-4846.

Rasul K, Seward C, Schuster I, Vollgraf R. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. Proc Mach Learn Res. 2021;139:8857-8868.

Rasul K, Sheikh AS, Schuster I, Bergmann U, Vollgraf R. Multivariate probabilistic time series forecasting via conditioned normalizing flows. In: International Conference on Learning Representations. 2021.

Shih SY, Sun FK, Lee HY. Temporal pattern attention for multivariate time series forecasting. Mach Learn. 2019;108:1421-1441.

Wen R, Torkkola K, Narayanaswamy B, Madeka D. A multi-horizon quantile recurrent forecaster. arXiv. 2017;1711.11053.

Published

2022-06-01