Deep Reinforcement Learning for Computation Offloading and Dynamic Resource Allocation in Mobile Edge Computing
Keywords:
Deep reinforcement learning, computation offloading, mobile edge computingAbstract
Mobile edge computing extends computational and storage capabilities toward the network edge to support latency-sensitive and resource-intensive applications running on mobile and Internet of Things devices. A central challenge is determining when computation should be processed locally, transferred to an edge server, or allocated across available network resources under rapidly changing wireless and computational conditions. Conventional optimization methods may become difficult to apply when system dynamics are unknown, high dimensional, or continuously varying. Deep reinforcement learning provides a framework in which agents learn sequential resource-management policies through interaction with the environment. This review examines deep-reinforcement-learning approaches to computation offloading and dynamic resource allocation available through 2021. Foundational value-based and actor-critic techniques - including deep Q-networks, asynchronous actor-critic methods, deterministic policy gradients, proximal policy optimization, and soft actor-critic - are reviewed in relation to edge-computing problems. Applications involving task offloading, wireless channel assignment, power management, vehicular edge computing, and multi-user resource allocation are subsequently examined. The review analyzes state-space design, reward formulation, discrete and continuous action spaces, single- versus multi-agent environments, convergence, energy consumption, service latency, and quality-of-service constraints. Persistent limitations include training instability, exploration cost, non-stationary network environments, generalization across deployment conditions, and the difficulty of validating learned policies in real infrastructure. Nevertheless, by 2021 deep reinforcement learning had emerged as an increasingly important mechanism for adaptive edge-resource orchestration.
References
Sutton RS, Barto AG. Reinforcement Learning: An Introduction. 2nd ed. Cambridge: MIT Press; 2018.
Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, et al. Human-level control through deep reinforcement learning. Nature. 2015;518(7540):529-533.
Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, Harley T, et al. Asynchronous methods for deep reinforcement learning. Proc Mach Learn Res. 2016;48:1928-1937.
Lillicrap TP, Hunt JJ, Pritzel A, Heess N, Erez T, Tassa Y, et al. Continuous control with deep reinforcement learning. In: International Conference on Learning Representations. 2016.
Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv. 2017;1707.06347.
Haarnoja T, Zhou A, Abbeel P, Levine S. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proc Mach Learn Res. 2018;80:1861-1870.
Mao Y, You C, Zhang J, Huang K, Letaief KB. A survey on mobile edge computing: the communication perspective. IEEE Commun Surv Tutor. 2017;19(4):2322-2358.
Mach P, Becvar Z. Mobile edge computing: a survey on architecture and computation offloading. IEEE Commun Surv Tutor. 2017;19(3):1628-1656.
Shi W, Cao J, Zhang Q, Li Y, Xu L. Edge computing: vision and challenges. IEEE Internet Things J. 2016;3(5):637-646.
Mao Y, Zhang J, Letaief KB. Dynamic computation offloading for mobile-edge computing with energy harvesting devices. IEEE J Sel Areas Commun. 2016;34(12):3590-3605.
Chen X, Jiao L, Li W, Fu X. Efficient multi-user computation offloading for mobile-edge cloud computing. IEEE/ACM Trans Netw. 2016;24(5):2795-2808.
Huang B, Li Z, Xu Y. Deep reinforcement learning for performance-aware adaptive resource allocation in mobile edge computing. Wirel Commun Mob Comput. 2020;2020:2765491.
Chen M, Wang T, Zhang S, Liu A. Deep reinforcement learning for computation offloading in mobile edge computing environment. Comput Commun. 2021;175:1-12.
Li D, Xu S, Li P. Deep reinforcement learning-empowered resource allocation for mobile edge computing in cellular V2X networks. Sensors. 2021;21(2):372.
Wang S, Liu H, Gomes PH, Krishnamachari B. Deep reinforcement learning for dynamic multichannel access in wireless networks. IEEE Trans Cogn Commun Netw. 2018;4(2):257-265.
Ye H, Li GY, Juang BHF. Deep reinforcement learning based resource allocation for V2V communications. IEEE Trans Veh Technol. 2019;68(4):3163-3173.
Huang L, Bi S, Zhang YJ. Deep reinforcement learning for online computation offloading in wireless powered mobile-edge computing networks. IEEE Trans Mob Comput. 2020;19(11):2581-2593.
Huang X, Yu R, Kang J, Zhang Y. Distributed reputation management for secure and efficient vehicular edge computing and networks. IEEE Access. 2017;5:25408-25420.
Zhang Y, Xia W, Yan F, Cheng H, Shen L. Multi-agent reinforcement learning for joint wireless and computational resource allocation in mobile edge computing system. In: Ad Hoc Networks. Cham: Springer; 2020. p. 156-169.
Xu J, Chen L, Zhou P. Joint service caching and task offloading for mobile edge computing in dense networks. In: IEEE INFOCOM 2018. Piscataway: IEEE; 2018. p. 207-215.