Tiny Machine Learning for Resource-Constrained IoT Devices: Efficient Architectures, Quantization, and Hardware-Aware Inference
Keywords:
embedded machine learning, MCUNet, neural architecture search, quantizationAbstract
The proliferation of low-power sensors and Internet of Things devices has created increasing demand for machine-learning inference directly on microcontrollers and other resource-constrained hardware. Conventional deep neural networks frequently exceed the memory, computation, latency, and energy budgets available on such platforms. Tiny machine learning, commonly termed TinyML, seeks to enable useful machine intelligence within kilobyte- to megabyte-scale memory environments while minimizing dependence on continuous cloud connectivity. This review examines the development of efficient embedded machine learning through 2021. Lightweight neural architectures including MobileNet, MobileNetV2, MobileNetV3, SqueezeNet, ShuffleNet, and EfficientNet are discussed together with model compression, pruning, low-precision inference, binary networks, quantization, neural architecture search, and hardware-aware optimization. Particular attention is given to MCUNet, which jointly designed neural architectures and an inference engine for microcontroller constraints, as well as Once-for-All and ProxylessNAS approaches for hardware-specialized deployment. Integer-only inference and mixed-precision quantization are evaluated as mechanisms for reducing memory and computational requirements. The review considers the trade-offs among model accuracy, SRAM consumption, flash storage, latency, energy efficiency, and device portability. Challenges include fragmented hardware ecosystems, limited benchmark standardization, memory bottlenecks, operator support, and differences between theoretical operation counts and real-device performance. The evidence available by 2021 suggested that hardware-software co-design would be central to making always-on artificial intelligence practical on very small IoT devices.
References
Warden P, Situnayake D. TinyML: Machine Learning with TensorFlow Lite on Arduino and Ultra-Low-Power Microcontrollers. Sebastopol: O'Reilly Media; 2019.
Lin J, Chen WM, Lin Y, Cohn J, Gan C, Han S. MCUNet: tiny deep learning on IoT devices. Adv Neural Inf Process Syst. 2020;33:11711-11722.
Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, et al. MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv. 2017;1704.04861.
Sandler M, Howard A, Zhu M, Zhmoginov A, Chen LC. MobileNetV2: inverted residuals and linear bottlenecks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 4510-4520.
Howard A, Sandler M, Chu G, Chen LC, Chen B, Tan M, et al. Searching for MobileNetV3. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. p. 1314-1324.
Tan M, Le QV. EfficientNet: rethinking model scaling for convolutional neural networks. Proc Mach Learn Res. 2019;97:6105-6114.
Han S, Mao H, Dally WJ. Deep compression: compressing deep neural networks with pruning, trained quantization and Huffman coding. In: International Conference on Learning Representations. 2016.
Jacob B, Kligys S, Chen B, Zhu M, Tang M, Howard A, et al. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 2704-2713.
Wang K, Liu Z, Lin Y, Lin J, Han S. HAQ: hardware-aware automated quantization with mixed precision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. p. 8612-8620.
Cai H, Zhu L, Han S. ProxylessNAS: direct neural architecture search on target task and hardware. In: International Conference on Learning Representations. 2019.
Cai H, Gan C, Wang T, Zhang Z, Han S. Once for all: train one network and specialize it for efficient deployment. In: International Conference on Learning Representations. 2020.
Tan M, Chen B, Pang R, Vasudevan V, Sandler M, Howard A, Le QV. MnasNet: platform-aware neural architecture search for mobile. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. p. 2820-2828.
Iandola FN, Han S, Moskewicz MW, Ashraf K, Dally WJ, Keutzer K. SqueezeNet: AlexNet-level accuracy with 50× fewer parameters and <0.5 MB model size. arXiv. 2016;1602.07360.
Zhang X, Zhou X, Lin M, Sun J. ShuffleNet: an extremely efficient convolutional neural network for mobile devices. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. p. 6848-6856.
Ma N, Zhang X, Zheng HT, Sun J. ShuffleNet V2: practical guidelines for efficient CNN architecture design. In: Ferrari V, Hebert M, Sminchisescu C, Weiss Y, editors. Computer Vision—ECCV 2018. Cham: Springer; 2018. p. 116-131.
Courbariaux M, Bengio Y, David JP. BinaryConnect: training deep neural networks with binary weights during propagations. Adv Neural Inf Process Syst. 2015;28:3123-3131.
Rastegari M, Ordonez V, Redmon J, Farhadi A. XNOR-Net: ImageNet classification using binary convolutional neural networks. In: Leibe B, Matas J, Sebe N, Welling M, editors. Computer Vision—ECCV 2016. Cham: Springer; 2016. p. 525-542.
Lai L, Suda N, Chandra V. CMSIS-NN: efficient neural network kernels for Arm Cortex-M CPUs. arXiv. 2018;1801.06601.
Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. arXiv. 2015;1503.02531.
Han S, Pool J, Tran J, Dally WJ. Learning both weights and connections for efficient neural networks. Adv Neural Inf Process Syst. 2015;28:1135-1143.