Adaptive Test-Time Reasoning in Large Language Models: Dynamic Compute Allocation, Overthinking Control, Verification, and Efficient Deliberation

Authors

  • Matthias Vogel Author

Keywords:

adaptive reasoning, reasoning models, Test-time scaling

Abstract

The development of reasoning-oriented large language models shifted the scaling problem from exclusively increasing training-time model capacity toward determining how much computation should be allocated to individual problems during inference. Early test-time scaling methods generally assumed that additional reasoning tokens, sampled solution paths, or verification steps would improve difficult-task performance. Evidence available by 2026, however, increasingly demonstrated that indiscriminate reasoning expansion can produce substantial inefficiency and may even reduce accuracy through overthinking, redundant exploration, or poor allocation of reasoning effort. This review examines adaptive test-time reasoning methods through September 2026. Chain-of-thought prompting, self-consistency, verifier-guided reasoning, Tree-of-Thoughts, iterative refinement, and process supervision are reviewed as foundations for inference-time deliberation. DeepSeek-R1 and related reinforcement-learned reasoning systems demonstrated that extended reasoning behavior could emerge through optimization targeted toward verifiable tasks. Subsequent work investigated compute-optimal scaling and controllable reasoning budgets. Plan-and-Budget introduced problem decomposition and adaptive token allocation to reduce both overthinking and underthinking, while OptimalThinkingBench explicitly evaluated whether models select an appropriate amount of reasoning for tasks of different difficulty. Research on training-data requirements further indicated that additional inference computation is useful only when prerequisite capabilities exist in the learned model. Asynchronous Test-Time Scaling examined latency-efficient parallel and sequential reasoning using speculative mechanisms and conformal prediction. The review concludes that the central challenge is no longer simply inducing longer reasoning traces, but learning or estimating when reasoning is required, how reasoning effort should be distributed, when verification is useful, and when computation should terminate.

 

References

Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, et al. Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst. 2022;35:24824-24837.

Wang X, Wei J, Schuurmans D, Le QV, Chi EH, Narang S, et al. Self-consistency improves chain of thought reasoning in language models. In: International Conference on Learning Representations. 2023.

Kojima T, Gu SS, Reid M, Matsuo Y, Iwasawa Y. Large language models are zero-shot reasoners. Adv Neural Inf Process Syst. 2022;35:22199-22213.

Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, et al. Training verifiers to solve math word problems. arXiv. 2021;2110.14168.

Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, et al. Let’s verify step by step. arXiv. 2023;2305.20050.

Yao S, Yu D, Zhao J, Shafran I, Griffiths TL, Cao Y, Narasimhan K. Tree of thoughts: deliberate problem solving with large language models. Adv Neural Inf Process Syst. 2023;36.

Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, et al. Self-Refine: iterative refinement with self-feedback. Adv Neural Inf Process Syst. 2023;36.

Zelikman E, Wu Y, Mu J, Goodman ND. STaR: bootstrapping reasoning with reasoning. Adv Neural Inf Process Syst. 2022;35:15476-15488.

Zelikman E, Harik G, Shao Y, Jayasiri V, Haber N, Goodman ND. Quiet-STaR: language models can teach themselves to think before speaking. arXiv. 2024;2403.09629.

Snell C, Lee J, Xu K, Kumar A. Scaling LLM test-time compute optimally can be more effective than scaling model parameters. In: International Conference on Learning Representations. 2025.

Guo D, Yang D, Zhang H, Song J, Zhang R, Xu R, et al. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. Nature. 2025;645:633-638.

Muennighoff N, Yang Z, Shi W, Li XL, Fei-Fei L, Hajishirzi H, et al. s1: simple test-time scaling. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.

Uesato J, Kushman N, Kumar R, Song F, Siegel N, Wang L, et al. Solving math word problems with process- and outcome-based feedback. arXiv. 2022;2211.14275.

Shao Z, Wang P, Zhu Q, Xu R, Song J, Bi X, et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv. 2024;2402.03300.

Lin J, Zeng X, Zhu J, Wang S, Shun J, Wu J, Zhou D. Plan and Budget: effective and efficient test-time scaling on reasoning large language models. In: International Conference on Learning Representations. 2026.

Aggarwal P, Kim S, Lanchantin J, Welleck S, Weston JE, Kulikov I, Saha S. OptimalThinkingBench: evaluating over and underthinking in LLMs. In: International Conference on Learning Representations. 2026.

Javanmard A, Mirzasoleiman B, Mirrokni V. Understanding the role of training data in test-time scaling. In: International Conference on Learning Representations. 2026.

Xiong J, Chen Q, Ye F, Wan Z, Zheng C, Zhao C, et al. ATTS: asynchronous test-time scaling via conformal prediction. In: International Conference on Learning Representations. 2026.

Bansal R, Zhang A, Tiwari R, Madaan L, Duvvuri VSSS, Khatri D, et al. Let’s (not) just put things in context: test-time training for long-context LLMs. In: International Conference on Learning Representations. 2026.

Rafailov R, Sharma A, Mitchell E, Manning CD, Ermon S, Finn C. Direct preference optimization: your language model is secretly a reward model. Adv Neural Inf Process Syst. 2023;36.

Published

2026-06-01