Test-Time Scaling and Reinforcement-Learned Reasoning in Large Language Models: Verifiers, Budget Forcing, and Emergent Deliberation

Authors

  • Niklas Hartmann Author

Keywords:

DeepSeek-R1, test-time scaling, large language models, Reasoning models

Abstract

The rapid development of reasoning-oriented large language models during 2024 and 2025 shifted attention from scaling only pretraining parameters toward allocating additional computation during inference. Test-time scaling allows a model to generate, evaluate, revise, or search among multiple reasoning trajectories before committing to a final answer. This review examines the foundations and principal developments in inference-time reasoning through 2025. Chain-of-thought prompting established that intermediate natural-language reasoning could substantially improve performance on complex tasks, while self-consistency demonstrated the value of sampling multiple reasoning trajectories and aggregating their answers. Self-Taught Reasoner and Quiet-STaR explored mechanisms through which models could learn to generate useful internal rationales. Process-supervision and verifier research further demonstrated the importance of evaluating intermediate reasoning rather than relying solely on final-answer correctness. Tree-of-Thoughts and iterative self-refinement extended reasoning toward explicit search and revision. Research on compute-optimal test-time scaling subsequently formalized the relationship among task difficulty, sampling budget, verifier quality, and inference computation. DeepSeek-R1 demonstrated that reinforcement learning could incentivize sophisticated reasoning behavior, including reflection and verification, while s1 showed that carefully selected training examples combined with budget forcing could produce strong reasoning and controllable test-time computation. This review examines outcome and process reward models, reinforcement learning, search, self-correction, budget forcing, inference latency, reasoning-token allocation, and distillation. Remaining challenges include excessive reasoning, unreliable self-verification, reward hacking, evaluation contamination, hidden computational cost, safety degradation, and the difficulty of establishing whether extended reasoning traces correspond to faithful internal computation.

References

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998-6008.

Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. Adv Neural Inf Process Syst. 2020;33:1877-1901.

Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, et al. Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst. 2022;35:27730-27744.

Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, et al. Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst. 2022;35:24824-24837.

Wang X, Wei J, Schuurmans D, Le QV, Chi EH, Narang S, et al. Self-consistency improves chain of thought reasoning in language models. In: International Conference on Learning Representations. 2023.

Zelikman E, Wu Y, Mu J, Goodman ND. STaR: bootstrapping reasoning with reasoning. Adv Neural Inf Process Syst. 2022;35:15476-15488.

Yao S, Yu D, Zhao J, Shafran I, Griffiths TL, Cao Y, Narasimhan K. Tree of thoughts: deliberate problem solving with large language models. Adv Neural Inf Process Syst. 2023;36.

Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, et al. Self-Refine: iterative refinement with self-feedback. Adv Neural Inf Process Syst. 2023;36.

Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, et al. Training verifiers to solve math word problems. arXiv. 2021;2110.14168.

Uesato J, Kushman N, Kumar R, Song F, Siegel N, Wang L, et al. Solving math word problems with process- and outcome-based feedback. arXiv. 2022;2211.14275.

Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, et al. Let’s verify step by step. arXiv. 2023;2305.20050.

Zelikman E, Harik G, Shao Y, Jayasiri V, Haber N, Goodman ND. Quiet-STaR: language models can teach themselves to think before speaking. arXiv. 2024;2403.09629.

Shao Z, Wang P, Zhu Q, Xu R, Song J, Bi X, et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv. 2024;2402.03300.

Yang A, Zhang B, Hui B, Gao B, Yu B, Li C, et al. Qwen2.5-Math technical report: toward mathematical expert model via self-improvement. arXiv. 2024;2409.12122.

Snell C, Lee J, Xu K, Kumar A. Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning. In: International Conference on Learning Representations. 2025.

Guo D, Yang D, Zhang H, Song J, Zhang R, Xu R, et al. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. Nature. 2025;645:633-638.

Muennighoff N, Yang Z, Shi W, Li XL, Fei-Fei L, Hajishirzi H, et al. s1: simple test-time scaling. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. p. 20275-20321.

Rafailov R, Sharma A, Mitchell E, Manning CD, Ermon S, Finn C. Direct preference optimization: your language model is secretly a reward model. Adv Neural Inf Process Syst. 2023;36.

Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv. 2017;1707.06347.

Bai Y, Kadavath S, Kundu S, Askell A, Kernion J, Jones A, et al. Constitutional AI: harmlessness from AI feedback. arXiv. 2022;2212.08073.

Published

2025-06-01