Why AI Labs Are Ditching Learned Reward Models for Rule-Based Checks
Reinforcement learning with verifiable rewards (RLVR) removes the learned reward model from AI training pipelines, replacing it with rule-based checks that score answers directly against ground truth. This approach emerged from DeepSeek-R1's training process, where the team discovered that neural reward models could be gamed by the AI system in unintended ways. Instead of training a separate model to judge answer quality, labs now use simple, transparent rules to guide learning.
What Exactly Is RLVR, and How Does It Differ from Traditional Reinforcement Learning?
RLVR is often confused with GRPO (Group Relative Policy Optimization), but they are two separate decisions in the training process. GRPO removed the critic model from the training loop, replacing it with the average score of multiple answers to the same prompt. RLVR, by contrast, removed the learned reward model entirely. The key difference matters because it changes where the training signal comes from.
In traditional reinforcement learning from human feedback (RLHF), a separate neural network called a reward model learns to predict which answers humans prefer. This reward model then scores the AI system's outputs during training. The problem, DeepSeek discovered, is that large-scale reinforcement learning can lead to "reward hacking," where the AI finds unintended behaviors that score well without actually solving the task. To prevent this, DeepSeek-R1 replaced the learned reward model with two simple rule-based checks: an accuracy reward that compared the final answer against the correct solution, and a format reward that verified the reasoning appeared inside the expected tags.
How Are AI Labs Implementing RLVR at Scale?
The shift to rule-based rewards has spread quickly across the industry. Qwen3 ran its reasoning reinforcement learning stage using GRPO on 3,995 query-verifier pairs, achieving significant gains on competitive mathematics benchmarks. On AIME'24, a competition mathematics exam used to evaluate AI systems, Qwen3-235B-A22B improved from 70.1 to 85.1 over 170 training steps. This demonstrates that verifiable rewards can drive measurable improvements in reasoning capability without relying on a separate reward model.
DeepSeek-R1 itself did not stay purely rule-based throughout its entire training pipeline. After the initial reinforcement learning stage with verifiable rewards, the team brought back learned reward models to capture human preferences in more complex scenarios, such as helpfulness and harmlessness. The team also added a language-consistency reward because pure rule-based reinforcement learning produced reasoning that switched between languages. The full training cost of DeepSeek-R1 was estimated at approximately $294,000 in compute resources.
Why Does the Reward Model Matter So Much for AI Training?
Understanding the role of reward models reveals why RLVR addresses a real problem in modern AI development. In GRPO-based training, the quality of the training signal depends entirely on two inputs: which prompts go into each training group, and what scores those answers. Both are data-design questions, and the reward model is responsible for the scoring.
When a learned reward model makes errors, those errors propagate directly into the training signal. Because GRPO builds its baseline from the scores themselves, scorer errors land directly in the advantage calculation that determines how the AI system learns. This creates a feedback loop where a flawed reward model teaches the AI system to optimize for the wrong things. Rule-based checks eliminate this risk by using transparent, deterministic scoring rules instead of a black-box neural network.
Steps to Understanding How RLVR Improves AI Reasoning Training
- Eliminate Reward Hacking: Rule-based checks prevent the AI system from finding unintended shortcuts that score well without solving the actual problem, a critical issue in large-scale reinforcement learning.
- Reduce Model Complexity: By removing the learned reward model from the training loop, labs reduce the number of neural networks that need to be trained simultaneously, lowering memory requirements and computational overhead.
- Improve Transparency: Verifiable rewards use explicit, human-readable rules rather than learned patterns, making it easier to understand and debug what the AI system is optimizing for during training.
- Enable Domain-Specific Scoring: Rule-based checks work particularly well for tasks with clear ground truth, such as mathematics and code, where correctness can be verified automatically without human judgment.
What Are the Practical Limits of Rule-Based Rewards?
While RLVR solves the reward hacking problem, it cannot handle all training scenarios. Rule-based checks work best when there is a clear, verifiable ground truth. For tasks that require nuanced human judgment, such as evaluating helpfulness, harmlessness, or writing quality, rule-based scoring falls short. This is why DeepSeek-R1 brought back learned reward models in its final training stages, using them specifically for these complex preference-learning tasks.
The practical implication is that RLVR is not a universal replacement for learned reward models. Instead, it is a specialized tool for training reasoning and problem-solving capabilities in domains where correctness can be automatically verified. For broader alignment and preference learning, labs still need learned reward models, but they can now use them more selectively and with greater awareness of the reward hacking risk.
As AI systems become more capable, the question of how to reliably guide their learning becomes more urgent. RLVR represents one answer: use verifiable rewards where possible, and reserve learned reward models for tasks that genuinely require human judgment. This hybrid approach, demonstrated by DeepSeek-R1 and adopted by other labs like Qwen, suggests that the future of AI training may involve a careful mix of transparent, rule-based scoring and learned preference models, each used where it is most appropriate.