As language models grow in complexity, and are increasingly used, so does the need for mechanisms that allow these models to learn to behave in useful, secure, and ethically aligned ways.

It is in this context that Deepseek, a young Chinese company founded in 2023, has taken a bold and promising step with the development of Deepseek-GRM-27B: a next-generation reward model that could redefine reinforcement learning in AI systems.

This article explores in depth Deepseek’s innovative approach, its key techniques, as well as its potential technical and strategic impact.

What are reward models and why do they matter?

In the field of reinforcement learning (RL), a reward model is the component that determines what actions are desirable for an AI agent.

This agent interacts with an environment, makes decisions and, depending on the reward model, receives positive signals (rewards) or negative signals (punishments), allowing it to learn to behave optimally over time.

For example, a robot can learn to pick up objects accurately by receiving points each time it does so correctly. In more sophisticated applications, such as conversational assistants, rewards are based on criteria such as usefulness, clarity, or user satisfaction.

The problem is that human preferences are subtle, changing and often subjective. Traditional methods, such as reinforcement learning with human feedback (RLHF), attempt to capture these preferences by manually evaluating responses.

While this has been useful (models like ChatGPT and Claude use RLHF), it is also an expensive, slow, and unscalable process.

Deepseek’s disruptive approach

Deepseek proposes a new architecture for reward models that seeks to eliminate over-reliance on human oversight without sacrificing alignment with human values. This is achieved thanks to three key technical pillars:

Generative Reward Modeling (GRM)

GRM is a technique by which the model learns to generate its own feedback principles. Instead of relying exclusively on human-labeled data, the model can analyze different responses to a task and deduce, on its own, which one is most appropriate.

For example, if a system is asked to answer a medical query, GRM can generate several options, evaluate them internally, and determine which best meets criteria such as clarity, accuracy, and usefulness. Thus, the model acts as its own judge, learning more autonomously and efficiently.

Self-Principled Critique Tuning (SPCT)

SPCT complements GRM with a self-criticism mechanism. This allows the model to dynamically adjust its own evaluation principles. If responses evaluated as “good” turn out to be inadequate in practice, the model detects these inconsistencies and improves its reward criteria.

This not only improves the adaptability of the system to new contexts, but also reduces the need for manual adjustments, achieving a more resilient and scalable system.

Scaling at inference time

A less common but very significant innovation in the Deepseek proposal is the inference time scaling approach that we previously talked about in “Advance in AI: An increasingly difficult and expensive path?”

Instead of concentrating all computational resources during the training phase, Deepseek proposes allocating them also during the use (inference) phase.

This allows the model to generate more accurate and contextualized responses without requiring retraining. This technique is especially useful for tasks that require real-time adaptability, such as content moderation or technical support.

Deepseek-GRM-27B: Results and performance

The Deepseek-GRM-27B model, with 27 billion parameters, represents the culmination of these innovations. According to the technical paper published on arXiv under the title “Inference-Time Scaling for Generalist Reward Modeling” (arXiv:2504.02495), this model outperforms several key competitors.

  • In tests with greedy decoding, it obtained a score of 69.9, surpassing the LLM-as-a-Judge model (67.0).
  • In voting@32 with MetaRM, it reached 72.8, compared to 69.0 for Deepseek-PairRM-27B.
  • In generalization tasks such as RMB BoN, the difference between different input types was less than 1%, showing its versatility.

Furthermore, compared to high-level models such as GPT-4o (version 2024-08-06), Deepseek-GRM-27B showed a slightly higher capacity to transfer reward principles to new tasks, which positions it as a highly reusable and adaptable model.

A dual strategy: Web, open source and community

Deepseek not only focuses on technical innovation, but also accessibility and collaboration.

The company has announced that GRM and SPCT functionalities are being gradually integrated into its web version, which will allow users and companies to access advanced reward models without requiring their own infrastructure.

Furthermore, Deepseek plans to release key components of its model as open-source. This decision opens the door for researchers, academic laboratories and startups to study, modify and take advantage of these technologies in their own projects.

In an environment dominated by proprietary initiatives of large technology companies, this commitment to open source represents a breath of fresh air for the scientific community.

Implications for the future of AI

Deepseek’s work is not just a technical advance: it is a redefinition of how we train and evaluate AI systems. Among its most relevant implications are:

  • Reduction of human dependency: By automating feedback, resources are freed and the training of larger models is facilitated.
  • Democratization of access to advanced AI: Thanks to its open-source and web approach, these advances will be available to actors outside the large centers of technological power.
  • Greater security and alignment: Models capable of criticizing and adjusting themselves reduce risks associated with biases, errors or unexpected behavior.
  • Acceleration of scientific research: With more aligned and adaptable systems, tasks such as data analysis or hypothesis formulation could benefit significantly.

Thanks to techniques such as GRM, SPCT and its innovative approach to inference, the company has not only outperformed its competitors in benchmarks, but has laid the foundation for more aligned, autonomous and accessible AI.

By integrating these capabilities into its web platform and betting on open-source, Deepseek not only seeks to lead technically, but also to foster a more collaborative and inclusive global community in the development of artificial intelligence.

This post is also available in: Español Français Русский Italiano