The Evolution of Reward Design: From RLHF to RLVR

Hyacehila

Previous articleExplain how reward is consumed by optimizers; this paper takes the question forward to see how reward is produced before it enters the training system.

If the last one is concerned, reward -> advantageThis article is concerned with how the objectives are first recast as monitorable. LLM post-training reward evolution, ostensibly from manual feedback to verifeier, judge, rubric and tornament; actual change closer to the subject of oversight itself: From the human comparison of preferences to the results of programmable validation; from the total score of the answer level to the process steps, the rubric critterion, to the relative ranking of open anent tracks.

This story will go down where the old interfaces are not enough. OpenAI's RLHF first turns human preferences into learningable substrate reform, and brings proxy Gap with optimization risks into the system; multi-target alignment quickly exposes a single score to helpful/harmles conflict; RLVR moves Learned process from core results signals to a valid mission; to open missions and angent scenes, rewarding is increasingly a layer of interface: first removes hard interfaces, core task signature, process signal and cost/security boundaries, and then decides whether to send signals to optimizer with pairwise, rubric tournament.

And we've seen it over and over again on the training line: the optimizer behind it, and it's not gonna save the incisive upstream signals. Many recent years have seen progress on the surface as changes in the loss, changes in the trainer, changes in the sampling, but changes often occur in the nature of the response: whether the subject of surveillance is the answer, sorting, steps, trajectory or future state, whether the source of oversight is the person, AI, the rule, the vifier, or the environmental consequences, and whether the signal organization is the pointwise, the signal, the listwise, the roubric, or the future. Now, it can no longer be understood as a line of fractions that come out of a single model, which is more like an interface link: the first one to recast the target as a monitorable object, the middle one to organize the object into a relatively stable signal, and the last one to consume is an optimizer.

Phase 1: RLHF for OpenAI, with manual feedback as a reward

Technically, the prototype of reward from references can be traced. Deep Reinforcement Learning from Human PreferencesI'm sorry. But if I pull out the LLM-era reward main line alone, I'd rather start with OpenAI, which is Learning to Summarize from Human Feedback and InstructGPTI'm sorry. The reason is simple, but it's a "preferred data--"> Incentive model--> The PPO is a specific group of jobs written as a reusable industrial stream line.

First look. 2020. Learning to Summarize from Human FeedbackI'm sorry. The problem it is going to solve is not at all abstract: it is hard to write the quality of the abstracts manually, but it is easy for humans to compare the two abstracts in relative terms. The paper then adopted a straightforward approach: to allow models to generate multiple abstract candidates for the same article, to collect artificial paper comparisons, to train reward Model with these comparative data, and then to put this reward Mode on the PPO, and to optimize the way that the policy is more easily available to high reward.

The detail worth retaining is not just “a reward model”, but rather a rewarded oversight object, which has changed from manual to preference ranking, and which is no longer equal to a given original human score, but rather a neuro-reward model for learning about preference relationships. The earlier question was, OpenAI was already in this job.Clearly observed that once over-traught the RM fraction, the output starts to deviate from the real human being. PreferI'm sorry. This suggests that the language task can indeed be included in the reference-based RL, but the problem Gap has emerged from the first generation.

Back up. InstructGPTAnd what it does is fix this set of three-part processes that are common to later:** use demontations to do SFTs, then train raked comparisons, and then continue to optimize the RM-led placy. ** This paper fixed the structures that many post-training recipe later inherited.

First, SFT is not a temporary warm-up, but a reward-optimised behavioral starting point, which starts with a model of basic answer style, task format and action legitimacy, and then RLs to fine-tune.

Second, the optimizer is never an isolated RM fraction, but more like a composite signal: $$ R(x, y) = r_\phi(x, y) - \beta \log \frac{\pi_\theta(y\mid x)}{\pi_{\text{ref}}(y\mid x)} $$

The power of forward movement is given by KL to not move too far, and the two decide where the policy can go.

Third, the classic result of the RLHF 1.3B model is that the original 175B GPT-3 is better than the original 175B GPS-3 and it is not that the parameters are not important, but rather that the reward signature and post-training lines may be more direct in the quality of user perception than in the continuation of the stack.

If you sum up this line with a single sentence, OpenAI does three things that bind each other:First, you change complex targets to comparable humans, then you learn these preferences into neuro-reward models, and then you optimize this one with the boundaries of reference policy.I'm sorry. And it's here, and all the subsequent expansions of reward have planted seeds because once you admit that reward is surogate and that surogate will be expluited, you have to continue to ask: how does this surogate really work?

Phase II: Towards multi-purpose and scalable monitoring

Once running in the language model, the question quickly becomes whether or not to make a single preference score like this is going to cover anything. The most important change in this phase is not what the trainer changed, but the reward object itself is beginning to be multidimensional.

Training a Helpful and Harmless Assistant with RLHF It's worth putting here because it's an early and systematic representation of the problem: helpful and harmless are not a dimension. If you only pursue more user-friendly answers, the model may become too responsive to hazardous requests; but if you only stress harmless, it will slide quickly to overconservative. The most important thing of this job is not to re-grave RLHF, but to make it official from a "single mass" to a "mixed goal." From the point of view of the design of reward, this is crucial because it shows that a single scalar reward is often a compression made to optimize convenience, and does not mean that the target itself has only one axis. And more and more systems are introducing fractions, rubric and hard construits, which can be traced here in a sense.

If HH RLHF is exposed to more than one target per se, then Constitutional AI and RLAIF It is not always human who are exposed to the source of oversight. The key approach of Constitual AI is to write first a set of references to self-critical, then to revise the model, and then to reconnect AI feedback to the waterline. RLAIF further refined this approach by replacing the upstream portion, which was entirely dependent on manual comparisons, with the more scalable AI feedback. They are not rewritten by who to label, but rather by the uppermost of the rewarding exercise began to be structured: principles themselves became a priori bound, the process of creating clitic became programmable, and the source of oversight became part of the design of the system instead of being simply manual.

After this phase, it's hard to imagine a clean single target again. It is more like a compression after a multi-purpose compromise, and the upstream production line itself is beginning to become the target of design: The way it is written, the way it is organized, the way it is sampled, changes the shape of the last. Once the target is multidimensional, continuing to press everything to the average point will only hide the conflict; and it is from here that the system is moving more and more naturally towards rubric, sub-points and hard construits.

Phase III: Structural diagnosis of proxy awards: can the preference for learning be fixed?

The first two phases expanded the target dimensions of the reward and the sources of oversight, but one more fundamental question has not been positively pursued:How reliable is it, as an agent?

This is important because RLHF (and later the Constitutory AI) has a structured choice to combine preference in a learning manner into a neuro-reward function, and to let the optimizer pursue it. If Proxy has structural defects in itself, then whatever changes are made to the trainer, to the loss, to the sampling, it is fixed on a foundation that is not built. A series of jobs around 2023 coincided with the discovery of the cracks in the foundations at different levels.

The information was lost before entering R.M.

Preference Ranking Optimization The question is not hard to understand: would you throw down the structure information if you had a natural sort of pool of candidates, tearing it down into a few separate pairwise winners and then training RRM? Pro's answer is yes, and it's very serious. A listwise contains a global sequence that is much richer than several pairwises that are removed from it.

GPO Another loss is exposed. It is closer to the extension of the DPO/ Preference Optimization, which is discussed here in the Preference Learning Context because it is equally concerned about how the preferences are compressed before they are optimized. Many true comparisons do not grow to be clean zero or one, but are more like “A slightly better than B”, “almost the same” or “stabilized differences between the different labelers”. But the tradition of RLHF is to make these greyscales a hard label. The GPO value is to keep the uncertainty in the preference as a matter of urgency labels, so that the optimizer does not over-optimise the already vague border sample.

The problem on this level is, in principle, removable - with the listless loss, with the soft label. But the models that they have shared are important:Every step of compression from the original preference to the RM training data is missing information and never comes back after the information that was lost.

Second layer of crack: reward modelling itself writing reward

A density estimation perspective on learning from pairwise human preferences and RLHF and IIA: Perverse Incentives The problem is pushed to the bottom. And one thing they all remind us is that learning is not natural to restore a stable global environment. You choose between Bradley-Terry or Thurstone, assuming that IA (not independent of options) is a technical model that directly shapes the geometric shape of the reward that is learned. In particular, IA, while making learning more effective, may induce distorted incentive structures and incentives.

This layer is deeper than the first: not just a problem with data formats, but rather a problem with data formats.The act of preference itself is a substitute for real preference.I'm sorry. Even if the data are perfect, the model family you chose is making your decision. Reward Mode is not a neutral statistical container, but is writing about it.

Third layer crack: even if all above is fixed, proxy is proxy

The first two cracks can be fixed at least within the preferred learning paradigm. But the third level of the problem is structural: no matter how faithful you are to organize data and to choose statistical assumptions, RM is ultimately a learning approximation, not the real goal itself.

This is not theoretically a concern. From Learning to Summarize At first, the reward over-optimization has been clearly observed: when CPR scores are pursued by the polity, the output starts to deviate from the true human preferences. The stronger the optimizer, the better it is to find the gap between the RM and the real target. This is not a bug for a particular RM, but a structural risk for suirogate resold:Under sufficiently high optimal pressure, the risk of expluit increases significantly.

The difference after the diagnosis.

These three layers are stacked together and point to a clear fork in the fork. For open missions (creative writing, open dialogue, complex advice), the ground truth is non-existent or unformatted, and proxy has to be used - but needs to be more structured, which is the problem behind LLM as Judge, Rubric reward and Tornament Ranking.

But the correctness of the answer can be directly verified for another type of task — the correctness of mathematical reasoning, the ability of codes to run, the conformity of the result of the tool call to expectations. If it's verifiable, why do you have to put a similar pattern in the middle? This is the starting point for the next stage of RLVR: instead of learning better about proxy, it is directly bypassing proxy over the verifiable part.

Phase 4: RLVR - a verifiable breakthrough in reward and mathematical reasoning

The third phase of the diagnosis has indicated the direction to be taken: For verifiable tasks, the preferred approach is not to learn more about proxy, but to allow the verifiable part to be given a central signal directly by the verifeier. And the most convincing breakthrough of the RLVR line is in the field of mathematical reasoning.

The math has a natural advantage: the answer to the question is programmable. That means you don't need a reward model to guess, but you can have the final results checked out first.DeepSeek-R1 R1-Zero sets this idea very far: in mathematical reasoning tasks, rule-based analysis is rewarded with basic formatting incentives, and then GRPO is doing RLs, with models showing long chain reasoning, self-correction and reflection. This result is important not only because it works well, but because it gives a strong engineering principle:When the core results of the task are validated, outside response can directly assume the main reward signal without default training first.

This is the core orientation of RLVR: the validated parts prioritize direct validation, minimizing the intervention in the core of the research path.

But immediately thereafter, two specific problems were encountered. First, pure outcome delivery only tells you the final answer, but not which one of the steps is wrong, which means credit attachment is extremely thin in long chain reasoning. Secondly, many tasks are not entirely verifiable, and only a partial structure is verifiable.

The first question is that of the future.Let's Verify Step by Step A project was proposed to change the oversight unit from an answer-level to step-level. It is necessary to distinguish that the PRM is usually still lost, not equal to certainty, and its value is to tear long chain reasoning to a more fine particle size, which allows process supervision, step diagnosis and candidate path screening.Math-Shepherd Use step-level validation models further on the test-time candidate and path filter, so that reward is not just a training signal, but also a system component for reasoning. It is understood that steps for validation can be determined to be given priority over those in the verifier, and it is difficult to determine the quality of the certification process to be expected that it may be necessary to learn PRM or judge.

The second question is:Crossing the Reward Bridge It provides a key methodological judgement: many tasks that appear to be based on judice scoring, but which are sufficiently detailed to be verified by some structural requirements, environmental outcomes or intermediate constraints. So, the key to whether you want to untangle the task and get the verifiable part out of it is that the scope of the application is much broader than the mathematical questions and codes.Beyond Outcome Verification This is supplemented by another perspective: Even if the final results can be verified, knowing at which point the model is deviating is still valuable, and the intermediate diagnosis itself can be used as a clear-cut test rather than reverting to pure learning score.

At this stage, the default of the reward is becoming clearer:If there is a verifiable structure in the mission, the core results incentives are given to the verifier; if the task is long chain reasoning, the split is used to provide as many intermediates as possible with verifiable basis; if there is also red lines such as security, format, protocol, they are individually made hard constrants, rather than mixed into average points. The key to RLVR is not just to continue with RLHF by changing its acronym, but to reward source as much as possible as a verifiable or auditable object; for tasks such as mathematical reasoning, the result has shown that this path has a strong training value.

Phase 5: LLM as Judge - from Park to the differential diagnosis

But real tasks are often less ideal. Many open missions have neither the single answer nor the full coverage of the verifiers, and it is difficult to return the manual label completely, because size is not allowed. That's it. LLM as Judge Get into the main line. It just fills the space between the verifier and the pure artificial.

The early sign of this line is G-Eval and Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaI'm sorry. G-Eval CoT + form-filling In a way that allows the model to think along the assessment logic and then to be judged by structured forms; MT-Bench and Chatbot Arena show that in the preference for open dialogue, strong LLM judge can be fairly close to human judgment, forming a scalable proxy evaluation layer. The change brought about by this group of work is not “no human need in the future”, but the emergence of an intermediate interface in open missions that can be broadly comparable to human preferences.

Once you put Judge back in the relief context, the first question is: whether Judge is better suited to direct scring, or pairwisse ranking? Direct scring is high, simple in form, easy to quantify, but it requires that Judge determine the scale of the fraction, which is usually not stable in open missions -- - What's the difference between points 7 and 8, and Judge himself may not be clear, and the scale across the prompt can easily drift. The pairwise raning let the judge go back to his best move: two better candidates.Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators It's clear that seeing the evaluation as a ranking problem, not a pointwise scring problem, is a lot of natural decline. From the design point of view of reward, this means that rewarding in an open mission is more like producing relative structures than creating stable points.

But the Judge's risk must be clear here. Its greatest danger is not “a little bias”, but the easy fact that it is wrong to be treated as truth for good behaviour.Large Language Models are not Fair Evaluators and Judging the Judges Noting that the choice of the choice is changed, the Judge ' s decision is likely to be reversed (position bias); MT-Bench / Chatbot Arena ' s research continues to expose the verbosity bias, and Judge tends to prefer longer, more answer-like texts;Self-Preference Bias in LLM-as-a-Judge It suggests that Judge may also systematically prefer an answer more like his own distribution; and JudgeBench But it is more acutely reminded — the same as human preferences is not the same as being able to rule on facts and logic.

Put these deviations back on the main line and the conclusion is very clear:PARK LLM-as-Judge could be the proxy review panel for an open mission, but in high-risk or factual missions it would not be appropriate to act as a single final reward source, preferably in conjunction with retrieval, tools, rules, manual tests. If the output of the Judge is a black box total, then it's essentially not fundamentally different from the three-phase diagnosis -- it's a approximation function that you can't see the internal structure. To stabilize the judge, not just to change a bigger model, but to give it a structure.

Phase 6: Structured Judge - Reward interface from total to multi-dimensional calibration

The core issues revealed during the previous phase can be summarized in one sentence:A total score is too black and too rough. Judge knows how to be good and bad, but let it just export a pointwise scalar, and the scale will drift, the dimensions will collapse, and the red line will be diluted on average. The idea of solving this problem is not to replace the judge, but to add a structured rating framework to it -- that is, rubric.

LLM-Rubric The first route to walk is to break the evaluation target into multiple independent dimensions, allowing LLM to distribute the probability of output along each dimension instead of a point estimate, and then map the scale of the true assessor with calibration layers. What it does is it turns the judge from a Black Box Ratter to a Structured Signal Ripper, plus a calibrated map layer. Once here, Rubric is not just a score sheet, but a design for a structured interface for reward. In engineering,Prometheus 2 Further, it is stated that Judge is not a small tool for a temporary set of prompts, but rather an independent system component that needs to be specially trained, disciplined and evaluated.

The real push of rubric-driving judge to the end of the reward interface, is Rubrics as RewardsI'm sorry. Its core judgment is clear: if the mission has no single answer, but can write sufficiently clear rubric, then reward can go straight along rubric.

It does three things technically: automatically generate rubric cliterion with a strong LLM conditioned on a gold reference answer; empowers cliterion by weight (Essential / Important / Optional / Pitfall); and selects the aggregation method, turning cliterion-level judgment into a reward scalar. The experiment consistently showed that hidden aggregates (all critterion given to the judge as a whole) were better than the apparent article-by-article weighted sum.

The more practical discovery is thatRubric guidance helps smaller judge close the gap with stronger direct scringI'm sorry. The paper discussed the jurisprudence of GPT-4o-mini and Qwen-32B/14B/7B, rather than simply betting on a single large model. This means a potential cost advantage for the scene of RL training, where there is a large number of calls to the judge. For full technical details of Rubrics as Rewards, see appendix.

This design addressed some of the core pains of Park Judge:Explanatory- Each score can be traced back to specific criterion;Red Lines to Insert— Security violations, factual errors can be judged as Essential or Pitfall alone, leaving red line items visible; if it is to be prevented from being offset by other dimensions, it will also require hard Gate, veto or lexicographic analysis;Cost control- Weak models with rubric are enough. Of course, the ceiling is clear:The rubric is poorly written, and the quality of the reward is immediately reduced. The focus of rewarding moving from training a better RM to writing a better rubric -- the latter is essentially a problem of specification, not an ML.

At this stage, the default for reward in the open task is clear:When the verifier does not cover, the stablest form of reward is not to let the judge give a total score, but to write the rubric definition of behavioral specifications, then to allow the judge to structure the output along the dimensions, and finally to aggregate the cliterion-level signatures into reward. But even if this is done, when the trajectory changes, candidate variants, differences are fine, pointwise ratings are still easy to show -- how much is uncertain, and who can feel better, and how much.

Phase 7: Arena RL - When structured scores are not enough, use relative ranking Bottom

If all the stages ahead are understood as how to make better points, then... ArenaRL The point of moving forward is that it starts to wonder whether the open anent should use pointswise points. Once the trajectory changes, candidate variants, differences become subtle, pointwise scalar reward – even if structured through rubric – is still easy to see discrimination: judge can feel who is better, but it is not sure how much, let alone cross-mission, cross-tracking sustains a uniform scale. We're talking about pairwise at the beginning of this paper, and now we're turning him over.

A three-tiered approach to the ArenaRL resolution project solves a specific problem.

  1. Process-aware pairwise evaluation: Not a single track, but a judge who compares the two tracks with who, while focusing on the logic of the reasoning chain, the call of tools and the steps in the middle, and a two-way rating to eliminate the location deviation.
  2. Seeded single-elimination tournament: Pre-order the race with anchor tracks to build the phase-out contest, and compress the order of the N-canner from the full cycle of O(N2) to the O(N) subcomparison.
  3. Quantile-based reward mapping: Converts the discrete ranking to a unipolar fraction, which is an advantage signal for policy objective, without having to cross the watch to maintain historical status.

For full technical details of the three-tier Pipeline (bi-directional rating mechanism, seed phaseout process, comparison of quantile vs. Elo vs. Bradley-Terry, experimental figures), see appendix.

The limitations of this package must also be clear: the curnament remains highly dependent on the quality of the judge, the cost is still higher than the direct pointwise scring, and the quantile mapping can only produce local relative signals and cannot give absolute quality judgement across the bat. But ArenaRl is very clear about the most important inspiration I have:In open-ended anent, the focus of rewarded design is no longer on writing a seemingly elegant flat score, but on how to get the minor but critical relative differences between candidates to the optimizers more reliably.

So, what exactly should we do about it? Out

And here, the old question is actually a different question: instead of asking what the total score is, the first question is what is on the mission to be monitored for. Single-wheel open answers, summaries and dialogues are naturally more appropriate to start with a pairwise or listwise preference, as it is relatively more stable than absolute scales; multi-target assistant missions should break down the dimensions of helpful, harmmes, truthfulness before deciding how to aggregate them; and mathematical, code and tool calls for such tasks, giving priority to the veritier-backed outcome distribution of credit in the chain, provided that core results are validated.

The problem of long chains is similar: only 0/1 is given to swallow the value of the intermediate process, but the learning PRM and the process verifier are not the same signals, the former being the quality of the learning process, and the latter being the verification step. Open angent tracks go further, with candidate differences often fine and insulated, and first a stable pointwise score is easily hallucinogenic; if the system has multiple candidate tracks, leaving the rancing structure and sending it to optimizers with pairwise or tournament, it is usually more stable than a premature one-line fraction. As for individualized or multi-user products, it is also recognized that different users do not necessarily share the same utility, rewarded features or context-condited rewards tend to be closer to real targets than a global average preference.

So, the sequence of reward design is actually a roll-over of risk. Pull out what cannot be offset by even-sum: illusions, the call for ultra vires tools, security violations, key formatting errors, costs and time-lapse borders should not be compensated by “good overall performance”. And look at what can be directly verified, what can be compared, which needs a judge or a rubric agent. The rule reward exposes the problem of under-specification in an open environment, and fixed RMs are dug out of OOD high-level subregions, LLM as Judge with long preferences, position preferences and self-prejudice; these risks are not only emerging after training, but are buried at the moment that reward is written as a target. A more systematic dismantling will be on the next page. Four high-end scenarios for Reward Hacking

So, reward should not be seen as a model for a line of fractions. It is more firmly understood that it is a production line: the Hard limits are responsible for securing the unassessed borders, and the core task signature is from the verifeier, pairwise ranking or rubric, procascination signatureal is responsible for the credit allocation in the long chain, with marginal amendments only for costs, time extensions, redundancy tools and readability. Multiple candidate sequencing, soft preferences, uncertainty and user differences are erased before entering reward, and the training is not brought back.

Back to the initial metaphor: reward is not a line of fractions that a model has thrown out, but a production line. This article went all the way from Openai's RLHF to Arena RL, which is really always the same thing: How upstream objects, mid- and downstream interfaces of this production line were rewritten step by step. After the structure of the reward production is clarified, the next one is a more dangerous question: And these rewards, once actively searched by optimizers, are hacked; and we continue to discuss how to consume these signals.

Appendix: Rubrics as Rewards Technical Details

Rubrics as Rewards Three key things were done technically, and here we go.

The first thing is the automatic generation of rubric. The paper does not ask you to write every critter, but to use a strong LLM (e.g. GPT-4o or o3-mini) to automatically generate 7-20 self-contained critters for each prompt, subject to a gold reference answer. Each criterion is essentially a binary verification check--"Did the answer clearly set out the core concept?""Are common logical traps avoided?""Has effective evidence of support been provided?"I'm sorry. Here's an important experiment:Rubric, anchored with a gold reference, has a significantly higher mass than unreferenced version, because reference answers provide a specific behavioural bar for the generation of the Cliterion, not for LLM to speculate"What should I look like?"。

The second thing is the hierarchy weight of the criterion. The resulting criterion is not equal, but is divided into four categories:Essential(Necessary, e.g. core logical correctness),Important(Immediate, e.g., sufficiency of arguments),Optional(plus sub-items, such as expression of grace) and Pitfall(Traps/tips, such as overdiagnosis, logical cycles). This classification directly determines the weight of each contribution in the final reward, and allows for the visible exposure of the red line items (Essential or Pitfall); if the red line items are to have hard binding effects, the hardgate, veto or dictionary rules need to be added to the polymer layer.

The third thing is the choice of a convergence approach. The paper explored two paths simultaneously.Visible Aggregation:Judge judge each critterion article by article (yes/no or fraction), then weight sum $R = \sum_i (s_i \times w_i) / \sum_i w_i$ Getting the final reward - the most interpretable but the weight-share is fragile.Invisible Aggregation: all criterion and weights are given together to the judge to export a 0-1 fraction of the whole after understanding all the criteria — more flexible and able to capture the interaction between criterion. The results of the experiment are consistent.The invisible polymers work better., the relative increase was 31% on HealthBench and 7% on GPQA-Diamund.

In integration with RL training, Rubrics as Rewards can use the rubric judge score as a group-based policy optimation reward: sample of each propt in the details of the the thesis training k=16 Candidates, rating each answer with rubric judge, then construct relative advantage with group statistics, access GRPO / PPO-crip-like strategies to update the target. The main difference between this process and the standard GRPO is that reward no longer comes from a lost RM or simple verifier, but from a structured assessment of the rubric judge.

Minimum viable physical path:

  1. And even if only part of the solution is a reference.
  2. Automatically generate rubric with a strong LLM conditioned on reference answers, 7-20 criterion per prompt.
  3. Empowering the classification of critterion by Essential / Important / Opportal / Pitfall.
  4. Select hidden aggregates (recommended) or visible aggregates.
  5. Access to GRPO training cycle, replacing traditional RM ratings with rubric judge ratings.

The whole process does not require any training in reward mode, rewarding engraining to the design and iterative of the rubric.

Appendix: Technical Details of Arena RL

ArenaRL The three-story Pipeline is being expanded as follows.

First tier: Process-Aware Pairwise Evaluation. Arena RL doesn't score a single track, but lets Judge go and compare two tracks to who's better. Its pairwise is more than just looking at the end result -- Judge, while at the same time concerned about the logical closeness of the CTT reasoning chain, the accuracy of the tools being used and the rationality of the intermediate steps, which is called processe-aware. This design is directed at a common problem in the open environment: The two tracks may have the same end results, but one arrived through reasonable planning, the other just happened to be right, and it is difficult to distinguish between pointswise, but more so. And more importantly, every track is made.Bi-directional Scorring: Put A in front of B and evaluate back and then exchange the order to eliminate the position deviation of LLM Judge.

Second floor: Feeded Single-Elimination Tournament. With the power of comparison, the next question is how to use it to handle the N-canoes. The simplest approach is to compare the Round-Robin with each other, but this requires O(N2) comparisons — costs are unacceptable. The solution for ArenaRL is Seeded Single-EliminationThe lided version of the paper needs to be available for the first round of the event. 2N-2 The next comparison is still O(N). Two steps: first.The ecstasy decodes to generate a baseline trajectory as anchor (anchor)This anchor is used to pre-sequence all candidates in a quick-sequence sequence, which prevents high-quality samples from colliding at an early stage and being eliminated; and then to build a fork tree based on seed sequencing, with two pairs of each round to be a pairwise evaluation and a winner to be promoted. On the Open-Travel benchmark, the Rund-Robin precision of the O(N) Tornament Program and O(N2) was close to 32.5% vs 32.9%, but costing was reduced by one to two orders of magnitude.

Third floor: Quantile-Based Reward Mapping. Tournament output is a discrete group of rankings, but RL optimizers require continuous advantage signatureal. ArenaRLBitmaps: First by ranking rank reward = 1 - Rank/(N-1), standardized to an advantage signal for clipped policy objective. Why did you choose Quantile / rank-based happening instead of Elo or Bradley-Terry? First, it is independent of each bat, and does not need to be a cross-batch to maintain historical status; second, relative ranking is better for assessing noise; and third, efficient calculation, sorting O (N log N) linear mapping is sufficient. Elo needs an iterative solution, sensitive to noise, and Bradley-Terry's data in the small bat is not reliable enough to match.

Serial three.resolution settlementresolution of conflictsAdded value to the calculation(O(N) vs O(N²));Quantile happening solver(Dispersion rankings change to continuous signals). The Open-Travel benchmark is 16.4 per cent for SFT baseline, 16.4 per cent for traditional GRPO and 41.8 per cent for Arena RL, an increase of 155 per cent.

Appendix: Can environmental consequences be directly monitored when reward is difficult to write

I still have this part in the appendix, instead of being part of the main line, because it's more like a reminder of the reward boundary than a necessary phase of the reward main line itself.Agent Learning via Early Experience The important thing is not to declare reward obsolete, but to remind us that reward is just a form of compression of environmental information, not the only way to monitor it.

The core approach of this line is to proceed from the status point of the actual projectory, to conduct controlled exploration of alternative actions, to actually execute them, and to use the grounded future status or monitoring. It is valuable that the subject of oversight is no longer just “how much is to be scored at this point”, but that it is to take back the environmental consequences themselves. This is why the information density of such methods tends to be higher than a single step by hand.

But I still put it in the appendix because it's more like telling us where the boundary of reward is: reward is never the target itself, nor is it the only way to encode environmental information. When reward is difficult to write, environmental consequences can certainly be intermediate layers of higher information density; but as long as we are discussing reward design, the main line remains concerned about how to turn the target into a secure signal interface as possible, rather than completely bypassing the issue of reward.

Appendix: When Judge has more than one - multi-smart debate and Evaluation Act

The LLM as Judge, which we're talking about on the main line, is basically a single intelligence body that is assessing whether it is simple, structured or structural. But... ChatEval(ICLR 2024) It is natural to ask: why cannot LLM assessments rely on multi-profiling collaboration?

The core approach of ChatEval is to build a multi-smart adjudicator. Several LLM parties have set different roles (e.g., ordinary readers, critics, news writers, psychologists, scientists) to debate the same text independently, exposing differences that a single perspective may ignore in the discussions and ultimately evaluating the results of the output by voting or by scoring the average. There are a few details of the design that deserve attention:

Role diversity simulates human-specified differences. The systemic differences between the different actors and the human beings who are marked in the real human scene are the same. Single Judge, even strong, can only represent one perspective; multi-role collaboration tries to simulate the human assessment process with AI"Diversity creates greatness."The characteristics of this are the idea of collective wisdom and cognitive synergy in sociology.

The design space for communication strategies. ChatEval also considered three multi-smart communication strategies: sequential chorus (agent speaking in turn, the latter seeing the full output of the former), a scramble of speeches, and chat history with a wrapper -- The latter has allowed follow-up angent to read only the summary of the previous discussion, not the full record, and to effectively control the length of the context. This design option is important in actual deployment because the cost of multi-wheel debate token will expand rapidly.

No consensus is required, but convergence. In multi-smart debate, all parties are not ultimately required to agree. The final results can be produced by voting or by a score of an average total number of participants. The design choice is pragmatic: enforcement of consensus may, on the contrary, reduce valuable divisive messages.

The results of the experiment also support this determination: on the TopicalChat benchmark, the multi-smart body discussion in ChatEval was 0.57 relevant to Kendall Tau, which is human judgment, while single GPT-4 evaluator was 0.52. Simple ensemble (albeit many times the average of independent ratings) does not significantly enhance the effect, which is largely due to the natural language interaction itself: the argument, challenge and addition in the debate provide additional signals.

ChatEval is placed in an appendix rather than in the main line because the multi-smart debate assesses more of the present than a direct component of the assessment methodology. It suggests that LLM as Judge may be an important middle ground between rule-based and human feedback, especially in the open generation of mission assessments.

From this perspective, the future of the assessment is not just a choice for a single contact debate and multi-smart body, but rather the construction of an evaluation ant with tools, multiple rounds of interaction and complex structures. ChatEval's multi-smart body debate was an early exploration of the road, but it also suggested another possibility: Multi-intellectual synergetics can mitigate the problem of a single judge ' s degeneration-of-thought, because the challenge and rebuttal in the debate naturally constitute a confrontational self-correction mechanism.

References

Core entrance

Extending reading

  • Title: The Evolution of Reward Design: From RLHF to RLVR
  • Author: Hyacehila
  • Created at : 2026-03-19 12:00:00
  • Link: https://hyacehila.github.io//blog/2026/03/19/reward-design-evolution-from-rlhf-to-rlvr/
  • License: This work is licensed under CC BY-NC-SA 4.0.
Comments