ML 20735 1
In multi-turn reinforcement learning (RL), your custom reward function decides what the model actually learns. A subtly wrong reward can quietly teach the wrong thing while every training curve looks healthy. Designing a reward that holds up over multi-turn, agentic tasks is one of the hardest parts of customizing Amazon Nova models. For multi-turn training, Amazon Nova Forge runs your reward logic in your own environment through its Bring Your Own Orchestration (BYOO) capability. You can focus on defining what a good outcome looks like while Nova Forge coordinates rollouts, message passing, and conversation state across turns. Nova Forge also offers a serverless multi-turn RL option, now generally available, for teams that prefer not to manage that environment. This post uses the BYOO path.
Amazon Nova offers multiple customization approaches, with reinforcement fine-tuning (RFT) standing out because it can teach models the behaviors you want through iterative feedback. RFT takes a different approach from supervised fine-tuning (SFT). Rather than requiring curated examples with annotated reasoning paths, it learns from evaluation signals on the model’s own outputs. Multi-turn RFT extends this to agents that act over a sequence of steps, such as calling tools, executing code, or recovering from a mistake. It optimizes cumulative reward across the whole trajectory rather than grading a single response. At the heart of RFT lies the reward function: the scoring mechanism that guides the model, and the part you design.
Figure 1 — Out-of-distribution (OOD) performance after equal-compute post-training from a shared checkpoint. RL improves OOD generalization across all task variants while SFT degrades. Adapted from Chu et al., 2025
This post focuses on the reward function itself: how to design a composite multi-turn reward that Group Relative Policy Optimization (GRPO) can learn from. This post also shows how to execute model-generated code safely inside the reward, and why to instrument each component so you can trust what training is learning. Part 1 of this series covers the Amazon SageMaker HyperPod and Nova Forge infrastructure. It also covers the training configuration that runs these rewards. We close with the pitfalls that can quietly collapse a reward, drawn from a real run where the highest-weighted component silently contributed no learning signal at all. We show how to catch them. The code throughout is illustrative. Use it as a starting point for your own reward implementation.
To follow along, you need the following:
RFT works by sampling completions from the current model and scoring them with a reward function. In Nova Forge, the reward function is a grader you write in code, and not a separately trained reward model. It can be a rule-based check that verifies the output (reinforcement learning with verifiable rewards), or it can call another large language model (LLM) to judge the response, an approach known as LLM-as-Judge.
RFT then adjusts the model weights to make higher-reward completions more likely. Nova Forge uses GRPO. For each conversation, GRPO uses the reward function to rank K model rollouts. GRPO uses the highest-ranked model completions to update the model according to the normalized reward (the advantage) of the batch. RFT with GRPO is a fundamental technique achieving noticeable performance gains over initial SFT.
A reward signal influences learning only through the variation it creates within a group. If a term takes the same value for every completion in a group, it contributes nothing to the advantage. It therefore contributes nothing to the gradient.
How your reward function runs with Nova Forge depends on the task. With single-turn RFT, you register the reward as an AWS Lambda function and point your recipe at it through reward_lambda_arn. Multi-turn tasks like the one in this post exceed what a single Lambda invocation supports. Multi-turn conversations and long-running scoring run past the 15-minute Lambda invocation limit. For these, Nova Forge uses BYOO. You set rollout.delegate: true and run your environment and reward logic in an environment container, for example on Amazon ECS. Nova Forge delegates each rollout to your environment. It then collects the completed episodes back for training. Your container manages the multi-turn interaction and conversation state: it runs the user simulator, executes code, and calls a verifier. It then returns an aggregate reward per sample (aggregate_reward_score), plus an optional list of per-component scores (metrics_list). Part 1 of this series covers this infrastructure and its AWS Cloud Development Kit (AWS CDK) deployment. This post focuses on the reward.
The training job generates candidate rollouts from the Nova model for each prompt. In a multi-turn task, a rollout is a full episode with a sequence of turns (a trajectory), not a single response. Your reward function receives each rollout and performs three steps:
metrics_list.aggregate_reward_score), which training turns into within-group advantages.Figure 2 — A single multi-turn rollout: Nova Forge delegates to your environment container, which asks the simulator or runs the committed code, then returns a reward score for GRPO
This cycle repeats over many training steps, progressively shaping the model to maximize cumulative reward across the whole sequence. The model optimizes toward whatever your reward actually rewards, which, as we show, is not always what you think you wrote.
Single scalar rewards are straightforward to game, and a single terminal reward is often too sparse to learn from in multi-turn tasks. Most production multi-turn rewards therefore combine three kinds of signal: outcome rewards, behavioral rewards, and penalties.
Episode-level (outcome) rewards capture whether the final artifact satisfied the goal. For example, did the unit tests pass, or did the workflow complete? They target the thing you ultimately care about, but they tend to be sparse and near-zero early in training.
Turn-level (behavioral) rewards capture whether the model exhibited the intermediate behavior you want, such as asking before acting, calling the right tool, or avoiding loops. They are best for shaping behavior the outcome reward is too sparse to teach, though they can be earned without real progress if not designed carefully. Penalties explicitly discourage a failure mode such as guessing, repeating, or stalling. They separate good and bad strategies so the optimizer sees a gradient.
Combine these so the model learns both the behavior and the outcome, without one component masking or starving the other. The rest of this post makes that concrete. We design a four-component reward for a real task and execute model-generated code safely inside it. Then we walk through the pitfalls that can collapse such a reward and how to fix them.
We built a multi-turn collaborative-coding task over 500 unique programming tasks. We trained Amazon Nova Lite 2.0 on it with multi-turn RFT, using GRPO with Low-Rank Adaptation (LoRA), on Amazon SageMaker HyperPod, implementing the reward inside a customer-managed environment container (the Nova Forge BYOO path).
The mechanics are as follows:
The design intent is that guessing produces wrong code, while asking surfaces the hidden detail and leads to correct code. “Ask first” should be forced by the task.
Make the target behavior directly and independently rewardable, and penalize the failure mode explicitly. For this task, the reward is a weighted sum of four components:
| Component | Weight | Definition |
correctness | 1.0 | fraction of hidden unit tests passing on the final code |
asked_before_coding | 0.6 | 1.0 if asked on turn 1 then committed; 0.6 if asked later then committed; else 0 (un-gated) |
guessed_immediately | 0.4 | penalty: -1.0 if the first turn is code with no question |
loop_penalty | 0.2 | -0.5 if the last two turns are more than 80% similar |
Two principles drive the design. First, un-gate the behavior you want: asked_before_coding is credited on its own, not conditioned on correctness, but it does require the model to eventually commit code, which closes the “ask forever, never answer” loophole. Second, penalize the failure mode: guessed_immediately makes guessing strictly worse than asking, which restores variation between strategies within a GRPO group, the variation the algorithm needs to produce a gradient.
Call these component scorers inside the reward handler in the environment container, and report each value through metrics_list:
The correctness component runs model-generated code against unit tests. Model output under RL is optimized through exploration, so treat it as not validated. The container runs in its own isolated execution environment, but you should still take precautions. Do not expose credentials or network to the generated code. Apply resource limits and run in a temporary directory. Use a per-run random sentinel so the model cannot forge the result by writing the expected marker to stderr. For execution that requires additional isolation, call a dedicated sandbox. This harness shows the pattern:
Also validate the number of tests actually run against the number expected, so the model cannot dilute the score with its own trivially-passing tests. For reward functions deployed in live environments, implement these security measures rather than treating them as optional.
Multi-turn reward design has a well-known set of failure modes. Reward hacking is where the model games a proxy instead of achieving the goal. Training instability is where updates diverge and entropy collapses or the Kullback-Leibler (KL) term blows up. Reward collapse is where the signal degenerates until within-group variation disappears and learning quietly stops. The first two usually announce themselves in transcripts or in loss and KL curves. Collapse is the dangerous one: aggregate reward, loss, and completion-length curves can all look healthy while a component you’re counting on contributes nothing. This section covers the two collapse failures that cost us the most time on this task, and how to catch them.
An earlier version of this reward gated the asking bonus behind correctness. You earned the asking reward only if the final code also passed. It also added an efficiency term that rewarded shorter conversations. Training collapsed. The model converged to guessing on turn one. The mean reward froze, and the GRPO advantage went to zero.
Two design errors caused it. First, the gate sat behind an unreachable condition. Correctness was near zero on these hard tasks, so the asking bonus almost never fired. The behavior we wanted to reward was invisible to the optimizer. Second, the efficiency term had a degenerate optimum. Fewer turns maximized it, so the policy collapsed onto a single, non-committal turn. Every completion looked alike, within-group variation vanished, and learning stopped.
The fix is the design in the previous section: un-gate the behavior you want, and penalize the failure mode explicitly. With both in place, distinct strategies keep producing distinct rewards within a group, which preserves the variance GRPO needs to learn.
When a reward component returns the same value for every completion in a GRPO group, its within-group variance is zero. As a result, it contributes nothing to the advantage or the gradient, even at the highest weight. The components that still vary keep aggregate reward, policy loss, advantage, and completion length looking healthy, so the curves never reveal it. One common cause in code rewards is a correctness scorer that returns 0 on every rollout because the harness never executes the model’s output. This can happen because of mismatched entry-point names, failed imports, or a setup error that makes every test fail before its assertions run. In our run, this is exactly what happened: the model’s clarifying-question rate rose from roughly 34–96 percent. Code correctness barely moved, because the correctness scorer was returning the same value on every rollout.
To catch a dead component, track each component’s within-group standard deviation, not the aggregate reward curve. Aggregate curves hide a dead channel behind the live ones. If that spread sits at or near zero, the component isn’t training, whatever its weight. The usual root cause in code rewards is a correctness scorer stuck at 0 because the harness never actually binds to and runs the model’s output. Fix that and confirm the spread becomes non-zero.
Several habits catch these failures, and would have caught ours on day one:
metrics_list, and track its mean and its within-group standard deviation. Any component with near-zero within-group variance contributes nothing to learning, regardless of its weight. You might dismiss a flat reward mean of 0.000 as “these tasks are just hard,” but a flat within-group variance is unambiguous. Automate this as a per-component advantage-variance panel so dead channels are flagged automatically, without manual inspection.correctness reward could not move the policy. If a behavioral shaping term dominates, the outcome term you care about may never get a gradient. Consider down-weighting a shaping term once it saturates, or up-weighting the outcome term.The training run and environment in this post use SageMaker HyperPod and Amazon ECS resources that incur cost while they run. When you finish experimenting, follow the teardown steps in Part 1 of this series to delete the SageMaker HyperPod cluster and the Amazon ECS environment, which stops the largest charges. Remove the rollout data and checkpoints from your Amazon S3 bucket if you no longer need them.
The reward function is the part of RFT you design, and it is where the subtle failures live. In your runs, the model may learn the behavior you train for while a term you care about contributes nothing to learning, with no aggregate metric revealing it. Better instrumentation, not a better algorithm, fixed the issue. Measure each component’s contribution to the advantage, read transcripts through the lens of the component you’re testing, and ablate what you claim is working. With a custom reward function on Amazon Nova Forge you have full control over the reward, which means the responsibility for getting it right is yours. For the infrastructure and AWS CDK deployment that make these runs reproducible, see Part 1 of this series.
Special thanks to Mahima Chaudhary for their review and contributions to this post.
My wife did this Cunk parody with a 3060 12gb and 32gb of system ram.…
In this article, you will learn what latent spaces are and how they serve three…
As concerns around data privacy in machine learning grow, the ability to unlearn, or remove,…
At a press conference outside Madison Square Garden, politicians, musicians, and privacy advocates argued for…
A tiny superconducting engine has successfully converted heat near absolute zero into useful work, demonstrating…
Among the many predictions about the future of artificial intelligence is that models will one…