GRPO-based LLM unlearning

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Reward success, RWKU forget scores, held-out completions, and training dynamics do not always tell the same story.

University of Valencia - September 2026

Pipeline overview for GRPO reward specification and evaluation

TL;DR

Getting a good unlearning score does not mean the model learned the behavior we wanted. Sometimes it forgets by refusing, sometimes it still leaks, and sometimes it finds weird ways to satisfy the reward. SFT warm-up helps larger models learn at all, but the reward still decides what kind of behavior they learn.

Abstract

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse.

We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, and training dynamics can point to different conclusions.

We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and a rubric reward that selects broad-topic answering with low semantic leakage under held-out evaluation.

Study Design

Question

What should an unlearned model do?

The target behavior is not just low leakage. For target-adjacent prompts, the model should avoid target-specific facts while still answering usefully at a broader topic level when possible.

Setup

Controlled GRPO comparisons

We run Qwen2.5-Instruct models at 0.5B, 1.5B, 3B, and 7B parameters. Across sizes, the study keeps the data source, LoRA pipeline, prompt splits, optimizer budget, and evaluation protocol as fixed as possible while changing reward specification and initialization.

Evaluation

Benchmark plus behavioral audit

RWKU metrics are interpreted alongside held-out completion audits and training dynamics.

Reward Specifications

Reward Intended endpoint Specification Cost
R0-Lex Suppress configured target patterns Lexical suppression baseline; relies on base-model behavior and KL regularization for usefulness. Low
R1-AntiRefusal Suppress leakage without simple refusal Adds explicit anti-refusal pressure to the R0-Lex lexical leakage signal. Medium
R2-Rubric Answer at a safe broader-topic level Uses a refusal-prefiltered rubric judge to prefer useful nearby-topic completions without target-specific leakage. High
R4-Refusal Contrastive refusal endpoint Rewards refusal-like behavior as a diagnostic contrast against useful-answer unlearning. Low

Main Findings

01

Reward proxies can be hacked into different endpoints.

R0-Lex can collapse to refusal-like behavior, while refusal-sensitive rewards can be satisfied by policy boilerplate or terse target-adjacent fragments. These are concrete optimized behaviors, not just noisy scores.

02

Forgetting scores can hide endpoint changes.

RWKU forget improvements can hide very different held-out behaviors: refusal collapse, residual leakage, classifier-aligned artifacts, or R2-Rubric broad-topic answering. This is clearest when benchmark deltas look comparable, but the behavioral audit separates the endpoints.

03

Utility and fluency separate the rewards.

R1-AntiRefusal and R2-Rubric most consistently preserve or improve fluency, while stronger suppressive or refusal-oriented endpoints can look better on forgetting but worse on useful generation.

04

SFT warm-up enables larger-model learning.

Without warm-up, the 3B and 7B runs often do not learn because rollout groups provide little useful reward variation. SFT expands support for non-target completions, enabling GRPO learning while leaving the endpoint reward-dependent.

Behavioral and Benchmark Diagnostics

Held-out Audit Cases

The held-out audit table combines the central R2-Rubric endpoint with three diagnostic phenomena: warm-start R0-Lex scale contrast, refusal collapse, and refusal-classifier hacking. Cells report median [Q1, Q3] rates across evaluated target entities; refusal/avoidance includes explicit refusal, inability claims, topic avoidance, and generic policy discussion.

What to look for. Read leakage together with prompt helpfulness, broad helpfulness, and refusal/avoidance: low leakage can mean useful broad answering, refusal collapse, or classifier-aligned shortcuts.

Model Init. Reward Lex. leak ↓ Sem. leak ↓ Prompt helpful ↑ Broad helpful ↑ Ref./avoid. Drift ↓ Neval
R2-Rubric broad-topic endpoint
0.5B Warm R2-Rubric 0.033 [0.033, 0.062] 0.008 [0.000, 0.029] 0.008 [0.000, 0.017] 0.983 [0.950, 1.000] 0.708 [0.633, 0.771] 0.000 [0.000, 0.000] 10
1.5B Warm R2-Rubric 0.058 [0.037, 0.117] 0.033 [0.017, 0.062] 0.025 [0.017, 0.067] 0.992 [0.958, 1.000] 0.608 [0.450, 0.779] 0.000 [0.000, 0.000] 10
3B Warm R2-Rubric 0.025 [0.000, 0.079] 0.017 [0.000, 0.046] 0.017 [0.000, 0.033] 0.983 [0.967, 0.996] 0.825 [0.762, 0.954] 0.000 [0.000, 0.000] 10
7B Warm R2-Rubric 0.017 [0.004, 0.017] 0.017 [0.004, 0.029] 0.008 [0.000, 0.046] 0.983 [0.933, 0.996] 0.775 [0.512, 0.863] 0.000 [0.000, 0.000] 10
Warm-start R0-Lex scale contrast
0.5B Warm R0-Lex 0.000 [0.000, 0.142] 0.025 [0.000, 0.217] 0.017 [0.000, 0.296] 0.058 [0.000, 0.442] 0.983 [0.504, 1.000] 0.000 [0.000, 0.000] 10
1.5B Warm R0-Lex 0.000 [0.000, 0.033] 0.008 [0.000, 0.058] 0.008 [0.000, 0.108] 0.017 [0.000, 0.425] 0.992 [0.604, 1.000] 0.000 [0.000, 0.000] 10
3B Warm R0-Lex 0.233 [0.092, 0.463] 0.325 [0.054, 0.388] 0.200 [0.050, 0.375] 0.600 [0.196, 0.846] 0.733 [0.517, 0.883] 0.000 [0.000, 0.000] 10
7B Warm R0-Lex 0.350 [0.100, 0.746] 0.500 [0.138, 0.621] 0.558 [0.104, 0.725] 0.958 [0.883, 1.000] 0.183 [0.075, 0.558] 0.000 [0.000, 0.000] 10
Refusal collapse
0.5B Cold R0-Lex 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 1.000 [1.000, 1.000] 0.000 [0.000, 0.000] 10
1.5B Cold R0-Lex 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 1.000 [1.000, 1.000] 0.000 [0.000, 0.000] 10
Refusal-classifier hacking
0.5B Cold R1-AntiRefusal 0.050 [0.021, 0.796] 0.000 [0.000, 0.767] 0.000 [0.000, 0.696] 0.000 [0.000, 0.496] 1.000 [0.250, 1.000] 0.000 [0.000, 0.000] 10
1.5B Cold R1-AntiRefusal 0.000 [0.000, 0.029] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 1.000 [1.000, 1.000] 0.000 [0.000, 0.033] 10
7B Warm R4-Refusal 0.408 [0.212, 0.617] 0.067 [0.037, 0.096] 0.008 [0.000, 0.017] 0.092 [0.042, 0.150] 0.950 [0.800, 0.983] 0.000 [0.000, 0.000] 10

Key takeaway. R2-Rubric is the clearest held-out broad-topic endpoint, while the diagnostic rows show how low leakage can also arise from scale-dependent R0 behavior, refusal collapse, or refusal-classifier shortcuts.

Selected RWKU and Held-out Endpoint Rows

This compact slice compares RWKU benchmark scores with matching held-out endpoints for 0.5B and 3B warm-start runs.

What to look for. The RWKU columns summarize benchmark forgetting and utility, while the held-out columns show whether the endpoint is broad answering, leakage, or refusal/avoidance.

Model Reward RWKU benchmark Held-out endpoint N
Δ Forget ↓ Δ Neighbor ↑ Factuality ↑ Fluency ↑ Broad helpful ↑ Sem. leak ↓ Ref./avoid.
0.5B R0-Lex -0.107 [-0.148, -0.051] -0.075 [-0.091, 0.002] 0.227 [0.194, 0.250] 5.787 [5.450, 6.125] 0.058 [0.000, 0.442] 0.025 [0.000, 0.217] 0.983 [0.504, 1.000] 10
0.5B R1-AntiRefusal -0.085 [-0.147, -0.051] -0.024 [-0.070, 0.073] 0.201 [0.191, 0.228] 6.677 [6.462, 6.907] 0.908 [0.842, 0.963] 0.575 [0.408, 0.783] 0.200 [0.079, 0.267] 10
0.5B R2-Rubric -0.082 [-0.143, -0.022] -0.055 [-0.081, -0.039] 0.176 [0.142, 0.202] 7.054 [6.802, 7.174] 0.983 [0.950, 1.000] 0.008 [0.000, 0.029] 0.708 [0.633, 0.771] 10
0.5B R4-Refusal -0.110 [-0.137, -0.087] -0.041 [-0.085, -0.013] 0.206 [0.189, 0.231] 3.718 [3.524, 3.844] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 1.000 [1.000, 1.000] 10
3B R0-Lex -0.316 [-0.376, -0.173] -0.073 [-0.139, -0.066] 0.193 [0.178, 0.208] 6.922 [6.812, 7.059] 0.600 [0.196, 0.846] 0.325 [0.054, 0.388] 0.733 [0.517, 0.883] 10
3B R1-AntiRefusal -0.217 [-0.270, -0.191] -0.057 [-0.100, -0.041] 0.153 [0.131, 0.177] 7.050 [6.950, 7.261] 0.950 [0.696, 0.996] 0.450 [0.183, 0.717] 0.325 [0.167, 0.588] 10
3B R2-Rubric -0.218 [-0.341, -0.178] -0.100 [-0.165, -0.033] 0.128 [0.124, 0.138] 7.154 [7.082, 7.248] 0.983 [0.967, 0.996] 0.017 [0.000, 0.046] 0.825 [0.762, 0.954] 10
3B R4-Refusal -0.261 [-0.336, -0.153] -0.129 [-0.197, -0.049] 0.238 [0.195, 0.273] 6.168 [6.050, 6.357] 0.067 [0.004, 0.217] 0.000 [0.000, 0.017] 1.000 [1.000, 1.000] 10

Key takeaway. RWKU captures benchmark forgetting and utility trade-offs, but it does not reveal whether the score came from broad-topic answering, residual leakage, or refusal-oriented behavior.

Training Diagnostics

Training diagnostics separate optimization failure from reward-endpoint failure.

What to look for. Reward variation, active groups, and KL show whether GRPO had room to update before the endpoint audits say what behavior was selected.

Training diagnostics for Qwen2.5 0.5B
Qwen2.5 0.5B training diagnostics.
Training diagnostics for Qwen2.5 1.5B
Qwen2.5 1.5B training diagnostics.
Training diagnostics for Qwen2.5 3B
Qwen2.5 3B training diagnostics.
Training diagnostics for Qwen2.5 7B
Qwen2.5 7B training diagnostics.

Key takeaway. SFT warm-up supplies a learning signal for larger models; after warm-up, the reward specification still determines which behavior is reinforced.

What This Means

In GRPO unlearning, the reward is not just an optimization detail. It defines the behavioral endpoint that can be reinforced from completions already in policy support.

Reliable evaluation should combine benchmark deltas, held-out audits, and training dynamics. Any one view can look healthy while another reveals refusal collapse, residual leakage, reward hacking, or missing learning signal.

BibTeX

@misc{balbastre2026empiricalstudyrewardspecification,
  title={An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning},
  author={Rubén Balbastre and Juan Manuel Orduña and Mariano Pérez},
  year={2026},
  eprint={2608.17804},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  doi={10.48550/arXiv.2608.17804},
  url={https://arxiv.org/abs/2608.17804}
}