benchmarks/bench_gspo.py script measures why fused GSPO on a SkyRL backend matters: it times the same frozen rollout batch on the same server two ways, once through the native fused path forward_backward(loss_fn="gspo") and once through the round-trip pattern hosted Tinker forces for losses it does not serve natively (forward to fetch logprobs, compute sequence-level GSPO weights on the client, then forward_backward(loss_fn="importance_sampling") with the reweighted batch).
What it measures
Path B reproduces the 2-pass wire pattern exactly. The client-side sequence-level clipped ratio computation is an approximation of GSPO gradients, but the quantity being measured (wall-clock per step) is the real cost of the 2-pass round-trip.
- A (native, 1 pass):
forward_backward(datums, loss_fn="gspo", loss_fn_config={"clip_low_threshold": 0.8, "clip_high_threshold": 1.2})thenoptim_step. - B (custom, 2 pass):
forward(datums, loss_fn="cross_entropy")to fetch per-token logprobs, compute sequence-level clipped importance ratios client-side, thenforward_backward(reweighted, loss_fn="importance_sampling")andoptim_step.
Latest result
Measured onQwen/Qwen3-4B-Instruct-2507 with --lora-rank 32, 16 sequences, 7,569 total target tokens per step, 8 timed steps after a warmup.
Fused GSPO cuts step time by 23.0% and lifts throughput by 29.8% on this configuration. The full JSON with per-step timings is written to
benchmarks/last_gspo_bench.json.
Run it yourself
1
Start a torch-backend server
Fused GSPO requires an FSDP or Megatron backend. See Backends & Losses for the loss matrix.
2
Run the benchmark
--steps (default 8), --max-tokens (default 512), --lora-rank (default 32).3
Read the summary
The script prints a
BENCH_RESULT line with the summary JSON and writes the full per-step record to benchmarks/last_gspo_bench.json.End-to-end learning: GSM8K with fused GSPO
Beyond wall-clock speed, fused GSPO on rlcli produces real learning on a single L4 GPU. The plot below isrlcli train rl --loss gspo on Qwen3-0.6B against GSM8K, GSM8K train accuracy per step.
Render both plots (bars and curve) from your own runs with benchmarks/plot_assets.py:
metrics.jsonl from the run directory (uses env/all/correct for the reward curve) and writes bench_gspo_bars.png plus curve_gsm8k_gspo.png.