Skip to main content
The benchmarks/bench_gspo.py script measures why fused GSPO on a SkyRL backend matters: it times the same frozen rollout batch on the same server two ways, once through the native fused path forward_backward(loss_fn="gspo") and once through the round-trip pattern hosted Tinker forces for losses it does not serve natively (forward to fetch logprobs, compute sequence-level GSPO weights on the client, then forward_backward(loss_fn="importance_sampling") with the reweighted batch).

What it measures

Path B reproduces the 2-pass wire pattern exactly. The client-side sequence-level clipped ratio computation is an approximation of GSPO gradients, but the quantity being measured (wall-clock per step) is the real cost of the 2-pass round-trip.
  • A (native, 1 pass): forward_backward(datums, loss_fn="gspo", loss_fn_config={"clip_low_threshold": 0.8, "clip_high_threshold": 1.2}) then optim_step.
  • B (custom, 2 pass): forward(datums, loss_fn="cross_entropy") to fetch per-token logprobs, compute sequence-level clipped importance ratios client-side, then forward_backward(reweighted, loss_fn="importance_sampling") and optim_step.
The batch is real rollout data: 16 GSM8K-style prompts sampled once from the model at startup with variable output lengths, then frozen and reused for every step of both paths so the comparison isolates the wire pattern.

Latest result

Measured on Qwen/Qwen3-4B-Instruct-2507 with --lora-rank 32, 16 sequences, 7,569 total target tokens per step, 8 timed steps after a warmup. Fused GSPO cuts step time by 23.0% and lifts throughput by 29.8% on this configuration. The full JSON with per-step timings is written to benchmarks/last_gspo_bench.json.

Run it yourself

1

Start a torch-backend server

Fused GSPO requires an FSDP or Megatron backend. See Backends & Losses for the loss matrix.
2

Run the benchmark

Optional flags: --steps (default 8), --max-tokens (default 512), --lora-rank (default 32).
3

Read the summary

The script prints a BENCH_RESULT line with the summary JSON and writes the full per-step record to benchmarks/last_gspo_bench.json.

End-to-end learning: GSM8K with fused GSPO

Beyond wall-clock speed, fused GSPO on rlcli produces real learning on a single L4 GPU. The plot below is rlcli train rl --loss gspo on Qwen3-0.6B against GSM8K, GSM8K train accuracy per step. Render both plots (bars and curve) from your own runs with benchmarks/plot_assets.py:
It reads metrics.jsonl from the run directory (uses env/all/correct for the reward curve) and writes bench_gspo_bars.png plus curve_gsm8k_gspo.png.
Single-GPU FSDP is verified clean for fused GSPO. Multi-GPU Megatron DP configs currently hit an upstream packing-order bug (see Backends & Losses).