Skip to main content
rlcli lets you run reinforcement learning (RL) training on the gsm8k benchmark using a variety of loss functions, including fused losses like gspo. This guide covers starting the right backend, selecting a supported loss, and launching a training run with config preview and loss guard validation.

Prerequisites

Fused losses such as gspo, dppo, and ppo_critic require a torch backend server running on Linux with CUDA. Use --backend fsdp or --backend megatron when starting the server. The JAX backend does not support these fused losses. For a full matrix of which losses work on which backends, see Backends & Losses.

Step-by-step

1

Start a torch backend server

Launch SkyRL with fsdp (or megatron) and allocate GPUs:
Verify the server is healthy before proceeding:
2

Preview the RL config

Use --dry-run to inspect the training configuration before committing GPU time. Include your chosen loss and any loss-specific config:
3

Run RL training

Remove --dry-run to start training:
rlcli validates your loss against the backend before training starts. If the loss is unsupported, you get a clear LossBackendError with a hint to switch backends.

Loss guard

rlcli calls ensure_loss_supported(loss, backend) before training begins. This fails fast with a clear error if your chosen loss is incompatible with the active backend.
  • If rlcli manages the server (local rlcli serve start), the guard knows the exact backend and rejects unsupported combinations immediately.
  • If you point at a remote or unmanaged --base-url, the guard only rejects entirely unknown losses. Provide --backend fsdp as a hint when using a remote server so the guard can still validate fused losses.

Important warnings

gspo multi-GPU / Megatron buggspo has a known upstream packing-order bug on multi-GPU and Megatron configurations (NovaSky-AI/SkyRL#2043). Single-GPU FSDP is verified clean. rlcli prints a warning to stderr when you use --loss gspo.
The only built-in dataset choice for rlcli train rl today is gsm8k. Use --eval-every N to run periodic evaluation during training.

Next steps