Prerequisites
Fused losses such asgspo, dppo, and ppo_critic require a torch backend server running on Linux with CUDA. Use --backend fsdp or --backend megatron when starting the server. The JAX backend does not support these fused losses.
For a full matrix of which losses work on which backends, see Backends & Losses.
Step-by-step
1
Start a torch backend server
Launch SkyRL with Verify the server is healthy before proceeding:
fsdp (or megatron) and allocate GPUs:2
Preview the RL config
Use
--dry-run to inspect the training configuration before committing GPU time. Include your chosen loss and any loss-specific config:3
Run RL training
Remove rlcli validates your loss against the backend before training starts. If the loss is unsupported, you get a clear
--dry-run to start training:LossBackendError with a hint to switch backends.Loss guard
rlcli callsensure_loss_supported(loss, backend) before training begins. This fails fast with a clear error if your chosen loss is incompatible with the active backend.
- If rlcli manages the server (local
rlcli serve start), the guard knows the exact backend and rejects unsupported combinations immediately. - If you point at a remote or unmanaged
--base-url, the guard only rejects entirely unknown losses. Provide--backend fsdpas a hint when using a remote server so the guard can still validate fused losses.
Important warnings
The only built-in dataset choice for
rlcli train rl today is gsm8k. Use --eval-every N to run periodic evaluation during training.Next steps
- See the full Backends & Losses compatibility matrix
- Try Harbor Agent RL for sandboxed agent tasks
- Review all
train rlCLI options