forward_backward call, backed by SkyRL. Hosted Tinker does not serve GSPO natively: reproducing it there costs a 2-pass round-trip (fetch logprobs, compute the loss client-side, ship the reweighted batch back). rlcli gives you the fused 1-pass path locally, and our benchmark measures a 23% step-time reduction and 29.8% throughput gain on Qwen3-4B-Instruct-2507.
It is also the missing front door for the whole stack: the official tinker CLI has no train verb, and SkyRL has no CLI. rlcli wires them together so you can serve a model, run SFT or RL, and sample from checkpoints without leaving your terminal.
This package is
rlcli on PyPI. Do not confuse it with rl-cli (Runloop’s CLI), which is a different tool.Why fused GSPO
1-pass vs 2-pass
23% faster step time and 29.8% higher throughput than the round-trip pattern.
Full loss set
GSPO, DPPO, PPO-Critic, CISPO, PPO, and importance sampling on FSDP or Megatron.
Get started
Quickstart
Go from install to your first local training run in minutes.
Installation
Install rlcli, extras, and the prerequisites for your backend.
Architecture
How rlcli, SkyRL, the Tinker API, and the tinker client fit together.
CLI Overview
Browse all commands: serve, train, import, sample, and Tinker passthrough.
What you can do
Supervised Fine-Tuning
Run SFT on a messages JSONL dataset with LoRA and configurable hyperparameters.
Reinforcement Learning
Train with RL losses like PPO, GSPO, and CISPO on GSM8K using your local server.
Harbor Agent RL
Train agentic models with Harbor tasks in a sandboxed Docker or Modal environment.
Import Traces
Normalize chat dumps from OpenAI, Anthropic, or messages JSONL into training-ready data.