Skip to main content
rlcli is a Python CLI that runs fused RL losses like GSPO on your own GPUs in a single forward_backward call, backed by SkyRL. Hosted Tinker does not serve GSPO natively: reproducing it there costs a 2-pass round-trip (fetch logprobs, compute the loss client-side, ship the reweighted batch back). rlcli gives you the fused 1-pass path locally, and our benchmark measures a 23% step-time reduction and 29.8% throughput gain on Qwen3-4B-Instruct-2507. It is also the missing front door for the whole stack: the official tinker CLI has no train verb, and SkyRL has no CLI. rlcli wires them together so you can serve a model, run SFT or RL, and sample from checkpoints without leaving your terminal.
This package is rlcli on PyPI. Do not confuse it with rl-cli (Runloop’s CLI), which is a different tool.

Why fused GSPO

1-pass vs 2-pass

23% faster step time and 29.8% higher throughput than the round-trip pattern.

Full loss set

GSPO, DPPO, PPO-Critic, CISPO, PPO, and importance sampling on FSDP or Megatron.

Get started

Quickstart

Go from install to your first local training run in minutes.

Installation

Install rlcli, extras, and the prerequisites for your backend.

Architecture

How rlcli, SkyRL, the Tinker API, and the tinker client fit together.

CLI Overview

Browse all commands: serve, train, import, sample, and Tinker passthrough.

What you can do

Supervised Fine-Tuning

Run SFT on a messages JSONL dataset with LoRA and configurable hyperparameters.

Reinforcement Learning

Train with RL losses like PPO, GSPO, and CISPO on GSM8K using your local server.

Harbor Agent RL

Train agentic models with Harbor tasks in a sandboxed Docker or Modal environment.

Import Traces

Normalize chat dumps from OpenAI, Anthropic, or messages JSONL into training-ready data.

One-liners

rlcli is open source under the Apache-2.0 license.