Skip to main content
This guide takes you from a fresh environment to a completed supervised fine-tuning run with rlcli. You will install the CLI, start a local SkyRL Tinker server, prepare a dataset, run training, and sample from the resulting checkpoint.
1

Install rlcli

Install rlcli with the [train] extras so the train and serve commands are available. You can use pip or uv:
You need Python 3.11 or later and uv installed for rlcli serve to work. See Installation for full details.
2

Start the local server

Start a SkyRL Tinker server on your local GPUs. This example uses the FSDP backend with 8 GPUs:
If you do not have CUDA or you are on macOS, use --backend jax instead. JAX runs on CPU and macOS, but it does not serve fused losses such as GSPO or DPPO. See Backends & Losses for the full matrix.
The command blocks until the server is ready (default timeout is 900 seconds). Once it returns, the server is running in the background.
3

Check server health

Verify the server is healthy before you submit training jobs:
You should see a JSON object that includes "healthy": true.
4

Prepare a dataset

Create a JSONL file where each line is an object with a messages array. For details, see Dataset Format.
conversations.jsonl
5

Run supervised fine-tuning

Train with the default SFT settings:
By default, logs are written to ~/.rlcli/runs/sl-<timestamp>/. The run uses LoRA rank 32, batch size 8, learning rate 1e-4, and max length 2048.
6

Sample from the checkpoint

After training saves checkpoints, list them and sample:
You can also sample from the base model directly with --model instead of --checkpoint.

Next steps

RL Training

Move from SFT to RL with PPO, GSPO, or CISPO on GSM8K.

Harbor Agent RL

Train agentic models with Harbor tasks in a sandboxed environment.

CLI Overview

Explore all commands, flags, and passthrough behavior.