Skip to main content
rlcli makes it easy to run supervised fine-tuning (SFT) on your own GPUs through a local SkyRL Tinker server. This guide walks you through starting the server, preparing your dataset, previewing the training configuration, and launching the actual run.

Prerequisites

Before you begin, make sure you have:
  • A running local server. See rlcli serve start for details on launching SkyRL.
  • A dataset in messages JSONL format. See Dataset Format for the expected structure.

Step-by-step

1

Start the server

Launch a local SkyRL server with your base model and backend. For GPU training, use a torch backend such as fsdp:
Wait for the server to become healthy. You can check status with rlcli serve status.
2

Prepare your dataset

Your dataset should be a JSONL file where each line contains a conversation object:
If you have raw chat dumps from OpenAI or Anthropic, use rlcli import to normalize them first:
You can also pipe imported traces directly into training:
3

Preview the training config

Use --dry-run to print the full training configuration without starting training. This helps you verify flags, paths, and renderer selection before consuming GPU time.
4

Run training

Remove --dry-run to start the actual supervised fine-tuning job:
Training logs are written to the default log path, ~/.rlcli/runs/sl-<timestamp>.
5

Inspect the run

After training, list available checkpoints:
You can sample from a checkpoint using its tinker://... path:

Renderer selection

rlcli auto-detects the recommended renderer for your model via tinker-cookbook’s model_info.get_recommended_renderer_name. If you are using an unknown or custom model, pass the renderer explicitly:
Available renderers include qwen3_instruct, llama3, and others provided by tinker-cookbook.

Next steps