Prerequisites
Before you begin, make sure you have:- A running local server. See
rlcli serve startfor details on launching SkyRL. - A dataset in messages JSONL format. See Dataset Format for the expected structure.
Step-by-step
1
Start the server
Launch a local SkyRL server with your base model and backend. For GPU training, use a torch backend such as Wait for the server to become healthy. You can check status with
fsdp:rlcli serve status.2
Prepare your dataset
Your dataset should be a JSONL file where each line contains a conversation object:If you have raw chat dumps from OpenAI or Anthropic, use You can also pipe imported traces directly into training:
rlcli import to normalize them first:3
Preview the training config
Use
--dry-run to print the full training configuration without starting training. This helps you verify flags, paths, and renderer selection before consuming GPU time.4
Run training
Remove Training logs are written to the default log path,
--dry-run to start the actual supervised fine-tuning job:~/.rlcli/runs/sl-<timestamp>.5
Inspect the run
After training, list available checkpoints:You can sample from a checkpoint using its
tinker://... path:Renderer selection
rlcli auto-detects the recommended renderer for your model via tinker-cookbook’smodel_info.get_recommended_renderer_name. If you are using an unknown or custom model, pass the renderer explicitly:
qwen3_instruct, llama3, and others provided by tinker-cookbook.
Next steps
- Learn about all
train sloptions - Convert chat dumps with
rlcli import - Explore reinforcement learning after SFT