Skip to main content
rlcli provides four training subcommands under train: sl for supervised fine-tuning, rl for reinforcement learning on math benchmarks, harbor for sandboxed agent RL, and opsd for on-policy self-distillation from a prompts file. Each subcommand sends requests to the Tinker API at --base-url and automatically sets TINKER_API_KEY=tml-dummy.

train sl

Supervised fine-tune a model on a messages JSONL dataset.

Flags

string
required
Path to a JSONL file, or - to read from stdin.
string
required
HuggingFace model name or path.
string
default:"http://localhost:8000"
Tinker API server URL.
string
Renderer to use. Defaults to the recommended renderer for the model.
int
default:"8"
Training batch size.
float
default:"1e-4"
Learning rate.
int
default:"2048"
Maximum sequence length for training examples.
int
default:"1"
Number of training epochs.
int
default:"32"
LoRA rank.
int
default:"20"
Save a checkpoint every N steps.
float
default:"0"
Fraction of data to hold out for evaluation.
string
default:"~/.rlcli/runs/sl-<timestamp>"
Directory for training logs.
flag
Print the training plan without executing.

Example

Common notes

  • TINKER_API_KEY is set to tml-dummy automatically if unset.
  • TINKER_BASE_URL is forwarded automatically.
  • The loss guard runs before training begins and fails fast with a clear message if the selected loss is unsupported by the backend.
gspo has a known upstream packing-order bug on multi-GPU and Megatron setups (NovaSky-AI/SkyRL#2043). Single-GPU FSDP is verified clean. rlcli prints a warning to stderr when --loss gspo is used.