rlcli train: one-shot SL, RL, Harbor agent, and self-distillation training
Use rlcli train sl, rl, harbor, and opsd to run supervised fine-tuning, RL on GSM8K, sandboxed Harbor agent RL, and on-policy self-distillation on local GPUs.
rlcli provides four training subcommands under train: sl for supervised fine-tuning, rl for reinforcement learning on math benchmarks, harbor for sandboxed agent RL, and opsd for on-policy self-distillation from a prompts file. Each subcommand sends requests to the Tinker API at --base-url and automatically sets TINKER_API_KEY=tml-dummy.
On-policy self-distillation on a prompts JSONL. Each step the student samples --group-size completions per prompt, a teacher scores the student’s own tokens, and the negative reverse KL (student || teacher) becomes the per-token advantage for an importance_sampling update. There is no reward: the teacher is the only signal. Drives the pinned cookbook’s distillation.train_on_policy recipe.
Path to a JSONL file, or - to read from stdin. Each line is {"prompt": "..."} or {"messages": [...]}; for messages, the last user turn is the prompt and earlier system/user/assistant text turns are context (trailing assistant turns are dropped).
Privileged text the teacher sees prepended to every prompt (hint, blank line, prompt). The student never sees it, so with the default teacher this is “same weights + hint” self-distillation.
rlcli train opsd --model Qwen/Qwen3-4B-Instruct-2507 --dataset prompts.jsonl \ --teacher-hint "Think step by step and double-check your arithmetic." \ --batch-size 16 --group-size 4 --steps 50
The teacher scores the student’s tokens as rendered with the student’s chat template; a --teacher with a different template is scored under the student’s rendering (the cookbook’s assumption too). With --teacher-hint, rlcli builds the hinted prompt with the same renderer and wraps the teacher client so its logprob calls see hint + prompt while the student’s sequence is unchanged.
TINKER_API_KEY is set to tml-dummy automatically if unset.
TINKER_BASE_URL is forwarded automatically.
The loss guard runs before training begins and fails fast with a clear message if the selected loss is unsupported by the backend.
gspo has a known upstream packing-order bug on multi-GPU and Megatron setups (NovaSky-AI/SkyRL#2043). Single-GPU FSDP is verified clean. rlcli prints a warning to stderr when --loss gspo is used.