uv run and exposes the Tinker HTTP API so rlcli can train and sample against it. Use rlcli serve to start, stop, inspect, and debug that server.
serve start
Start the managed SkyRL Tinker server.Flags
string
required
HuggingFace model name or local path to load.
choice
default:"jax"
Training backend. Choices:
jax, fsdp, megatron.jax: runs anywhere, including CPU and macOS. Servescross_entropy,importance_sampling,ppo, andcispo.fsdp/megatron: require Linux and CUDA. Serve the full loss set includinggspo,cispo,dppo, andppo_critic.
int
default:"8000"
HTTP port for the Tinker API.
int
GPUs per node. Also configures one vLLM engine per GPU. Only valid for torch backends (
fsdp, megatron).int
Number of training nodes. Only valid for torch backends.
int
Tensor-parallel size per inference engine. Only valid for torch backends. The number of engines becomes
max(1, gpus // tp).string
Base directory for checkpoints. Defaults to the server’s default.
int
default:"16384"
Cap the vLLM context length on torch backends. Set to
0 to use the model’s native context length. Be careful: uncapped native contexts on large models like Qwen3-4B-Instruct-2507 (262K) can segfault vLLM. Any explicit non-default value is trusted as-is.JSON string
Raw SkyRL-Train configuration overrides, merged last. Nested dictionaries like
engine_init_kwargs are merged rather than overwritten. User keys win over rlcli defaults.int
default:"900"
Startup timeout in seconds.
Example
serve stop
Stop the managed server.SIGTERM to the server process group. If the process does not exit within 30 seconds, it sends SIGKILL. The server state file is removed on success.
serve status
Print the current managed server state and a health check.~/.rlcli/server.json plus a healthy boolean from GET /docs. If no managed server is running, the command prints:
serve logs
Tail the server log.Flags
int
default:"40"
Number of trailing lines to print from
~/.rlcli/server.log.Example
Server paths and environment
- State file:
~/.rlcli/server.json - Log file:
~/.rlcli/server.log - SkyRL checkout:
~/.rlcli/skyrl-src
RLCLI_SKYRL_SOURCE environment variable, or change the entire home directory with RLCLI_HOME.
Errors
ServerError is raised in these situations:
- Server exits early during startup
- Startup exceeds the
--waittimeout - An unknown backend is specified
- A duplicate managed server is already running