rlcli train sl expects a JSON Lines file where each line is one conversation. The format is the standard messages array used by OpenAI, Anthropic, and most modern chat APIs. Getting the shape right lets you pipe imported traces directly into training without extra preprocessing.
Line format
Each line is a single JSON object with amessages key. Every message has a role and a content string.
- One JSON object per line. No outer array.
rolemust beuser,assistant, orsystem.contentis a plain string. No content-block lists at this stage.
How rlcli import normalizes data
rlcli import reads chat dumps and writes messages JSONL to stdout (or -o). It applies two rules before emitting a line.
Role normalization
Content flattening
- Strings are kept as-is.
- Anthropic-style content-block lists join only
type:"text"blocks with newlines. tool_use,tool_result,image, andthinkingblocks are dropped.- If flattening leaves a message with no content, the turn is skipped.
--min-messages (default 2) messages, and at least one assistant message must remain after filtering. Lines that fail these checks are dropped silently.
Import formats
rlcli import supports three source shapes:
messages(default): each line is already{"messages":[...]}.openai: chat-completions dumps.tool_callsturns are dropped.anthropic: Messages API dumps. Content blocks are flattened as above, and a top-levelsystemfield is kept as a system message.
rlcli import raises ImportFormatError with the line number so you can fix the source file.
Pipe into training
Becauserlcli import writes to stdout, you can stream cleaned data straight into supervised fine-tuning:
--dataset - flag tells rlcli train sl to read from stdin.
SFT versus RL datasets
Imported traces are for supervised fine-tuning and distillation. They provide (prompt, response) pairs that the model learns to imitate. Reinforcement learning needs an environment and a reward function, not a static messages file. Userlcli train rl for gsm8k math RL, or rlcli train harbor for sandboxed agent tasks. See the RL guides for details.
Next steps
- Import and clean chat dumps:
/guides/importing-traces - Run supervised fine-tuning:
/guides/supervised-fine-tuning - Explore RL with an environment:
/guides/reinforcement-learning