Skip to main content
rlcli train sl expects a JSON Lines file where each line is one conversation. The format is the standard messages array used by OpenAI, Anthropic, and most modern chat APIs. Getting the shape right lets you pipe imported traces directly into training without extra preprocessing.

Line format

Each line is a single JSON object with a messages key. Every message has a role and a content string.
  • One JSON object per line. No outer array.
  • role must be user, assistant, or system.
  • content is a plain string. No content-block lists at this stage.

How rlcli import normalizes data

rlcli import reads chat dumps and writes messages JSONL to stdout (or -o). It applies two rules before emitting a line. Role normalization Content flattening
  • Strings are kept as-is.
  • Anthropic-style content-block lists join only type:"text" blocks with newlines.
  • tool_use, tool_result, image, and thinking blocks are dropped.
  • If flattening leaves a message with no content, the turn is skipped.
A conversation must pass two filters to be emitted: it needs at least --min-messages (default 2) messages, and at least one assistant message must remain after filtering. Lines that fail these checks are dropped silently.

Import formats

rlcli import supports three source shapes:
  • messages (default): each line is already {"messages":[...]}.
  • openai: chat-completions dumps. tool_calls turns are dropped.
  • anthropic: Messages API dumps. Content blocks are flattened as above, and a top-level system field is kept as a system message.
If the input contains unknown formats or malformed JSON, rlcli import raises ImportFormatError with the line number so you can fix the source file.

Pipe into training

Because rlcli import writes to stdout, you can stream cleaned data straight into supervised fine-tuning:
The --dataset - flag tells rlcli train sl to read from stdin.

SFT versus RL datasets

Imported traces are for supervised fine-tuning and distillation. They provide (prompt, response) pairs that the model learns to imitate. Reinforcement learning needs an environment and a reward function, not a static messages file. Use rlcli train rl for gsm8k math RL, or rlcli train harbor for sandboxed agent tasks. See the RL guides for details.

Next steps