import reads raw chat dumps and writes normalized messages JSONL suitable for supervised fine-tuning. It supports OpenAI chat-completions dumps, Anthropic Messages API dumps, and pre-shaped messages JSONL. This guide covers formats, role normalization, filtering, and piping into training.
Supported formats
messages (default)
Each line is already a
{"messages": [...]} object. Passed through with validation.openai
Chat-completions API dumps. Tool call turns are dropped because their flattened content is empty.
anthropic
Messages API dumps. Content blocks are flattened and the top-level
system message is kept as a system message.Step-by-step
1
Point at your raw dumps
Identify the format of your JSONL file. For example,
prod-traces.jsonl might contain Anthropic Messages API responses.2
Convert to messages JSONL
Run For OpenAI dumps:
rlcli import with the correct format flag and output path:3
Pipe directly into training
You can stream imported traces straight into supervised fine-tuning without writing an intermediate file:
Role normalization
rlcli normalizes roles to the standarduser, assistant, and system set. Unsupported roles are dropped.
Content flattening
- String content is kept as-is.
- Anthropic-style content blocks:
type: "text"blocks are joined with newlines. tool_use,tool_result,image, andthinkingblocks are dropped.- Empty flattened messages are skipped.
Filtering options
Failure modes
rlcli fails loud with clear errors and line numbers:--format messageson OpenAI dumps:ImportFormatErrorwith line number, because OpenAI dumps lack amessageskey- Invalid JSON: error includes the offending line number
- Unknown format flag: clear usage error listing supported formats
Imported traces are for supervised fine-tuning and distillation. For reinforcement learning, you need an environment and reward function. See RL Training and Harbor Agent RL.
Next steps
- See all
rlcli importoptions - Review the expected Dataset Format
- Start supervised fine-tuning with your cleaned data