Skip to main content
rlcli train harbor runs reinforcement learning on Harbor terminal tasks, rewarding agents based on tests/test.sh results. By default, it uses local Docker sandboxes, so you do not need a cloud account. This guide covers setup, dataset prefetching, task filtering, and launching a training run.

Prerequisites

  • Docker daemon running locally
  • A torch backend server for fused losses (see Backends & Losses)
  • For cloud sandboxes: a Modal account and the Modal CLI configured

Step-by-step

1

Start the server

Launch SkyRL with a torch backend. FSDP on a single GPU is recommended for stability with fused losses:
2

Prefetch the dataset (optional)

The default dataset is [email protected]. rlcli auto-downloads it on first use via uvx harbor datasets download, but you can prefetch it manually:
3

Filter tasks

Narrow the task set with substring filtering and a limit:
If no tasks match your filter, rlcli exits with a clear usage error.
4

Preview the config

Verify your setup before consuming GPU time:
5

Run Harbor agent RL

Remove --dry-run to start training:
Training logs are written to ~/.rlcli/runs/harbor-<timestamp> by default.

Sandbox options

Docker (default)

Runs tasks in local Docker containers. No cloud account required. Ensure your Docker daemon is running.

Modal

Runs tasks in Modal cloud sandboxes. Pass --sandbox modal if you have Modal configured.

Tokenizer stability

rlcli installs a locked tokenizer to avoid “Already borrowed” panics that can occur when concurrent Harbor environments render with a shared fast tokenizer.

Next steps