train harbor runs reinforcement learning on Harbor terminal tasks, rewarding agents based on tests/test.sh results. By default, it uses local Docker sandboxes, so you do not need a cloud account. This guide covers setup, dataset prefetching, task filtering, and launching a training run.
Prerequisites
- Docker daemon running locally
- A torch backend server for fused losses (see Backends & Losses)
- For cloud sandboxes: a Modal account and the Modal CLI configured
Step-by-step
1
Start the server
Launch SkyRL with a torch backend. FSDP on a single GPU is recommended for stability with fused losses:
2
Prefetch the dataset (optional)
The default dataset is
[email protected]. rlcli auto-downloads it on first use via uvx harbor datasets download, but you can prefetch it manually:3
Filter tasks
Narrow the task set with substring filtering and a limit:If no tasks match your filter, rlcli exits with a clear usage error.
4
Preview the config
Verify your setup before consuming GPU time:
5
Run Harbor agent RL
Remove Training logs are written to
--dry-run to start training:~/.rlcli/runs/harbor-<timestamp> by default.Sandbox options
Docker (default)
Runs tasks in local Docker containers. No cloud account required. Ensure your Docker daemon is running.
Modal
Runs tasks in Modal cloud sandboxes. Pass
--sandbox modal if you have Modal configured.Tokenizer stability
rlcli installs a locked tokenizer to avoid “Already borrowed” panics that can occur when concurrent Harbor environments render with a shared fast tokenizer.Next steps
- Review all
train harborCLI options - Learn about Backends & Losses
- Explore supervised fine-tuning to distill agent traces