Skip to main content
Multi-steps tasks provide a way to iterleave verification through the agent run and measure an agent’s ability to continue work from a prior session. They are helpful for long-horizon tasks with early stopping conditions and for measuring memory and continual learning ability.

Format

Multi-step tasks have a different format than a typical Harbor task:
A step can contain the following files and folders:
  • instruction.md (required)
  • tests/ (optional)
  • solution/ (optional)
  • workdir/ (optional)

Configuration

Declare steps in the root task.toml, in execution order. Each name matches a directory under steps/. Other task settings use the standard task configuration.
Here, step-1 gets a 300-second agent timeout; step-2 inherits 600 seconds. step-2 runs only if step-1 earns a reward of at least 1.0.

Steps schema

Each [[steps]] entry accepts:
string
required
Unique, portable directory name under steps/.
AgentConfig
Per-step agent settings, using the task’s [agent] schema.
VerifierConfig
Per-step verifier settings, using the task’s [verifier] schema, including separate environments.
number | object | null
default:"null"
Reward threshold required to continue. See Early stopping.
HealthcheckConfig | null
default:"null"
Additional healthcheck after step setup, before the agent runs. Uses the environment healthcheck schema.
list[string | ArtifactConfig]
default:"[]"
Extra artifacts collected after verification into steps/<name>/artifacts/, alongside task- and trial-level artifacts.

Trial reward

Set multi_step_reward_strategy at the top level of task.toml:
  • "mean" (default): average each reward across steps with verifier results; missing keys count as zero.
  • "final": use the last executed step’s verifier result.

Early stopping

Set min_reward on a step to skip remaining steps when its score is too low. A number checks the reward key; an object requires every named reward to meet its threshold:
Missing reward keys fail the check; equality passes. Without min_reward, low scores do not stop execution. Thresholds are ignored when verification is disabled, but a step error without a verifier result still stops the task. Skipped steps do not contribute to the trial reward. "final" uses the last executed step; "mean" averages the available verifier results.

Test helpers

Put shared helpers in the base tests/. Harbor copies them to /tests/, then overlays the step’s tests, overwriting matching files. Separate verifiers use bundled tests with dedicated verifier images, or the same upload behavior when falling back to the agent environment.

Adding step-specific environment files

Sometimes, you want to upload files to the agent’s workspace at the start of a step. To do this, place files in steps/<step-name>/workdir/. Harbor copies them into the agent’s working directory before the step, overwriting matching paths.
An optional setup.sh runs with Bash after copying. Filesystem changes persist between steps.

Resuming an agent session

Each step starts a fresh conversation by default. To continue the previous step’s session, use:
Requires an agent with native resume support. The environment persists regardless of this setting.

Starting from a specific step

Starting from a specific step is coming soon. To enable this feature, make sure prior steps contain solution/solve.sh files.