> ## Documentation Index
> Fetch the complete documentation index at: https://docs.harborframework.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Multi-step

> Evaluate agent performance across multiple milestones.

Multi-steps tasks provide a way to iterleave verification through the agent run and measure an agent's ability to continue work from a prior session. They are helpful for long-horizon tasks with early stopping conditions and for measuring memory and continual learning ability.

## Format

Multi-step tasks have a different format than a typical Harbor task:

```bash theme={"system"}
task.toml
environment/
├── Dockerfile  # or some other spec
└── ...
tests/
├── helpers.py  # optional shared grading utilities
└── ...
steps/
├── step-1/
│   ├── instruction.md
│   ├── tests/
│   │   ├── test.sh
│   │   └── ...
│   ├── solution/
│   │   ├── solve.sh
│   │   └── ...
│   └── workdir/
│       ├── setup.sh  # optional
│       └── ...
└── step-2/
    └── ...
```

A step can contain the following files and folders:

* `instruction.md` (required)
* `tests/` (optional)
* `solution/` (optional)
* `workdir/` (optional)

## Configuration

Declare steps in the root `task.toml`, in execution order. Each `name` matches a directory under `steps/`. Other task settings use the standard [task configuration](/tasks/configuration).

```toml theme={"system"}
multi_step_reward_strategy = "mean"

[agent]
timeout_sec = 600

[[steps]]
name = "step-1"
min_reward = 1.0

[steps.agent]
timeout_sec = 300

[[steps]]
name = "step-2"

[steps.verifier]
timeout_sec = 120
```

Here, `step-1` gets a 300-second agent timeout; `step-2` inherits 600 seconds. `step-2` runs only if `step-1` earns a `reward` of at least `1.0`.

### Steps schema

Each `[[steps]]` entry accepts:

<ParamField body="name" type="string" required>
  Unique, portable directory name under `steps/`.
</ParamField>

<ParamField body="agent" type="AgentConfig">
  Per-step agent settings, using the task's `[agent]` schema.
</ParamField>

<ParamField body="verifier" type="VerifierConfig">
  Per-step verifier settings, using the task's `[verifier]` schema, including [separate environments](/tasks/separate-verifier).
</ParamField>

<ParamField body="min_reward" type="number | object | null" default="null">
  Reward threshold required to continue. See [Early stopping](#early-stopping).
</ParamField>

<ParamField body="healthcheck" type="HealthcheckConfig | null" default="null">
  Additional healthcheck after step setup, before the agent runs. Uses the environment healthcheck schema.
</ParamField>

<ParamField body="artifacts" type="list[string | ArtifactConfig]" default="[]">
  Extra [artifacts](/tasks/configuration#artifacts) collected after verification into `steps/<name>/artifacts/`, alongside task- and trial-level artifacts.
</ParamField>

### Trial reward

Set `multi_step_reward_strategy` at the top level of `task.toml`:

* `"mean"` (default): average each reward across steps with verifier results; missing keys count as zero.
* `"final"`: use the last executed step's verifier result.

## Early stopping

Set `min_reward` on a step to skip remaining steps when its score is too low. A number checks the `reward` key; an object requires every named reward to meet its threshold:

```toml theme={"system"}
[[steps]]
name = "step-1"
min_reward = { accuracy = 0.9, safety = 1.0 }

[[steps]]
name = "step-2"
```

Missing reward keys fail the check; equality passes. Without `min_reward`, low scores do not stop execution. Thresholds are ignored when verification is disabled, but a step error without a verifier result still stops the task.

Skipped steps do not contribute to the trial reward. `"final"` uses the last executed step; `"mean"` averages the available verifier results.

## Test helpers

Put shared helpers in the base `tests/`. Harbor copies them to `/tests/`, then overlays the step's tests, overwriting matching files.

[Separate verifiers](/tasks/separate-verifier#image-selection) use bundled tests with dedicated verifier images, or the same upload behavior when falling back to the agent environment.

## Adding step-specific environment files

Sometimes, you want to upload files to the agent's workspace at the start of a step. To do this, place files in `steps/<step-name>/workdir/`. Harbor copies them into the agent's working directory before the step, overwriting matching paths.

```bash theme={"system"}
steps/step-2/workdir/
├── input.csv
└── setup.sh
```

An optional `setup.sh` runs with Bash after copying. Filesystem changes persist between steps.

## Resuming an agent session

Each step starts a fresh conversation by default. To continue the previous step's session, use:

<Tabs>
  <Tab title="CLI">
    ```bash theme={"system"}
    harbor run \
      --path "<task-path>" \
      --agent "<agent>" \
      --model "<model>" \
      --resume-trajectory
    ```
  </Tab>

  <Tab title="Config">
    ```yaml theme={"system"}
    tasks:
      - path: "<task-path>"
    agents:
      - name: "<agent>"
        model_name: "<model>"
        resume_trajectory: true
    ```
  </Tab>
</Tabs>

Requires an agent with native resume support. The environment persists regardless of this setting.

## Starting from a specific step

<Note>
  Starting from a specific step is coming soon. To enable this feature, make sure prior steps contain `solution/solve.sh` files.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.