> ## Documentation Index
> Fetch the complete documentation index at: https://docs.harborframework.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Overview

> A Harbor task is an instruction, environment, and test script.

A Harbor task is an instruction, environment, and test script. The task is defined as a directory of files with the following structure:

```bash theme={"system"}
my-task/
├── instruction.md
├── task.toml
├── environment/
│   ├── Dockerfile  # or some other spec
│   └── ...
├── solution/
│   ├── solve.sh
│   └── ...
└── tests/
    ├── test.sh
    └── ...
```

Environments can be defined using any spec, as long as the consumer of the task (e.g. Harbor framework) supports it. Popular specs include `Dockerfile`, `docker-compose.yaml` (for multi-container environments), and `Apptainer.def` (for HPCs).

## Trials

A Harbor trial is an agent's attempt at completing a task. The diagram below shows how the components of a task are used in a trial.

```mermaid theme={"system"}
%%{init: {"flowchart": {"curve": "basis", "nodeSpacing": 60, "rankSpacing": 70}, "look": "handDrawn"}}%%
flowchart LR
  Task -.-> SellShare[Share]
  Task[/Task/] --> Start[Start<br/>Environment]
  Start --> C1[Environment]
  C1 --> C2[Environment]
  C2 --> Stop[Stop<br/>Environment]

  C1 --> Agent[Agent]
  Agent --> C1
  Agent --> Trajectory[/Trajectory/]
  C2 --> Verifier[Verifier]
  Verifier --> C2
  Verifier --> Reward[/Reward/]
  Reward -.-> Evals[Aggregate<br/>across eval]
  Reward -.-> Optimize[RL & context<br/>optimization]
  Trajectory -.-> Optimize
  Trajectory -..-> SFT[SFT data &<br/>failure analysis]
```

## Rewards

A task defines the logic to verify the completion of the instruction and produce a reward in the `tests/test.sh` script. The script must produce a numerical reward in the environment at `/logs/verifier/reward.txt` or `/logs/verifier/reward.json` (for multi-dimensional or labeled rewards).

## Independent, isolated, and reproducible

Harbor tasks are independent, isolated, reproducible pieces of code. Harbor tasks have no dependency on the Harbor framework. They can easily be plugged into any framework that supports the Harbor format. However, the Harbor framework is a simple way to start running tasks at scale.

Think of tasks like software packages that should be self-contained, verisioned, maintained, and developed over time.

## Task components

* [`instruction.md`](/tasks/instruction)
* [`task.toml`](/tasks/configuration)
* [`environment/`](/tasks/environment)
* [`solution/`](/tasks/solution)
* [`tests/`](/tasks/verifier)

## Multi-step tasks

Harbor tasks can be split into multiple steps, each with its own instruction, tests, and solution. Multi-step tasks are great for defining milestones in long-horizon tasks, testing continual learning methods like memory, and observing an agent's ability to build on its prior work.

See [Multi-step](/tasks/multi-step) for a detailed overview.

```bash theme={"system"}
my-task/
├── task.toml
├── environment/
│   ├── Dockerfile  # or some other spec
│   └── ...
├── tests/
│   ├── helpers.py
│   └── ...
└── steps/
    ├── step-1/
    │   ├── instruction.md
    │   ├── tests/
    │   │   ├── test.sh
    │   │   └── ...
    │   ├── solution/
    │   │   ├── solve.sh
    │   │   └── ...
    │   └── workdir/
    │       ├── setup.sh
    │       └── ...
    └── step-2/
        └── ...
```

`Dockerfile` can be replaced with another supported environment spec. Shared grading utilities in `tests/helpers.py` and per-step `workdir/setup.sh` scripts are optional. Directories can contain additional supporting files.

Each step's test script can use the shared utilities. A base `tests/test.sh` is also supported as a fallback for steps without their own test script; a step-specific `test.sh` overrides it.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.