> ## Documentation Index
> Fetch the complete documentation index at: https://docs.harborframework.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Quick start

> Define and run verifiers that produce reward scores.

Rewardkit lets you define verifiers against an agent's workspace and trajectory. It runs criteria in parallel and writes their scores to JSON. Criteria can be:

* **Programmatic:** Python functions that inspect files, run commands, or evaluate outputs
* **Judge-based:** LLM or agent judges configured with reusable TOML files

<Info>
  Rewardkit is a self-contained Python package. It is most powerful when you use
  it with Harbor but it works just fine on its own.
</Info>

## Installation

```bash theme={"system"}
uv tool install harbor-rewardkit
```

For criteria or judges that read images or common document files such as PDF, DOCX, PPTX, and XLSX, install the extras:

```bash theme={"system"}
uv tool install harbor-rewardkit[all]
```

## Using with Harbor

Harbor copies a task's `tests/` directory to `/tests` and runs `test.sh`. Put your criteria files alongside it:

```bash theme={"system"}
tests/
├── files.py
├── quality.toml
└── test.sh
```

Programmatic criteria are implemented in Python files, and judges are specified in `.toml` files. Rewardkit picks up both from the tests directory.

```bash title="tests/test.sh" theme={"system"}
#!/bin/bash
uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
```

Rewardkit discovers the criteria in `/tests`, runs them against the workspace at `/app`, and writes the result to `/logs/verifier/reward.json`:

```json title="/logs/verifier/reward.json" theme={"system"}
{ "reward": 0.75 }
```

All defaults match Harbor's conventions.

<Card title="Working example" icon="github" href="https://github.com/harbor-framework/harbor/tree/main/examples/tasks/reward-kit-example">
  A complete task that uses Rewardkit, available as a starting point.
</Card>

<Info>
  Want your coding agent to help design verifiers with Rewardkit? Install the `rewardkit` skill:

  ```bash theme={"system"}
  npx skills add harbor-framework/harbor --skill rewardkit
  ```
</Info>

## Programmatic criteria

### Built-in criteria

Rewardkit includes common criteria used across popular benchmarks. Call them from any Python file in the tests directory:

```python title="tests/files.py" theme={"system"}
import rewardkit as rk

rk.file_exists("output.txt")
rk.file_contains("output.txt", "hello")
rk.command_succeeds("python main.py")
```

There are 20+ built-in criteria for files, commands, JSON, CSV, HTTP, images, and agent trajectories. See [Built-in criteria](/rewardkit/built-in-criteria) for the full list.

### Custom criteria

When you need logic specific to your task, define a function with the `@criterion` decorator. The first parameter is always `workspace: Path`, and the function returns a `bool`, a `float`, or a dict with a `score`:

```python title="tests/output.py" theme={"system"}
from pathlib import Path

from rewardkit import criterion


@criterion
def has_valid_output(workspace: Path) -> bool:
    output = (workspace / "output.txt").read_text()
    return len(output.splitlines()) >= 10
```

A criterion with no arguments besides `workspace` is called automatically. Criteria with additional arguments must be called through `rewardkit`:

```python title="tests/output.py" theme={"system"}
from pathlib import Path

import rewardkit as rk
from rewardkit import criterion


@criterion(description="output has at least {n} lines")
def has_n_lines(workspace: Path, n: int) -> bool:
    output = (workspace / "output.txt").read_text()
    return len(output.splitlines()) >= n


rk.has_n_lines(10, weight=2.0)
rk.has_n_lines(50, weight=1.0)
```

A criterion can also return a dict. `reasoning`, `confidence`, and `model` are optional and are written to `reward-details.json`:

```python title="tests/output.py" theme={"system"}
@criterion
def report_is_relevant(workspace: Path) -> dict:
    probability = classify((workspace / "report.md").read_text())
    return {
        "score": probability >= 0.5,
        "reasoning": "classified as on-topic",
        "confidence": max(probability, 1 - probability),
        "model": "relevance-classifier-v1",
    }
```

## Judge criteria

Some things are easier to grade with an LLM, such as code quality, edge-case handling, or readability. Define these criteria in a TOML file:

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
files = ["/app/main.py"]

[[criterion]]
description = "Is the code correct?"
type = "binary"

[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
```

Python criteria and judge TOMLs can live in the same directory. See [Judge criteria](/rewardkit/judge-criteria) for agent judges, individual evaluation, MCP servers, and the complete TOML reference.

Judges need API keys. In a Harbor task, pass them through `[verifier.env]` in `task.toml`:

```toml title="task.toml" theme={"system"}
[verifier]
timeout_sec = 300.0

[verifier.env]
ANTHROPIC_API_KEY = "${ANTHROPIC_API_KEY}"
```

See [Provider routing](/rewardkit/judge-criteria#provider-routing) to swap judge providers without editing a rubric.

## How scores are combined

Every criterion accepts an optional `weight`; the default is `1.0`:

```python theme={"system"}
rk.file_exists("output.txt", weight=3.0)
rk.file_exists("readme.md", weight=1.0)
```

Each Python file that registers criteria produces one score, and so does each judge TOML. The name is the filename without `.py` or `.toml`: `files.py` becomes `files`, and `quality.toml` becomes `quality`. Files that only contain imports or shared criterion definitions and register no criteria are ignored.

When the verifier files live directly under `tests/`, their scores are combined with equal weight into `reward`. To change either level, add `reward.toml` beside the criteria files:

```toml title="tests/reward.toml" theme={"system"}
# Require every criterion registered by files.py to pass.
[scoring.files]
aggregation = "all-pass"

# Give files.py twice the weight of quality.toml.
[[reward]]
name = "reward"
aggregation = "weighted-mean"
weights = { files = 2.0, quality = 1.0 }
```

`[scoring.files]` controls how criteria from `files.py` are combined. The named `[[reward]]` table controls how `files`, `quality`, and any other scores in the directory are combined. Without this table, Rewardkit creates `reward` using a weighted average, where a judge's share comes from `weight` in its `[judge]` section.

The available aggregation modes are `weighted-mean`, `weighted-sum`, `all-pass`, `any-pass`, `threshold`, and `required-pass`. `weighted-sum` computes `sum(value × weight)` without normalization. It is the only mode that accepts negative weights, and its result may fall outside `[0, 1]`. Set `threshold` in the same table when using threshold aggregation; its default is `0.5`. For programmatic criteria, `required-pass` is the same as `all-pass`. Judge TOMLs use their own `[scoring]` section to combine criteria within that judge.

## Multi-reward tasks

When you want separate scores for dimensions such as correctness, structure, and quality, organize criteria into subdirectories. Each subdirectory becomes one reward:

```bash theme={"system"}
tests/
├── test.sh
├── correctness/
│   ├── files.py
│   └── behavior.py
├── structure/
│   └── files.py
└── quality/
    └── judge.toml
```

This produces separate scores:

```json title="/logs/verifier/reward.json" theme={"system"}
{ "correctness": 0.75, "structure": 1.0, "quality": 0.6 }
```

Within a dimension, Python files, judge TOMLs, and subdirectories are combined with equal weight. Add a `reward.toml` inside the dimension to change that behavior:

```toml title="tests/correctness/reward.toml" theme={"system"}
[[reward]]
aggregation = "weighted-mean"
weights = { files = 2.0, behavior = 1.0 }
```

Python files and judge TOMLs are referenced by filename stem and subdirectories by directory name. Inputs left out of `weights` keep their own weight. Omit `name`, because the directory name defines the score's name.

Python files and judge TOMLs placed directly under `tests/` next to the dimension directories become top-level rewards named after their filename stems. Files that only define shared criteria, described below, produce no score.

### Nested groups

Subdirectories can be nested to group related criteria within a dimension:

```bash theme={"system"}
tests/
└── correctness/
    ├── reward.toml
    ├── files.py
    └── behavior/
        ├── runtime.py
        └── edge_cases.py
```

`behavior/` combines its criteria into an internal `behavior` score. Its parent can reference that score by directory name:

```toml title="tests/correctness/reward.toml" theme={"system"}
[[reward]]
weights = { files = 1.0, behavior = 2.0 }
```

The weights inside `behavior/` never leak into its parent. Each directory exports one score.

To add a main reward to the output across the existing dimensions, create a root-level `tests/reward.toml` with a named `[[reward]]` table:

```toml title="tests/reward.toml" theme={"system"}
[[reward]]
name = "reward"
aggregation = "weighted-mean"
weights = { correctness = 2.0, structure = 1.0, quality = 1.0 }
```

```json title="/logs/verifier/reward.json" theme={"system"}
{
  "correctness": 0.75,
  "structure": 1.0,
  "quality": 0.6,
  "reward": 0.775
}
```

The dimension scores remain in the output. Without a root aggregation, a multi-reward task has no implicit `reward`. Harbor reads the `reward` key as the task's main score. Additional named `[[reward]]` tables can add other aggregate scores; a name may not collide with a dimension.

To reuse custom criteria across dimensions, define them at the tests root with `shared=True`:

```python title="tests/shared.py" theme={"system"}
from pathlib import Path

from rewardkit import criterion


@criterion(shared=True)
def word_count_correct(workspace: Path) -> float:
    # Criterion logic ...
    return score
```

```python title="tests/correctness/behavior.py" theme={"system"}
import rewardkit as rk

rk.word_count_correct(weight=3.0)
```

## Isolation

Some criteria run commands that modify the workspace. To prevent one criterion from affecting another, run it in isolation. Rewardkit mounts the workspace as read-only with overlayfs and discards that criterion's changes after it finishes.

For programmatic criteria, pass `isolated=True`:

```python theme={"system"}
rk.command_succeeds("python main.py", isolated=True)
```

For agent judges, set `isolated = true` in the `[judge]` section:

```toml theme={"system"}
[judge]
judge = "claude-code"
isolated = true
```

Agent judges share the workspace, so Rewardkit raises when a task makes more than one agent run (several agent judges, several samples, or individual mode with several criteria) unless each agent judge sets `isolated = true`.

### Isolation in a Harbor task

Mounting needs more privileges than a default container has. Give them to a [separate verifier environment](/tasks/separate-verifier), never to the agent environment. The verifier image installs `fuse-overlayfs`:

```dockerfile title="tests/Dockerfile" theme={"system"}
FROM ubuntu:24.04

RUN apt-get update \
    && apt-get install -y fuse-overlayfs \
    && rm -rf /var/lib/apt/lists/*
COPY --from=ghcr.io/astral-sh/uv:latest /uv /uvx /bin/

COPY . /tests/
```

A compose file lets the verifier container mount filesystems:

```yaml title="tests/docker-compose.yaml" theme={"system"}
services:
  main:
    cap_add:
      - SYS_ADMIN
    devices:
      - /dev/fuse
    security_opt:
      - apparmor:unconfined
```

`task.toml` turns on the separate verifier and lists the files to grade as artifacts:

```toml title="task.toml" theme={"system"}
artifacts = ["/app/main.py"]

[verifier]
environment_mode = "separate"
```

This works on the Docker environment. Other sandbox providers differ in whether they allow mounts.

## Output

Rewardkit writes two files side by side:

* `reward.json`: the reward scores
* `reward-details.json`: individual criterion scores, judge reasoning, errors, and warnings

`harbor view` renders the details under **Verifier Logs → Rewards** as a collapsible tree. Agent judge details also include normalized token `usage`, paths to copied native JSONL `judge_logs`, and text logs with stdout and stderr for failed Codex attempts.

## Comparing verifiers

Pass multiple test directories to compare verifier designs side by side:

```bash theme={"system"}
rewardkit /tests/v1 /tests/v2
```

Rewardkit runs each directory independently, prints a comparison table for matching reward names, and writes namespaced keys such as `v1/correctness` and `v2/correctness` to `reward.json`.

## CLI

```bash theme={"system"}
rewardkit <tests_dirs...> \
    --workspace /app \
    --output /logs/verifier/reward.json \
    --max-concurrent-programmatic 8 \
    --max-concurrent-llm 8 \
    --max-concurrent-agent 2
```

All flags are optional. Passing multiple test directories runs each one independently and prints a comparison.

## Python API

```python theme={"system"}
import rewardkit as rk

# Run and get scores
scores = rk.run("/tests", workspace="/app")

# Inspect discovered rewards without running them
rewards = rk.discover("/tests", workspace="/app")
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.