Skip to main content
Rewardkit lets you define verifiers against an agent’s workspace and trajectory. It runs criteria in parallel and writes their scores to JSON. Criteria can be:
  • Programmatic: Python functions that inspect files, run commands, or evaluate outputs
  • Judge-based: LLM or agent judges configured with reusable TOML files
Rewardkit is a self-contained Python package. It is most powerful when you use it with Harbor but it works just fine on its own.

Installation

For criteria or judges that read images or common document files such as PDF, DOCX, PPTX, and XLSX, install the extras:

Using with Harbor

Harbor copies a task’s tests/ directory to /tests and runs test.sh. Put your criteria files alongside it:
Programmatic criteria are implemented in Python files, and judges are specified in .toml files. Rewardkit picks up both from the tests directory.
tests/test.sh
Rewardkit discovers the criteria in /tests, runs them against the workspace at /app, and writes the result to /logs/verifier/reward.json:
/logs/verifier/reward.json
All defaults match Harbor’s conventions.

Working example

A complete task that uses Rewardkit, available as a starting point.
Want your coding agent to help design verifiers with Rewardkit? Install the rewardkit skill:

Programmatic criteria

Built-in criteria

Rewardkit includes common criteria used across popular benchmarks. Call them from any Python file in the tests directory:
tests/files.py
There are 20+ built-in criteria for files, commands, JSON, CSV, HTTP, images, and agent trajectories. See Built-in criteria for the full list.

Custom criteria

When you need logic specific to your task, define a function with the @criterion decorator. The first parameter is always workspace: Path, and the function returns a bool, a float, or a dict with a score:
tests/output.py
A criterion with no arguments besides workspace is called automatically. Criteria with additional arguments must be called through rewardkit:
tests/output.py
A criterion can also return a dict. reasoning, confidence, and model are optional and are written to reward-details.json:
tests/output.py

Judge criteria

Some things are easier to grade with an LLM, such as code quality, edge-case handling, or readability. Define these criteria in a TOML file:
tests/quality.toml
Python criteria and judge TOMLs can live in the same directory. See Judge criteria for agent judges, individual evaluation, MCP servers, and the complete TOML reference. Judges need API keys. In a Harbor task, pass them through [verifier.env] in task.toml:
task.toml
See Provider routing to swap judge providers without editing a rubric.

How scores are combined

Every criterion accepts an optional weight; the default is 1.0:
Each Python file that registers criteria produces one score, and so does each judge TOML. The name is the filename without .py or .toml: files.py becomes files, and quality.toml becomes quality. Files that only contain imports or shared criterion definitions and register no criteria are ignored. When the verifier files live directly under tests/, their scores are combined with equal weight into reward. To change either level, add reward.toml beside the criteria files:
tests/reward.toml
[scoring.files] controls how criteria from files.py are combined. The named [[reward]] table controls how files, quality, and any other scores in the directory are combined. Without this table, Rewardkit creates reward using a weighted average, where a judge’s share comes from weight in its [judge] section. The available aggregation modes are weighted-mean, weighted-sum, all-pass, any-pass, threshold, and required-pass. weighted-sum computes sum(value × weight) without normalization. It is the only mode that accepts negative weights, and its result may fall outside [0, 1]. Set threshold in the same table when using threshold aggregation; its default is 0.5. For programmatic criteria, required-pass is the same as all-pass. Judge TOMLs use their own [scoring] section to combine criteria within that judge.

Multi-reward tasks

When you want separate scores for dimensions such as correctness, structure, and quality, organize criteria into subdirectories. Each subdirectory becomes one reward:
This produces separate scores:
/logs/verifier/reward.json
Within a dimension, Python files, judge TOMLs, and subdirectories are combined with equal weight. Add a reward.toml inside the dimension to change that behavior:
tests/correctness/reward.toml
Python files and judge TOMLs are referenced by filename stem and subdirectories by directory name. Inputs left out of weights keep their own weight. Omit name, because the directory name defines the score’s name. Python files and judge TOMLs placed directly under tests/ next to the dimension directories become top-level rewards named after their filename stems. Files that only define shared criteria, described below, produce no score.

Nested groups

Subdirectories can be nested to group related criteria within a dimension:
behavior/ combines its criteria into an internal behavior score. Its parent can reference that score by directory name:
tests/correctness/reward.toml
The weights inside behavior/ never leak into its parent. Each directory exports one score. To add a main reward to the output across the existing dimensions, create a root-level tests/reward.toml with a named [[reward]] table:
tests/reward.toml
/logs/verifier/reward.json
The dimension scores remain in the output. Without a root aggregation, a multi-reward task has no implicit reward. Harbor reads the reward key as the task’s main score. Additional named [[reward]] tables can add other aggregate scores; a name may not collide with a dimension. To reuse custom criteria across dimensions, define them at the tests root with shared=True:
tests/shared.py
tests/correctness/behavior.py

Isolation

Some criteria run commands that modify the workspace. To prevent one criterion from affecting another, run it in isolation. Rewardkit mounts the workspace as read-only with overlayfs and discards that criterion’s changes after it finishes. For programmatic criteria, pass isolated=True:
For agent judges, set isolated = true in the [judge] section:
Agent judges share the workspace, so Rewardkit raises when a task makes more than one agent run (several agent judges, several samples, or individual mode with several criteria) unless each agent judge sets isolated = true.

Isolation in a Harbor task

Mounting needs more privileges than a default container has. Give them to a separate verifier environment, never to the agent environment. The verifier image installs fuse-overlayfs:
tests/Dockerfile
A compose file lets the verifier container mount filesystems:
tests/docker-compose.yaml
task.toml turns on the separate verifier and lists the files to grade as artifacts:
task.toml
This works on the Docker environment. Other sandbox providers differ in whether they allow mounts.

Output

Rewardkit writes two files side by side:
  • reward.json: the reward scores
  • reward-details.json: individual criterion scores, judge reasoning, errors, and warnings
harbor view renders the details under Verifier Logs → Rewards as a collapsible tree. Agent judge details also include normalized token usage, paths to copied native JSONL judge_logs, and text logs with stdout and stderr for failed Codex attempts.

Comparing verifiers

Pass multiple test directories to compare verifier designs side by side:
Rewardkit runs each directory independently, prints a comparison table for matching reward names, and writes namespaced keys such as v1/correctness and v2/correctness to reward.json.

CLI

All flags are optional. Passing multiple test directories runs each one independently and prints a comparison.

Python API