- Programmatic: Python functions that inspect files, run commands, or evaluate outputs
- Judge-based: LLM or agent judges configured with reusable TOML files
Rewardkit is a self-contained Python package. It is most powerful when you use
it with Harbor but it works just fine on its own.
Installation
Using with Harbor
Harbor copies a task’stests/ directory to /tests and runs test.sh. Put your criteria files alongside it:
.toml files. Rewardkit picks up both from the tests directory.
tests/test.sh
/tests, runs them against the workspace at /app, and writes the result to /logs/verifier/reward.json:
/logs/verifier/reward.json
Working example
A complete task that uses Rewardkit, available as a starting point.
Want your coding agent to help design verifiers with Rewardkit? Install the
rewardkit skill:Programmatic criteria
Built-in criteria
Rewardkit includes common criteria used across popular benchmarks. Call them from any Python file in the tests directory:tests/files.py
Custom criteria
When you need logic specific to your task, define a function with the@criterion decorator. The first parameter is always workspace: Path, and the function returns a bool, a float, or a dict with a score:
tests/output.py
workspace is called automatically. Criteria with additional arguments must be called through rewardkit:
tests/output.py
reasoning, confidence, and model are optional and are written to reward-details.json:
tests/output.py
Judge criteria
Some things are easier to grade with an LLM, such as code quality, edge-case handling, or readability. Define these criteria in a TOML file:tests/quality.toml
[verifier.env] in task.toml:
task.toml
How scores are combined
Every criterion accepts an optionalweight; the default is 1.0:
.py or .toml: files.py becomes files, and quality.toml becomes quality. Files that only contain imports or shared criterion definitions and register no criteria are ignored.
When the verifier files live directly under tests/, their scores are combined with equal weight into reward. To change either level, add reward.toml beside the criteria files:
tests/reward.toml
[scoring.files] controls how criteria from files.py are combined. The named [[reward]] table controls how files, quality, and any other scores in the directory are combined. Without this table, Rewardkit creates reward using a weighted average, where a judge’s share comes from weight in its [judge] section.
The available aggregation modes are weighted-mean, weighted-sum, all-pass, any-pass, threshold, and required-pass. weighted-sum computes sum(value × weight) without normalization. It is the only mode that accepts negative weights, and its result may fall outside [0, 1]. Set threshold in the same table when using threshold aggregation; its default is 0.5. For programmatic criteria, required-pass is the same as all-pass. Judge TOMLs use their own [scoring] section to combine criteria within that judge.
Multi-reward tasks
When you want separate scores for dimensions such as correctness, structure, and quality, organize criteria into subdirectories. Each subdirectory becomes one reward:/logs/verifier/reward.json
reward.toml inside the dimension to change that behavior:
tests/correctness/reward.toml
weights keep their own weight. Omit name, because the directory name defines the score’s name.
Python files and judge TOMLs placed directly under tests/ next to the dimension directories become top-level rewards named after their filename stems. Files that only define shared criteria, described below, produce no score.
Nested groups
Subdirectories can be nested to group related criteria within a dimension:behavior/ combines its criteria into an internal behavior score. Its parent can reference that score by directory name:
tests/correctness/reward.toml
behavior/ never leak into its parent. Each directory exports one score.
To add a main reward to the output across the existing dimensions, create a root-level tests/reward.toml with a named [[reward]] table:
tests/reward.toml
/logs/verifier/reward.json
reward. Harbor reads the reward key as the task’s main score. Additional named [[reward]] tables can add other aggregate scores; a name may not collide with a dimension.
To reuse custom criteria across dimensions, define them at the tests root with shared=True:
tests/shared.py
tests/correctness/behavior.py
Isolation
Some criteria run commands that modify the workspace. To prevent one criterion from affecting another, run it in isolation. Rewardkit mounts the workspace as read-only with overlayfs and discards that criterion’s changes after it finishes. For programmatic criteria, passisolated=True:
isolated = true in the [judge] section:
isolated = true.
Isolation in a Harbor task
Mounting needs more privileges than a default container has. Give them to a separate verifier environment, never to the agent environment. The verifier image installsfuse-overlayfs:
tests/Dockerfile
tests/docker-compose.yaml
task.toml turns on the separate verifier and lists the files to grade as artifacts:
task.toml
Output
Rewardkit writes two files side by side:reward.json: the reward scoresreward-details.json: individual criterion scores, judge reasoning, errors, and warnings
harbor view renders the details under Verifier Logs → Rewards as a collapsible tree. Agent judge details also include normalized token usage, paths to copied native JSONL judge_logs, and text logs with stdout and stderr for failed Codex attempts.
Comparing verifiers
Pass multiple test directories to compare verifier designs side by side:v1/correctness and v2/correctness to reward.json.

