Skip to main content
This tutorial builds a small task and grades it with Rewardkit. The agent writes a text statistics module and an analysis script. The verifier scores structure and correctness separately and combines them into a main reward. The finished task is reward-kit-example in the Harbor repository.

Step 1: Create the task

Install Harbor, then create a task directory:
The finished task looks like this:

Step 2: Write the instruction and environment

textstats/instruction.md
The environment needs Python, uv, and the sample text:
textstats/environment/Dockerfile
textstats/environment/sample.txt
Leave the generated task.toml as is.

Step 3: Write the solution

The Oracle agent runs this script to confirm the task is solvable:
textstats/solution/solve.sh

Step 4: Run Rewardkit from the test script

Harbor copies tests/ to /tests and runs test.sh. The script only invokes Rewardkit, which discovers every criterion under /tests, runs it against /app, and writes /logs/verifier/reward.json:
textstats/tests/test.sh

Step 5: Add structure criteria

Each subdirectory of tests/ becomes one reward named after the directory, and each Python file inside it is one score named after the file. structure checks that the required files and functions exist, using built-in criteria. Paths are relative to /app:
textstats/tests/structure/files_exist.py
textstats/tests/structure/functions_defined.py

Step 6: Add correctness criteria

Checking return values needs custom logic. Define it once at the tests root with shared=True so any dimension can call it. Returning a float gives partial credit:
textstats/tests/criteria.py
Call the shared criteria from correctness, one per file, and put the end-to-end pipeline check in a nested pipeline/ group so it can be weighted separately:
textstats/tests/correctness/word_count.py
textstats/tests/correctness/most_common.py
textstats/tests/correctness/pipeline/pipeline_runs.py

Step 7: Configure the scoring

By default, the criteria in a Python file are averaged by weight, the files in a dimension count equally, and each dimension is written to reward.json on its own. Three reward.toml files adjust this. structure passes only when every file and function is present. [scoring.<stem>] configures the criteria within one Python file, and the unnamed [[reward]] combines the files:
textstats/tests/structure/reward.toml
In correctness, the two function checks and the nested pipeline group are weighted by name:
textstats/tests/correctness/reward.toml
A root reward.toml adds the main reward that Harbor reads. It is 1.0 only when both dimensions score 1.0:
textstats/tests/reward.toml
The verifier now writes:
/logs/verifier/reward.json
See How scores are combined for the aggregation modes.

Step 8 (Optional): Add an LLM judge

Add a quality dimension with a judge TOML and pass the API key through task.toml:
textstats/tests/quality/quality.toml
textstats/task.toml
quality becomes a fourth score, and the root all-pass aggregation includes it. The criterion is binary because all-pass needs a score of 1.0, which a Likert judge rarely gives. See Judge criteria for agent judges and the full TOML reference.

Step 9: Run and inspect

Run the solution against the verifier, then open the viewer:
Every dimension should score 1. In the trial’s Verifier Logs tab, the Rewards section renders reward-details.json as a tree with the score, weight, and any error per criterion.

reward-kit-example

The complete task built in this tutorial.