structure and correctness separately and combines them into a main reward. The finished task is reward-kit-example in the Harbor repository.
Step 1: Create the task
Install Harbor, then create a task directory:Step 2: Write the instruction and environment
textstats/instruction.md
uv, and the sample text:
textstats/environment/Dockerfile
textstats/environment/sample.txt
task.toml as is.
Step 3: Write the solution
The Oracle agent runs this script to confirm the task is solvable:textstats/solution/solve.sh
Step 4: Run Rewardkit from the test script
Harbor copiestests/ to /tests and runs test.sh. The script only invokes Rewardkit, which discovers every criterion under /tests, runs it against /app, and writes /logs/verifier/reward.json:
textstats/tests/test.sh
Step 5: Add structure criteria
Each subdirectory oftests/ becomes one reward named after the directory, and each Python file inside it is one score named after the file. structure checks that the required files and functions exist, using built-in criteria. Paths are relative to /app:
textstats/tests/structure/files_exist.py
textstats/tests/structure/functions_defined.py
Step 6: Add correctness criteria
Checking return values needs custom logic. Define it once at the tests root withshared=True so any dimension can call it. Returning a float gives partial credit:
textstats/tests/criteria.py
correctness, one per file, and put the end-to-end pipeline check in a nested pipeline/ group so it can be weighted separately:
textstats/tests/correctness/word_count.py
textstats/tests/correctness/most_common.py
textstats/tests/correctness/pipeline/pipeline_runs.py
Step 7: Configure the scoring
By default, the criteria in a Python file are averaged by weight, the files in a dimension count equally, and each dimension is written toreward.json on its own. Three reward.toml files adjust this.
structure passes only when every file and function is present. [scoring.<stem>] configures the criteria within one Python file, and the unnamed [[reward]] combines the files:
textstats/tests/structure/reward.toml
correctness, the two function checks and the nested pipeline group are weighted by name:
textstats/tests/correctness/reward.toml
reward.toml adds the main reward that Harbor reads. It is 1.0 only when both dimensions score 1.0:
textstats/tests/reward.toml
/logs/verifier/reward.json
Step 8 (Optional): Add an LLM judge
Add aquality dimension with a judge TOML and pass the API key through task.toml:
textstats/tests/quality/quality.toml
textstats/task.toml
quality becomes a fourth score, and the root all-pass aggregation includes it. The criterion is binary because all-pass needs a score of 1.0, which a Likert judge rarely gives. See Judge criteria for agent judges and the full TOML reference.
Step 9: Run and inspect
Run the solution against the verifier, then open the viewer:reward-details.json as a tree with the score, weight, and any error per criterion.
reward-kit-example
The complete task built in this tutorial.

