Motivation
When creating new tasks for Harbor, writing the verifier is often the most tedious part. We looked at verifier implementations across popular benchmarks and found that they are often complicated scripts that are hard to read, extend, and reuse across tasks. This leads to copy-and-pasting of boilerplate code between tasks, which causes subtle errors in verifiers during evals. Rewardkit solves this. It packages common patterns into reusable components that can be shared and reused across tasks. Rewardkit gets rid of boilerplate code by enabling LLM- and agent-as-a-judge evaluation and partial credit. Because independent reward components are defined through the directory structure, verifiers are easy to understand at a glance even without reading the code. Evaluating multiple criteria in parallel and in isolation is natively supported in Rewardkit.Design principles
- Simplicity: Verifiers are defined by directory structure, making them easy to read at a glance. Common criteria are one-liners. Custom logic is a decorated function. Judges are reusable TOML files.
- Reuse and sharing: Reward criteria are plain files in a directory. Share them across tasks and version them in git.
- Zero boilerplate: 20+ built-in criteria for files, commands, JSON, CSV, HTTP, images, and trajectories. LLM- and agent-as-a-judge are supported natively.
- Isolation: Criteria can run in isolated filesystem snapshots so they cannot interfere with each other or corrupt the agent’s workspace.

