Skip to main content
Judge criteria let you use an LLM or agent judge to grade a task. They are configured with TOML files. This makes it easy to reuse and share rubrics between tasks.

LLM judge

tests/quality.toml
The judge field accepts any LiteLLM model string. It can be overridden at invocation time without editing the rubric. See Provider routing.

Agent judge

Agent judges can explore the filesystem and run commands to grade the task.
tests/review.toml

MCP servers

Each [[judge.mcp_servers]] entry matches a Harbor task’s [[environment.mcp_servers]]. Per-server allowed_tools lists the tools the judge may call; omit it to allow all of the server’s tools.
tests/review.toml
Codex does not support sse servers.

Individual mode

Set mode = "individual" to grade one criterion per call instead of batching them all into one. LLM judges make one request per criterion; agent judges run one turn per criterion, sequentially. For LLM judges, each criterion can also scope its own files:
tests/quality.toml
Criteria without files fall back to [judge].files. If a judge call times out, Rewardkit records the affected criteria as 0.0 with an error and warning in reward-details.json.

Sampling the judge

Set samples to run the judge several times and score each criterion with its median sample. reward-details.json lists every sample’s answer and an agreement between 0 and 1.
tests/quality.toml
An agent judge with several samples must be isolated. Pass --yolo to run without isolation.

Guarding the judge

guard protects the judge from prompt injection: an agent writing instructions for the judge into the files being graded. The judge is told not to follow instructions in the graded files and reports whether it found any under guard in reward-details.json. flag only reports it; penalize also scores a flagged submission 0.
tests/quality.toml

Provider routing

Judges call LiteLLM, which reads credentials from environment variables. You can change providers at invocation time without editing the rubric:
  • --je KEY=VALUE sets an environment variable for the run and can be repeated.
  • --judge MODEL_OR_AGENT overrides [judge].judge. The equivalent environment variable is REWARDKIT_JUDGE.
  • --model MODEL overrides [judge].model for an agent judge. The equivalent environment variable is REWARDKIT_MODEL.
Harbor users can pass the same environment variables with --ve.
The LiteLLM provider docs list the environment variables used by each provider.

Subscription authentication

For Anthropic LLM judges, Rewardkit uses a Claude subscription token when it is the only Anthropic credential present. Create one with claude setup-token and set CLAUDE_CODE_OAUTH_TOKEN. If ANTHROPIC_API_KEY is also set, the API key has priority. Set REWARDKIT_FORCE_SUBSCRIPTION=1 to require the subscription token. For the codex agent judge, set OPENAI_API_KEY or pass the contents of a ChatGPT-authenticated Codex auth.json as CODEX_AUTH_JSON. The API key has priority unless REWARDKIT_FORCE_SUBSCRIPTION=1 is set.

Configuration reference

Judge TOMLs are validated when the tests directory is scanned, before any judge runs. Unknown keys and unknown values raise an error.

[judge] section

string
default:"\"anthropic/claude-opus-5-5\""
LiteLLM model name (e.g. "anthropic/claude-opus-5-5"), agent judge name ("claude-code", "codex", "fx"), or "jev".
string | null
default:"null"
For agent judges, the LLM the agent should use. For the JEV judge, the JEV model.
list[string]
default:"[]"
Workspace file paths to include in the judge prompt.
"batched" | "individual"
default:"\"batched\""
Whether to grade all criteria in one call (batched) or one call per criterion (individual).
integer
default:"300"
Seconds to wait for the judge response.
"none" | "minimal" | "low" | "medium" | "high" | "xhigh"
default:"\"medium\""
Effort level for LLM judges. Accepts the levels LiteLLM supports; a given model may not support all of them.
boolean
default:"false"
For agent judges, mount the workspace as read-only via overlayfs.
string | null
default:"null"
For agent judges, the working directory the agent runs in. With isolated, it must be inside the workspace.
list[table]
default:"[]"
For agent judges, MCP servers to configure before running. Each entry matches a Harbor task’s [[environment.mcp_servers]], plus a per-server allowed_tools allowlist. Codex does not support sse servers.
string | null
default:"null"
Path to a reference solution file for comparison.
string | null
default:"null"
Path to an ATIF trajectory JSON to include in the prompt.
number
default:"1.0"
Weight of this judge’s score when combined with the other scores in the same directory.
string | null
default:"null"
Custom prompt template (.md or .txt). Must contain a {criteria} placeholder.
integer
default:"1"
Number of times to run the judge.
"off" | "flag" | "penalize"
default:"\"off\""
Protect the judge from instructions the agent wrote into the graded files. penalize also scores a flagged submission 0.

[[criterion]] entries

string
required
What to grade. This text is sent to the judge.
"binary" | "likert" | "numeric" | "rubric"
default:"\"binary\""
Output format.
string | null
default:"null"
Identifier for this criterion. Auto-generated from description if omitted.
string | null
default:"null"
Stable rubric identifier carried through to reward-details.json for provenance (e.g. "1.1", "2.3"). Independent of name; survives rewording the description.
integer
default:"5"
Scale size for the likert type.
number
default:"0.0"
Minimum value for the numeric type.
number
default:"1.0"
Maximum value for the numeric type.
list[string]
default:"[]"
Level descriptions for the rubric type, ordered from lowest to highest. 2 to 10 entries.
number
default:"1.0"
Importance multiplier for score aggregation. Negative weights support penalty criteria and require aggregation = "weighted-sum".
list[string]
default:"[]"
Files scoped to this criterion. Requires [judge].mode = "individual". Falls back to [judge].files when omitted.
boolean
default:"false"
Invert the normalized score. Use for a criterion describing behavior the answer should not exhibit. See Negated criteria.
boolean
default:"false"
Exempt this criterion from gating under aggregation = "required-pass".

[scoring] section

Controls how the judge’s criteria are aggregated into a single score. This only affects criteria within this TOML file. It does not change how programmatic and judge scores are combined across the directory.
required-pass returns 1.0 only when every non-optional criterion passes (value > 0); optional criteria never gate. With no non-optional criteria it warns and scores 0.0. weighted-sum computes sum(value × weight) without normalization. It is the only aggregation that accepts negative weights, and its result may fall outside [0, 1].

Score normalization

  • Binary: yes/true/1 → 1.0, anything else → 0.0
  • Likert: normalized to [0, 1] as (raw - 1) / (points - 1)
  • Numeric: normalized to [0, 1] as (raw - min) / (max - min)
  • Rubric: levels are numbered from 0, normalized to [0, 1] as raw / (levels - 1)

Negated criteria

Set negate = true for a criterion describing behavior the answer should not exhibit. The judge scores the criteria as usual, then the score is inverted (value → 1 - value): present → 0.0, absent → 1.0. The raw judge answer is kept in reward-details.json so the flip is auditable.

Trajectory evaluation

To grade the agent’s process rather than just its output, point the judge at the trajectory file:
The trajectory content is truncated proportionally to fit within the model’s context window while preserving all steps.

Custom prompt templates

You can provide your own prompt instead of the built-in one:
The template must contain a {criteria} placeholder where criterion descriptions get injected.

JEV judge

JEV is a new type of language model from TypeSafe. It returns a probability or a rubric score per criterion and no reasoning, so it is fast and cheap. It requires the jev extra (uv tool install harbor-rewardkit[jev]) and a TYPESAFE_API_KEY:
tests/quality.toml
Binary criteria score 1.0 at a probability of 0.5 or higher. Rubric levels must run from worst to best, because the position in the list sets the score. The raw probability or score is kept in reward-details.json. JEV grades text files only and supports binary and rubric criteria. atif-trajectory and prompt_template are not supported. The files and the longest criterion must fit in 32k tokens, and the task image needs CA certificates (ca-certificates on Debian and Ubuntu). To route through a LiteLLM proxy or Vercel AI Gateway, set TYPESAFE_BASE_URL to the gateway’s TypeSafe endpoint and TYPESAFE_API_KEY to the gateway key. Vercel also needs model = "typesafe-ai/jev".