> ## Documentation Index
> Fetch the complete documentation index at: https://docs.harborframework.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Judge criteria

> Grade tasks with LLM or agent judges configured in TOML.

Judge criteria let you use an LLM or agent judge to grade a task. They are configured with TOML files. This makes it easy to reuse and share rubrics between tasks.

## LLM judge

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
files = ["/app/main.py", "/app/utils.py"]

[[criterion]]
description = "Is the code correct?"
type = "binary"

[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
weight = 2.0

[[criterion]]
description = "Rate the test coverage on a scale from 0 to 100"
type = "numeric"
min = 0
max = 100
```

The `judge` field accepts any [LiteLLM model string](https://docs.litellm.ai/docs/providers). It can be overridden at invocation time without editing the rubric. See [Provider routing](#provider-routing).

## Agent judge

Agent judges can explore the filesystem and run commands to grade the task.

```toml title="tests/review.toml" theme={"system"}
[judge]
judge = "claude-code"
model = "anthropic/claude-opus-5-5"
isolated = true

[[criterion]]
description = "Does the solution handle edge cases?"
type = "binary"
```

### MCP servers

Each `[[judge.mcp_servers]]` entry matches a Harbor task's `[[environment.mcp_servers]]`. Per-server `allowed_tools` lists the tools the judge may call; omit it to allow all of the server's tools.

```toml title="tests/review.toml" theme={"system"}
[judge]
judge = "claude-code"

[[judge.mcp_servers]]
name = "playwright"
transport = "stdio"
command = "npx"
args = ["@playwright/mcp@latest", "--headless", "--isolated"]
allowed_tools = ["navigate", "click"]

[[criterion]]
description = "Does the rendered page match the spec?"
type = "binary"
```

<Note>Codex does not support `sse` servers.</Note>

## Individual mode

Set `mode = "individual"` to grade one criterion per call instead of batching them all into one. LLM judges make one request per criterion; agent judges run one turn per criterion, sequentially. For LLM judges, each criterion can also scope its own `files`:

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
mode = "individual"

[[criterion]]
description = "Is the analysis correct?"
files = ["/app/analysis.pdf"]

[[criterion]]
description = "Is the spreadsheet well-structured?"
files = ["/app/data.xlsx"]
```

Criteria without `files` fall back to `[judge].files`.

If a judge call times out, Rewardkit records the affected criteria as `0.0` with an error and warning in `reward-details.json`.

## Sampling the judge

Set `samples` to run the judge several times and score each criterion with its median sample. `reward-details.json` lists every sample's answer and an `agreement` between 0 and 1.

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
samples = 5
```

An agent judge with several samples must be [isolated](/rewardkit/quick-start#isolation). Pass `--yolo` to run without isolation.

## Guarding the judge

`guard` protects the judge from prompt injection: an agent writing instructions for the judge into the files being graded. The judge is told not to follow instructions in the graded files and reports whether it found any under `guard` in `reward-details.json`. `flag` only reports it; `penalize` also scores a flagged submission 0.

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
guard = "penalize"
```

## Provider routing

Judges call LiteLLM, which reads credentials from environment variables. You can change providers at invocation time without editing the rubric:

* `--je KEY=VALUE` sets an environment variable for the run and can be repeated.
* `--judge MODEL_OR_AGENT` overrides `[judge].judge`. The equivalent environment variable is `REWARDKIT_JUDGE`.
* `--model MODEL` overrides `[judge].model` for an agent judge. The equivalent environment variable is `REWARDKIT_MODEL`.

Harbor users can pass the same environment variables with `--ve`.

```bash theme={"system"}
rewardkit /tests \
  --judge bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0 \
  --je AWS_ACCESS_KEY_ID=$AWS_ACCESS_KEY_ID \
  --je AWS_REGION_NAME=us-east-1

rewardkit /tests \
  --judge claude-code \
  --model anthropic/claude-opus-5-5
```

The [LiteLLM provider docs](https://docs.litellm.ai/docs/providers) list the environment variables used by each provider.

### Subscription authentication

For Anthropic LLM judges, Rewardkit uses a Claude subscription token when it is the only Anthropic credential present. Create one with `claude setup-token` and set `CLAUDE_CODE_OAUTH_TOKEN`. If `ANTHROPIC_API_KEY` is also set, the API key has priority. Set `REWARDKIT_FORCE_SUBSCRIPTION=1` to require the subscription token.

For the `codex` agent judge, set `OPENAI_API_KEY` or pass the contents of a ChatGPT-authenticated Codex `auth.json` as `CODEX_AUTH_JSON`. The API key has priority unless `REWARDKIT_FORCE_SUBSCRIPTION=1` is set.

## Configuration reference

Judge TOMLs are validated when the tests directory is scanned, before any judge runs. Unknown keys and unknown values raise an error.

### `[judge]` section

<ParamField body="judge" type="string" default={'"anthropic/claude-opus-5-5"'}>
  LiteLLM model name (e.g. `"anthropic/claude-opus-5-5"`), agent judge name
  (`"claude-code"`, `"codex"`, `"fx"`), or `"jev"`.
</ParamField>

<ParamField body="model" type="string | null" default="null">
  For agent judges, the LLM the agent should use. For the JEV judge, the JEV
  model.
</ParamField>

<ParamField body="files" type="list[string]" default="[]">
  Workspace file paths to include in the judge prompt.
</ParamField>

<ParamField body="mode" type={'"batched" | "individual"'} default={'"batched"'}>
  Whether to grade all criteria in one call (`batched`) or one call per
  criterion (`individual`).
</ParamField>

<ParamField body="timeout" type="integer" default="300">
  Seconds to wait for the judge response.
</ParamField>

<ParamField body="reasoning_effort" type={'"none" | "minimal" | "low" | "medium" | "high" | "xhigh"'} default={'"medium"'}>
  Effort level for LLM judges. Accepts the levels LiteLLM supports; a given
  model may not support all of them.
</ParamField>

<ParamField body="isolated" type="boolean" default="false">
  For agent judges, mount the workspace as read-only via overlayfs.
</ParamField>

<ParamField body="cwd" type="string | null" default="null">
  For agent judges, the working directory the agent runs in. With `isolated`,
  it must be inside the workspace.
</ParamField>

<ParamField body="mcp_servers" type="list[table]" default="[]">
  For agent judges, MCP servers to configure before running. Each entry matches
  a Harbor task's `[[environment.mcp_servers]]`, plus a per-server
  `allowed_tools` allowlist. Codex does not support `sse` servers.
</ParamField>

<ParamField body="reference" type="string | null" default="null">
  Path to a reference solution file for comparison.
</ParamField>

<ParamField body="atif-trajectory" type="string | null" default="null">
  Path to an ATIF trajectory JSON to include in the prompt.
</ParamField>

<ParamField body="weight" type="number" default="1.0">
  Weight of this judge's score when combined with the other scores in the same
  directory.
</ParamField>

<ParamField body="prompt_template" type="string | null" default="null">
  Custom prompt template (`.md` or `.txt`). Must contain a `{criteria}`
  placeholder.
</ParamField>

<ParamField body="samples" type="integer" default="1">
  Number of times to run the judge.
</ParamField>

<ParamField body="guard" type={'"off" | "flag" | "penalize"'} default={'"off"'}>
  Protect the judge from instructions the agent wrote into the graded files.
  `penalize` also scores a flagged submission 0.
</ParamField>

### `[[criterion]]` entries

<ParamField body="description" type="string" required>
  What to grade. This text is sent to the judge.
</ParamField>

<ParamField body="type" type={'"binary" | "likert" | "numeric" | "rubric"'} default={'"binary"'}>
  Output format.
</ParamField>

<ParamField body="name" type="string | null" default="null">
  Identifier for this criterion. Auto-generated from `description` if omitted.
</ParamField>

<ParamField body="id" type="string | null" default="null">
  Stable rubric identifier carried through to `reward-details.json` for
  provenance (e.g. `"1.1"`, `"2.3"`). Independent of `name`; survives rewording
  the description.
</ParamField>

<ParamField body="points" type="integer" default="5">
  Scale size for the `likert` type.
</ParamField>

<ParamField body="min" type="number" default="0.0">
  Minimum value for the `numeric` type.
</ParamField>

<ParamField body="max" type="number" default="1.0">
  Maximum value for the `numeric` type.
</ParamField>

<ParamField body="levels" type="list[string]" default="[]">
  Level descriptions for the `rubric` type, ordered from lowest to highest. 2 to
  10 entries.
</ParamField>

<ParamField body="weight" type="number" default="1.0">
  Importance multiplier for score aggregation. Negative weights support penalty
  criteria and require `aggregation = "weighted-sum"`.
</ParamField>

<ParamField body="files" type="list[string]" default="[]">
  Files scoped to this criterion. Requires `[judge].mode = "individual"`. Falls
  back to `[judge].files` when omitted.
</ParamField>

<ParamField body="negate" type="boolean" default="false">
  Invert the normalized score. Use for a criterion describing behavior the
  answer should not exhibit. See [Negated criteria](#negated-criteria).
</ParamField>

<ParamField body="optional" type="boolean" default="false">
  Exempt this criterion from gating under `aggregation = "required-pass"`.
</ParamField>

### `[scoring]` section

Controls how the judge's criteria are aggregated into a single score. This only affects criteria within this TOML file. It does not change how programmatic and judge scores are combined across the directory.

```toml theme={"system"}
[scoring]
aggregation = "all-pass"  # weighted-mean | weighted-sum | all-pass | any-pass | threshold | required-pass
threshold = 0.7           # only used with "threshold" aggregation
```

`required-pass` returns `1.0` only when every non-`optional` criterion passes (`value > 0`); `optional` criteria never gate. With no non-optional criteria it warns and scores `0.0`.

`weighted-sum` computes `sum(value × weight)` without normalization. It is the only aggregation that accepts negative weights, and its result may fall outside `[0, 1]`.

## Score normalization

* **Binary**: yes/true/1 → 1.0, anything else → 0.0
* **Likert**: normalized to \[0, 1] as `(raw - 1) / (points - 1)`
* **Numeric**: normalized to \[0, 1] as `(raw - min) / (max - min)`
* **Rubric**: levels are numbered from 0, normalized to \[0, 1] as `raw / (levels - 1)`

## Negated criteria

Set `negate = true` for a criterion describing behavior the answer should **not** exhibit. The judge scores the criteria as usual, then the score is inverted (`value → 1 - value`): present → 0.0, absent → 1.0. The raw judge answer is kept in `reward-details.json` so the flip is auditable.

```toml theme={"system"}
[[criterion]]
description = "States there is no task execution history tracking in the database"
type = "binary"
negate = true   # the answer should NOT make this (false) claim
```

## Trajectory evaluation

To grade the agent's process rather than just its output, point the judge at the trajectory file:

```toml theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
atif-trajectory = "/logs/agent/trajectory.json"
files = ["/app/main.py"]

[[criterion]]
description = "Did the agent take an efficient approach?"
type = "likert"
points = 5
```

The trajectory content is truncated proportionally to fit within the model's context window while preserving all steps.

## Custom prompt templates

You can provide your own prompt instead of the built-in one:

```toml theme={"system"}
[judge]
judge = "anthropic/claude-opus-5-5"
prompt_template = "my_prompt.md"
```

The template must contain a `{criteria}` placeholder where criterion descriptions get injected.

## JEV judge

[JEV](https://docs.typesafe.ai) is a new type of language model from TypeSafe. It returns a probability or a rubric score per criterion and no reasoning, so it is fast and cheap. It requires the `jev` extra (`uv tool install harbor-rewardkit[jev]`) and a `TYPESAFE_API_KEY`:

```toml title="tests/quality.toml" theme={"system"}
[judge]
judge = "jev"
files = ["/app/answer.md"]

[[criterion]]
description = "Does the answer address the requested task?"
type = "binary"

[[criterion]]
description = "How complete is the answer?"
type = "rubric"
levels = [
  "Omits the requested information",
  "Provides some requested information but misses important details",
  "Provides all requested information",
]
```

Binary criteria score 1.0 at a probability of 0.5 or higher. Rubric `levels` must run from worst to best, because the position in the list sets the score. The raw probability or score is kept in `reward-details.json`.

JEV grades text files only and supports `binary` and `rubric` criteria. `atif-trajectory` and `prompt_template` are not supported. The files and the longest criterion must fit in 32k tokens, and the task image needs CA certificates (`ca-certificates` on Debian and Ubuntu).

To route through a LiteLLM proxy or Vercel AI Gateway, set `TYPESAFE_BASE_URL` to the gateway's TypeSafe endpoint and `TYPESAFE_API_KEY` to the gateway key. Vercel also needs `model = "typesafe-ai/jev"`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.