Skip to main content
Metrics define how to aggregate rewards across tasks in a job or dataset. By default, Harbor averages rewards across tasks and treats missing rewards as 0. However, you may want to define custom logic or handle missing rewards differently. You can do this by creating a metric.py file. If you are working with a dataset.toml, you can run the following command to create a metric.py file and automatically add it to the [[files]] section of your dataset.toml:
If you are running a local dataset, it will automatically use the metric.py if it’s present in the dataset directory.

Default behavior

By default, Harbor averages rewards across tasks and treats missing rewards as 0.
Rewards
Aggregate
For single-dimension rewards, Harbor names the aggregate after the metric (mean). For multi-dimensional rewards, it aggregates each dimension independently. Missing rewards and missing dimensions count as 0.

Custom metrics with metric.py

A metric.py script must accept the following arguments:
  • -i, --input-path: rewards JSONL file
  • -o, --output-path: output JSON file with computed metrics

Example

metric.py

Considerations

When implementing a custom metric, be sure to account for
  1. Missing or invalid reward keys
  2. Invalid reward values
  3. Null rows in the JSONL input
  4. Default aggregate reward keys

Publishing custom metrics

When publishing a dataset, Harbor automatically publishes the metric.py file in the same directory:

Output format

metric.py output should be a JSON object. Multiple metric values are allowed: