Skip to main content
A leaderboard is a ranked table of results for a dataset. You define its columns and ranking rules, add rows with metadata and scores, and optionally link each row to the trials behind it.
Harbor Hub intentionally does not calculate row scores from linked trials, to enable maximally flexible leaderboard construction.
Public leaderboards can be read without signing in. Private leaderboards are visible to members of the owning organization. Creating a leaderboard requires owner access to the dataset’s organization. In the commands below, <leaderboard> is either a UUID or an org/dataset/leaderboard slug.

General approach

  1. Choose a published dataset and the versions the leaderboard will cover.
  2. Generate a configuration with init, then define the metadata, metrics, columns, and ranking rules.
  3. Prepare your results as rows, calculating the scores yourself. Include trial IDs if you want to link each result to its runs.
  4. Create the leaderboard with its definition and rows files. Keep it private while reviewing the results.
  5. Inspect the leaderboard, check the rankings and trial links, then make it public when it is ready.

Browse

list optionally filters by dataset slug or ID.

Create

Generate a configuration file, edit it, then create the leaderboard:
New leaderboards are private. Set visibility: public in the file or pass --visibility public to create.

Configuration schema

The definition file passed to create --config accepts the fields below. Use either package or package_id. Rows go in a separate file passed with --rows.
string
Dataset slug, such as acme/my-dataset. Required unless package_id is provided.
string (UUID)
Dataset package UUID, as an alternative to package.
string
required
Leaderboard slug, up to 100 characters. Starts with a lowercase letter or digit; may also contain ., _, and -.
string
required
Display title, from 1 to 200 characters.
string | null
Optional description, up to 5,000 characters.
string
default:"private"
public or private.
object
default:"{}"
Schema for each row’s metadata, such as the agent name and configuration.
object
default:"{}"
Schema for each row’s metrics, such as scores and costs.
object[]
default:"[]"
Display columns, in left-to-right order.
object[]
default:"[]"
Ranking rules, evaluated in order to break ties.
string[]
Dataset version refs to associate, such as latest or a version tag. Resolved to fixed UUIDs at creation.
string (UUID)[]
Dataset version UUIDs to associate. Can be combined with refs. Omit both fields to associate all existing versions; set both to [] for none.

Columns and ranking

The generated file includes an example with an agent name and reward score:
This displays an agent column and a reward column, with the highest reward ranked first.

Dataset versions

By default, a new leaderboard is associated with every dataset version that exists at creation time. Later versions are not added automatically. To select versions explicitly, add dataset_version_refs or dataset_version_ids to the configuration:
Refs resolve to fixed version UUIDs. Both fields can be combined; setting both to [] associates no versions.

Add rows

Create rows.yaml with the values your columns expect:
Then add the rows:

Row file schema

The file passed to row create --config or create --rows has this structure:
object[]
required
Between 1 and 500 new rows. Trial IDs must be unique across the rows in the request.
To create the leaderboard and its rows together, pass the separate rows file to create:

Edit a leaderboard

Change its title, description, or visibility directly:
Use --description to change the description or --visibility private to make the board private. To replace dataset-version associations, repeat --dataset-version-ref or --dataset-version-id. Omit these fields to preserve existing associations. For columns, schemas, or ranking changes, export the definition, edit it, then apply it:
Flags override corresponding values in the file. The package and leaderboard name cannot be changed through update.

Edit rows

List rows to find their IDs, then inspect or edit one:
Row updates change metadata, metrics, or status. To hide, restore, or permanently delete rows:
Deleting rows also removes their trial links. Omit --yes for an interactive confirmation. The CLI has no command to delete an entire leaderboard.

Batch edits

Export all visible rows, edit them, then apply the file:
If a schema change requires row changes, apply both files together so they succeed or fail as one transaction:
Export both files before editing. Their timestamps protect against overwriting newer changes. --dry-run requires both a definition change and row updates. Trial links record which runs support a row. set replaces all links; add and remove change only the specified links.
Repeat --trial-id for multiple trials. Alternatively, replace the links from a YAML or JSON file containing a non-empty trial_ids list:
Use the file option on its own; use --clear to remove all links.

Files and scripting

Configuration files accept YAML or JSON. Exports require a .yaml, .yml, or .json extension; init also accepts --format json. Use --force to overwrite an existing output file. Read and mutation commands support --json. Use list --quiet for leaderboard slugs, row list --quiet for row IDs, and row trial list --quiet for trial IDs. Row and trial lists support --limit (default 50, maximum 1,000) and --page (starting at 1). JSON returns one page. Quiet or piped output streams all pages unless --page is supplied; --no-headers removes piped table headers.

Display a leaderboard on your own website

Fetch a public leaderboard without authentication to display it on your website:
The response contains the leaderboard definition, ranked rows, and pagination. Fetch through pagination.total_pages for all rows. You can also select a board by leaderboard_id instead of package and name. For private boards, authenticate from your server; keep API keys out of browser code. See terminal-bench-2-1 for an example pipeline and tbench.ai for custom visuals backed by Hub. Links to trials and versioned datasets let readers audit results and reproduce evaluations.