Harbor Hub intentionally does not calculate row scores from linked trials, to enable maximally flexible leaderboard construction.
<leaderboard> is either a UUID or an org/dataset/leaderboard slug.
General approach
- Choose a published dataset and the versions the leaderboard will cover.
- Generate a configuration with
init, then define the metadata, metrics, columns, and ranking rules. - Prepare your results as rows, calculating the scores yourself. Include trial IDs if you want to link each result to its runs.
- Create the leaderboard with its definition and rows files. Keep it private while reviewing the results.
- Inspect the leaderboard, check the rankings and trial links, then make it public when it is ready.
Browse
list optionally filters by dataset slug or ID.
Create
Generate a configuration file, edit it, then create the leaderboard:visibility: public in the file or pass --visibility public to create.
Configuration schema
The definition file passed tocreate --config accepts the fields below. Use either package or package_id. Rows go in a separate file passed with --rows.
string
Dataset slug, such as
acme/my-dataset. Required unless package_id is provided.string (UUID)
Dataset package UUID, as an alternative to
package.string
required
Leaderboard slug, up to 100 characters. Starts with a lowercase letter or digit; may also contain
., _, and -.string
required
Display title, from 1 to 200 characters.
string | null
Optional description, up to 5,000 characters.
string
default:"private"
public or private.object
default:"{}"
Schema for each row’s metadata, such as the agent name and configuration.
object
default:"{}"
Schema for each row’s metrics, such as scores and costs.
object[]
default:"[]"
Display columns, in left-to-right order.
object[]
default:"[]"
Ranking rules, evaluated in order to break ties.
string[]
Dataset version refs to associate, such as
latest or a version tag. Resolved to fixed UUIDs at creation.string (UUID)[]
Dataset version UUIDs to associate. Can be combined with refs. Omit both fields to associate all existing versions; set both to
[] for none.Columns and ranking
The generated file includes an example with an agent name and reward score:Dataset versions
By default, a new leaderboard is associated with every dataset version that exists at creation time. Later versions are not added automatically. To select versions explicitly, adddataset_version_refs or dataset_version_ids to the configuration:
[] associates no versions.
Add rows
Createrows.yaml with the values your columns expect:
Row file schema
The file passed torow create --config or create --rows has this structure:
object[]
required
Between 1 and 500 new rows. Trial IDs must be unique across the rows in the request.
create:
Edit a leaderboard
Change its title, description, or visibility directly:--description to change the description or --visibility private to make the board private. To replace dataset-version associations, repeat --dataset-version-ref or --dataset-version-id. Omit these fields to preserve existing associations.
For columns, schemas, or ranking changes, export the definition, edit it, then apply it:
update.
Edit rows
List rows to find their IDs, then inspect or edit one:metadata, metrics, or status. To hide, restore, or permanently delete rows:
--yes for an interactive confirmation. The CLI has no command to delete an entire leaderboard.
Batch edits
Export all visible rows, edit them, then apply the file:--dry-run requires both a definition change and row updates.
Link trials
Trial links record which runs support a row.set replaces all links; add and remove change only the specified links.
--trial-id for multiple trials. Alternatively, replace the links from a YAML or JSON file containing a non-empty trial_ids list:
--clear to remove all links.
Files and scripting
Configuration files accept YAML or JSON. Exports require a.yaml, .yml, or .json extension; init also accepts --format json. Use --force to overwrite an existing output file.
Read and mutation commands support --json. Use list --quiet for leaderboard slugs, row list --quiet for row IDs, and row trial list --quiet for trial IDs.
Row and trial lists support --limit (default 50, maximum 1,000) and --page (starting at 1). JSON returns one page. Quiet or piped output streams all pages unless --page is supplied; --no-headers removes piped table headers.
Display a leaderboard on your own website
Fetch a public leaderboard without authentication to display it on your website:leaderboard definition, ranked rows, and pagination. Fetch through pagination.total_pages for all rows. You can also select a board by leaderboard_id instead of package and name.
For private boards, authenticate from your server; keep API keys out of browser code.
See terminal-bench-2-1 for an example pipeline and tbench.ai for custom visuals backed by Hub. Links to trials and versioned datasets let readers audit results and reproduce evaluations.
