Benchmark Specifications
Benchmark specs define runnable benchmark suites (datasets, metrics, and difficulty metadata).
Each benchmark must declare one or more evaluator dependencies in the
evaluators field.
Required Fields​
idversionnamedescriptioncategorytask_countmetricevaluators
Common Fields​
difficultylanguagesdataset_sourcesupports_live_monitoringsupports_experiment_comparisonevaluator_shapesrecommended_windowstrace_integrationdataset_editabilitysdk_support
Notes​
- Benchmarks are referenced from agentspecs and runtime workflows.
- Benchmark generation validates that each evaluator reference exists in the eval catalog.
- Prefer explicit
id:versionreferences.