.honeyhive/ and apply them to a project with the HoneyHive CLI. For the directory layout and schema discovery, see Config as Code.
Use a project API key (HH_PROJECT_API_KEY), created under Settings > Project > API Keys on the Project tab. It reaches one project.
Define an evaluator
Python evaluators are functions that score events on the server. To define one as a file, capture the same fields you would set in the Evaluators UI:.honeyhive/evaluators/keyword-check.yaml
criteria, and the model is selected with model_provider and model_name. return_type is float for numeric scores, boolean for pass/fail, string for free-form, or categorical when paired with a categories list:
.honeyhive/evaluators/relevance-llm.yaml
Every
return_type: float evaluator needs scale, the top of the rating range. Here it is 5, because the prompt rates 1 to 5. The JSON Schema marks scale as optional, but the API rejects a float evaluator without it with a 400.Apply a file
Pass--filename (or -f) to send the whole file as the request body. The file can be .yaml, .yml, .json, or .jsonc.
To update a resource later, you need its ID. Pick one way to keep IDs and use it across your repo:
- Embedded ID: write the ID into the YAML file. This is the simplest option. See the example below.
- Lockfile: keep IDs in
.honeyhive/state.jsonand leave the YAML files unchanged. See A simple sync script.
metric_id in the file and call metrics update:
.honeyhive/evaluators/keyword-check.yaml
--filename flow works for datasets create, datasets update, datapoints create, datapoints update, and every other namespace. Run --show-file-schema on the command you want to use to see the exact shape.
Define a dataset
A dataset definition only needs a name, an optional description, and the datapoint IDs it should include:.honeyhive/datasets/qa-eval-set.yaml
.honeyhive/datapoints/q1.yaml
Link from the datapoint side (
linked_datasets), as shown above. It keeps the dataset YAML stable across runs.This page covers the shape of your dataset as code. For keeping the contents of a dataset in sync with an external source (S3, a database, an internal tool), see Sync from External Sources.
A simple sync script
The CLI has noapply command. This script implements the lockfile pattern from Apply a file: it upserts every evaluator under .honeyhive/evaluators/ and tracks IDs in .honeyhive/state.json. Commit the lockfile so collaborators and CI share IDs.
To apply the same files to several projects, run the script once per project with that project’s key and its own lockfile, for example STATE_FILE=.honeyhive/state-staging.json HH_PROJECT_API_KEY=... bash sync-evaluators.sh.
sync-evaluators.sh
Run from CI
A minimal GitHub Actions job:.github/workflows/honeyhive-sync.yml
HH_PROJECT_API_KEY and HH_DATA_PLANE_URL values per environment without changing any file under .honeyhive/.
Validate before applying
Check your files before you push. Save the command’s schema, then run a JSON Schema validator such ascheck-jsonschema. Files without an ID use the create schema. Files with an embedded ID use the update schema:
Related references
Sync Datasets from External Sources
Keep dataset contents in sync with S3, databases, or internal tools.
Python and LLM Evaluators
Background on the evaluator model the YAML files describe.