Skip to main content
Define evaluators, datasets, and datapoints as files under .honeyhive/ and apply them to a project with the HoneyHive CLI. For the directory layout and schema discovery, see Config as Code. Use a project API key (HH_PROJECT_API_KEY), created under Settings > Project > API Keys on the Project tab. It reaches one project.

Define an evaluator

Python evaluators are functions that score events on the server. To define one as a file, capture the same fields you would set in the Evaluators UI:
.honeyhive/evaluators/keyword-check.yaml
LLM evaluators follow the same pattern. The prompt goes in criteria, and the model is selected with model_provider and model_name. return_type is float for numeric scores, boolean for pass/fail, string for free-form, or categorical when paired with a categories list:
.honeyhive/evaluators/relevance-llm.yaml
Every return_type: float evaluator needs scale, the top of the rating range. Here it is 5, because the prompt rates 1 to 5. The JSON Schema marks scale as optional, but the API rejects a float evaluator without it with a 400.
To see every field the API accepts (categories, thresholds, event filters, etc.), inspect the schema:

Apply a file

Pass --filename (or -f) to send the whole file as the request body. The file can be .yaml, .yml, .json, or .jsonc.
Each resource returns its new ID in a different place: To update a resource later, you need its ID. Pick one way to keep IDs and use it across your repo:
  • Embedded ID: write the ID into the YAML file. This is the simplest option. See the example below.
  • Lockfile: keep IDs in .honeyhive/state.json and leave the YAML files unchanged. See A simple sync script.
To update an existing evaluator, include metric_id in the file and call metrics update:
.honeyhive/evaluators/keyword-check.yaml
The same --filename flow works for datasets create, datasets update, datapoints create, datapoints update, and every other namespace. Run --show-file-schema on the command you want to use to see the exact shape.

Define a dataset

A dataset definition only needs a name, an optional description, and the datapoint IDs it should include:
.honeyhive/datasets/qa-eval-set.yaml
Apply it the same way:
For each datapoint, define inputs and ground truth and link it to the dataset on create:
.honeyhive/datapoints/q1.yaml
Link from the datapoint side (linked_datasets), as shown above. It keeps the dataset YAML stable across runs.
This page covers the shape of your dataset as code. For keeping the contents of a dataset in sync with an external source (S3, a database, an internal tool), see Sync from External Sources.

A simple sync script

The CLI has no apply command. This script implements the lockfile pattern from Apply a file: it upserts every evaluator under .honeyhive/evaluators/ and tracks IDs in .honeyhive/state.json. Commit the lockfile so collaborators and CI share IDs. To apply the same files to several projects, run the script once per project with that project’s key and its own lockfile, for example STATE_FILE=.honeyhive/state-staging.json HH_PROJECT_API_KEY=... bash sync-evaluators.sh.
sync-evaluators.sh
Use --verbose (or HH_VERBOSE=true) when debugging to log the resolved API URL and masked key for each invocation. This makes it obvious whether a CI job is hitting staging or production.

Run from CI

A minimal GitHub Actions job:
.github/workflows/honeyhive-sync.yml
For staging/production parity, pass different HH_PROJECT_API_KEY and HH_DATA_PLANE_URL values per environment without changing any file under .honeyhive/.

Validate before applying

Check your files before you push. Save the command’s schema, then run a JSON Schema validator such as check-jsonschema. Files without an ID use the create schema. Files with an embedded ID use the update schema:

Sync Datasets from External Sources

Keep dataset contents in sync with S3, databases, or internal tools.

Python and LLM Evaluators

Background on the evaluator model the YAML files describe.