> ## Documentation Index
> Fetch the complete documentation index at: https://docs.honeyhive.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluators and Datasets as Code

> Define HoneyHive evaluators, datasets, and datapoints as YAML files, apply them with the CLI, and keep them in sync from a script or CI.

Define evaluators, datasets, and datapoints as files under `.honeyhive/` and apply them to a project with the [HoneyHive CLI](/v2/cli-reference/getting-started). For the directory layout and schema discovery, see [Config as Code](/v2/cli-reference/config-as-code).

Use a project API key (`HH_PROJECT_API_KEY`), created under **Settings > Project > API Keys** on the **Project** tab. It reaches one project.

## Define an evaluator

[Python evaluators](/v2/evaluators/python) are functions that score events on the server. To define one as a file, capture the same fields you would set in the [Evaluators UI](https://app.us.honeyhive.ai/metrics):

```yaml .honeyhive/evaluators/keyword-check.yaml theme={null}
name: keyword-check
type: PYTHON
return_type: boolean
description: Checks whether the response mentions the word "honey".
criteria: |
  def keyword_check():
      return "honey" in outputs["content"].lower()
```

[LLM evaluators](/v2/evaluators/llm) follow the same pattern. The prompt goes in `criteria`, and the model is selected with `model_provider` and `model_name`. `return_type` is `float` for numeric scores, `boolean` for pass/fail, `string` for free-form, or `categorical` when paired with a `categories` list:

```yaml .honeyhive/evaluators/relevance-llm.yaml theme={null}
name: relevance-llm
type: LLM
return_type: float
scale: 5
model_provider: openai
model_name: gpt-4o
sampling_percentage: 25
description: Rates how well the answer addresses the question, 1-5.
criteria: |
  [Instruction]
  Rate the assistant's answer for relevance to the question on a scale of 1 to 5.

  [Question]
  {{ inputs.question }}

  [Answer]
  {{ outputs.content }}

  [Evaluation]
  Rating: [[X]]
```

<Note>
  Every `return_type: float` evaluator needs `scale`, the top of the rating range. Here it is `5`, because the prompt rates 1 to 5. The JSON Schema marks `scale` as optional, but the API rejects a float evaluator without it with a `400`.
</Note>

To see every field the API accepts (categories, thresholds, event filters, etc.), inspect the schema:

```bash theme={null}
honeyhive metrics create --show-file-schema | jq '.properties | keys'
```

## Apply a file

Pass `--filename` (or `-f`) to send the whole file as the request body. The file can be `.yaml`, `.yml`, `.json`, or `.jsonc`.

```bash theme={null}
# Create the evaluator
honeyhive metrics create --filename .honeyhive/evaluators/keyword-check.yaml
# {
#   "inserted": true,
#   "metric_id": "01KRJB6SX9YA4J51NRFT6M27RC"
# }
```

Each resource returns its new ID in a different place:

| Resource | Get the new ID | An update file needs |
| - | - | - |
| `metrics` | `jq -r '.metric_id'` | `metric_id` |
| `datasets` | `jq -r '.result.insertedId'` | `dataset_id` |
| `datapoints` | `jq -r '.result.insertedIds[0]'` | `datapoint_id` |

To update a resource later, you need its ID. Pick one way to keep IDs and use it across your repo:

* **Embedded ID**: write the ID into the YAML file. This is the simplest option. See the example below.
* **Lockfile**: keep IDs in `.honeyhive/state.json` and leave the YAML files unchanged. See [A simple sync script](#a-simple-sync-script).

To update an existing evaluator, include `metric_id` in the file and call `metrics update`:

```yaml .honeyhive/evaluators/keyword-check.yaml theme={null}
metric_id: 01KRJB6SX9YA4J51NRFT6M27RC
name: keyword-check
type: PYTHON
return_type: boolean
description: Checks whether the response mentions the word "honey" or "hive".
criteria: |
  def keyword_check():
      content = outputs["content"].lower()
      return "honey" in content or "hive" in content
```

```bash theme={null}
honeyhive metrics update --filename .honeyhive/evaluators/keyword-check.yaml
```

The same `--filename` flow works for `datasets create`, `datasets update`, `datapoints create`, `datapoints update`, and every other namespace. Run `--show-file-schema` on the command you want to use to see the exact shape.

## Define a dataset

A dataset definition only needs a name, an optional description, and the datapoint IDs it should include:

```yaml .honeyhive/datasets/qa-eval-set.yaml theme={null}
name: qa-eval-set
description: Question/answer pairs used for the relevance evaluator.
datapoints: []
```

Apply it the same way:

```bash theme={null}
DATASET_ID=$(honeyhive datasets create \
  --filename .honeyhive/datasets/qa-eval-set.yaml \
  | jq -r '.result.insertedId')
```

For each datapoint, define inputs and ground truth and link it to the dataset on create:

```yaml .honeyhive/datapoints/q1.yaml theme={null}
inputs:
  question: What is the capital of France?
ground_truth:
  answer: Paris
metadata:
  external_id: q1
linked_datasets:
  - "01KRJB7WD8E2H4M9X3K2Y7Q1A5"  # paste the dataset_id printed by the create above
```

```bash theme={null}
honeyhive datapoints create --filename .honeyhive/datapoints/q1.yaml
```

<Note>
  Link from the datapoint side (`linked_datasets`), as shown above. It keeps the dataset YAML stable across runs.
</Note>

<Note>
  This page covers the **shape** of your dataset as code. For keeping the **contents** of a dataset in sync with an external source (S3, a database, an internal tool), see [Sync from External Sources](/v2/datasets/sync).
</Note>

## A simple sync script

The CLI has no `apply` command. This script implements the lockfile pattern from [Apply a file](#apply-a-file): it upserts every evaluator under `.honeyhive/evaluators/` and tracks IDs in `.honeyhive/state.json`. Commit the lockfile so collaborators and CI share IDs.

To apply the same files to several projects, run the script once per project with that project's key and its own lockfile, for example `STATE_FILE=.honeyhive/state-staging.json HH_PROJECT_API_KEY=... bash sync-evaluators.sh`.

```bash sync-evaluators.sh theme={null}
#!/usr/bin/env bash
# Usage: HH_PROJECT_API_KEY=... bash sync-evaluators.sh
set -euo pipefail

STATE_FILE="${STATE_FILE:-.honeyhive/state.json}"
[ -f "$STATE_FILE" ] || echo '{"evaluators": {}}' > "$STATE_FILE"

for file in .honeyhive/evaluators/*.yaml; do
  name=$(yq '.name' "$file")
  existing_id=$(jq -r --arg n "$name" '.evaluators[$n] // ""' "$STATE_FILE")

  if [ -n "$existing_id" ]; then
    # Update in place; CLI reads metric_id from the file body.
    # mktemp + .yaml suffix appended manually so this works on macOS (BSD
    # mktemp does not support --suffix).
    tmp="$(mktemp).yaml"
    yq ". + {\"metric_id\": \"$existing_id\"}" "$file" > "$tmp"
    honeyhive metrics update --filename "$tmp" > /dev/null
    rm -f "$tmp"
    echo "Updated $name ($existing_id)"
  else
    new_id=$(honeyhive metrics create --filename "$file" | jq -r '.metric_id')
    tmp=$(mktemp)
    jq --arg n "$name" --arg id "$new_id" \
      '.evaluators[$n] = $id' "$STATE_FILE" > "$tmp"
    mv "$tmp" "$STATE_FILE"
    echo "Created $name ($new_id)"
  fi
done
```

<Tip>
  Use `--verbose` (or `HH_VERBOSE=true`) when debugging to log the resolved API URL and masked key for each invocation. This makes it obvious whether a CI job is hitting staging or production.
</Tip>

## Run from CI

A minimal GitHub Actions job:

```yaml .github/workflows/honeyhive-sync.yml theme={null}
name: Sync HoneyHive resources

on:
  push:
    branches: [main]
    paths: [".honeyhive/**"]

jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Install HoneyHive CLI
        # Pin to a specific release so CI runs are reproducible.
        run: curl -fsSL https://github.com/honeyhiveai/honeyhive-cli/releases/download/v1.8.0/install.sh | sh
      - name: Apply evaluators
        env:
          HH_PROJECT_API_KEY: ${{ secrets.HH_PROJECT_API_KEY }}
        run: bash sync-evaluators.sh
```

For staging/production parity, pass different `HH_PROJECT_API_KEY` and `HH_DATA_PLANE_URL` values per environment without changing any file under `.honeyhive/`.

## Validate before applying

Check your files before you push. Save the command's schema, then run a JSON Schema validator such as `check-jsonschema`. Files without an ID use the `create` schema. Files with an embedded ID use the `update` schema:

<CodeGroup>
  ```bash Lockfile pattern theme={null}
  # Files have no metric_id; validate against the create schema.
  honeyhive metrics create --show-file-schema > /tmp/metric-create.schema.json
  check-jsonschema --schemafile /tmp/metric-create.schema.json .honeyhive/evaluators/*.yaml
  ```

  ```bash Embedded-ID pattern theme={null}
  # Files have metric_id embedded; validate against the update schema.
  honeyhive metrics update --show-file-schema > /tmp/metric-update.schema.json
  check-jsonschema --schemafile /tmp/metric-update.schema.json .honeyhive/evaluators/*.yaml
  ```
</CodeGroup>

## Related references

<CardGroup cols={2}>
  <Card title="Sync Datasets from External Sources" icon="rotate" href="/v2/datasets/sync">
    Keep dataset contents in sync with S3, databases, or internal tools.
  </Card>

  <Card title="Python and LLM Evaluators" icon="flask" href="/v2/evaluators/introduction">
    Background on the evaluator model the YAML files describe.
  </Card>
</CardGroup>


## Related topics

- [How to migrate from Langfuse to HoneyHive](/v2/tracing/migrate-from-langfuse.md)
- [Introduction](/v2/evaluators/introduction.md)
- [Introducing HoneyHive](/v2/introduction/what-is-hhai.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.