Adding Metrics to Traces
Compute scores in your application code and attach them to traces for monitoring and analysis. Use cases: Format validation, safety checks, PII detection, latency tracking, relevance scores.For complete documentation on adding metrics to traces, see Custom Metrics.
Evaluator Functions for Experiments
Define scoring functions that run locally duringevaluate() to score outputs against expected results.
Writing an Evaluator
Evaluators receive three arguments and return a score:Use
ground_truth (singular) for both the datapoint field and the evaluator argument. ground_truths was the pre-1.0 SDK name.Running Evaluators
Pass evaluator functions toevaluate():
For a complete tutorial with real examples, see Run Your First Experiment.
Evaluating Multi-Step Pipelines
For pipelines with multiple steps, combine both approaches:- Session-level: Pass evaluators to
evaluate()for overall scoring - Span-level: Use
enrich_span()within traced functions for step-specific metrics
answer_qualityscores at the session levelretrieval_score,num_docs,answer_lengthat the span level
Next Steps
Custom Metrics
Full guide to adding metrics to traces
Run Your First Experiment
Complete tutorial with real examples
Sync Offline Evaluations
Upload results you scored yourself, with or without the SDK
Server-Side Evaluators
Run evaluators on HoneyHive infrastructure
LLM-as-Judge
Use LLMs to evaluate outputs

