5a558eb09e
TypeScript SDK Compatibility V1.x E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / TypeScript SDK Compatibility V1.x E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
TypeScript SDK E2E Tests / TypeScript SDK E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
Python SDK E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK E2E Tests / Python SDK E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK E2E Tests / build-opik (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Python SDK Compatibility V1.x E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK E2E Tests / build-opik (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer E2E Tests Python ${{matrix.python_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer Integration Smoke Tests (push) Has been cancelled
🐙 Code Quality / detect (push) Has been cancelled
🐙 Code Quality / lint (${{ matrix.leg.name }}) (push) Has been cancelled
🐙 Code Quality / summary (push) Has been cancelled
TypeScript SDK Library Integration Tests / Check Secrets (push) Has been cancelled
TypeScript SDK Library Integration Tests / opik-vercel (Vercel AI SDK / eve) (push) Has been cancelled
SDK Library Integration Tests Runner / Check Secrets (push) Has been cancelled
SDK Library Integration Tests Runner / Missed OpenAI API Key Warning (push) Has been cancelled
SDK Library Integration Tests Runner / Build (push) Has been cancelled
SDK Library Integration Tests Runner / openai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_legacy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / llama_index_tests (push) Has been cancelled
SDK Library Integration Tests Runner / anthropic_tests (push) Has been cancelled
SDK Library Integration Tests Runner / mistral_tests (push) Has been cancelled
SDK Library Integration Tests Runner / groq_tests (push) Has been cancelled
SDK Library Integration Tests Runner / aisuite_tests (push) Has been cancelled
SDK Library Integration Tests Runner / haystack_tests (push) Has been cancelled
SDK Library Integration Tests Runner / dspy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v1_tests (push) Has been cancelled
SDK Library Integration Tests Runner / genai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_legacy_1_3_0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / evaluation_metrics_tests (push) Has been cancelled
SDK Library Integration Tests Runner / bedrock_tests (push) Has been cancelled
SDK Library Integration Tests Runner / litellm_tests (push) Has been cancelled
SDK Library Integration Tests Runner / harbor_tests (push) Has been cancelled
SDK Library Integration Tests Runner / Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / render-equality (push) Has been cancelled
Opik Optimizer - Unit Tests / Opik Optimizer Unit Tests Python ${{matrix.python_version}} (push) Has been cancelled
Python BE E2E Tests / Python BE E2E (push) Has been cancelled
Python Backend Tests / run-python-backend-tests (push) Has been cancelled
Python SDK Unit Tests / Python SDK Unit Tests ${{matrix.python_version}} (push) Has been cancelled
Release Drafter / update_release_draft (push) Has been cancelled
SDK E2E Libraries Integration Tests / Check Secrets (push) Has been cancelled
SDK E2E Libraries Integration Tests / Missed OpenAI API Key Warning (push) Has been cancelled
SDK E2E Libraries Integration Tests / build-opik (push) Has been cancelled
SDK E2E Libraries Integration Tests / E2E Lib Integration Python ${{matrix.python_version}} (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-gemini) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-langchain) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-openai) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-otel) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-vercel) (push) Has been cancelled
TypeScript SDK Build & Publish / build-and-publish (push) Has been cancelled
TypeScript SDK Unit Tests / Test on Node ${{ matrix.node-version }} (push) Has been cancelled
Backend Tests / discover-tests (push) Has been cancelled
Backend Tests / ${{ matrix.name }} (push) Has been cancelled
Build and Publish SDK / build-and-publish (push) Has been cancelled
Build Opik Docker Images / set-version (push) Has been cancelled
Build Opik Docker Images / build-backend (push) Has been cancelled
Build Opik Docker Images / build-sandbox-executor-python (push) Has been cancelled
Build Opik Docker Images / build-python-backend (push) Has been cancelled
Build Opik Docker Images / build-frontend (push) Has been cancelled
Build Opik Docker Images / create-git-tag (push) Has been cancelled
ClickHouse Migration Cluster Check / validate-clickhouse-migrations (push) Has been cancelled
Docs - Publish / run (push) Has been cancelled
E2E Tests - Post Merge (v2) / 🧪 E2E v2 Tests (${{ github.event.inputs.tier || 't1' }}) (push) Has been cancelled
E2E Tests - Post Merge (v2) / 📢 Slack Notification (push) Has been cancelled
Frontend Unit Tests / Test on Node 20 (push) Has been cancelled
Guardrails E2E Tests / Select Python version matrix (push) Has been cancelled
Guardrails E2E Tests / Guardrails E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Guardrails E2E Tests / 📢 Slack Notification (push) Has been cancelled
Guardrails Backend Unit Tests / Guardrails Backend Unit Tests (push) Has been cancelled
Guardrails Backend Unit Tests / 📢 Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v3.21.0) (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v4.2.0) (push) Has been cancelled
Lint Opik Helm Chart / unittest-helm-chart (push) Has been cancelled
137 lines
9.6 KiB
Plaintext
137 lines
9.6 KiB
Plaintext
---
|
||
description: Describes all the built-in evaluation metrics provided by Opik
|
||
headline: Overview | Opik Documentation
|
||
og:description: Explore Opik's evaluation metrics to assess LLM behavior with heuristic
|
||
and LLM as a Judge methods for precise analysis.
|
||
og:site_name: Opik Documentation
|
||
og:title: Evaluation Metrics Overview - Opik
|
||
title: Overview
|
||
canonical-url: https://www.comet.com/docs/opik/evaluation/metrics/overview
|
||
---
|
||
|
||
# Overview
|
||
|
||
Opik provides a set of built-in evaluation metrics that you can mix and match to evaluate LLM behaviour. These metrics are broken down into two main categories:
|
||
|
||
1. **Heuristic metrics** – deterministic checks that rely on rules, statistics, or classical NLP algorithms.
|
||
2. **LLM as a Judge metrics** – delegate scoring to an LLM so you can capture semantic, task-specific, or conversation-level quality signals.
|
||
|
||
Heuristic metrics are ideal when you need reproducible checks such as exact matching, regex validation, or similarity scores against a reference. LLM as a Judge metrics are useful when you want richer qualitative feedback (hallucination detection, helpfulness, summarisation quality, regulatory risk, etc.).
|
||
|
||
## Built-in metrics
|
||
|
||
### Heuristic metrics
|
||
|
||
| Metric | Description | Documentation |
|
||
| --- | --- | --- |
|
||
| BERTScore | Contextual embedding similarity score | [BERTScore](/v1/evaluation/metrics/heuristic_metrics#bertscore) |
|
||
| ChrF | Character n-gram F-score (chrF / chrF++) | [ChrF](/v1/evaluation/metrics/heuristic_metrics#chrf) |
|
||
| Contains | Checks whether the output contains a specific substring | [Contains](/v1/evaluation/metrics/heuristic_metrics#contains) |
|
||
| Corpus BLEU | Computes corpus-level BLEU across multiple outputs | [CorpusBLEU](/v1/evaluation/metrics/heuristic_metrics#bleu) |
|
||
| Equals | Checks if the output exactly matches an expected string | [Equals](/v1/evaluation/metrics/heuristic_metrics#equals) |
|
||
| GLEU | Estimates grammatical fluency for candidate sentences | [GLEU](/v1/evaluation/metrics/heuristic_metrics#gleu) |
|
||
| IsJson | Validates that the output can be parsed as JSON | [IsJson](/v1/evaluation/metrics/heuristic_metrics#isjson) |
|
||
| JSDivergence | Jensen–Shannon similarity between token distributions | [JSDivergence](/v1/evaluation/metrics/heuristic_metrics#jsdivergence) |
|
||
| JSDistance | Raw Jensen–Shannon divergence | [JSDistance](/v1/evaluation/metrics/heuristic_metrics#jsdistance) |
|
||
| KLDivergence | Kullback–Leibler divergence with smoothing | [KLDivergence](/v1/evaluation/metrics/heuristic_metrics#kldivergence) |
|
||
| Language Adherence | Verifies output language code | [Language Adherence](/v1/evaluation/metrics/heuristic_metrics#language-adherence) |
|
||
| Levenshtein | Calculates the normalized Levenshtein distance between output and reference | [Levenshtein](/v1/evaluation/metrics/heuristic_metrics#levenshteinratio) |
|
||
| Readability | Reports Flesch Reading Ease and FK grade | [Readability](/v1/evaluation/metrics/heuristic_metrics#readability) |
|
||
| RegexMatch | Checks if the output matches a specified regular expression pattern | [RegexMatch](/v1/evaluation/metrics/heuristic_metrics#regexmatch) |
|
||
| ROUGE | Calculates ROUGE variants (rouge1/2/L/Lsum/W) | [ROUGE](/v1/evaluation/metrics/heuristic_metrics#rouge) |
|
||
| Sentence BLEU | Computes a BLEU score for a single output against one or more references | [SentenceBLEU](/v1/evaluation/metrics/heuristic_metrics#bleu) |
|
||
| Sentiment | Scores sentiment using VADER | [Sentiment](/v1/evaluation/metrics/heuristic_metrics#sentiment) |
|
||
| Spearman Ranking | Spearman's rank correlation | [Spearman Ranking](/v1/evaluation/metrics/heuristic_metrics#spearman-ranking) |
|
||
| Tone | Flags tone issues such as shouting or negativity | [Tone](/v1/evaluation/metrics/heuristic_metrics#tone) |
|
||
|
||
### Conversation heuristic metrics
|
||
|
||
| Metric | Description | Documentation |
|
||
| --- | --- | --- |
|
||
| DegenerationC | Detects repetition and degeneration patterns over a conversation | [DegenerationC](/v1/evaluation/metrics/conversation_threads_metrics#conversation-degeneration-metric) |
|
||
| Knowledge Retention | Checks whether the last assistant reply preserves user facts from earlier turns | [Knowledge Retention](/v1/evaluation/metrics/conversation_threads_metrics#knowledge-retention-metric) |
|
||
|
||
### LLM as a Judge metrics
|
||
|
||
| Metric | Description | Documentation |
|
||
| --- | --- | --- |
|
||
| Agent Task Completion Judge | Checks whether an agent fulfilled its assigned task | [Agent Task Completion](/v1/evaluation/metrics/agent_task_completion) |
|
||
| Agent Tool Correctness Judge | Evaluates whether an agent used tools correctly | [Agent Tool Correctness](/v1/evaluation/metrics/agent_tool_correctness) |
|
||
| Answer Relevance | Checks whether the answer stays on-topic with the question | [Answer Relevance](/v1/evaluation/metrics/answer_relevance) |
|
||
| Compliance Risk Judge | Identifies non-compliant or high-risk statements | [Compliance Risk](/v1/evaluation/metrics/compliance_risk) |
|
||
| Context Precision | Ensures the answer only uses relevant context | [Context Precision](/v1/evaluation/metrics/context_precision) |
|
||
| Context Recall | Measures how well the answer recalls supporting context | [Context Recall](/v1/evaluation/metrics/context_recall) |
|
||
| Dialogue Helpfulness Judge | Evaluates how helpful an assistant reply is in a dialogue | [Dialogue Helpfulness](/v1/evaluation/metrics/dialogue_helpfulness) |
|
||
| G-Eval | Task-agnostic judge configurable with custom instructions | [G-Eval](/v1/evaluation/metrics/g_eval) |
|
||
| Hallucination | Detects unsupported or hallucinated claims using an LLM judge | [Hallucination](/v1/evaluation/metrics/hallucination) |
|
||
| LLM Juries Judge | Averages scores from multiple judge metrics for ensemble scoring | [LLM Juries](/v1/evaluation/metrics/llm_juries) |
|
||
| Meaning Match | Evaluates semantic equivalence between output and ground truth | [Meaning Match](/v1/evaluation/metrics/meaning_match) |
|
||
| Moderation | Flags safety or policy violations in assistant responses | [Moderation](/v1/evaluation/metrics/moderation) |
|
||
| Prompt Uncertainty Judge | Detects ambiguity in prompts that may confuse LLMs | [Prompt Diagnostics](/v1/evaluation/metrics/prompt_diagnostics) |
|
||
| QA Relevance Judge | Determines whether an answer directly addresses the user question | [QA Relevance](/v1/evaluation/metrics/g_eval#qa-relevance-judge) |
|
||
| Structured Output Compliance | Checks JSON or schema adherence for structured responses | [Structured Output](/v1/evaluation/metrics/structure_output_compliance) |
|
||
| Summarization Coherence Judge | Rates the structure and coherence of a summary | [Summarization Coherence](/v1/evaluation/metrics/summarization_coherence) |
|
||
| Summarization Consistency Judge | Checks if a summary stays faithful to the source | [Summarization Consistency](/v1/evaluation/metrics/summarization_consistency) |
|
||
| Trajectory Accuracy | Scores how closely agent trajectories follow expected steps | [Trajectory Accuracy](/v1/evaluation/metrics/trajectory_accuracy) |
|
||
| Usefulness | Rates how useful the answer is to the user | [Usefulness](/v1/evaluation/metrics/usefulness) |
|
||
|
||
### Conversation LLM as a Judge metrics
|
||
|
||
| Metric | Description | Documentation |
|
||
| --- | --- | --- |
|
||
| Conversational Coherence | Evaluates coherence across sliding windows of a dialogue | [Conversational Coherence](/v1/evaluation/metrics/conversation_threads_metrics#conversationalcoherencemetric) |
|
||
| Session Completeness Quality | Checks whether user goals were satisfied during the session | [Session Completeness](/v1/evaluation/metrics/conversation_threads_metrics#sessioncompletenessquality) |
|
||
| User Frustration | Estimates the likelihood a user was frustrated | [User Frustration](/v1/evaluation/metrics/conversation_threads_metrics#userfrustrationmetric) |
|
||
|
||
|
||
## Customizing LLM as a Judge metrics
|
||
|
||
By default, Opik uses GPT-5-nano from OpenAI as the LLM to evaluate the output of other LLMs. However, you can easily switch to another LLM provider by specifying a different `model` parameter.
|
||
|
||
<CodeBlocks>
|
||
```python title="Python" language="python"
|
||
from opik.evaluation.metrics import Hallucination
|
||
|
||
metric = Hallucination(model="bedrock/anthropic.claude-3-sonnet-20240229-v1:0")
|
||
|
||
metric.score(
|
||
input="What is the capital of France?",
|
||
output="The capital of France is Paris. It is famous for its iconic Eiffel Tower and rich cultural heritage.",
|
||
)
|
||
|
||
````
|
||
|
||
```typescript title="TypeScript" language="typescript"
|
||
import { Hallucination } from 'opik';
|
||
import { openai } from '@ai-sdk/openai';
|
||
|
||
// Using model ID string (simplest approach)
|
||
const metric1 = new Hallucination({ model: 'gpt-4o' });
|
||
const metric2 = new Hallucination({ model: 'claude-3-5-sonnet-latest' });
|
||
const metric3 = new Hallucination({ model: 'gemini-2.0-flash' });
|
||
|
||
// With generation parameters (temperature, seed, maxTokens)
|
||
const metric4 = new Hallucination({
|
||
model: 'gpt-4o',
|
||
temperature: 0.3,
|
||
seed: 42
|
||
});
|
||
|
||
// Using custom LanguageModel instance for provider-specific configuration
|
||
const customModel = openai('gpt-4o', {
|
||
structuredOutputs: true
|
||
});
|
||
const metric5 = new Hallucination({ model: customModel });
|
||
|
||
// Score using the metric
|
||
await metric4.score({
|
||
input: "What is the capital of France?",
|
||
output: "The capital of France is Paris. It is famous for its iconic Eiffel Tower and rich cultural heritage.",
|
||
});
|
||
````
|
||
|
||
</CodeBlocks>
|
||
|
||
For **Python**, this functionality is based on LiteLLM framework. You can find a full list of supported LLM providers and how to configure them in the [LiteLLM Providers](https://docs.litellm.ai/docs/providers) guide.
|
||
|
||
For **TypeScript**, the SDK integrates with the Vercel AI SDK. You can use model ID strings for simplicity or LanguageModel instances for advanced configuration. See the [Models documentation](/reference/typescript-sdk/evaluation/models) for more details. |