Files
wehub-resource-sync 5a558eb09e
TypeScript SDK Compatibility V1.x E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / TypeScript SDK Compatibility V1.x E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
TypeScript SDK E2E Tests / TypeScript SDK E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
Python SDK E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK E2E Tests / Python SDK E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK E2E Tests / build-opik (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Python SDK Compatibility V1.x E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK E2E Tests / build-opik (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer E2E Tests Python ${{matrix.python_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer Integration Smoke Tests (push) Has been cancelled
🐙 Code Quality / detect (push) Has been cancelled
🐙 Code Quality / lint (${{ matrix.leg.name }}) (push) Has been cancelled
🐙 Code Quality / summary (push) Has been cancelled
TypeScript SDK Library Integration Tests / Check Secrets (push) Has been cancelled
TypeScript SDK Library Integration Tests / opik-vercel (Vercel AI SDK / eve) (push) Has been cancelled
SDK Library Integration Tests Runner / Check Secrets (push) Has been cancelled
SDK Library Integration Tests Runner / Missed OpenAI API Key Warning (push) Has been cancelled
SDK Library Integration Tests Runner / Build (push) Has been cancelled
SDK Library Integration Tests Runner / openai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_legacy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / llama_index_tests (push) Has been cancelled
SDK Library Integration Tests Runner / anthropic_tests (push) Has been cancelled
SDK Library Integration Tests Runner / mistral_tests (push) Has been cancelled
SDK Library Integration Tests Runner / groq_tests (push) Has been cancelled
SDK Library Integration Tests Runner / aisuite_tests (push) Has been cancelled
SDK Library Integration Tests Runner / haystack_tests (push) Has been cancelled
SDK Library Integration Tests Runner / dspy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v1_tests (push) Has been cancelled
SDK Library Integration Tests Runner / genai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_legacy_1_3_0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / evaluation_metrics_tests (push) Has been cancelled
SDK Library Integration Tests Runner / bedrock_tests (push) Has been cancelled
SDK Library Integration Tests Runner / litellm_tests (push) Has been cancelled
SDK Library Integration Tests Runner / harbor_tests (push) Has been cancelled
SDK Library Integration Tests Runner / Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / render-equality (push) Has been cancelled
Opik Optimizer - Unit Tests / Opik Optimizer Unit Tests Python ${{matrix.python_version}} (push) Has been cancelled
Python BE E2E Tests / Python BE E2E (push) Has been cancelled
Python Backend Tests / run-python-backend-tests (push) Has been cancelled
Python SDK Unit Tests / Python SDK Unit Tests ${{matrix.python_version}} (push) Has been cancelled
Release Drafter / update_release_draft (push) Has been cancelled
SDK E2E Libraries Integration Tests / Check Secrets (push) Has been cancelled
SDK E2E Libraries Integration Tests / Missed OpenAI API Key Warning (push) Has been cancelled
SDK E2E Libraries Integration Tests / build-opik (push) Has been cancelled
SDK E2E Libraries Integration Tests / E2E Lib Integration Python ${{matrix.python_version}} (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-gemini) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-langchain) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-openai) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-otel) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-vercel) (push) Has been cancelled
TypeScript SDK Build & Publish / build-and-publish (push) Has been cancelled
TypeScript SDK Unit Tests / Test on Node ${{ matrix.node-version }} (push) Has been cancelled
Backend Tests / discover-tests (push) Has been cancelled
Backend Tests / ${{ matrix.name }} (push) Has been cancelled
Build and Publish SDK / build-and-publish (push) Has been cancelled
Build Opik Docker Images / set-version (push) Has been cancelled
Build Opik Docker Images / build-backend (push) Has been cancelled
Build Opik Docker Images / build-sandbox-executor-python (push) Has been cancelled
Build Opik Docker Images / build-python-backend (push) Has been cancelled
Build Opik Docker Images / build-frontend (push) Has been cancelled
Build Opik Docker Images / create-git-tag (push) Has been cancelled
ClickHouse Migration Cluster Check / validate-clickhouse-migrations (push) Has been cancelled
Docs - Publish / run (push) Has been cancelled
E2E Tests - Post Merge (v2) / 🧪 E2E v2 Tests (${{ github.event.inputs.tier || 't1' }}) (push) Has been cancelled
E2E Tests - Post Merge (v2) / 📢 Slack Notification (push) Has been cancelled
Frontend Unit Tests / Test on Node 20 (push) Has been cancelled
Guardrails E2E Tests / Select Python version matrix (push) Has been cancelled
Guardrails E2E Tests / Guardrails E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Guardrails E2E Tests / 📢 Slack Notification (push) Has been cancelled
Guardrails Backend Unit Tests / Guardrails Backend Unit Tests (push) Has been cancelled
Guardrails Backend Unit Tests / 📢 Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v3.21.0) (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v4.2.0) (push) Has been cancelled
Lint Opik Helm Chart / unittest-helm-chart (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 13:25:44 +08:00

382 lines
16 KiB
Plaintext
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
description: Describes Opik's built-in G-Eval metric which is a task agnostic LLM
as a Judge metric
headline: G-Eval | Opik Documentation
og:description: Evaluate tasks using G-Eval, a LLM-as-a-judge metric. Specify tasks
and criteria to receive scores from 0 to 1 efficiently.
og:site_name: Opik Documentation
og:title: G-Eval Metrics - Opik for Task Evaluation
title: G-Eval
canonical-url: https://www.comet.com/docs/opik/evaluation/metrics/g_eval
---
G-Eval is a task-agnostic LLM-as-a-judge metric that allows you to specify a task description and evaluation criteria. The model first drafts step-by-step evaluation instructions and then produces a score between 0 and 1. You can learn more about G-Eval in the [original paper](https://arxiv.org/abs/2303.16634).
To use G-Eval, supply two pieces of information:
1. A task introduction describing what should be evaluated.
2. Evaluation criteria outlining what “good” looks like.
The judge responds with an **integer between 0 and 10**. Opik divides that value by 10 so callers receive a score between 0.0 and 1.0. We recommend packaging the full scenario (prompt, context, answer, etc.) inside a single string and passing it via the `output` argument; any other keyword arguments are ignored by the metric interface.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import GEval
metric = GEval(
task_introduction="You are an expert judge tasked with evaluating the faithfulness of an AI-generated answer to the given context.",
evaluation_criteria="In the provided text the OUTPUT must not introduce new information beyond what's provided in the CONTEXT.",
)
payload = """INPUT: What is the capital of France?
CONTEXT: France is a country in Western Europe. Its capital is Paris, which is known for landmarks like the Eiffel Tower.
OUTPUT: Paris is the capital of France.
"""
metric.score(output=payload)
````
```typescript title="TypeScript" language="typescript"
import { GEval } from "opik/evaluation/metrics";
const metric = new GEval({
taskIntroduction: "You are an expert judge tasked with evaluating the faithfulness of an AI-generated answer to the given context.",
evaluationCriteria: "In the provided text the OUTPUT must not introduce new information beyond what's provided in the CONTEXT.",
});
const payload = `INPUT: What is the capital of France?
CONTEXT: France is a country in Western Europe. Its capital is Paris, which is known for landmarks like the Eiffel Tower.
OUTPUT: Paris is the capital of France.
`;
await metric.score({ output: payload });
````
</CodeBlocks>
## How it works
G-Eval first expands your task description into a step-by-step Chain of Thought (CoT). This CoT becomes the rubric the judge will follow when scoring the provided answer. The model then evaluates the answer, returning a score in the 010 range which Opik normalises to 01.
By default, the `gpt-5-nano` model is used, but you can change this to any model supported by [LiteLLM](https://docs.litellm.ai/docs/providers) via the `model` parameter. Learn more in the [custom model guide](/v1/evaluation/metrics/custom_model).
<Note>
To make the metric more robust, Opik requests the top 20 log probabilities from the LLM and computes a weighted average of the scores, as recommended by the original paper. The evaluator always returns an **integer between 0 and 10**; Opik divides that value by 10 before exposing it so callers see numbers in the [0,1] range. Newer models in the GPT-5 family and other providers may not expose log probabilities, so scores can vary when switching models.
</Note>
## Built-in G-Eval judges
Opik ships opinionated presets for common evaluation needs. Each class inherits from `GEval` and exposes the same constructor parameters (`model`, `track`, `temperature`, etc.).
### Compliance Risk Judge
Flags statements that may be non-factual, non-compliant, or risky (e.g. finance, healthcare, legal). This judge is useful when you need an automated review step before customer-facing responses are sent, or when auditing historical conversations for policy breaches.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import ComplianceRiskJudge
metric = ComplianceRiskJudge(model="gpt-4o-mini")
payload = """INPUT: Customer asked about wire-transfer reversal policies.
OUTPUT: Just reverse it whenever the customer asks.
"""
score = metric.score(output=payload)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { ComplianceRiskJudge } from "opik/evaluation/metrics";
const metric = new ComplianceRiskJudge({ model: "gpt-4o-mini" });
const payload = `INPUT: Customer asked about wire-transfer reversal policies.
OUTPUT: Just reverse it whenever the customer asks.
`;
const score = await metric.score({ output: payload });
console.log(score.value, score.reason);
````
</CodeBlocks>
Inspect `score.reason` for granular rationales and route risky cases accordingly. The raw 010 judgement is divided by 10 in the returned value.
### Prompt Uncertainty Judge
`PromptUncertaintyJudge` estimates how ambiguous a user prompt is before it reaches your model. Run it on raw user messages to prioritise agent hand-offs or to warn users when the request is ill-posed.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import PromptUncertaintyJudge
prompt = "Summarise the attached 400 page contract in one sentence and guarantee there are no mistakes."
uncertainty = PromptUncertaintyJudge().score(prompt=prompt)
print(uncertainty.value)
````
```typescript title="TypeScript" language="typescript"
import { PromptUncertaintyJudge } from "opik/evaluation/metrics";
const prompt = "Summarise the attached 400 page contract in one sentence and guarantee there are no mistakes.";
const uncertainty = await new PromptUncertaintyJudge().score({ output: prompt });
console.log(uncertainty.value);
````
</CodeBlocks>
Use the score to highlight prompts that may confuse downstream models; the judge emits an integer from 0 (best) to 10 (worst) before normalisation.
### Summarization Consistency Judge
Checks whether a generated summary is faithful to the source material. This is the right choice when a downstream workflow consumes summaries and you need to enforce factual alignment with the source document.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import SummarizationConsistencyJudge
metric = SummarizationConsistencyJudge(model="gpt-4o")
payload = """CONTEXT: ...long article text...
SUMMARY: The article confirms new safety protocols but misstates the deadline.
"""
score = metric.score(output=payload)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { SummarizationConsistencyJudge } from "opik/evaluation/metrics";
const metric = new SummarizationConsistencyJudge({ model: "gpt-4o" });
const payload = `CONTEXT: ...long article text...
SUMMARY: The article confirms new safety protocols but misstates the deadline.
`;
const score = await metric.score({ output: payload });
console.log(score.value, score.reason);
````
</CodeBlocks>
Pair this metric with alerts or automated rollbacks when the score drops below a threshold; the evaluator still returns raw integers in 010 before Opik scales them.
### Summarization Coherence Judge
Scores the structure, clarity, and organisation of a summary. Use it when you optimise for human readability or want to catch summaries that are factually right but poorly written.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import SummarizationCoherenceJudge
metric = SummarizationCoherenceJudge()
score = metric.score(output="""SUMMARY: First... Secondly... Finally...""")
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { SummarizationCoherenceJudge } from "opik/evaluation/metrics";
const metric = new SummarizationCoherenceJudge();
const score = await metric.score({ output: "SUMMARY: First... Secondly... Finally..." });
console.log(score.value, score.reason);
````
</CodeBlocks>
High scores correlate with summaries that maintain logical ordering and concise transitions between ideas. A perfect 10 becomes 1.0 after Opik normalisation.
### Dialogue Helpfulness Judge
Examines how helpful an assistant reply is in the context of the preceding dialogue. Helpful for agent tuning or support chat routing where you want to surface conversations that require escalation.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import DialogueHelpfulnessJudge
transcript = """USER: How do I reset my password?
ASSISTANT: Visit settings and click reset.
USER: I cannot see that option.
ASSISTANT: Please contact support.
"""
score = DialogueHelpfulnessJudge().score(output=transcript)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { DialogueHelpfulnessJudge } from "opik/evaluation/metrics";
const transcript = `USER: How do I reset my password?
ASSISTANT: Visit settings and click reset.
USER: I cannot see that option.
ASSISTANT: Please contact support.
`;
const score = await new DialogueHelpfulnessJudge().score({ output: transcript });
console.log(score.value, score.reason);
````
</CodeBlocks>
Low scores typically indicate the assistant ignored prior context or refused to offer actionable steps. The normalised value originates from an integer between 0 and 10.
### QA Relevance Judge
Determines whether an answer directly addresses the users question. Ideal for dataset regression tests where each sample has a clear question/answer pair.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import QARelevanceJudge
metric = QARelevanceJudge()
payload = """QUESTION: What causes rainbows?
ANSWER: The capital of France is Paris.
"""
score = metric.score(output=payload)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { QARelevanceJudge } from "opik/evaluation/metrics";
const metric = new QARelevanceJudge();
const payload = `QUESTION: What causes rainbows?
ANSWER: The capital of France is Paris.
`;
const score = await metric.score({ output: payload });
console.log(score.value, score.reason);
````
</CodeBlocks>
Combine with hallucination metrics to distinguish totally off-topic answers from confident but wrong responses; the judge still works on a 010 scale internally.
### Agent Task Completion Judge
Evaluates if an agent fulfilled its assigned high-level task. Works well for long-running workflows where success is defined by end-state rather than a single response.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import AgentTaskCompletionJudge
trace_summary = "Agent gathered quotes, compared options, and booked travel."
score = AgentTaskCompletionJudge().score(output=trace_summary)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { AgentTaskCompletionJudge } from "opik/evaluation/metrics";
const traceSummary = "Agent gathered quotes, compared options, and booked travel.";
const score = await new AgentTaskCompletionJudge().score({ output: traceSummary });
console.log(score.value, score.reason);
````
</CodeBlocks>
Use the reason text to inspect which sub-goals the judge believed were satisfied; a raw 010 verdict is divided by 10 in the returned value.
### Agent Tool Correctness Judge
Assesses whether an agent invoked tools appropriately and interpreted outputs correctly. Especially useful for production agents integrating external APIs.
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import AgentToolCorrectnessJudge
call_trace = "Tool weather_api called with city='Paris' but response ignored."
score = AgentToolCorrectnessJudge().score(output=call_trace)
print(score.value, score.reason)
````
```typescript title="TypeScript" language="typescript"
import { AgentToolCorrectnessJudge } from "opik/evaluation/metrics";
const callTrace = "Tool weather_api called with city='Paris' but response ignored.";
const score = await new AgentToolCorrectnessJudge().score({ output: callTrace });
console.log(score.value, score.reason);
````
</CodeBlocks>
Lower scores suggest the agent mis-handled tool results or skipped required invocations. Raw values remain in the 010 range before normalisation.
### Trajectory Accuracy
Scores whether an agents trajectory (series of states or actions) matches the expected path. Use it to audit reinforcement-learning agents or scripted flows that should follow specific checkpoints.
```python title="Trajectory accuracy"
from opik.evaluation.metrics import TrajectoryAccuracy
expected = ["start", "search_docs", "summarise", "respond"]
actual = ["start", "search_docs", "respond"]
score = TrajectoryAccuracy(expected_path=expected).score(output=actual)
print(score.value, score.reason)
```
This metric highlights missing or out-of-order actions so you can tighten guardrails around multi-step agents.
## LLM Juries Judge
`LLMJuriesJudge` is an ensemble wrapper that averages the outputs of multiple judge metrics. This is useful when you want to combine bespoke criteria—e.g. take the mean of hallucination, helpfulness, and compliance scores.
```python
from opik.evaluation.metrics import LLMJuriesJudge, Hallucination, ComplianceRiskJudge
jury = LLMJuriesJudge([
Hallucination(model="gpt-4o-mini"),
ComplianceRiskJudge(model="gpt-4o-mini"),
])
payload = """INPUT: Summarise compliance requirements for fintech onboarding.
OUTPUT: No need for KYC; just accept the payment.
"""
result = jury.score(output=payload)
print(result.value, result.metadata["judge_scores"])
```
## Conversation adapters
Need to apply G-Eval-based judges to full conversations? Use the conversation adapters in `opik.evaluation.metrics.conversation.llm_judges.g_eval_wrappers`, exposed via `Conversation*` classes. They focus on the last assistant turn (or full transcript for summaries) and keep the original GEval reasoning.
Refer to [Conversation-level GEval Metrics](/v1/evaluation/metrics/g_eval_conversation_metrics) for available adapters and usage examples.
## Customising models
All GEval-derived metrics expose the `model` parameter so you can switch the underlying LLM. For example:
<CodeBlocks>
```python title="Python" language="python"
from opik.evaluation.metrics import ComplianceRiskJudge
metric = ComplianceRiskJudge(model="bedrock/anthropic.claude-3-sonnet-20240229-v1:0")
payload = """INPUT: What is the capital of France?
OUTPUT: The capital of France is Paris. It is famous for its iconic Eiffel Tower and rich cultural heritage.
"""
score = metric.score(output=payload)
````
```typescript title="TypeScript" language="typescript"
import { ComplianceRiskJudge } from "opik/evaluation/metrics";
import { anthropic } from "@ai-sdk/anthropic";
const metric = new ComplianceRiskJudge({
model: anthropic("claude-3-5-sonnet-latest")
});
const payload = `INPUT: What is the capital of France?
OUTPUT: The capital of France is Paris. It is famous for its iconic Eiffel Tower and rich cultural heritage.
`;
const score = await metric.score({ output: payload });
````
</CodeBlocks>
In Python, this functionality relies on LiteLLM. See the [LiteLLM Providers](https://docs.litellm.ai/docs/providers) guide for a full list of supported providers and model identifiers.
In TypeScript, the SDK uses the [Vercel AI SDK](https://sdk.vercel.ai/providers) for model integration. See the [Models documentation](/reference/typescript-sdk/evaluation/models) for configuration details.