5a558eb09e
TypeScript SDK Compatibility V1.x E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / TypeScript SDK Compatibility V1.x E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
TypeScript SDK E2E Tests / TypeScript SDK E2E Tests Node ${{matrix.node_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
Python SDK E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK E2E Tests / Python SDK E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK E2E Tests / build-opik (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Select Python version matrix (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / Python SDK Compatibility V1.x E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Python SDK Compatibility V1.x E2E Tests / build-opik (push) Has been cancelled
TypeScript SDK E2E Tests / Select Node version matrix (push) Has been cancelled
TypeScript SDK E2E Tests / build-opik (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer E2E Tests Python ${{matrix.python_version}} (push) Has been cancelled
Opik Optimizer - E2E Tests / Opik Optimizer Integration Smoke Tests (push) Has been cancelled
🐙 Code Quality / detect (push) Has been cancelled
🐙 Code Quality / lint (${{ matrix.leg.name }}) (push) Has been cancelled
🐙 Code Quality / summary (push) Has been cancelled
TypeScript SDK Library Integration Tests / Check Secrets (push) Has been cancelled
TypeScript SDK Library Integration Tests / opik-vercel (Vercel AI SDK / eve) (push) Has been cancelled
SDK Library Integration Tests Runner / Check Secrets (push) Has been cancelled
SDK Library Integration Tests Runner / Missed OpenAI API Key Warning (push) Has been cancelled
SDK Library Integration Tests Runner / Build (push) Has been cancelled
SDK Library Integration Tests Runner / openai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_tests (push) Has been cancelled
SDK Library Integration Tests Runner / langchain_legacy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / llama_index_tests (push) Has been cancelled
SDK Library Integration Tests Runner / anthropic_tests (push) Has been cancelled
SDK Library Integration Tests Runner / mistral_tests (push) Has been cancelled
SDK Library Integration Tests Runner / groq_tests (push) Has been cancelled
SDK Library Integration Tests Runner / aisuite_tests (push) Has been cancelled
SDK Library Integration Tests Runner / haystack_tests (push) Has been cancelled
SDK Library Integration Tests Runner / dspy_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / crewai_v1_tests (push) Has been cancelled
SDK Library Integration Tests Runner / genai_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_tests (push) Has been cancelled
SDK Library Integration Tests Runner / adk_legacy_1_3_0_tests (push) Has been cancelled
SDK Library Integration Tests Runner / evaluation_metrics_tests (push) Has been cancelled
SDK Library Integration Tests Runner / bedrock_tests (push) Has been cancelled
SDK Library Integration Tests Runner / litellm_tests (push) Has been cancelled
SDK Library Integration Tests Runner / harbor_tests (push) Has been cancelled
SDK Library Integration Tests Runner / Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / render-equality (push) Has been cancelled
Opik Optimizer - Unit Tests / Opik Optimizer Unit Tests Python ${{matrix.python_version}} (push) Has been cancelled
Python BE E2E Tests / Python BE E2E (push) Has been cancelled
Python Backend Tests / run-python-backend-tests (push) Has been cancelled
Python SDK Unit Tests / Python SDK Unit Tests ${{matrix.python_version}} (push) Has been cancelled
Release Drafter / update_release_draft (push) Has been cancelled
SDK E2E Libraries Integration Tests / Check Secrets (push) Has been cancelled
SDK E2E Libraries Integration Tests / Missed OpenAI API Key Warning (push) Has been cancelled
SDK E2E Libraries Integration Tests / build-opik (push) Has been cancelled
SDK E2E Libraries Integration Tests / E2E Lib Integration Python ${{matrix.python_version}} (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-gemini) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-langchain) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-openai) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-otel) (push) Has been cancelled
TypeScript SDK Integration Build & Publish / build-and-publish (opik-vercel) (push) Has been cancelled
TypeScript SDK Build & Publish / build-and-publish (push) Has been cancelled
TypeScript SDK Unit Tests / Test on Node ${{ matrix.node-version }} (push) Has been cancelled
Backend Tests / discover-tests (push) Has been cancelled
Backend Tests / ${{ matrix.name }} (push) Has been cancelled
Build and Publish SDK / build-and-publish (push) Has been cancelled
Build Opik Docker Images / set-version (push) Has been cancelled
Build Opik Docker Images / build-backend (push) Has been cancelled
Build Opik Docker Images / build-sandbox-executor-python (push) Has been cancelled
Build Opik Docker Images / build-python-backend (push) Has been cancelled
Build Opik Docker Images / build-frontend (push) Has been cancelled
Build Opik Docker Images / create-git-tag (push) Has been cancelled
ClickHouse Migration Cluster Check / validate-clickhouse-migrations (push) Has been cancelled
Docs - Publish / run (push) Has been cancelled
E2E Tests - Post Merge (v2) / 🧪 E2E v2 Tests (${{ github.event.inputs.tier || 't1' }}) (push) Has been cancelled
E2E Tests - Post Merge (v2) / 📢 Slack Notification (push) Has been cancelled
Frontend Unit Tests / Test on Node 20 (push) Has been cancelled
Guardrails E2E Tests / Select Python version matrix (push) Has been cancelled
Guardrails E2E Tests / Guardrails E2E Tests ${{matrix.python_version}} (push) Has been cancelled
Guardrails E2E Tests / 📢 Slack Notification (push) Has been cancelled
Guardrails Backend Unit Tests / Guardrails Backend Unit Tests (push) Has been cancelled
Guardrails Backend Unit Tests / 📢 Slack Notification (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v3.21.0) (push) Has been cancelled
Lint Opik Helm Chart / lint-helm-chart (Helm v4.2.0) (push) Has been cancelled
Lint Opik Helm Chart / unittest-helm-chart (push) Has been cancelled
106 lines
5.4 KiB
Plaintext
106 lines
5.4 KiB
Plaintext
---
|
|
description: Describes the Usefulness metric
|
|
headline: Usefulness | Opik Documentation
|
|
og:description: Evaluate the usefulness of LLM responses with Opik's metric, scoring
|
|
from 0.0 to 1.0 for precise assessments.
|
|
og:site_name: Opik Documentation
|
|
og:title: Evaluate Usefulness with Opik Metrics
|
|
title: Usefulness
|
|
canonical-url: https://www.comet.com/docs/opik/evaluation/metrics/usefulness
|
|
---
|
|
|
|
# Usefulness
|
|
|
|
The usefulness metric allows you to evaluate how useful an LLM response is given an input. It uses a language model to assess the usefulness and provides a score between 0.0 and 1.0, where higher values indicate higher usefulness. Along with the score, it provides a detailed explanation of why that score was assigned.
|
|
|
|
## How to use the Usefulness metric
|
|
|
|
You can use the `Usefulness` metric as follows:
|
|
|
|
<CodeBlocks>
|
|
```python title="Python" language="python"
|
|
from opik.evaluation.metrics import Usefulness
|
|
|
|
metric = Usefulness()
|
|
|
|
result = metric.score(
|
|
input="How can I optimize the performance of my Python web application?",
|
|
output="To optimize your Python web application's performance, focus on these key areas:\n1. Database optimization: Use connection pooling, index frequently queried fields, and cache common queries\n2. Caching strategy: Implement Redis or Memcached for session data and frequently accessed content\n3. Asynchronous operations: Use async/await for I/O-bound operations to handle more concurrent requests\n4. Code profiling: Use tools like cProfile to identify bottlenecks in your application\n5. Load balancing: Distribute traffic across multiple server instances for better scalability",
|
|
)
|
|
|
|
print(result.value) # A float between 0.0 and 1.0
|
|
print(result.reason) # Explanation for the score
|
|
|
|
````
|
|
|
|
```typescript title="TypeScript" language="typescript"
|
|
import { Usefulness } from 'opik';
|
|
|
|
const metric = new Usefulness();
|
|
|
|
const result = await metric.score({
|
|
input: "How can I optimize the performance of my Python web application?",
|
|
output: "To optimize your Python web application's performance, focus on these key areas:\n1. Database optimization: Use connection pooling, index frequently queried fields, and cache common queries\n2. Caching strategy: Implement Redis or Memcached for session data and frequently accessed content\n3. Asynchronous operations: Use async/await for I/O-bound operations to handle more concurrent requests\n4. Code profiling: Use tools like cProfile to identify bottlenecks in your application\n5. Load balancing: Distribute traffic across multiple server instances for better scalability",
|
|
});
|
|
|
|
console.log(result.value); // A float between 0.0 and 1.0
|
|
console.log(result.reason); // Explanation for the score
|
|
````
|
|
|
|
</CodeBlocks>
|
|
|
|
Asynchronous scoring is also supported with the `ascore` method in Python and `score` method in TypeScript (which is always async).
|
|
|
|
## Understanding the scores
|
|
|
|
The usefulness score ranges from 0.0 to 1.0:
|
|
|
|
- Scores closer to 1.0 indicate that the response is highly useful, directly addressing the input query with relevant and accurate information
|
|
- Scores closer to 0.0 indicate that the response is less useful, possibly being off-topic, incomplete, or not addressing the input query effectively
|
|
|
|
Each score comes with a detailed explanation (`result.reason`) that helps understand why that particular score was assigned.
|
|
|
|
## Usefulness Prompt
|
|
|
|
Opik uses an LLM as a Judge to evaluate usefulness, for this we have a prompt template that is used to generate the prompt for the LLM. By default, the `gpt-4o` model is used to evaluate responses but you can change this to any model supported by [LiteLLM](https://docs.litellm.ai/docs/providers) by setting the `model` parameter. You can learn more about customizing models in the [Customize models for LLM as a Judge metrics](/v1/evaluation/metrics/custom_model) section.
|
|
|
|
The template is as follows:
|
|
|
|
```
|
|
You are an impartial judge tasked with evaluating the quality and usefulness of AI-generated responses.
|
|
|
|
Your evaluation should consider the following key factors:
|
|
- Helpfulness: How well does it solve the user's problem?
|
|
- Relevance: How well does it address the specific question?
|
|
- Accuracy: Is the information correct and reliable?
|
|
- Depth: Does it provide sufficient detail and explanation?
|
|
- Creativity: Does it offer innovative or insightful perspectives when appropriate?
|
|
- Level of detail: Is the amount of detail appropriate for the question?
|
|
|
|
###EVALUATION PROCESS###
|
|
|
|
1. **ANALYZE** the user's question and the AI's response carefully
|
|
2. **EVALUATE** how well the response meets each of the criteria above
|
|
3. **CONSIDER** the overall effectiveness and usefulness of the response
|
|
4. **PROVIDE** a clear, objective explanation for your evaluation
|
|
5. **SCORE** the response on a scale from 0.0 to 1.0:
|
|
- 1.0: Exceptional response that excels in all criteria
|
|
- 0.8: Excellent response with minor room for improvement
|
|
- 0.6: Good response that adequately addresses the question
|
|
- 0.4: Fair response with significant room for improvement
|
|
- 0.2: Poor response that barely addresses the question
|
|
- 0.0: Completely inadequate or irrelevant response
|
|
|
|
###OUTPUT FORMAT###
|
|
|
|
Your evaluation must be provided as a JSON object with exactly two fields:
|
|
- "score": A float between 0.0 and 1.0
|
|
- "reason": A brief, objective explanation justifying your score based on the criteria above
|
|
|
|
Now, please evaluate the following:
|
|
|
|
User Question: {input}
|
|
AI Response: {output}
|
|
|
|
Provide your evaluation in the specified JSON format.
|
|
``` |