chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,71 @@
|
||||
# eval-f-score (F-Score HuggingFace Dataset Sentiment Analysis Eval)
|
||||
|
||||
You can run this example with:
|
||||
|
||||
```bash
|
||||
npx promptfoo@latest init --example eval-f-score
|
||||
cd eval-f-score
|
||||
```
|
||||
|
||||
This project evaluates GPT-4o-mini's zero-shot performance on IMDB movie review sentiment analysis using promptfoo. Each model response includes:
|
||||
|
||||
- Sentiment classification
|
||||
- Confidence score (1-10)
|
||||
- Reasoning for the classification
|
||||
|
||||
## Quick Start
|
||||
|
||||
Set your OpenAI API key and run the evaluation:
|
||||
|
||||
```bash
|
||||
promptfoo eval
|
||||
```
|
||||
|
||||
## Dataset
|
||||
|
||||
The evaluation uses the IMDB dataset from HuggingFace's datasets library, sampled to 100 reviews. The dataset is preprocessed into a CSV with two columns:
|
||||
|
||||
- `text`: The movie review content
|
||||
- `sentiment`: The label ("positive" or "negative")
|
||||
|
||||
To modify the sample size or generate a new dataset, you can use `prepare_data.py`. First, install the Python dependencies:
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
Then run the preparation script:
|
||||
|
||||
```bash
|
||||
python prepare_data.py
|
||||
```
|
||||
|
||||
## Metrics Overview
|
||||
|
||||
The evaluation implements F-score and related metrics using promptfoo's assertion system:
|
||||
|
||||
1. **Base Metrics** calculated for each test case using JavaScript assertions:
|
||||
|
||||
```yaml
|
||||
- type: javascript
|
||||
value: "output.sentiment === 'positive' && context.vars.sentiment === 'positive' ? 1 : 0"
|
||||
metric: true_positives
|
||||
```
|
||||
|
||||
2. **Derived Metrics** calculated from base metrics after the evaluation completes:
|
||||
|
||||
```yaml
|
||||
- name: precision
|
||||
value: true_positives / (true_positives + false_positives)
|
||||
|
||||
- name: f1_score
|
||||
value: 2 * true_positives / (2 * true_positives + false_positives + false_negatives)
|
||||
```
|
||||
|
||||
The evaluation tracks:
|
||||
|
||||
- **True/False Positives/Negatives**: Base metrics for classification
|
||||
- **Precision**: TP / (TP + FP)
|
||||
- **Recall**: TP / (TP + FN)
|
||||
- **F1 Score**: 2 × (precision × recall) / (precision + recall)
|
||||
- **Accuracy**: (TP + TN) / Total
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,49 @@
|
||||
from typing import Dict, List
|
||||
|
||||
import pandas as pd
|
||||
from datasets import Dataset, load_dataset
|
||||
|
||||
|
||||
def prepare_imdb_data() -> None:
|
||||
"""
|
||||
Prepare and sample IMDB dataset for sentiment analysis evaluation.
|
||||
Loads data from HuggingFace, converts to DataFrame, and saves a sample to CSV.
|
||||
"""
|
||||
# Load the IMDB dataset
|
||||
print("Loading IMDB dataset...")
|
||||
imdb: Dataset = load_dataset("imdb") # type: ignore
|
||||
|
||||
# Convert labels to more readable format
|
||||
label_map: Dict[int, str] = {0: "negative", 1: "positive"}
|
||||
|
||||
# Create dataframe from test set (we'll use this for zero-shot evaluation)
|
||||
texts: List[str] = imdb["test"]["text"] # type: ignore
|
||||
labels: List[int] = imdb["test"]["label"] # type: ignore
|
||||
|
||||
eval_df: pd.DataFrame = pd.DataFrame(
|
||||
{
|
||||
"text": texts,
|
||||
"sentiment": [label_map[label] for label in labels],
|
||||
}
|
||||
)
|
||||
|
||||
# Take a small sample for evaluation
|
||||
eval_sample: pd.DataFrame = eval_df.sample(n=100, random_state=0)
|
||||
|
||||
# Save to CSV file
|
||||
print("Saving sample to CSV...")
|
||||
eval_sample.to_csv("imdb_eval_sample.csv", index=False)
|
||||
|
||||
print(f"Saved {len(eval_sample)} examples for evaluation")
|
||||
|
||||
# Print some statistics
|
||||
print("\nLabel distribution in evaluation set:")
|
||||
print(eval_sample["sentiment"].value_counts())
|
||||
|
||||
print("\nSample review:")
|
||||
print("Text:", eval_sample["text"].iloc[0][:200], "...")
|
||||
print("Sentiment:", eval_sample["sentiment"].iloc[0])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
prepare_imdb_data()
|
||||
@@ -0,0 +1,87 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: IMDB Review Sentiment Analysis
|
||||
|
||||
providers:
|
||||
- openai:gpt-4.1-mini
|
||||
|
||||
prompts:
|
||||
- label: Sentiment Analysis
|
||||
raw: |
|
||||
Analyze the sentiment of the following movie review. Classify it as either positive or negative.
|
||||
|
||||
Review: "{{text}}"
|
||||
|
||||
Respond with a JSON object in the following format:
|
||||
{
|
||||
"sentiment": "positive" or "negative",
|
||||
"confidence": number between 1-10,
|
||||
"reasoning": "brief explanation"
|
||||
}
|
||||
config:
|
||||
response_format:
|
||||
type: json_schema
|
||||
json_schema:
|
||||
name: MovieReviewSentiment
|
||||
schema:
|
||||
type: object
|
||||
properties:
|
||||
sentiment:
|
||||
type: string
|
||||
enum: ['positive', 'negative']
|
||||
confidence:
|
||||
type: integer
|
||||
minimum: 1
|
||||
maximum: 10
|
||||
reasoning:
|
||||
type: string
|
||||
required: ['sentiment', 'confidence', 'reasoning']
|
||||
|
||||
defaultTest:
|
||||
assert:
|
||||
# Basic JSON validation
|
||||
- type: is-json
|
||||
|
||||
# Track binary classification metrics
|
||||
- type: javascript
|
||||
value: 'output.sentiment === context.vars.sentiment'
|
||||
metric: accuracy
|
||||
|
||||
# For F1 score components (treating 'positive' as the positive class)
|
||||
- type: javascript
|
||||
value: "output.sentiment === 'positive' && context.vars.sentiment === 'positive' ? 1 : 0"
|
||||
metric: true_positives
|
||||
weight: 0
|
||||
|
||||
- type: javascript
|
||||
value: "output.sentiment === 'positive' && context.vars.sentiment === 'negative' ? 1 : 0"
|
||||
metric: false_positives
|
||||
weight: 0
|
||||
|
||||
- type: javascript
|
||||
value: "output.sentiment === 'negative' && context.vars.sentiment === 'positive' ? 1 : 0"
|
||||
metric: false_negatives
|
||||
weight: 0
|
||||
|
||||
- type: javascript
|
||||
value: "output.sentiment === 'negative' && context.vars.sentiment === 'negative' ? 1 : 0"
|
||||
metric: true_negatives
|
||||
weight: 0
|
||||
|
||||
derivedMetrics:
|
||||
# Precision = TP / (TP + FP)
|
||||
- name: precision
|
||||
value: true_positives / (true_positives + false_positives)
|
||||
|
||||
# Recall = TP / (TP + FN)
|
||||
- name: recall
|
||||
value: true_positives / (true_positives + false_negatives)
|
||||
|
||||
# F1 Score = 2 * (precision * recall) / (precision + recall)
|
||||
- name: f1_score
|
||||
value: 2 * true_positives / (2 * true_positives + false_positives + false_negatives)
|
||||
|
||||
# Accuracy = (TP + TN) / (TP + TN + FP + FN)
|
||||
- name: accuracy_score
|
||||
value: (true_positives + true_negatives) / (true_positives + true_negatives + false_positives + false_negatives)
|
||||
|
||||
tests: file://imdb_eval_sample.csv
|
||||
@@ -0,0 +1,2 @@
|
||||
datasets==4.2.0
|
||||
pandas==2.3.3
|
||||
Reference in New Issue
Block a user