chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled

This commit is contained in:
wehub-resource-sync
2026-07-13 13:24:08 +08:00
commit 0d3cb498a3
5438 changed files with 1316560 additions and 0 deletions
+71
View File
@@ -0,0 +1,71 @@
# eval-f-score (F-Score HuggingFace Dataset Sentiment Analysis Eval)
You can run this example with:
```bash
npx promptfoo@latest init --example eval-f-score
cd eval-f-score
```
This project evaluates GPT-4o-mini's zero-shot performance on IMDB movie review sentiment analysis using promptfoo. Each model response includes:
- Sentiment classification
- Confidence score (1-10)
- Reasoning for the classification
## Quick Start
Set your OpenAI API key and run the evaluation:
```bash
promptfoo eval
```
## Dataset
The evaluation uses the IMDB dataset from HuggingFace's datasets library, sampled to 100 reviews. The dataset is preprocessed into a CSV with two columns:
- `text`: The movie review content
- `sentiment`: The label ("positive" or "negative")
To modify the sample size or generate a new dataset, you can use `prepare_data.py`. First, install the Python dependencies:
```bash
pip install -r requirements.txt
```
Then run the preparation script:
```bash
python prepare_data.py
```
## Metrics Overview
The evaluation implements F-score and related metrics using promptfoo's assertion system:
1. **Base Metrics** calculated for each test case using JavaScript assertions:
```yaml
- type: javascript
value: "output.sentiment === 'positive' && context.vars.sentiment === 'positive' ? 1 : 0"
metric: true_positives
```
2. **Derived Metrics** calculated from base metrics after the evaluation completes:
```yaml
- name: precision
value: true_positives / (true_positives + false_positives)
- name: f1_score
value: 2 * true_positives / (2 * true_positives + false_positives + false_negatives)
```
The evaluation tracks:
- **True/False Positives/Negatives**: Base metrics for classification
- **Precision**: TP / (TP + FP)
- **Recall**: TP / (TP + FN)
- **F1 Score**: 2 × (precision × recall) / (precision + recall)
- **Accuracy**: (TP + TN) / Total
File diff suppressed because one or more lines are too long
+49
View File
@@ -0,0 +1,49 @@
from typing import Dict, List
import pandas as pd
from datasets import Dataset, load_dataset
def prepare_imdb_data() -> None:
"""
Prepare and sample IMDB dataset for sentiment analysis evaluation.
Loads data from HuggingFace, converts to DataFrame, and saves a sample to CSV.
"""
# Load the IMDB dataset
print("Loading IMDB dataset...")
imdb: Dataset = load_dataset("imdb") # type: ignore
# Convert labels to more readable format
label_map: Dict[int, str] = {0: "negative", 1: "positive"}
# Create dataframe from test set (we'll use this for zero-shot evaluation)
texts: List[str] = imdb["test"]["text"] # type: ignore
labels: List[int] = imdb["test"]["label"] # type: ignore
eval_df: pd.DataFrame = pd.DataFrame(
{
"text": texts,
"sentiment": [label_map[label] for label in labels],
}
)
# Take a small sample for evaluation
eval_sample: pd.DataFrame = eval_df.sample(n=100, random_state=0)
# Save to CSV file
print("Saving sample to CSV...")
eval_sample.to_csv("imdb_eval_sample.csv", index=False)
print(f"Saved {len(eval_sample)} examples for evaluation")
# Print some statistics
print("\nLabel distribution in evaluation set:")
print(eval_sample["sentiment"].value_counts())
print("\nSample review:")
print("Text:", eval_sample["text"].iloc[0][:200], "...")
print("Sentiment:", eval_sample["sentiment"].iloc[0])
if __name__ == "__main__":
prepare_imdb_data()
@@ -0,0 +1,87 @@
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: IMDB Review Sentiment Analysis
providers:
- openai:gpt-4.1-mini
prompts:
- label: Sentiment Analysis
raw: |
Analyze the sentiment of the following movie review. Classify it as either positive or negative.
Review: "{{text}}"
Respond with a JSON object in the following format:
{
"sentiment": "positive" or "negative",
"confidence": number between 1-10,
"reasoning": "brief explanation"
}
config:
response_format:
type: json_schema
json_schema:
name: MovieReviewSentiment
schema:
type: object
properties:
sentiment:
type: string
enum: ['positive', 'negative']
confidence:
type: integer
minimum: 1
maximum: 10
reasoning:
type: string
required: ['sentiment', 'confidence', 'reasoning']
defaultTest:
assert:
# Basic JSON validation
- type: is-json
# Track binary classification metrics
- type: javascript
value: 'output.sentiment === context.vars.sentiment'
metric: accuracy
# For F1 score components (treating 'positive' as the positive class)
- type: javascript
value: "output.sentiment === 'positive' && context.vars.sentiment === 'positive' ? 1 : 0"
metric: true_positives
weight: 0
- type: javascript
value: "output.sentiment === 'positive' && context.vars.sentiment === 'negative' ? 1 : 0"
metric: false_positives
weight: 0
- type: javascript
value: "output.sentiment === 'negative' && context.vars.sentiment === 'positive' ? 1 : 0"
metric: false_negatives
weight: 0
- type: javascript
value: "output.sentiment === 'negative' && context.vars.sentiment === 'negative' ? 1 : 0"
metric: true_negatives
weight: 0
derivedMetrics:
# Precision = TP / (TP + FP)
- name: precision
value: true_positives / (true_positives + false_positives)
# Recall = TP / (TP + FN)
- name: recall
value: true_positives / (true_positives + false_negatives)
# F1 Score = 2 * (precision * recall) / (precision + recall)
- name: f1_score
value: 2 * true_positives / (2 * true_positives + false_positives + false_negatives)
# Accuracy = (TP + TN) / (TP + TN + FP + FN)
- name: accuracy_score
value: (true_positives + true_negatives) / (true_positives + true_negatives + false_positives + false_negatives)
tests: file://imdb_eval_sample.csv
+2
View File
@@ -0,0 +1,2 @@
datasets==4.2.0
pandas==2.3.3