Files
wehub-resource-sync 0d3cb498a3
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 13:24:08 +08:00
..

compare-gpt-model-tiers-mmlu-pro (GPT Model Tiers MMLU-Pro Comparison)

You can run this example with:

npx promptfoo@latest init --example compare-gpt-model-tiers-mmlu-pro
cd compare-gpt-model-tiers-mmlu-pro

This example demonstrates how to benchmark full, mini, and nano OpenAI GPT model tiers using MMLU-Pro, a more challenging successor to MMLU with up to 10 answer options per question.

Prerequisites

  • promptfoo CLI installed (npm install -g promptfoo or brew install promptfoo)
  • OpenAI API key set as OPENAI_API_KEY
  • Hugging Face account and access token (optional for public MMLU-Pro data, useful for higher rate limits)

Hugging Face Authentication

For higher rate limits or private datasets, authenticate with Hugging Face:

  1. Create a Hugging Face account at huggingface.co if you don't have one

  2. Generate an access token at huggingface.co/settings/tokens

  3. Set your token as an environment variable:

    export HF_TOKEN=your_token_here
    

    Or add it to your .env file:

    HF_TOKEN=your_token_here
    

Running the Eval

  1. Get a local copy of the promptfooconfig:

    npx promptfoo@latest init --example compare-gpt-model-tiers-mmlu-pro
    cd compare-gpt-model-tiers-mmlu-pro
    
  2. Run the evaluation:

    npx promptfoo@latest eval
    
  3. View the results:

    npx promptfoo@latest view
    

What's Being Tested

This comparison evaluates all three models on 100 MMLU-Pro questions spanning many subject categories.

Test Structure

The configuration in promptfooconfig.yaml includes:

  1. Prompt Template: Renders all available MMLU-Pro options dynamically and asks for a final answer in a fixed format
  2. Quality Checks:
    • 60-second timeout per question
    • Required final answer format (Therefore, the answer is X)
    • Deterministic JavaScript scoring that compares the parsed final letter against answer
  3. Model Configuration:
    • 1200 max completion tokens for concise reasoning plus the final answer

Customizing

You can modify the test by editing promptfooconfig.yaml:

  1. Evaluate more MMLU-Pro questions:

    tests:
      - huggingface://datasets/TIGER-Lab/MMLU-Pro?split=test&config=default&limit=250
    
  2. Change the number of questions:

    tests:
      - huggingface://datasets/TIGER-Lab/MMLU-Pro?split=test&config=default&limit=200
    
  3. Adjust model parameters:

    providers:
      - id: openai:chat:gpt-5.4
        config:
          max_completion_tokens: 1500
    

Additional Resources