chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,111 @@
|
||||
# xai/chat (xAI Grok Models Evaluation)
|
||||
|
||||
This example demonstrates how to evaluate xAI's Grok models across their main capabilities: text generation with reasoning, image creation, and server-side search tools.
|
||||
|
||||
You can run this example with:
|
||||
|
||||
```bash
|
||||
npx promptfoo@latest init --example xai/chat
|
||||
cd xai/chat
|
||||
```
|
||||
|
||||
## Environment Variables
|
||||
|
||||
This example requires the following environment variable:
|
||||
|
||||
- `XAI_API_KEY` - Your xAI API key. You can obtain this from the [xAI Console](https://console.x.ai/)
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Set your API key
|
||||
export XAI_API_KEY=your_api_key_here
|
||||
|
||||
# Run the main evaluation
|
||||
promptfoo eval
|
||||
|
||||
# View results in the web interface
|
||||
promptfoo view
|
||||
```
|
||||
|
||||
## What's Tested
|
||||
|
||||
This example includes configurations to test different Grok capabilities:
|
||||
|
||||
- **Text Generation** (`promptfooconfig.yaml`) - Mathematical reasoning with the current Grok 4.3 and Grok 4.20 families
|
||||
- **Image Generation** (`promptfooconfig.images.yaml`) - Artistic image creation using Grok's image models
|
||||
- **Search Tools** (`promptfooconfig.search.yaml`) - Real-time web and X search using the Responses API
|
||||
- **Agent Tools (Responses API)** (`promptfooconfig.responses.yaml`) - Autonomous web and X search using Agent Tools
|
||||
- **Search Demo** (`promptfooconfig.promptfoo-search.yaml`) - Responses API search with assertions example
|
||||
|
||||
## Run Individual Tests
|
||||
|
||||
```bash
|
||||
# Text generation with mathematical reasoning
|
||||
promptfoo eval -c promptfooconfig.yaml
|
||||
|
||||
# Image generation with artistic styles
|
||||
promptfoo eval -c promptfooconfig.images.yaml
|
||||
|
||||
# Search tools with web and X sources
|
||||
promptfoo eval -c promptfooconfig.search.yaml
|
||||
|
||||
# Agent Tools with Responses API (recommended)
|
||||
promptfoo eval -c promptfooconfig.responses.yaml
|
||||
|
||||
# Search demo with assertions
|
||||
promptfoo eval -c promptfooconfig.promptfoo-search.yaml
|
||||
```
|
||||
|
||||
## Featured Models
|
||||
|
||||
### Grok 4.3
|
||||
|
||||
The recommended starting point for general text workflows:
|
||||
|
||||
- `xai:grok-4.3` - General-purpose reasoning model
|
||||
- `reasoning_effort` - Supports `none`, `low`, `medium`, and `high` in chat configs
|
||||
- `xai:responses:grok-4.3` - Recommended form for server-side tools
|
||||
|
||||
### Grok 4.20
|
||||
|
||||
- `xai:grok-4.20-reasoning` - Reasoning model
|
||||
- `xai:grok-4.20-non-reasoning` - Non-reasoning model
|
||||
- `xai:grok-4.20-multi-agent` - Multi-agent variant
|
||||
|
||||
### Legacy Model Note
|
||||
|
||||
xAI periodically retires older model slugs and may keep them working through redirects to newer replacements. This example uses Grok 4.3 plus alias-style Grok 4.20 family IDs, matching xAI's guidance for configs that should track the current release within a family.
|
||||
|
||||
### Agent Tools (Responses API)
|
||||
|
||||
Enable autonomous tool execution via the Responses API:
|
||||
|
||||
```yaml
|
||||
providers:
|
||||
- id: xai:responses:grok-4.3
|
||||
config:
|
||||
tools:
|
||||
- type: web_search
|
||||
- type: x_search
|
||||
- type: code_interpreter
|
||||
```
|
||||
|
||||
### Search Tools
|
||||
|
||||
Enable real-time search via the Responses API:
|
||||
|
||||
```yaml
|
||||
providers:
|
||||
- id: xai:responses:grok-4.3
|
||||
config:
|
||||
tools:
|
||||
- type: web_search
|
||||
- type: x_search
|
||||
```
|
||||
|
||||
## Expected Results
|
||||
|
||||
- **Text Generation**: Grok will provide step-by-step mathematical solutions with clear reasoning
|
||||
- **Image Generation**: Generated images in the requested artistic styles
|
||||
- **Search Tools**: Current information from web and X with source citations
|
||||
@@ -0,0 +1,6 @@
|
||||
- role: system
|
||||
content: |
|
||||
You are a helpful assistant skilled at solving problems through careful reasoning and analysis.
|
||||
Approach problems step-by-step, showing your thought process clearly.
|
||||
- role: user
|
||||
content: '{{problem}}'
|
||||
@@ -0,0 +1,15 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: xAI image generation evaluation
|
||||
prompts:
|
||||
- 'In the style of {{artist}}: {{subject}}'
|
||||
|
||||
providers:
|
||||
- xai:image:grok-imagine-image
|
||||
|
||||
tests:
|
||||
- vars:
|
||||
artist: Van Gogh
|
||||
subject: sunset over mountains
|
||||
- vars:
|
||||
artist: Dali
|
||||
subject: melting clocks
|
||||
@@ -0,0 +1,70 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: xAI search for promptfoo info
|
||||
|
||||
prompts:
|
||||
- |
|
||||
{{question}}
|
||||
|
||||
Search both web and X/Twitter for the most relevant and recent information.
|
||||
Include specific examples, quotes from posts, or citations where available.
|
||||
|
||||
providers:
|
||||
# Grok 4.3 with server-side search tools
|
||||
- id: xai:responses:grok-4.3
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_output_tokens: 2048
|
||||
tools:
|
||||
- type: web_search
|
||||
- type: x_search
|
||||
|
||||
tests:
|
||||
# Developer love and adoption
|
||||
- vars:
|
||||
question: Why do developers love promptfoo? What features make it stand out as the go-to open-source tool for LLM evaluation and testing?
|
||||
assert:
|
||||
- type: contains
|
||||
value: promptfoo
|
||||
- type: llm-rubric
|
||||
value: The response should highlight positive aspects and strengths of promptfoo
|
||||
|
||||
# Developer adoption and community
|
||||
- vars:
|
||||
question: How many developers are using promptfoo? What does the community say about it on GitHub and developer forums?
|
||||
assert:
|
||||
- type: contains
|
||||
value: promptfoo
|
||||
- type: llm-rubric
|
||||
value: The response should mention developer adoption numbers or positive community feedback about promptfoo
|
||||
|
||||
# Red teaming leadership
|
||||
- vars:
|
||||
question: How has promptfoo become the leading open-source solution for AI red teaming and LLM security testing? What vulnerabilities can it detect?
|
||||
assert:
|
||||
- type: contains-any
|
||||
value:
|
||||
- red team
|
||||
- security
|
||||
- vulnerabilities
|
||||
- jailbreak
|
||||
- prompt injection
|
||||
- type: llm-rubric
|
||||
value: The response should explain promptfoo's red teaming and security testing capabilities
|
||||
|
||||
# Recent momentum and growth
|
||||
- vars:
|
||||
question: What's driving promptfoo's rapid growth and developer adoption? What recent milestones or funding has it achieved?
|
||||
assert:
|
||||
- type: contains
|
||||
value: promptfoo
|
||||
- type: llm-rubric
|
||||
value: The response should mention growth, funding, or adoption milestones for promptfoo
|
||||
|
||||
# Unique value proposition
|
||||
- vars:
|
||||
question: What makes promptfoo unique for testing LLM applications? Why do teams choose it for evaluating prompts, agents, and RAG systems?
|
||||
assert:
|
||||
- type: contains
|
||||
value: promptfoo
|
||||
- type: llm-rubric
|
||||
value: The response should explain what makes promptfoo valuable and differentiated
|
||||
@@ -0,0 +1,39 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: xAI Responses API with Agent Tools
|
||||
|
||||
prompts:
|
||||
- '{{question}}'
|
||||
|
||||
providers:
|
||||
# Grok 4.3 with Agent Tools (web + X search)
|
||||
- id: xai:responses:grok-4.3
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_output_tokens: 2048
|
||||
tools:
|
||||
- type: web_search
|
||||
- type: x_search
|
||||
|
||||
tests:
|
||||
# Test web search capability
|
||||
- vars:
|
||||
question: What is the current weather in San Francisco? Search the web for the latest information.
|
||||
assert:
|
||||
- type: contains-any
|
||||
value:
|
||||
- San Francisco
|
||||
- weather
|
||||
- temperature
|
||||
- degrees
|
||||
|
||||
# Test X search capability
|
||||
- vars:
|
||||
question: What are people saying about AI on X/Twitter today? Search X for recent posts.
|
||||
assert:
|
||||
- type: contains-any
|
||||
value:
|
||||
- AI
|
||||
- artificial intelligence
|
||||
- posts
|
||||
- X
|
||||
- Twitter
|
||||
@@ -0,0 +1,29 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: xAI Responses API search demo
|
||||
|
||||
prompts:
|
||||
- 'What is the latest news about {{topic}}?'
|
||||
|
||||
providers:
|
||||
# Grok 4.3 with web + X search
|
||||
- id: xai:responses:grok-4.3
|
||||
label: grok-4.3-web-and-x-search
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_output_tokens: 2048
|
||||
tools:
|
||||
- type: web_search
|
||||
- type: x_search
|
||||
|
||||
# Grok 4.3 with web search only
|
||||
- id: xai:responses:grok-4.3
|
||||
label: grok-4.3-web-search
|
||||
config:
|
||||
tools:
|
||||
- type: web_search
|
||||
|
||||
tests:
|
||||
- vars:
|
||||
topic: xAI
|
||||
- vars:
|
||||
topic: artificial intelligence breakthroughs
|
||||
@@ -0,0 +1,41 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: Testing xAI Grok models
|
||||
|
||||
prompts:
|
||||
- What are the key differences between machine learning and deep learning?
|
||||
- file://./prompt.yaml
|
||||
|
||||
providers:
|
||||
# Grok 4.3 (recommended general-purpose model)
|
||||
- id: xai:grok-4.3
|
||||
config:
|
||||
temperature: 0.7
|
||||
reasoning_effort: high
|
||||
max_completion_tokens: 4096
|
||||
|
||||
# Grok 4.20 reasoning family
|
||||
- id: xai:grok-4.20-reasoning
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_completion_tokens: 4096
|
||||
|
||||
# Grok 4.20 non-reasoning family
|
||||
- id: xai:grok-4.20-non-reasoning
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_completion_tokens: 4096
|
||||
|
||||
# Grok 4.20 multi-agent family
|
||||
- id: xai:grok-4.20-multi-agent
|
||||
config:
|
||||
temperature: 0.7
|
||||
max_completion_tokens: 4096
|
||||
|
||||
defaultTest:
|
||||
assert:
|
||||
- type: javascript
|
||||
value: 'output.length >= 100'
|
||||
|
||||
tests:
|
||||
- vars:
|
||||
problem: 'A baker has 3 boxes of cookies. Each box contains 15 cookies. He sells 12 cookies. How many cookies does he have left?'
|
||||
Reference in New Issue
Block a user