chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,54 @@
|
||||
# anthropic/opus-4-8-coding (Claude Opus 4.8 Advanced Coding)
|
||||
|
||||
This example exercises Claude Opus 4.8 on hard coding tasks using the `xhigh` effort level with adaptive thinking.
|
||||
|
||||
You can run this example with:
|
||||
|
||||
```bash
|
||||
npx promptfoo@latest init --example anthropic/opus-4-8-coding
|
||||
cd opus-4-8-coding
|
||||
```
|
||||
|
||||
## What This Tests
|
||||
|
||||
Claude Opus 4.8 is Anthropic's most capable model for complex reasoning and agentic coding. This example evaluates:
|
||||
|
||||
- **Bug diagnosis** across multiple system boundaries
|
||||
- **Production-quality code generation** with proper error handling
|
||||
- **Code review** with nuanced, prioritized feedback
|
||||
|
||||
## Working with Opus 4.8
|
||||
|
||||
- **Builds on Opus 4.7.** Opus 4.8 supports the same feature set as 4.7 (no breaking API changes) and improves capability on complex reasoning and long-horizon agentic coding.
|
||||
- **Adaptive thinking is opt-in.** Set `thinking: { type: adaptive }` (as this example does) to let the model decide when and how much to reason per request. Without an explicit `thinking` block the model runs **without** extended thinking, even at high effort.
|
||||
- **`effort` defaults to `high`; `xhigh` is available.** Setting `effort: high` behaves the same as omitting it. Start with `xhigh` for coding and agentic work, and pair high effort with a large `max_tokens`.
|
||||
- **Sampling controls are managed for you.** Opus 4.8 rejects `temperature`, `top_p`, and `top_k` at the model level; promptfoo omits them automatically (don't set them in config).
|
||||
|
||||
## Running the Example
|
||||
|
||||
```bash
|
||||
# Set your API key
|
||||
export ANTHROPIC_API_KEY=your_api_key_here
|
||||
|
||||
# Run the evaluation
|
||||
npx promptfoo@latest eval
|
||||
|
||||
# View results
|
||||
npx promptfoo@latest view
|
||||
```
|
||||
|
||||
## Other providers
|
||||
|
||||
Opus 4.8 is also reachable through:
|
||||
|
||||
- AWS Bedrock — `bedrock:us.anthropic.claude-opus-4-8` (or `bedrock:converse:us.anthropic.claude-opus-4-8`)
|
||||
- Google Vertex — `vertex:claude-opus-4-8` with `config.region: global`
|
||||
- Azure AI Foundry — point `anthropic:messages:claude-opus-4-8` at `https://<resource>.services.ai.azure.com/anthropic` via `apiBaseUrl`
|
||||
|
||||
Across all four providers, promptfoo automatically omits the unsupported sampling parameters (`temperature`, `top_p`, `top_k`) for Opus 4.8. The Anthropic Messages provider also logs a one-time warning if you set them explicitly; the Bedrock, Vertex, and Azure paths omit them silently.
|
||||
|
||||
## Learn More
|
||||
|
||||
- [Claude Opus 4.8 announcement](https://www.anthropic.com/news/claude-opus-4-8)
|
||||
- [Anthropic documentation](https://docs.anthropic.com)
|
||||
- [Promptfoo Anthropic provider docs](https://promptfoo.dev/docs/providers/anthropic)
|
||||
@@ -0,0 +1,124 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: Claude Opus 4.8 advanced coding with xhigh effort and adaptive thinking
|
||||
|
||||
prompts:
|
||||
- |
|
||||
{{task}}
|
||||
|
||||
providers:
|
||||
- id: anthropic:messages:claude-opus-4-8
|
||||
config:
|
||||
# Opus 4.8 deprecates manual sampling controls (temperature/top_p/top_k) at
|
||||
# the model level — promptfoo omits them automatically, so don't set them here.
|
||||
#
|
||||
# Adaptive thinking is opt-in: without an explicit `thinking` block the model
|
||||
# runs WITHOUT extended thinking even at high effort. Set it to let the model
|
||||
# decide when and how much to reason per request.
|
||||
thinking:
|
||||
type: adaptive
|
||||
effort: xhigh # Recommended starting point for coding/agentic work (between high and max)
|
||||
max_tokens: 8000
|
||||
|
||||
tests:
|
||||
# Complex bug diagnosis across multiple systems
|
||||
- vars:
|
||||
task: |
|
||||
You're debugging a production issue where users can't log in. Here's what you know:
|
||||
|
||||
1. The frontend shows "Authentication failed" after username/password submission
|
||||
2. Backend logs show successful JWT generation
|
||||
3. Redis cache is returning stale session data
|
||||
4. Database shows correct user credentials
|
||||
5. The issue only affects 10% of login attempts
|
||||
6. It started after deploying a load balancer configuration change
|
||||
|
||||
Diagnose the root cause and propose a fix. Explain your reasoning about what's causing the intermittent nature of the bug.
|
||||
assert:
|
||||
- type: contains-any
|
||||
value: ['load balancer', 'session', 'sticky', 'affinity', 'routing']
|
||||
reason: Should identify load balancer session routing as the issue
|
||||
- type: llm-rubric
|
||||
value: |
|
||||
The response should:
|
||||
1. Identify the root cause (likely session affinity/sticky sessions issue with load balancer)
|
||||
2. Explain why it's intermittent (different backend servers, inconsistent session state)
|
||||
3. Propose concrete fixes (enable sticky sessions, shared session store, stateless tokens)
|
||||
4. Show reasoning about the tradeoffs of different solutions
|
||||
|
||||
# Production-quality code generation with error handling
|
||||
- vars:
|
||||
task: |
|
||||
Write a Python function that:
|
||||
1. Fetches user data from a REST API (may timeout or return errors)
|
||||
2. Caches results in Redis with 5-minute TTL
|
||||
3. Falls back to database if cache miss
|
||||
4. Returns user object or raises appropriate exception
|
||||
|
||||
Include proper error handling, typing, and comments explaining design decisions.
|
||||
assert:
|
||||
- type: contains
|
||||
value: 'def'
|
||||
reason: Should include Python function definition
|
||||
- type: contains-any
|
||||
value: ['try', 'except', 'raise', 'error']
|
||||
reason: Should include error handling
|
||||
- type: contains-any
|
||||
value: ['cache', 'redis', 'ttl']
|
||||
reason: Should implement caching logic
|
||||
- type: llm-rubric
|
||||
value: |
|
||||
The code should:
|
||||
1. Include proper type hints (from typing import ...)
|
||||
2. Handle network timeouts and API errors gracefully
|
||||
3. Implement cache-aside pattern correctly
|
||||
4. Include docstrings and comments explaining design decisions
|
||||
5. Use appropriate exception types
|
||||
6. Be production-ready (not a toy example)
|
||||
|
||||
# Code review with nuanced feedback
|
||||
- vars:
|
||||
task: |
|
||||
Review this React component and provide feedback:
|
||||
|
||||
```jsx
|
||||
function UserList() {
|
||||
const [users, setUsers] = useState([]);
|
||||
|
||||
useEffect(() => {
|
||||
fetch('/api/users')
|
||||
.then(res => res.json())
|
||||
.then(data => setUsers(data));
|
||||
}, []);
|
||||
|
||||
return (
|
||||
<div>
|
||||
{users.map(user => (
|
||||
<div key={user.id}>
|
||||
<h3>{user.name}</h3>
|
||||
<p>{user.email}</p>
|
||||
</div>
|
||||
))}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
```
|
||||
|
||||
Identify issues, suggest improvements, and explain the reasoning behind each suggestion.
|
||||
assert:
|
||||
- type: contains-any
|
||||
value: ['error', 'loading', 'state', 'async']
|
||||
reason: Should identify missing error and loading states
|
||||
- type: llm-rubric
|
||||
value: |
|
||||
The review should identify multiple issues:
|
||||
1. No error handling for failed fetch
|
||||
2. No loading state
|
||||
3. No cleanup for fetch in useEffect
|
||||
4. Missing dependencies might cause issues in strict mode
|
||||
5. No null/empty checks for users array
|
||||
|
||||
For each issue, it should:
|
||||
- Explain why it's a problem
|
||||
- Suggest specific improvements
|
||||
- Provide example code where helpful
|
||||
- Prioritize issues by severity
|
||||
Reference in New Issue
Block a user