Files
promptfoo--promptfoo/examples/amazon-bedrock/models/promptfooconfig.inference-profiles-simple.yaml
T
wehub-resource-sync 0d3cb498a3
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 13:24:08 +08:00

65 lines
2.2 KiB
YAML

# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: 'Simple Inference Profile Example - High Availability Setup'
# This example demonstrates a typical production setup using inference profiles
# for high availability and automatic failover across regions
prompts:
- 'Answer this customer support question professionally: {{question}}'
providers:
# Production inference profile with automatic failover
# This profile might route to us-east-1 primarily, with failover to us-west-2
- id: bedrock:arn:aws:bedrock:us-east-1:YOUR_ACCOUNT_ID:application-inference-profile/prod-claude-ha
label: 'Production Claude (HA)'
config:
inferenceModelType: 'claude' # Required for inference profiles
region: 'us-east-1'
temperature: 0.3 # Lower temperature for consistent customer support
max_tokens: 500
anthropic_version: 'bedrock-2023-05-31'
# Direct model for comparison (no failover)
- id: bedrock:us.anthropic.claude-sonnet-4-6
label: 'Direct Claude (Single Region)'
config:
region: 'us-east-1'
temperature: 0.3
max_tokens: 500
tests:
- vars:
question: 'How do I reset my password?'
assert:
- type: contains
value: 'password'
- type: llm-rubric
value: 'Response should be helpful, professional, and provide clear steps'
- vars:
question: "My order hasn't arrived yet, and it's been 2 weeks. What should I do?"
assert:
- type: llm-rubric
value: 'Response should be empathetic and provide actionable next steps'
- vars:
question: 'Can I change my subscription plan mid-cycle?'
assert:
- type: llm-rubric
value: 'Response should clearly explain the policy and any potential charges'
# Use an inference profile for consistent grading across regions
defaultTest:
options:
provider:
id: bedrock:arn:aws:bedrock:us-east-1:YOUR_ACCOUNT_ID:application-inference-profile/grading-claude
config:
inferenceModelType: 'claude'
temperature: 0 # Zero temperature for consistent grading
max_tokens: 256
assert:
- type: not-contains
value: 'I cannot'
- type: min-length
value: 50