chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled

This commit is contained in:
wehub-resource-sync
2026-07-13 13:24:08 +08:00
commit 0d3cb498a3
5438 changed files with 1316560 additions and 0 deletions
+48
View File
@@ -0,0 +1,48 @@
# openai-agents-advanced (Sessions, Tracing, and Sandbox Agents)
This example exercises the OpenAI Agents SDK TypeScript features that matter once you move beyond a single-turn agent:
- persistent `MemorySession` history
- Promptfoo vars forwarded as SDK local run context
- file-exported SDK tools
- Promptfoo trajectory assertions over SDK traces
- `SandboxAgent` execution with the SDK local sandbox client
- sandbox skills loaded through SDK capability objects
## Prerequisites
- `OPENAI_API_KEY`
## Installation
```bash
npx promptfoo@latest init --example openai-agents-advanced
cd openai-agents-advanced
```
Or, from a cloned repository:
```bash
cd examples/openai-agents-advanced
npm install
```
## Run the session and tracing eval
```bash
npx promptfoo eval -c promptfooconfig.yaml --no-cache -j 1
```
The second test depends on the first test's remembered code word, so run this config with `-j 1`.
For stateful red-team strategies, use a session factory keyed by a per-test `sessionId` rather than one shared inline session. The [OpenAI Agents provider docs](https://www.promptfoo.dev/docs/providers/openai-agents/#stateful-red-team-runs) show the `transformVars` plus session-factory pattern that keeps turns together without sharing history across unrelated tests; the [multi-turn strategy docs](https://www.promptfoo.dev/docs/red-team/strategies/multi-turn/) explain when `stateful: true` is appropriate.
## Run the sandbox and skill eval
```bash
npx promptfoo eval -c promptfooconfig.sandbox.yaml --no-cache
```
The sandbox eval mounts a synthetic `task.md`, asks the agent to use the `ticket-summary` skill, and asserts on traced shell activity plus the final answer.
See also the [Tracing docs](https://www.promptfoo.dev/docs/tracing/) for trajectory assertions and the [OpenAI Agents provider docs](https://www.promptfoo.dev/docs/providers/openai-agents/) for the full JavaScript SDK configuration surface.
@@ -0,0 +1,12 @@
import { Agent } from '@openai/agents';
export default new Agent({
name: 'Support Agent',
model: 'gpt-5-mini',
instructions: `You are a concise support agent.
- Remember short code words the user gives you.
- When the user asks about an order, use lookup_order before answering.
- When the user asks about their customer tier, use lookup_customer_context before answering.
- Reply with the remembered code word when the user asks for it later.`,
});
@@ -0,0 +1,45 @@
import { fileURLToPath } from 'node:url';
import { localFile, Manifest, SandboxAgent, shell, skills } from '@openai/agents/sandbox';
const taskFilePath = fileURLToPath(new URL('../workspace/task.md', import.meta.url));
export default new SandboxAgent({
name: 'Workspace Reviewer',
model: 'gpt-5-mini',
modelSettings: {
toolChoice: 'required',
},
instructions: `You are reviewing a small sandbox workspace.
- Use the ticket-summary skill before answering.
- Inspect task.md through the available sandbox tools.
- Always run exec_command to read task.md before answering, even if you think you already know its contents.
- Reply with the ticket ID followed by one concise sentence.`,
defaultManifest: new Manifest({
entries: {
'task.md': localFile({
src: taskFilePath,
}),
},
}),
capabilities: [
shell(),
skills({
skills: [
{
name: 'ticket-summary',
description: 'Summarize a ticket file into one sentence.',
content: `---
name: ticket-summary
description: Summarize a ticket file into one sentence.
---
1. Read the requested ticket file before answering.
2. Preserve the ticket ID exactly.
3. Return one concise sentence after the ID.`,
},
],
}),
],
});
@@ -0,0 +1,12 @@
{
"name": "@promptfoo/openai-agents-advanced-example",
"version": "1.0.0",
"license": "MIT",
"private": true,
"description": "Advanced OpenAI Agents SDK TypeScript example with sessions, sandbox agents, skills, and tracing",
"type": "module",
"dependencies": {
"@openai/agents": "^0.11.3",
"zod": "^4.3.6"
}
}
@@ -0,0 +1,34 @@
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: OpenAI Agents SDK sandbox and skill features
prompts:
- '{{query}}'
tracing:
enabled: true
otlp:
http:
enabled: true
port: 4318
providers:
- id: openai:agents:workspace-agent
config:
agent: file://./agents/workspace-agent.ts
sandbox:
type: unix-local
tracing: true
maxTurns: 12
tests:
- description: Inspect sandbox workspace through the skill workflow
vars:
query: Use the ticket-summary skill, run exec_command to read task.md, and summarize the ticket.
assert:
- type: contains
value: PF-42
- type: trajectory:step-count
value:
type: command
pattern: '*task.md*'
min: 1
@@ -0,0 +1,57 @@
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: OpenAI Agents SDK advanced TypeScript features
prompts:
- '{{query}}'
tracing:
enabled: true
otlp:
http:
enabled: true
port: 4318
providers:
- id: openai:agents:support-agent
config:
agent: file://./agents/support-agent.ts
tools: file://./tools/support-tools.ts
session:
type: memory
sessionId: support-demo
tracing: true
maxTurns: 8
tests:
- description: Store a remembered code word
vars:
query: 'Remember this code word exactly: violet.'
assert:
- type: contains
value: violet
- description: Reuse session history and traced tool calls
vars:
query: What code word did I tell you earlier? Then look up order 123.
assert:
- type: contains
value: violet
- type: contains
value: shipped
- type: trajectory:tool-used
value: lookup_order
- type: trajectory:tool-args-match
value:
name: lookup_order
args:
order_id: '123'
- description: Read Promptfoo vars through SDK local context
vars:
customer_tier: gold
query: What customer tier am I on?
assert:
- type: contains
value: gold
- type: trajectory:tool-used
value: lookup_customer_context
@@ -0,0 +1,30 @@
import { tool } from '@openai/agents';
import { z } from 'zod';
export const lookupOrder = tool({
name: 'lookup_order',
description: 'Look up the shipping status for an order.',
parameters: z.object({
order_id: z.string(),
}),
execute: async ({ order_id }) => ({
order_id,
status: 'shipped',
}),
});
export const lookupCustomerContext = tool({
name: 'lookup_customer_context',
description: 'Read the current customer tier from local run context.',
parameters: z.object({}),
execute: async (_args, runContext) => ({
customer_tier:
typeof runContext?.context === 'object' &&
runContext.context !== null &&
'customer_tier' in runContext.context
? runContext.context.customer_tier
: 'unknown',
}),
});
export default [lookupOrder, lookupCustomerContext];
@@ -0,0 +1,3 @@
Ticket PF-42: update the release note for the Agents SDK workflow launch.
Mention that the workflow now supports persistent sessions, local context, tracing, and sandbox execution.