chore: import upstream snapshot with attribution
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
CI / Shell Format Check (push) Has been cancelled
CI / Check Ruby (3.4) (push) Has been cancelled
CI / CI Config (push) Has been cancelled
CI / Test on Node ${{ matrix.node }} and ${{ matrix.os }}${{ matrix.shard && format(' (shard {0}/3)', matrix.shard) || '' }} (push) Has been cancelled
CI / Build on Node ${{ matrix.node }} (push) Has been cancelled
CI / Style Check (push) Has been cancelled
CI / Generate Assets (push) Has been cancelled
CI / Check Python (3.14) (push) Has been cancelled
CI / Check Python (3.9) (push) Has been cancelled
CI / Build Docs (push) Has been cancelled
CI / Code Scan Action (push) Has been cancelled
CI / Site tests (push) Has been cancelled
CI / webui tests (push) Has been cancelled
CI / Run Integration Tests (push) Has been cancelled
CI / Run Smoke Tests (push) Has been cancelled
CI / Go Tests (push) Has been cancelled
CI / Share Test (push) Has been cancelled
CI / Redteam (Production API) (push) Has been cancelled
CI / Redteam (Staging API) (push) Has been cancelled
CI / GitHub Actions Lint (push) Has been cancelled
CI / Check Ruby (3.0) (push) Has been cancelled
release-please / release-please (push) Has been cancelled
release-please / build (push) Has been cancelled
release-please / publish-npm (push) Has been cancelled
release-please / publish-npm-backfill (push) Has been cancelled
release-please / docker (push) Has been cancelled
release-please / publish-code-scan-action (push) Has been cancelled
release-please / attest-code-scan-action (push) Has been cancelled
Deploy local.promptfoo.app / Deploy to Cloudflare Pages (push) Has been cancelled
Test and Publish Multi-arch Docker Image / test (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-amd64 platform:linux/amd64 runner:ubuntu-latest]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / build-docker-and-push-digests (map[digest-suffix:linux-arm64 platform:linux/arm64 runner:ubuntu-24.04-arm]) (push) Has been cancelled
Test and Publish Multi-arch Docker Image / merge-docker-digests (push) Has been cancelled
Test and Publish Multi-arch Docker Image / Attest Multi-arch Image (push) Has been cancelled
Validate Renovate Config / Validate Renovate Configuration (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# openai-agents-advanced (Sessions, Tracing, and Sandbox Agents)
|
||||
|
||||
This example exercises the OpenAI Agents SDK TypeScript features that matter once you move beyond a single-turn agent:
|
||||
|
||||
- persistent `MemorySession` history
|
||||
- Promptfoo vars forwarded as SDK local run context
|
||||
- file-exported SDK tools
|
||||
- Promptfoo trajectory assertions over SDK traces
|
||||
- `SandboxAgent` execution with the SDK local sandbox client
|
||||
- sandbox skills loaded through SDK capability objects
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- `OPENAI_API_KEY`
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
npx promptfoo@latest init --example openai-agents-advanced
|
||||
cd openai-agents-advanced
|
||||
```
|
||||
|
||||
Or, from a cloned repository:
|
||||
|
||||
```bash
|
||||
cd examples/openai-agents-advanced
|
||||
npm install
|
||||
```
|
||||
|
||||
## Run the session and tracing eval
|
||||
|
||||
```bash
|
||||
npx promptfoo eval -c promptfooconfig.yaml --no-cache -j 1
|
||||
```
|
||||
|
||||
The second test depends on the first test's remembered code word, so run this config with `-j 1`.
|
||||
|
||||
For stateful red-team strategies, use a session factory keyed by a per-test `sessionId` rather than one shared inline session. The [OpenAI Agents provider docs](https://www.promptfoo.dev/docs/providers/openai-agents/#stateful-red-team-runs) show the `transformVars` plus session-factory pattern that keeps turns together without sharing history across unrelated tests; the [multi-turn strategy docs](https://www.promptfoo.dev/docs/red-team/strategies/multi-turn/) explain when `stateful: true` is appropriate.
|
||||
|
||||
## Run the sandbox and skill eval
|
||||
|
||||
```bash
|
||||
npx promptfoo eval -c promptfooconfig.sandbox.yaml --no-cache
|
||||
```
|
||||
|
||||
The sandbox eval mounts a synthetic `task.md`, asks the agent to use the `ticket-summary` skill, and asserts on traced shell activity plus the final answer.
|
||||
|
||||
See also the [Tracing docs](https://www.promptfoo.dev/docs/tracing/) for trajectory assertions and the [OpenAI Agents provider docs](https://www.promptfoo.dev/docs/providers/openai-agents/) for the full JavaScript SDK configuration surface.
|
||||
@@ -0,0 +1,12 @@
|
||||
import { Agent } from '@openai/agents';
|
||||
|
||||
export default new Agent({
|
||||
name: 'Support Agent',
|
||||
model: 'gpt-5-mini',
|
||||
instructions: `You are a concise support agent.
|
||||
|
||||
- Remember short code words the user gives you.
|
||||
- When the user asks about an order, use lookup_order before answering.
|
||||
- When the user asks about their customer tier, use lookup_customer_context before answering.
|
||||
- Reply with the remembered code word when the user asks for it later.`,
|
||||
});
|
||||
@@ -0,0 +1,45 @@
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
import { localFile, Manifest, SandboxAgent, shell, skills } from '@openai/agents/sandbox';
|
||||
|
||||
const taskFilePath = fileURLToPath(new URL('../workspace/task.md', import.meta.url));
|
||||
|
||||
export default new SandboxAgent({
|
||||
name: 'Workspace Reviewer',
|
||||
model: 'gpt-5-mini',
|
||||
modelSettings: {
|
||||
toolChoice: 'required',
|
||||
},
|
||||
instructions: `You are reviewing a small sandbox workspace.
|
||||
|
||||
- Use the ticket-summary skill before answering.
|
||||
- Inspect task.md through the available sandbox tools.
|
||||
- Always run exec_command to read task.md before answering, even if you think you already know its contents.
|
||||
- Reply with the ticket ID followed by one concise sentence.`,
|
||||
defaultManifest: new Manifest({
|
||||
entries: {
|
||||
'task.md': localFile({
|
||||
src: taskFilePath,
|
||||
}),
|
||||
},
|
||||
}),
|
||||
capabilities: [
|
||||
shell(),
|
||||
skills({
|
||||
skills: [
|
||||
{
|
||||
name: 'ticket-summary',
|
||||
description: 'Summarize a ticket file into one sentence.',
|
||||
content: `---
|
||||
name: ticket-summary
|
||||
description: Summarize a ticket file into one sentence.
|
||||
---
|
||||
|
||||
1. Read the requested ticket file before answering.
|
||||
2. Preserve the ticket ID exactly.
|
||||
3. Return one concise sentence after the ID.`,
|
||||
},
|
||||
],
|
||||
}),
|
||||
],
|
||||
});
|
||||
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"name": "@promptfoo/openai-agents-advanced-example",
|
||||
"version": "1.0.0",
|
||||
"license": "MIT",
|
||||
"private": true,
|
||||
"description": "Advanced OpenAI Agents SDK TypeScript example with sessions, sandbox agents, skills, and tracing",
|
||||
"type": "module",
|
||||
"dependencies": {
|
||||
"@openai/agents": "^0.11.3",
|
||||
"zod": "^4.3.6"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,34 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: OpenAI Agents SDK sandbox and skill features
|
||||
|
||||
prompts:
|
||||
- '{{query}}'
|
||||
|
||||
tracing:
|
||||
enabled: true
|
||||
otlp:
|
||||
http:
|
||||
enabled: true
|
||||
port: 4318
|
||||
|
||||
providers:
|
||||
- id: openai:agents:workspace-agent
|
||||
config:
|
||||
agent: file://./agents/workspace-agent.ts
|
||||
sandbox:
|
||||
type: unix-local
|
||||
tracing: true
|
||||
maxTurns: 12
|
||||
|
||||
tests:
|
||||
- description: Inspect sandbox workspace through the skill workflow
|
||||
vars:
|
||||
query: Use the ticket-summary skill, run exec_command to read task.md, and summarize the ticket.
|
||||
assert:
|
||||
- type: contains
|
||||
value: PF-42
|
||||
- type: trajectory:step-count
|
||||
value:
|
||||
type: command
|
||||
pattern: '*task.md*'
|
||||
min: 1
|
||||
@@ -0,0 +1,57 @@
|
||||
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
|
||||
description: OpenAI Agents SDK advanced TypeScript features
|
||||
|
||||
prompts:
|
||||
- '{{query}}'
|
||||
|
||||
tracing:
|
||||
enabled: true
|
||||
otlp:
|
||||
http:
|
||||
enabled: true
|
||||
port: 4318
|
||||
|
||||
providers:
|
||||
- id: openai:agents:support-agent
|
||||
config:
|
||||
agent: file://./agents/support-agent.ts
|
||||
tools: file://./tools/support-tools.ts
|
||||
session:
|
||||
type: memory
|
||||
sessionId: support-demo
|
||||
tracing: true
|
||||
maxTurns: 8
|
||||
|
||||
tests:
|
||||
- description: Store a remembered code word
|
||||
vars:
|
||||
query: 'Remember this code word exactly: violet.'
|
||||
assert:
|
||||
- type: contains
|
||||
value: violet
|
||||
|
||||
- description: Reuse session history and traced tool calls
|
||||
vars:
|
||||
query: What code word did I tell you earlier? Then look up order 123.
|
||||
assert:
|
||||
- type: contains
|
||||
value: violet
|
||||
- type: contains
|
||||
value: shipped
|
||||
- type: trajectory:tool-used
|
||||
value: lookup_order
|
||||
- type: trajectory:tool-args-match
|
||||
value:
|
||||
name: lookup_order
|
||||
args:
|
||||
order_id: '123'
|
||||
|
||||
- description: Read Promptfoo vars through SDK local context
|
||||
vars:
|
||||
customer_tier: gold
|
||||
query: What customer tier am I on?
|
||||
assert:
|
||||
- type: contains
|
||||
value: gold
|
||||
- type: trajectory:tool-used
|
||||
value: lookup_customer_context
|
||||
@@ -0,0 +1,30 @@
|
||||
import { tool } from '@openai/agents';
|
||||
import { z } from 'zod';
|
||||
|
||||
export const lookupOrder = tool({
|
||||
name: 'lookup_order',
|
||||
description: 'Look up the shipping status for an order.',
|
||||
parameters: z.object({
|
||||
order_id: z.string(),
|
||||
}),
|
||||
execute: async ({ order_id }) => ({
|
||||
order_id,
|
||||
status: 'shipped',
|
||||
}),
|
||||
});
|
||||
|
||||
export const lookupCustomerContext = tool({
|
||||
name: 'lookup_customer_context',
|
||||
description: 'Read the current customer tier from local run context.',
|
||||
parameters: z.object({}),
|
||||
execute: async (_args, runContext) => ({
|
||||
customer_tier:
|
||||
typeof runContext?.context === 'object' &&
|
||||
runContext.context !== null &&
|
||||
'customer_tier' in runContext.context
|
||||
? runContext.context.customer_tier
|
||||
: 'unknown',
|
||||
}),
|
||||
});
|
||||
|
||||
export default [lookupOrder, lookupCustomerContext];
|
||||
@@ -0,0 +1,3 @@
|
||||
Ticket PF-42: update the release note for the Agents SDK workflow launch.
|
||||
|
||||
Mention that the workflow now supports persistent sessions, local context, tracing, and sandbox execution.
|
||||
Reference in New Issue
Block a user