426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
136 lines
4.7 KiB
Python
136 lines
4.7 KiB
Python
#!/usr/bin/env python3
|
|
"""
|
|
BFCL Benchmark Runner Script
|
|
|
|
Runs the BFCL benchmark against ElizaOS Python and generates results.
|
|
|
|
Usage:
|
|
python benchmarks/bfcl/scripts/run_benchmark.py [OPTIONS]
|
|
|
|
Options:
|
|
--sample N Run only N sample tests
|
|
--mock Use mock agent (for testing)
|
|
--output DIR Output directory
|
|
"""
|
|
|
|
import asyncio
|
|
import json
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
# Add project root to path
|
|
project_root = Path(__file__).parent.parent.parent.parent
|
|
sys.path.insert(0, str(project_root))
|
|
|
|
|
|
async def main():
|
|
"""Run the BFCL benchmark."""
|
|
import argparse
|
|
|
|
parser = argparse.ArgumentParser(description="Run BFCL Benchmark")
|
|
parser.add_argument("--sample", type=int, help="Number of sample tests")
|
|
parser.add_argument("--mock", action="store_true", help="Use mock agent")
|
|
parser.add_argument("--output", default="./benchmark_results/bfcl",
|
|
help="Output directory")
|
|
parser.add_argument("--categories", help="Comma-separated categories")
|
|
parser.add_argument("--expand-scenarios", action="store_true",
|
|
help="Run each selected BFCL case with 10 edge variants")
|
|
parser.add_argument("--count-scenarios", action="store_true",
|
|
help="Print base/edge/total scenario counts and exit")
|
|
parser.add_argument("--validate-scenarios", action="store_true",
|
|
help="Validate selected base/expanded BFCL cases and exit")
|
|
args = parser.parse_args()
|
|
|
|
from benchmarks.bfcl import BFCLRunner, BFCLConfig, BFCLCategory
|
|
from benchmarks.bfcl.dataset import BFCLDataset, expand_test_cases, validate_test_cases
|
|
from benchmarks.bfcl.reporting import print_results
|
|
|
|
# Configure benchmark
|
|
config = BFCLConfig(
|
|
output_dir=args.output,
|
|
generate_report=True,
|
|
compare_baselines=True,
|
|
include_edge_scenarios=bool(args.expand_scenarios),
|
|
)
|
|
|
|
if args.categories:
|
|
config.categories = [
|
|
BFCLCategory(c.strip())
|
|
for c in args.categories.split(",")
|
|
]
|
|
|
|
if args.count_scenarios or args.validate_scenarios:
|
|
dataset = BFCLDataset(config)
|
|
await dataset.load()
|
|
base_cases = list(dataset)
|
|
if args.sample:
|
|
base_cases = base_cases[: args.sample]
|
|
cases = expand_test_cases(base_cases) if config.include_edge_scenarios else base_cases
|
|
if args.validate_scenarios:
|
|
validate_test_cases(cases)
|
|
print(json.dumps({
|
|
"base": len(base_cases),
|
|
"edge": len(cases) - len(base_cases),
|
|
"total": len(cases),
|
|
}, sort_keys=True))
|
|
return 0
|
|
|
|
# Create runner
|
|
runner = BFCLRunner(config, use_mock_agent=args.mock)
|
|
|
|
print("\n" + "=" * 60)
|
|
print("BFCL BENCHMARK - ElizaOS Python")
|
|
print("=" * 60)
|
|
|
|
try:
|
|
if args.sample:
|
|
print(f"\nRunning sample of {args.sample} tests...\n")
|
|
results = await runner.run_sample(n=args.sample)
|
|
else:
|
|
print("\nRunning full benchmark...\n")
|
|
results = await runner.run()
|
|
|
|
# Print results
|
|
print_results(results)
|
|
|
|
# Save results summary
|
|
summary_path = Path(args.output) / "RESULTS.md"
|
|
summary_path.parent.mkdir(parents=True, exist_ok=True)
|
|
|
|
with open(summary_path, "w") as f:
|
|
f.write("# BFCL Benchmark Results - ElizaOS Python\n\n")
|
|
f.write(f"## Overall Score: {results.metrics.overall_score:.2%}\n\n")
|
|
f.write("| Metric | Score |\n")
|
|
f.write("|--------|-------|\n")
|
|
f.write(f"| AST Accuracy | {results.metrics.ast_accuracy:.2%} |\n")
|
|
f.write(f"| Exec Accuracy | {results.metrics.exec_accuracy:.2%} |\n")
|
|
f.write(f"| Relevance Accuracy | {results.metrics.relevance_accuracy:.2%} |\n")
|
|
f.write(f"\nTotal Tests: {results.metrics.total_tests}\n")
|
|
f.write(f"Passed: {results.metrics.passed_tests}\n")
|
|
f.write(f"Failed: {results.metrics.failed_tests}\n")
|
|
|
|
if results.baseline_comparison:
|
|
f.write("\n## Baseline Comparison\n\n")
|
|
f.write("| Model | Difference |\n")
|
|
f.write("|-------|------------|\n")
|
|
for model, diff in sorted(
|
|
results.baseline_comparison.items(),
|
|
key=lambda x: x[1],
|
|
reverse=True,
|
|
):
|
|
sign = "+" if diff > 0 else ""
|
|
f.write(f"| {model} | {sign}{diff:.2%} |\n")
|
|
|
|
print(f"\nResults saved to {summary_path}")
|
|
return 0
|
|
|
|
except Exception as e:
|
|
print(f"\n❌ Benchmark failed: {e}")
|
|
import traceback
|
|
traceback.print_exc()
|
|
return 1
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(asyncio.run(main()))
|