426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
189 lines
5.4 KiB
Python
189 lines
5.4 KiB
Python
"""Shared CLI scaffolding for standard benchmark adapters.
|
|
|
|
Each adapter calls ``build_parser`` to add its specific args on top of
|
|
the shared base, then ``run_cli`` to execute the runner. This eliminates
|
|
~40 lines of duplicated argparse boilerplate per adapter.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import logging
|
|
import sys
|
|
from pathlib import Path
|
|
from typing import Sequence
|
|
|
|
from ._base import (
|
|
BenchmarkResult,
|
|
BenchmarkRunner,
|
|
TrajectoryRecordingClient,
|
|
make_client,
|
|
resolve_api_key,
|
|
resolve_endpoint,
|
|
)
|
|
|
|
|
|
def build_parser(*, prog: str, description: str) -> argparse.ArgumentParser:
|
|
parser = argparse.ArgumentParser(prog=prog, description=description)
|
|
parser.add_argument(
|
|
"--model-endpoint",
|
|
default=None,
|
|
help="OpenAI-compatible chat-completion endpoint URL (e.g. http://localhost:8000/v1)",
|
|
)
|
|
parser.add_argument(
|
|
"--provider",
|
|
default=None,
|
|
help="Shortcut for --model-endpoint (one of: openai, groq, openrouter, cerebras, vllm, ollama, elizacloud)",
|
|
)
|
|
parser.add_argument(
|
|
"--model",
|
|
required=False,
|
|
default="gpt-4o-mini",
|
|
help="Model id to send in the chat-completion request",
|
|
)
|
|
parser.add_argument(
|
|
"--api-key-env",
|
|
default="OPENAI_API_KEY",
|
|
help="Environment variable name that holds the API key",
|
|
)
|
|
parser.add_argument(
|
|
"--output",
|
|
required=True,
|
|
help="Output directory for the results JSON",
|
|
)
|
|
parser.add_argument(
|
|
"--limit",
|
|
type=int,
|
|
default=None,
|
|
help="Optional cap on number of evaluated examples",
|
|
)
|
|
parser.add_argument(
|
|
"--mock",
|
|
action="store_true",
|
|
help="Run with a deterministic mock client (no network).",
|
|
)
|
|
parser.add_argument("--expand-scenarios", action="store_true")
|
|
parser.add_argument("--count-scenarios", action="store_true")
|
|
parser.add_argument("--validate-scenarios", action="store_true")
|
|
parser.add_argument(
|
|
"--log-level",
|
|
default="INFO",
|
|
help="Logging level (DEBUG, INFO, WARNING, ERROR)",
|
|
)
|
|
return parser
|
|
|
|
|
|
def run_cli(
|
|
*,
|
|
runner_factory: "RunnerFactory",
|
|
output_filename: str,
|
|
argv: Sequence[str] | None = None,
|
|
) -> int:
|
|
"""Drive the shared CLI flow:
|
|
|
|
1. Parse args + setup logging.
|
|
2. Resolve endpoint / api key / mock-vs-real client.
|
|
3. Invoke the runner.
|
|
4. Persist results to ``<output>/<output_filename>``.
|
|
|
|
``runner_factory`` is a callable taking the parsed args and returning
|
|
a ``BenchmarkRunner`` plus an optional list of mock responses.
|
|
"""
|
|
|
|
parser = build_parser(prog=runner_factory.prog, description=runner_factory.description)
|
|
runner_factory.augment_parser(parser)
|
|
args = parser.parse_args(argv)
|
|
|
|
logging.basicConfig(
|
|
level=getattr(logging, str(args.log_level).upper(), logging.INFO),
|
|
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
|
|
)
|
|
|
|
output_dir = Path(args.output)
|
|
output_dir.mkdir(parents=True, exist_ok=True)
|
|
|
|
runner, mock_responses = runner_factory.build(args)
|
|
if args.count_scenarios or args.validate_scenarios:
|
|
if not hasattr(runner, "scenario_counts"):
|
|
raise RuntimeError(f"{runner.benchmark_id} does not expose scenario_counts")
|
|
counts = runner.scenario_counts(limit=args.limit) # type: ignore[attr-defined]
|
|
if args.validate_scenarios:
|
|
print("Scenario validation: ok")
|
|
if args.count_scenarios:
|
|
print(json.dumps(counts, sort_keys=True))
|
|
return 0
|
|
|
|
endpoint = (
|
|
"mock://standard-benchmark"
|
|
if args.mock
|
|
else resolve_endpoint(
|
|
model_endpoint=args.model_endpoint,
|
|
provider=args.provider,
|
|
)
|
|
)
|
|
api_key = resolve_api_key(args.api_key_env)
|
|
client = make_client(
|
|
endpoint=endpoint,
|
|
api_key=api_key,
|
|
mock_responses=mock_responses if args.mock else None,
|
|
)
|
|
client = TrajectoryRecordingClient(
|
|
client,
|
|
output_path=output_dir / "trajectories.jsonl",
|
|
benchmark_id=runner.benchmark_id,
|
|
model=args.model,
|
|
)
|
|
|
|
result: BenchmarkResult = runner.run(
|
|
client=client,
|
|
model=args.model,
|
|
endpoint=endpoint,
|
|
output_dir=output_dir,
|
|
limit=args.limit,
|
|
)
|
|
out_path = result.write(output_dir / output_filename)
|
|
print(json.dumps({"output": str(out_path), "metrics": result.metrics}, indent=2))
|
|
return 0
|
|
|
|
|
|
class RunnerFactory:
|
|
"""Per-adapter factory that the shared CLI driver calls into.
|
|
|
|
Adapters subclass this to provide their parser-augmentation and
|
|
runner construction without depending on argparse internals.
|
|
"""
|
|
|
|
prog: str = "benchmarks.standard"
|
|
description: str = ""
|
|
|
|
def augment_parser(self, parser: argparse.ArgumentParser) -> None:
|
|
"""Override to add adapter-specific args."""
|
|
|
|
def build(
|
|
self,
|
|
args: argparse.Namespace,
|
|
) -> tuple[BenchmarkRunner, Sequence[str] | None]:
|
|
raise NotImplementedError
|
|
|
|
|
|
def main_entry(
|
|
runner_factory: RunnerFactory,
|
|
*,
|
|
output_filename: str,
|
|
argv: Sequence[str] | None = None,
|
|
) -> int:
|
|
return run_cli(
|
|
runner_factory=runner_factory,
|
|
output_filename=output_filename,
|
|
argv=argv,
|
|
)
|
|
|
|
|
|
def cli_dispatch(
|
|
runner_factory: RunnerFactory,
|
|
*,
|
|
output_filename: str,
|
|
) -> None:
|
|
sys.exit(main_entry(runner_factory, output_filename=output_filename))
|