Files
elizaos--eliza/packages/benchmarks/visualwebbench
wehub-resource-sync 426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:43:05 +08:00
..

VisualWebBench Benchmark for ElizaOS

A faithful implementation of VisualWebBench (Apache-2.0), a seven-subtask multimodal web understanding and grounding benchmark.

The package routes through the real Eliza adapter by default and downloads the dataset (with screenshots) lazily from Hugging Face.

Subtasks and metrics

Subtask Metric
web_caption ROUGE-1 / ROUGE-2 / ROUGE-L F1 (headline = ROUGE-L)
webqa ROUGE-1 F1, best of reference list
heading_ocr ROUGE-1 / ROUGE-2 / ROUGE-L F1
element_ocr ROUGE-1 / ROUGE-2 / ROUGE-L F1
element_ground MCQ accuracy
action_prediction MCQ accuracy
action_ground MCQ accuracy

These mirror the upstream scorers in VisualWebBench/utils/eval_utils.py. The ROUGE implementation is in-tree (lcs- and ngram-based F1) and matches the reference rouge package within rounding.

Quick start

Run the seven subtasks against the live HF dataset using the Eliza adapter:

pip install -e "packages/benchmarks/visualwebbench[hf]"
PYTHONPATH=packages:packages/benchmarks/eliza-adapter \
  python -m benchmarks.visualwebbench --max-tasks 70

Screenshots are cached as PNG under ~/.cache/elizaos/visualwebbench/images/ the first time each task is encountered.

Offline / CI mode

A 7-row labeled JSONL fixture is bundled at fixtures/smoke.jsonl (one row per subtask, no images). It is only a metric-plumbing helper — scores from it are not comparable to upstream. Combine with --mock to short-circuit the agent entirely:

PYTHONPATH=packages:packages/benchmarks/eliza-adapter \
  python -m benchmarks.visualwebbench --use-sample-tasks --mock --max-tasks 7

--mock reads task.answer and echoes a well-formed response, so it always scores 100. It is gated to this flag — every other run path uses the real agent.

Outputs

  • visualwebbench-results.json — full per-task records with per-subtask metrics
  • summary.md — headline table plus per-subtask breakdown
  • traces/<task-id>.json — one trace per task

CLI flags worth knowing

Flag Purpose
--mock Use the offline oracle (CI only)
--use-sample-tasks Use the bundled labeled JSONL helper
--max-tasks N Cap total tasks (divided across subtasks)
--task-types a,b,c Restrict to a subset of subtasks
--image-cache-dir P Override the on-disk image cache
--no-image-cache Keep image bytes in memory only

Hugging Face details

  • Repo: visualwebbench/VisualWebBench (Apache-2.0)
  • Splits: test only
  • Each of the seven subtasks is its own HF config
  • Images: PIL Image cells, decoded lazily; written to PNG on disk by default
  • Sizes: ~1.5k rows total, several hundred MB of screenshots — fetched lazily so capping --max-tasks keeps downloads small.