426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
64 lines
3.0 KiB
Markdown
64 lines
3.0 KiB
Markdown
# Deterministic real-agent app-build e2e
|
|
|
|
Proves the **orchestrator → real eliza-code coding agent → plan/tool/file-write →
|
|
task_complete** pipeline end-to-end, deterministically, with **no live LLM**.
|
|
|
|
This is a *real* agent, not a fake/stub one. The only thing mocked is the model
|
|
**provider** (an OpenAI-compatible HTTP endpoint): the agent's provider is
|
|
pointed at a local record/replay proxy that serves a recorded "ideal"
|
|
gemma-4-31b session. Everything else is the production code path — the
|
|
orchestrator's `AcpService` spawns the same `src/acp.ts` ACP agent the live bot
|
|
uses (`codingOnly`), the agent plans, calls `fs/write_text_file`, and the
|
|
orchestrator executes those writes into a real workspace.
|
|
|
|
## Files
|
|
|
|
- `deterministic-app-build-replay.mjs` — the driver. Spawns the real agent via
|
|
`AcpService`, sends the exact recorded prompt, asserts it builds a sane
|
|
`index.html` with `stopReason === "end_turn"`.
|
|
- `llm-record-replay-proxy.mjs` — OpenAI-compatible record/replay proxy. In
|
|
`record` it forwards to Cerebras and captures raw response bodies (handles SSE
|
|
streaming) while absorbing 429 TPM rate-limits with backoff; in `replay` it
|
|
serves recorded responses keyed by a volatile-normalized hash of
|
|
`(model, messages, tools)`, with sequential fallback.
|
|
- `fixtures/random-color-gemma-session.json` — the recorded gemma-4-31b session.
|
|
|
|
## Run
|
|
|
|
Prerequisite (once per checkout): `bun install` at the repo root, which runs the
|
|
core codegen this from-source run needs. If you skipped the postinstall codegen,
|
|
run it explicitly: `node packages/shared/scripts/generate-keywords.mjs --target ts`.
|
|
|
|
Replay (default — **keyless, no live LLM, deterministic**, safe for CI):
|
|
|
|
```bash
|
|
bun run --cwd packages/examples/code e2e:deterministic-replay
|
|
# equivalently, from packages/examples/code:
|
|
bun --conditions eliza-source --tsconfig-override ../../../tsconfig.json \
|
|
tests/e2e/deterministic-app-build-replay.mjs
|
|
```
|
|
|
|
Re-record the fixture against live Cerebras gemma-4-31b (needs a key):
|
|
|
|
```bash
|
|
LLM_MODE=record CEREBRAS_API_KEY=csk-... \
|
|
bun --conditions eliza-source --tsconfig-override ../../../tsconfig.json \
|
|
tests/e2e/deterministic-app-build-replay.mjs
|
|
```
|
|
|
|
## Why the fixed workspace
|
|
|
|
Record and replay use a **fixed** workspace path (reset clean each run), not a
|
|
per-run temp dir. Replaying an agentic loop is only deterministic when the
|
|
filesystem context is identical: the agent's tool results (file reads/writes,
|
|
git state) feed back into the conversation, so a different workspace would make
|
|
requests drift off the recorded turn sequence. The proxy additionally normalizes
|
|
volatile tokens (paths, UUIDs, timestamps) out of the match key.
|
|
|
|
The replay proxy also rewrites recorded workspace paths inside model responses
|
|
to the current fixed workspace. That keeps a fixture recorded on macOS usable on
|
|
Linux CI, where `/tmp/eliza-det-replay-workspace` is the real write target.
|
|
|
|
Re-record whenever the agent's prompt scaffolding or the orchestrator's ACP
|
|
event mapping changes in a way that alters the request sequence.
|