426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
129 lines
4.5 KiB
Python
129 lines
4.5 KiB
Python
"""Purge the literal default-thought strings from packed train.jsonl.
|
|
|
|
The dataset adapters were updated to use a varied thought-phrasing pool, but
|
|
existing `data/normalized/*.jsonl` files were generated before the fix and
|
|
still contain literal strings like `"Reply to the user."` injected as the
|
|
model's reasoning. Re-normalizing 75 GB to fix this is wasteful when we can
|
|
post-process the packed train.jsonl directly.
|
|
|
|
This script walks train.jsonl line by line. For each record where
|
|
`expectedResponse` starts with `thought: <literal default>`, replaces just
|
|
that line with a varied phrasing picked deterministically from the
|
|
adapter's pool (seeded by the user message).
|
|
|
|
Output: rewrites train.jsonl (and val.jsonl) in place via atomic temp file.
|
|
|
|
Usage:
|
|
.venv/bin/python scripts/transform_purge_default_thoughts.py
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import logging
|
|
import os
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
ROOT = Path(__file__).resolve().parent.parent
|
|
sys.path.insert(0, str(ROOT / "scripts"))
|
|
|
|
from lib.adapters import ( # noqa: E402
|
|
_picked_thought,
|
|
_REPLY_THOUGHT_POOL,
|
|
_TOOL_THOUGHT_POOL,
|
|
_SHELL_THOUGHT_POOL,
|
|
_IGNORE_THOUGHT_POOL,
|
|
_AGENT_TRACE_THOUGHT_POOL,
|
|
)
|
|
from lib.eliza_record import DEFAULT_THOUGHT_LEAKS # noqa: E402
|
|
|
|
assert "Reply to the user." in DEFAULT_THOUGHT_LEAKS, (
|
|
"DEFAULT_THOUGHT_LEAKS lost its canonical entry — see lib/eliza_record.py"
|
|
)
|
|
|
|
logging.basicConfig(level=logging.INFO,
|
|
format="%(asctime)s [%(levelname)s] %(message)s")
|
|
log = logging.getLogger("purge")
|
|
|
|
# Map from old literal → pool to use for replacement
|
|
_LEAK_PATTERNS: list[tuple[str, tuple[str, ...]]] = [
|
|
("thought: Reply to the user.", _REPLY_THOUGHT_POOL),
|
|
("thought: Call the tool to satisfy the request.", _TOOL_THOUGHT_POOL),
|
|
("thought: Run the command to satisfy the request.", _SHELL_THOUGHT_POOL),
|
|
('thought: "This message is not for me — ignore."', _IGNORE_THOUGHT_POOL),
|
|
("thought: This message is not for me — ignore.", _IGNORE_THOUGHT_POOL),
|
|
("thought: Continue the agent task.", _AGENT_TRACE_THOUGHT_POOL),
|
|
]
|
|
|
|
|
|
def rewrite_expected_response(er: str, seed: str) -> tuple[str, bool]:
|
|
"""Return (rewritten, was_leaked). Replaces the first thought line if
|
|
it's a known literal default."""
|
|
if not er:
|
|
return er, False
|
|
lines = er.split("\n", 1)
|
|
first = lines[0]
|
|
rest = lines[1] if len(lines) > 1 else ""
|
|
for prefix, pool in _LEAK_PATTERNS:
|
|
if first.startswith(prefix):
|
|
varied = _picked_thought(pool, seed)
|
|
new_first = f"thought: {varied}"
|
|
return f"{new_first}\n{rest}" if rest else new_first, True
|
|
return er, False
|
|
|
|
|
|
def process_file(path: Path) -> tuple[int, int]:
|
|
"""Returns (records, rewritten)."""
|
|
if not path.exists():
|
|
log.warning("missing %s; skipping", path)
|
|
return 0, 0
|
|
tmp = path.with_suffix(".jsonl.purged.tmp")
|
|
n_total = n_rewritten = 0
|
|
with path.open("r", encoding="utf-8") as f, \
|
|
tmp.open("w", encoding="utf-8") as g:
|
|
for line in f:
|
|
line = line.rstrip("\n")
|
|
if not line:
|
|
continue
|
|
try:
|
|
rec = json.loads(line)
|
|
except json.JSONDecodeError:
|
|
continue
|
|
n_total += 1
|
|
er = rec.get("expectedResponse") or ""
|
|
seed = (rec.get("currentMessage") or {}).get("content", "") or ""
|
|
new_er, was_leaked = rewrite_expected_response(er, seed)
|
|
if was_leaked:
|
|
rec["expectedResponse"] = new_er
|
|
n_rewritten += 1
|
|
g.write(json.dumps(rec, ensure_ascii=False) + "\n")
|
|
os.replace(tmp, path)
|
|
return n_total, n_rewritten
|
|
|
|
|
|
def main() -> int:
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument("--final-dir", type=Path,
|
|
default=ROOT / "data" / "final")
|
|
args = ap.parse_args()
|
|
|
|
grand_total = grand_rewritten = 0
|
|
for split in ("train", "val", "test"):
|
|
path = args.final_dir / f"{split}.jsonl"
|
|
log.info("processing %s", path)
|
|
n, r = process_file(path)
|
|
grand_total += n
|
|
grand_rewritten += r
|
|
pct = 100 * r / max(1, n)
|
|
log.info(" %s: %d records, %d rewritten (%.2f%%)",
|
|
split, n, r, pct)
|
|
log.info("DONE: %d/%d records rewritten across all splits (%.2f%%)",
|
|
grand_rewritten, grand_total,
|
|
100 * grand_rewritten / max(1, grand_total))
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|