Files
elizaos--eliza/packages/training/scripts/transform_purge_default_thoughts.py
wehub-resource-sync 426e9eeabd
Voice Workbench / headless workbench (mocked backends) (push) Has been cancelled
Voice Workbench / real acoustic lane (nightly, provisioned only) (push) Has been cancelled
ci / test (push) Has been cancelled
ci / lint-and-format (push) Has been cancelled
ci / build (push) Has been cancelled
ci / dev-startup (push) Has been cancelled
gitleaks / gitleaks (push) Has been cancelled
Markdown Links / Relative Markdown Links (push) Has been cancelled
Quality (Extended) / Homepage Build (PR smoke) (push) Has been cancelled
Quality (Extended) / Comment-only diff guard (push) Has been cancelled
Quality (Extended) / Format + Type Safety Ratchet (push) Has been cancelled
Quality (Extended) / Develop Gate (secret scan + UI determinism) (push) Has been cancelled
Quality (Extended) / Develop Gate (lint) (push) Has been cancelled
Chat shell gestures / Chat shell gesture + parity e2e (push) Has been cancelled
Cloud Gateway Discord / Test (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx @biomejs/biome check packages/lifeops-bench/src, benchmark-lint) (push) Has been cancelled
Benchmark Bridge Tests / benchmark (bunx vitest run --config packages/lifeops-bench/vitest.config.ts --root packages/lifeops-bench --passWithNoTests, benchmark-tests) (push) Has been cancelled
Build Agent Image / build-and-push (push) Has been cancelled
Dev Smoke / bun run dev onboarding chat (push) Has been cancelled
Dev Smoke / Vite HMR dependency-level smoke (push) Has been cancelled
Electrobun Submodule Guard / electrobun gitlink is fetchable (push) Has been cancelled
Publish @elizaos/example-code / check_npm (push) Has been cancelled
Publish @elizaos/example-code / publish_npm (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / verify_version (push) Has been cancelled
Publish @elizaos/plugin-elizacloud / publish_npm (push) Has been cancelled
Sandbox Live Smoke / Sandbox live smoke (push) Has been cancelled
Snap Build & Test / Build Snap (amd64) (push) Has been cancelled
Snap Build & Test / Build Snap (arm64) (push) Has been cancelled
Test Packaging / elizaos CLI global-install smoke (node + bun) (push) Has been cancelled
Cloud Gateway Webhook / Test (push) Has been cancelled
Cloud Tests / lint-and-types (push) Has been cancelled
Cloud Tests / unit-tests (push) Has been cancelled
Cloud Tests / integration-tests (push) Has been cancelled
Cloud Tests / e2e-tests (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Deploy Apps Worker (Product 2) / Determine environment (push) Has been cancelled
Deploy Apps Worker (Product 2) / Deploy apps worker to apps-control host (${{ needs.determine-env.outputs.environment }}) (push) Has been cancelled
Deploy Eliza Provisioning Worker / Determine environment (push) Has been cancelled
Deploy Eliza Provisioning Worker / Deploy worker to Hetzner host (${{ needs.determine-env.outputs.environment }} @ ${{ needs.determine-env.outputs.deployment_sha }}) (push) Has been cancelled
Dev Smoke / Classify changed paths (push) Has been cancelled
supply-chain / sbom (push) Has been cancelled
supply-chain / vulnerability-scan (push) Has been cancelled
Build, Push & Deploy to Phala Cloud / build-and-push (push) Has been cancelled
Test Packaging / Validate Packaging Configs (push) Has been cancelled
Test Packaging / Build & Test PyPI Package (push) Has been cancelled
Test Packaging / PyPI on Python ${{ matrix.python }} (push) Has been cancelled
Test Packaging / Pack & Test JS Tarballs (push) Has been cancelled
UI Fixture E2E / ui-fixture-e2e (push) Has been cancelled
UI Fixture E2E / fixture-e2e (push) Has been cancelled
UI Story Gate / story-gate (push) Has been cancelled
vault-ci / test (macos-latest) (push) Has been cancelled
vault-ci / test (ubuntu-latest) (push) Has been cancelled
vault-ci / test (windows-latest) (push) Has been cancelled
vault-ci / app-core wiring tests (push) Has been cancelled
verify-patches / verify patches/CHECKSUMS.sha256 (push) Has been cancelled
Voice Benchmark Smoke / voice-emotion fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voiceagentbench fixture smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench-quality unit smoke (push) Has been cancelled
Voice Benchmark Smoke / voicebench TypeScript unit (no audio) (push) Has been cancelled
Voice Benchmark Smoke / voice bench smoke summary (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/app-core test bun run --cwd packages/elizaos test bun run --cwd packages/cloud/shared test], app-and-cli) (push) Has been cancelled
Windows CI / windows ([bun run --cwd packages/scenario-runner test bun run --cwd packages/vault test bun run --cwd packages/security test bun run --cwd plugins/plugin-coding-tools test], framework-packages) (push) Has been cancelled
Windows CI / windows ([bun run --cwd plugins/plugin-elizacloud test bun run --cwd plugins/plugin-discord test bun run --cwd plugins/plugin-anthropic test bun run --cwd plugins/plugin-openai test bun run --cwd plugins/plugin-app-control test bun run --cwd plugins/pl… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run build --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/agent --concurrency=4 node packages/scripts/run-bash-linux-only.mjs scripts/verify-riscv64-buildpaths.sh node packages/scripts/run… (push) Has been cancelled
Windows CI / windows ([node packages/scripts/run-turbo.mjs run typecheck --filter=@elizaos/core --filter=@elizaos/shared --filter=@elizaos/cloud-shared --concurrency=4 bun run --cwd packages/core test bun run --cwd packages/shared test], core-runtime, 75) (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:43:05 +08:00

129 lines
4.5 KiB
Python

"""Purge the literal default-thought strings from packed train.jsonl.
The dataset adapters were updated to use a varied thought-phrasing pool, but
existing `data/normalized/*.jsonl` files were generated before the fix and
still contain literal strings like `"Reply to the user."` injected as the
model's reasoning. Re-normalizing 75 GB to fix this is wasteful when we can
post-process the packed train.jsonl directly.
This script walks train.jsonl line by line. For each record where
`expectedResponse` starts with `thought: <literal default>`, replaces just
that line with a varied phrasing picked deterministically from the
adapter's pool (seeded by the user message).
Output: rewrites train.jsonl (and val.jsonl) in place via atomic temp file.
Usage:
.venv/bin/python scripts/transform_purge_default_thoughts.py
"""
from __future__ import annotations
import argparse
import json
import logging
import os
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(ROOT / "scripts"))
from lib.adapters import ( # noqa: E402
_picked_thought,
_REPLY_THOUGHT_POOL,
_TOOL_THOUGHT_POOL,
_SHELL_THOUGHT_POOL,
_IGNORE_THOUGHT_POOL,
_AGENT_TRACE_THOUGHT_POOL,
)
from lib.eliza_record import DEFAULT_THOUGHT_LEAKS # noqa: E402
assert "Reply to the user." in DEFAULT_THOUGHT_LEAKS, (
"DEFAULT_THOUGHT_LEAKS lost its canonical entry — see lib/eliza_record.py"
)
logging.basicConfig(level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(message)s")
log = logging.getLogger("purge")
# Map from old literal → pool to use for replacement
_LEAK_PATTERNS: list[tuple[str, tuple[str, ...]]] = [
("thought: Reply to the user.", _REPLY_THOUGHT_POOL),
("thought: Call the tool to satisfy the request.", _TOOL_THOUGHT_POOL),
("thought: Run the command to satisfy the request.", _SHELL_THOUGHT_POOL),
('thought: "This message is not for me — ignore."', _IGNORE_THOUGHT_POOL),
("thought: This message is not for me — ignore.", _IGNORE_THOUGHT_POOL),
("thought: Continue the agent task.", _AGENT_TRACE_THOUGHT_POOL),
]
def rewrite_expected_response(er: str, seed: str) -> tuple[str, bool]:
"""Return (rewritten, was_leaked). Replaces the first thought line if
it's a known literal default."""
if not er:
return er, False
lines = er.split("\n", 1)
first = lines[0]
rest = lines[1] if len(lines) > 1 else ""
for prefix, pool in _LEAK_PATTERNS:
if first.startswith(prefix):
varied = _picked_thought(pool, seed)
new_first = f"thought: {varied}"
return f"{new_first}\n{rest}" if rest else new_first, True
return er, False
def process_file(path: Path) -> tuple[int, int]:
"""Returns (records, rewritten)."""
if not path.exists():
log.warning("missing %s; skipping", path)
return 0, 0
tmp = path.with_suffix(".jsonl.purged.tmp")
n_total = n_rewritten = 0
with path.open("r", encoding="utf-8") as f, \
tmp.open("w", encoding="utf-8") as g:
for line in f:
line = line.rstrip("\n")
if not line:
continue
try:
rec = json.loads(line)
except json.JSONDecodeError:
continue
n_total += 1
er = rec.get("expectedResponse") or ""
seed = (rec.get("currentMessage") or {}).get("content", "") or ""
new_er, was_leaked = rewrite_expected_response(er, seed)
if was_leaked:
rec["expectedResponse"] = new_er
n_rewritten += 1
g.write(json.dumps(rec, ensure_ascii=False) + "\n")
os.replace(tmp, path)
return n_total, n_rewritten
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--final-dir", type=Path,
default=ROOT / "data" / "final")
args = ap.parse_args()
grand_total = grand_rewritten = 0
for split in ("train", "val", "test"):
path = args.final_dir / f"{split}.jsonl"
log.info("processing %s", path)
n, r = process_file(path)
grand_total += n
grand_rewritten += r
pct = 100 * r / max(1, n)
log.info(" %s: %d records, %d rewritten (%.2f%%)",
split, n, r, pct)
log.info("DONE: %d/%d records rewritten across all splits (%.2f%%)",
grand_rewritten, grand_total,
100 * grand_rewritten / max(1, grand_total))
return 0
if __name__ == "__main__":
sys.exit(main())