Files
yvgude--lean-ctx/bench/agent-task/swebench_harness/collect.py
T
wehub-resource-sync 26382a7ac6
CI / Clippy (push) Failing after 15m13s
CI / Test (ubuntu-latest) (push) Failing after 16m1s
CI / Test (macos-latest) (push) Has been cancelled
CI / Test (windows-latest) (push) Has been cancelled
CI / Build (no embeddings / no ORT) (push) Has been cancelled
CI / Format (push) Has been cancelled
CI / Cookbook (Node) (push) Has been cancelled
CI / Pi Extension (Node) (push) Has been cancelled
CI / Rust SDK (lean-ctx-client) (push) Has been cancelled
CI / Embed SDK (lean-ctx-sdk) (push) Has been cancelled
CI / Python SDK (leanctx) (push) Has been cancelled
CI / Hermes Plugin (Python) (push) Has been cancelled
CI / SDK Conformance Matrix (push) Has been cancelled
CI / Coverage (push) Has been cancelled
CI / cargo-deny (push) Has been cancelled
CI / Adversarial Safety (push) Has been cancelled
CI / Benchmarks (push) Has been cancelled
CI / Output-Quality Gate (eval A/B) (push) Has been cancelled
CI / Documentation (push) Has been cancelled
CI / CI Green (push) Has been cancelled
JetBrains Plugin / Actionlint (push) Has been cancelled
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (rust) (push) Has been cancelled
JetBrains Plugin / Validation (push) Has been cancelled
JetBrains Plugin / Build (push) Has been cancelled
JetBrains Plugin / Test (push) Has been cancelled
Security Check / Security Scan (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:35:30 +08:00

56 lines
1.7 KiB
Python

"""Assemble per-arm predictions.jsonl for the official SWE-bench evaluation.
Usage:
python -m swebench_harness.collect --run-id v1
Then evaluate each arm (docker required):
python -m swebench.harness.run_evaluation \
--dataset_name princeton-nlp/SWE-bench_Verified \
--predictions_path runs/v1/predictions-native.jsonl \
--run_id v1-native --max_workers 4
"""
from __future__ import annotations
import argparse
import json
from pathlib import Path
from . import ARMS, BENCH_ROOT, load_config, load_tasks_lock
def collect_arm(run_root: Path, arm: str, instances: list) -> Path:
out_path = run_root / f"predictions-{arm}.jsonl"
rows, missing = [], []
for inst in instances:
iid = inst["instance_id"]
patch_file = run_root / iid / arm / "model_patch.diff"
if not patch_file.exists():
missing.append(iid)
continue
rows.append({
"instance_id": iid,
"model_name_or_path": f"claude-code-{arm}",
"model_patch": patch_file.read_text(),
})
with out_path.open("w") as fh:
for row in rows:
fh.write(json.dumps(row) + "\n")
print(f"{arm}: {len(rows)} predictions -> {out_path}" + (f" (missing: {', '.join(missing)})" if missing else ""))
return out_path
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--run-id", required=True)
args = ap.parse_args()
cfg = load_config()
run_root = BENCH_ROOT / cfg["runs_dir"] / args.run_id
instances = load_tasks_lock()
for arm in ARMS:
collect_arm(run_root, arm, instances)
if __name__ == "__main__":
main()