Files
wehub-resource-sync 23f7624596
ADR-166 MCP Bridge Security Lock / Static-source security lock (push) Failing after 0s
ADR-166 MCP Bridge Security Lock / Compose default binds loopback + Mongo has auth (push) Failing after 2s
CodeQL Advanced / Analyze (rust) (push) Failing after 0s
ADR-166 MCP Bridge Security Lock / plugin-agent-federation bindHost default (push) Failing after 1s
ADR-166 MCP Bridge Security Lock / Runtime behavior — 401 + terminal gate + fail-closed (push) Failing after 4s
business-pods-smoke / smoke (push) Failing after 1s
all-plugins-smoke / smoke-all (push) Failing after 2s
CI/CD Pipeline / Security & Code Quality (push) Failing after 1s
CI/CD Pipeline / Test Suite (ubuntu-latest) (push) Failing after 1s
CI/CD Pipeline / Build & Package (macos-latest) (push) Has been skipped
CI/CD Pipeline / Build & Package (ubuntu-latest) (push) Has been skipped
CI/CD Pipeline / Build & Package (windows-latest) (push) Has been skipped
CI/CD Pipeline / Documentation & Examples (push) Failing after 1s
Clone Tracker (14-day rolling) / Snapshot clones for ruflo ecosystem (push) Failing after 1s
CodeQL Advanced / Analyze (actions) (push) Failing after 1s
CodeQL Advanced / Analyze (javascript-typescript) (push) Failing after 1s
federation-peer-rust / stable-noop (push) Failing after 1s
metaharness-ci / score (push) Failing after 1s
metaharness-ci / router-compat (push) Failing after 0s
metaharness-ci / similarity-tests (push) Failing after 0s
no-agentbbs-smoke / smoke-without-agentbbs (push) Failing after 1s
V3 CI/CD Pipeline / Build V3 (windows-latest) (push) Has been skipped
codex-integration-audit / Codex integration audit (push) Failing after 1s
helpers-manifest-guard / guard (push) Failing after 1s
🔗 Cross-Agent Integration Tests / 🤝 Agent Coordination Tests (push) Has been skipped
🔗 Cross-Agent Integration Tests / 🧠 Memory Sharing Integration (push) Has been skipped
🔗 Cross-Agent Integration Tests / 🛡️ Fault Tolerance Tests (push) Has been skipped
🔗 Cross-Agent Integration Tests / ⚡ Performance Integration Tests (push) Has been skipped
metaharness-ci / mcp-scan (push) Failing after 1s
metaharness-ci / eject-dryrun (push) Failing after 1s
metaharness-ci / metaharness-real-data (push) Failing after 0s
no-cli-optdep-bloat-2561 / guard (push) Failing after 1s
no-metaharness-smoke / smoke-without-metaharness (push) Failing after 1s
no-phantom-agentic-flow-subpath / guard (push) Failing after 1s
🔄 Automated Rollback Manager / 🚨 Failure Detection (push) Failing after 1s
V3 CI/CD Pipeline / Plugin hooks smoke / ubuntu-latest / Node 22 (push) Failing after 1s
V3 CI/CD Pipeline / ruflo-graph-intelligence build + test smoke (#2044, ADR-123) (push) Failing after 1s
CVE Audit Gate / Audit root (critical-blocking) (push) Failing after 2s
cost-tracker-smoke / smoke (push) Failing after 3s
oia-audit-weekly / audit (push) Failing after 2s
ruflo-agent-smoke / ruflo-agent structural smoke (push) Failing after 1s
📊 Status Badges Update / 📊 Update Status Badges (push) Failing after 1s
V3 CI/CD Pipeline / Static regression guards (#2267 YAML + (push) Failing after 1s
V3 CI/CD Pipeline / Test V3 Packages (push) Failing after 0s
V3 CI/CD Pipeline / agent_execute provider routing smoke (#2042) (push) Failing after 0s
CVE Audit Gate / Audit v3 (critical-blocking) (push) Failing after 1s
federation-peer-rust / stable-native (push) Failing after 2s
🔗 Cross-Agent Integration Tests / 🚀 Integration Test Setup (push) Failing after 2s
neural-trader-smoke / runtime-smoke (push) Failing after 1s
V3 CI/CD Pipeline / Build V3 (macos-latest) (push) Has been skipped
V3 CI/CD Pipeline / Build V3 (ubuntu-latest) (push) Has been skipped
V3 CI/CD Pipeline / Type Check V3 (push) Failing after 1s
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / ubuntu-latest / Node 24 (push) Failing after 1s
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / ubuntu-latest / Node 22 (push) Failing after 2s
V3 CI/CD Pipeline / browser rvf create flag smoke (#2015) (push) Failing after 0s
V3 CI/CD Pipeline / Dependency review (#2046) (push) Has been skipped
V3 CI/CD Pipeline / Supply-chain audit (#2046) (push) Failing after 0s
V3 CI/CD Pipeline / witness marker drift smoke (#2021) (push) Failing after 1s
V3 CI/CD Pipeline / neural-trader portfolio CG smoke (#2068, ADR-126 Phase 3) (push) Failing after 1s
V3 CI/CD Pipeline / neural-trader backtest signing smoke (#2068, ADR-126 Phase 4) (push) Failing after 1s
V3 CI/CD Pipeline / kg-extract type-import classification smoke (#2049) (push) Failing after 0s
V3 CI/CD Pipeline / witness verify precondition smoke (#1880) (push) Failing after 2s
V3 CI/CD Pipeline / neural-trader pipeline risk-gate smoke (#2068, ADR-126 Phase 5) (push) Failing after 0s
V3 CI/CD Pipeline / neural-trader feature attribution smoke (#2068, ADR-126 Phase 6) (push) Failing after 0s
V3 CI/CD Pipeline / plugin-registry signature verification smoke (#1922, CWE-347) (push) Failing after 4s
V3 CI/CD Pipeline / memory stats legacy-DB smoke (#2120) (push) Failing after 4s
V3 CI/CD Pipeline / github deprecated actions smoke (#2089, ADR-127 Phase 3) (push) Failing after 1s
V3 CI/CD Pipeline / graph query + pathfinder smoke (ADR-130 P2+P5) (push) Has been skipped
V3 CI/CD Pipeline / graph trajectory hooks smoke (ADR-130 P3) (push) Has been skipped
V3 CI/CD Pipeline / graph plugin adapter smoke (ADR-130 P4) (push) Has been skipped
V3 CI/CD Pipeline / graph benchmark (ADR-130 P6) (push) Has been skipped
V3 CI/CD Pipeline / statusline generator delegation smoke (#2195) (push) Failing after 1s
V3 CI/CD Pipeline / wizard init regression guard (#2206 (push) Failing after 1s
V3 CI/CD Pipeline / memory no-stray-db smoke (ADR-125 P7) (push) Failing after 1s
V3 CI/CD Pipeline / github-safe injection smoke (#2089, ADR-127 Phase 1) (push) Failing after 1s
V3 CI/CD Pipeline / github actions pin smoke (#2089, ADR-127 Phase 1) (push) Failing after 1s
V3 CI/CD Pipeline / github attribution opt-in smoke (#2089, ADR-127 Phase 4) (push) Failing after 1s
V3 CI/CD Pipeline / pre-bash hook safety smoke (#2017) (push) Failing after 1s
V3 CI/CD Pipeline / Memory import smoke / ubuntu-latest (push) Failing after 0s
V3 CI/CD Pipeline / MCP protocol smoke / ubuntu-latest (push) Failing after 2s
V3 CI/CD Pipeline / ruvllm WASM auto-init smoke (#2086) (push) Failing after 4s
V3 CI/CD Pipeline / MCP paired-tool round-trip smoke (#1889) (push) Failing after 1s
V3 CI/CD Pipeline / Plugin package install-safety (#1902/#1903/#1904) (push) Failing after 1s
V3 CI/CD Pipeline / Tool description discoverability (ADR-112) (push) Failing after 3s
V3 CI/CD Pipeline / CLI npx-install smoke (#1147 / (22) (push) Failing after 1s
V3 CI/CD Pipeline / CLI npx-install smoke (#1147 / (24) (push) Failing after 1s
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / ubuntu-latest (push) Failing after 2s
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / ubuntu-latest (push) Failing after 1s
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / ubuntu-latest (push) Failing after 1s
V3 CI/CD Pipeline / Vector-index dimension audit (#1947) (push) Failing after 0s
V3 CI/CD Pipeline / Hook-command install safety (#1921) (push) Failing after 1s
V3 CI/CD Pipeline / ToolOutputGuardrail smoke (ADR-131, (push) Failing after 1s
V3 CI/CD Pipeline / init-bundle invariants smoke (#2095, ADR-128 Phase 5) (push) Failing after 1s
V3 CI/CD Pipeline / wasm provider bridge smoke (ADR-129 P1) (push) Failing after 2s
V3 CI/CD Pipeline / wasm gallery CRUD smoke (ADR-129 P3) (push) Failing after 1s
V3 CI/CD Pipeline / wasm plugin bridge smoke (ADR-129 P4) (push) Failing after 0s
V3 CI/CD Pipeline / wasm compose smoke (ADR-129 P2) (push) Failing after 4s
V3 CI/CD Pipeline / graph schema smoke (ADR-130 P1) (push) Failing after 0s
Validate Marketplace / validate (push) Failing after 1s
🔍 Verification Pipeline / 🚀 Setup Verification (push) Failing after 1s
🔍 Verification Pipeline / 🛡️ Security Verification (push) Has been skipped
🔍 Verification Pipeline / 📝 Code Quality (push) Has been skipped
🔍 Verification Pipeline / 🧪 Test Verification (${{ matrix.os }}, Node ${{ matrix.node }}) (push) Has been skipped
🔍 Verification Pipeline / 🏗️ Build Verification (push) Has been skipped
🔍 Verification Pipeline / 📚 Documentation Verification (push) Has been skipped
CVE Audit Gate / High-severity report (warn only) (push) Has been cancelled
🔄 Automated Rollback Manager / 🔄 Execute Rollback (push) Has been cancelled
🔄 Automated Rollback Manager / ✅ Post-Rollback Verification (push) Has been cancelled
🔄 Automated Rollback Manager / 📊 Rollback Monitoring (push) Has been cancelled
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / windows-latest (push) Has been cancelled
🔄 Automated Rollback Manager / ⏳ Manual Rollback Approval (push) Has been cancelled
V3 CI/CD Pipeline / MCP protocol smoke / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Memory import smoke / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / ubuntu-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Publish to npm (alpha) (push) Has been cancelled
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / macos-latest / Node 22 (push) Has been cancelled
V3 CI/CD Pipeline / Plugin hooks smoke / macos-latest / Node 22 (push) Has been cancelled
CI/CD Pipeline / Deploy & Release (push) Has been cancelled
CI/CD Pipeline / CI Status (push) Has been cancelled
🔗 Cross-Agent Integration Tests / 📊 Integration Test Report (push) Has been cancelled
🔄 Automated Rollback Manager / 🔍 Pre-Rollback Validation (push) Has been cancelled
🔍 Verification Pipeline / ⚡ Performance Verification (push) Has been cancelled
🔍 Verification Pipeline / 📊 Verification Report (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:02:19 +08:00

170 lines
7.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Capabilities Historical Reference
Per-OS performance baseline + capability status. Tracks how key user-visible operations evolve across releases. Each measurement appends one line to `verification/<os>/performance.jsonl` so regressions are catchable the same way `manifest.md.json` catches a documented fix disappearing.
> See [README.md](README.md) for the witness manifest layer (presence). This doc covers the **performance** layer (speed) — they're complementary.
---
## What's tracked
Each entry in `verification/<os>/performance.jsonl` records one capability×measurement at one git commit:
```jsonc
{
"v": 1,
"commit": "<full sha>",
"issuedAt": "2026-05-09T15:00:00.000Z",
"os": "macos",
"capability": "install_pack",
"durationMs": 373,
"baselineMs": 410, // present when --baseline flag set
"deltaPct": -9 // negative = faster than rolling median
}
```
## Capabilities
| Capability | What it measures | Why it matters |
|---|---|---|
| `install_pack` | Time for `pnpm pack @claude-flow/memory` | Catch regressions in package size / pack-time pipeline |
| `install_no_optional` | `npm install <tarball> --omit=optional` end-to-end | The user-visible "fresh install on a platform without prebuilds" — this is what was 152s on Node 26 before #1867 fix; now ~5s on a clean dir |
| `memory_load` | Cold `import('@claude-flow/memory')` in a fresh node process | Catches accidentally-eager imports of heavy native modules |
| `memory_round_trip` | `createDatabase(auto) → store → get → shutdown` | End-to-end runtime behaviour of the auto-fallback path |
| `witness_verify` | `verify.mjs --manifest <os>/manifest.md.json` | The witness verification itself — should stay sub-second even at 100+ fixes |
Add capabilities by extending the `runners` map in `plugins/ruflo-core/scripts/witness/perf.mjs`. The framework supports any synchronous benchmark that throws on failure.
---
## Reference baselines (macOS, M-class hardware, Node 22.22.1, post-warmup)
Recorded 2026-05-09 against commit `5372f83`. Treat as "should not regress beyond ~3×" — anything larger is signal.
| Capability | Median ms | P95 ms | Notes |
|---|---:|---:|---|
| `install_pack` | 370 | 450 | pnpm-pack pipeline; rewrites workspace:* → resolved versions |
| `install_no_optional` | ~5,000 | 8,000 | Network-bound (npm registry); flaps with cache state |
| `memory_load` | 18 | 35 | Cold module load; sub-50ms = no eager native imports |
| `memory_round_trip` | ~80 | 120 | Backend selection + RVF fallback + open + write + read + close |
| `witness_verify` | 53 | 90 | 82 markers × file read + sha256; @noble/ed25519 sig verify |
Linux + Windows baselines populate as CI runs the perf job on those runners. Median across the rolling-5 window is the comparison baseline; a single slow run doesn't trigger a regression.
---
## Historical reference for key incidents
### #1867 — Node 26 install failure (2026-05-08)
| Phase | install_no_optional (median ms) | Notes |
|---|---:|---|
| Pre-fix (3.7.0-alpha.17) | **fails** | `node-gyp` cannot rebuild `better-sqlite3@^11` on Node 26; install never completes |
| Post-fix (3.7.0-alpha.18+) | ~5,000 | `better-sqlite3` moved to `optionalDependencies`; `--omit=optional` makes it skipable; runtime falls back to RVF/sql.js |
Captured in `verification.md.json` fix `#1867` (marker: `(await import('better-sqlite3')).default` — guards against re-introduction of a static import).
### #1859 + #1862 — Plugin/CLI flag drift (2026-05-08)
| Phase | hooks/post-edit handler | Result |
|---|---|---|
| Pre-fix | `cat | jq | tr | xargs -0 -I {} npx ... post-edit --file '{}' --format true` | `[ERROR] Invalid value for --format: true` on every Edit/Write |
| Post-fix | `bash -c '...; npx ... post-edit -f "$FILE" -s true'` | Records correct file path |
Captured in `verification.md.json` fixes `#1862` (marker: `hooks post-edit -f \"$FILE\" -s true`) and `#1859` (CLI parser swap, marker: `ctx.flags.file || ctx.args[0]`).
### #1608 — bcrypt → bcryptjs migration (PR #1818)
| Phase | dependencies | Notes |
|---|---|---|
| Pre-migration | `bcrypt@6.0.0` (native, brings tar CVE chain) | 6 HIGH CVEs in transitive `tar` |
| Post-migration | `bcryptjs@^3.0.3` (pure-JS) | No native dep, no tar; same `$2a$` hash compatibility |
Captured in `verification.md.json` fix `#1608` (marker: `bcryptjs`). Briefly regressed in early sessions (dist not rebuilt against migrated source); witness-verify caught it as `markerVerified: false` and a rebuild restored it to `pass`.
### Memory backend fallback chain (ADR-009)
Auto-selection priority (highest first, falls through on failure):
1. **RVF** — pure-TS HNSW; always available. Default in CI.
2. **better-sqlite3** — native SQLite; fastest. Available when prebuild is fetched.
3. **sql.js** — WASM SQLite. Pure-JS fallback for restricted environments.
4. **JSON** — last-ditch flat file. Never used in practice.
Verified by `memory_round_trip` capability — the round-trip succeeds on whichever backend was selected, so a regression in fallback selection shows as a runtime error on platforms where the preferred backend is unavailable.
---
## Daily workflow
```bash
# Run all benchmarks now and append to verification/<os>/performance.jsonl
node plugins/ruflo-core/scripts/witness/perf.mjs
# Run with baseline comparison (median of last 5 entries per capability)
node plugins/ruflo-core/scripts/witness/perf.mjs --baseline
# Run a subset
node plugins/ruflo-core/scripts/witness/perf.mjs \
--capabilities install_pack,memory_load \
--json
```
For CI, gate on regressions exceeding a threshold:
```yaml
- name: Performance verification
run: |
node plugins/ruflo-core/scripts/witness/perf.mjs --baseline --json > /tmp/perf.json
node -e "
const r = require('/tmp/perf.json');
const regressed = r.results.filter(x => x.deltaPct != null && x.deltaPct > 200);
if (regressed.length) {
console.error('regressions (>200% slower than baseline):');
for (const x of regressed) console.error(\` \${x.capability}: \${x.durationMs}ms vs \${x.baselineMs}ms baseline\`);
process.exit(1);
}
"
```
---
## What's not tracked yet (and why)
- **HNSW search latency** — depends on dataset size; needs a fixture, follow-up.
- **CLI startup time** — `ruflo --version` is the obvious metric, but currently dominated by node startup + module graph; not stable enough as a regression signal until the cli-core split (PR #1764) lands.
- **Memory growth over long-running processes** — needs an instrumented harness; out of scope for snapshot-style verification.
---
## Schema
### `verification/<os>/performance.jsonl` (one entry per line)
```jsonc
{
"v": 1, // schema version
"commit": "<full sha>",
"issuedAt": "<ISO timestamp>",
"os": "linux" | "macos" | "windows",
"capability": "<runner name>",
"durationMs": 373, // null if measurement errored
"error": "...", // present iff measurement errored
"metadata": { /* free-form per-capability */ },
"baselineMs": 410, // optional; present when --baseline used
"deltaPct": -9 // optional; (durationMs - baselineMs) / baselineMs * 100
}
```
The file is append-only and OS-specific. Cross-OS comparison happens by reading the three files and joining on capability — different OSes have different native code paths, so absolute numbers don't compare directly, but **trends do**.
---
## References
- [README.md](README.md) — the witness manifest layer (fix presence)
- [witness-fixes.json](witness-fixes.json) — fix list (input to manifest regen)
- [results.md](results.md) — last verification run report
- [`plugins/ruflo-core/scripts/witness/perf.mjs`](../plugins/ruflo-core/scripts/witness/perf.mjs) — benchmark runner
- [ADR-103](../v3/docs/adr/ADR-103-witness-temporal-history.md) — temporal history pattern (presence) that perf.mjs mirrors for measurements