意图画布
intent confidence 100/100; Intent is clear enough to package the first routeable version.
diff --git a/references/review-studio-method.md b/references/review-studio-method.md index da67a4c..2c52a40 100644 --- a/references/review-studio-method.md +++ b/references/review-studio-method.md @@ -40,7 +40,7 @@ Production, library, and governed reviews should also show a blind A/B review pa When `reports/output_execution_runs.json` exists, Review Studio should show the number of variant runs, command-executed runs, model-executed runs, recorded fixtures, timing-observed runs, and token-estimated runs. Recorded fixtures are valid reproducibility evidence, but they must not be described as model-executed output evidence. -When `reports/output_review_adjudication.json` exists, Review Studio should show reviewed pairs and pending pairs. Pending reviewer decisions are acceptable as an explicit state, but they must not be counted as agreement or human review evidence. Invalid adjudication records should block release because they make the blind review audit untrustworthy. +When `reports/output_review_adjudication.json` exists, Review Studio should show reviewed pairs and pending pairs. Pending reviewer decisions are acceptable as an explicit state, but they must not be counted as agreement or human review evidence. For production, library, and governed packages, pending reviewer decisions should keep the Output Lab in `warn` until reviewer decisions are recorded or the warning is explicitly accepted in the waiver ledger. Invalid adjudication records should block release because they make the blind review audit untrustworthy. The Operations Loop must never display raw telemetry logs. It should link only to `reports/adoption_drift_report.md`; privacy or schema violations are blockers. diff --git a/registry/index.json b/registry/index.json index fef4b96..661ecf0 100644 --- a/registry/index.json +++ b/registry/index.json @@ -15,7 +15,7 @@ "agent-skills-compatible" ], "package_metadata": "registry/packages/yao-meta-skill.json", - "package_sha256": "d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5" + "package_sha256": "af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068" } ] } diff --git a/registry/packages/yao-meta-skill.json b/registry/packages/yao-meta-skill.json index 622bec3..d4c2f84 100644 --- a/registry/packages/yao-meta-skill.json +++ b/registry/packages/yao-meta-skill.json @@ -15,8 +15,8 @@ "trust_level": "local", "license": "MIT", "checksums": { - "package_sha256": "d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5", - "archive_sha256": "6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451" + "package_sha256": "af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068", + "archive_sha256": "9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b" }, "compatibility": { "openai": "pass", @@ -47,7 +47,7 @@ }, "distribution": { "archive_verified": true, - "archive_sha256": "6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451", + "archive_sha256": "9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b", "package_verification": "reports/package_verification.json", "install_simulated": true, "install_simulation": "reports/install_simulation.json" diff --git a/reports/adoption_drift_report.json b/reports/adoption_drift_report.json index 453be36..d88f91a 100644 --- a/reports/adoption_drift_report.json +++ b/reports/adoption_drift_report.json @@ -1,7 +1,7 @@ { "ok": true, "schema_version": "2.0", - "generated_at": "2026-06-13T11:16:31Z", + "generated_at": "2026-06-13T11:23:46Z", "skill_dir": ".", "privacy_contract": { "storage": "local-first", diff --git a/reports/context_budget.json b/reports/context_budget.json index 3639880..9b5216c 100644 --- a/reports/context_budget.json +++ b/reports/context_budget.json @@ -6,9 +6,9 @@ "context_budget_tier": "production", "context_budget_limit": 1000, "skill_body_tokens": 811, - "other_text_tokens": 871624, + "other_text_tokens": 872858, "estimated_initial_load_tokens": 987, - "estimated_total_text_tokens": 872435, + "estimated_total_text_tokens": 873669, "relevant_file_count": 368, "unused_resource_dirs": [], "quality_signal_points": 130, diff --git a/reports/output_execution_runs.json b/reports/output_execution_runs.json index 153ce5d..ca2c305 100644 --- a/reports/output_execution_runs.json +++ b/reports/output_execution_runs.json @@ -34,7 +34,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 32.51, + "duration_ms": 26.38, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -62,7 +62,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 32.64, + "duration_ms": 27.27, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -85,7 +85,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 38.43, + "duration_ms": 26.19, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -113,7 +113,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 31.51, + "duration_ms": 26.45, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -136,7 +136,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 31.34, + "duration_ms": 26.16, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -164,7 +164,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 30.94, + "duration_ms": 26.14, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -187,7 +187,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 30.5, + "duration_ms": 27.24, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -214,7 +214,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 31.42, + "duration_ms": 28.91, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -237,7 +237,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 31.18, + "duration_ms": 27.35, "provider": "local-output-eval-runner", "model": "", "usage": { @@ -266,7 +266,7 @@ "execution_mode": "command", "model_executed": false, "command_executed": true, - "duration_ms": 30.97, + "duration_ms": 26.7, "provider": "local-output-eval-runner", "model": "", "usage": { diff --git a/reports/output_execution_runs.md b/reports/output_execution_runs.md index 58e706c..2bb6b47 100644 --- a/reports/output_execution_runs.md +++ b/reports/output_execution_runs.md @@ -23,16 +23,16 @@ Command runner evidence is present. This proves the eval harness executed an ext | Case | Variant | Mode | Model | Duration ms | Tokens | Score | Status | | --- | --- | --- | --- | ---: | ---: | ---: | --- | -| skill-package-contract | baseline | command | local-output-eval-runner | 32.51 | 33 | 0.0 | pass | -| skill-package-contract | with_skill | command | local-output-eval-runner | 32.64 | 73 | 100.0 | pass | -| output-eval-expectation | baseline | command | local-output-eval-runner | 38.43 | 36 | 0.0 | pass | -| output-eval-expectation | with_skill | command | local-output-eval-runner | 31.51 | 80 | 100.0 | pass | -| ir-before-packaging | baseline | command | local-output-eval-runner | 31.34 | 33 | 0.0 | pass | -| ir-before-packaging | with_skill | command | local-output-eval-runner | 30.94 | 80 | 100.0 | pass | -| near-neighbor-boundary | baseline | command | local-output-eval-runner | 30.5 | 36 | 0.0 | pass | -| near-neighbor-boundary | with_skill | command | local-output-eval-runner | 31.42 | 65 | 100.0 | pass | -| file-backed-governed-package | baseline | command | local-output-eval-runner | 31.18 | 37 | 0.0 | pass | -| file-backed-governed-package | with_skill | command | local-output-eval-runner | 30.97 | 98 | 100.0 | pass | +| skill-package-contract | baseline | command | local-output-eval-runner | 26.38 | 33 | 0.0 | pass | +| skill-package-contract | with_skill | command | local-output-eval-runner | 27.27 | 73 | 100.0 | pass | +| output-eval-expectation | baseline | command | local-output-eval-runner | 26.19 | 36 | 0.0 | pass | +| output-eval-expectation | with_skill | command | local-output-eval-runner | 26.45 | 80 | 100.0 | pass | +| ir-before-packaging | baseline | command | local-output-eval-runner | 26.16 | 33 | 0.0 | pass | +| ir-before-packaging | with_skill | command | local-output-eval-runner | 26.14 | 80 | 100.0 | pass | +| near-neighbor-boundary | baseline | command | local-output-eval-runner | 27.24 | 36 | 0.0 | pass | +| near-neighbor-boundary | with_skill | command | local-output-eval-runner | 28.91 | 65 | 100.0 | pass | +| file-backed-governed-package | baseline | command | local-output-eval-runner | 27.35 | 37 | 0.0 | pass | +| file-backed-governed-package | with_skill | command | local-output-eval-runner | 26.7 | 98 | 100.0 | pass | ## Next Fixes diff --git a/reports/package_verification.json b/reports/package_verification.json index 56c45c8..e9e6abf 100644 --- a/reports/package_verification.json +++ b/reports/package_verification.json @@ -8,7 +8,7 @@ "target_count": 3, "adapter_count": 3, "archive_present": true, - "archive_sha256": "6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451", + "archive_sha256": "9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b", "archive_entry_count": 490, "failure_count": 0, "warning_count": 0 diff --git a/reports/package_verification.md b/reports/package_verification.md index caa1f02..a83053c 100644 --- a/reports/package_verification.md +++ b/reports/package_verification.md @@ -4,7 +4,7 @@ - Package directory: `dist` - Targets: `3 / 3` adapters present - Archive present: `True` -- Archive SHA256: `6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451` +- Archive SHA256: `9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b` - Failures: `0` - Warnings: `0` diff --git a/reports/registry_audit.json b/reports/registry_audit.json index 608ace9..69c7d2c 100644 --- a/reports/registry_audit.json +++ b/reports/registry_audit.json @@ -20,8 +20,8 @@ "trust_level": "local", "license": "MIT", "checksums": { - "package_sha256": "d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5", - "archive_sha256": "6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451" + "package_sha256": "af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068", + "archive_sha256": "9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b" }, "compatibility": { "openai": "pass", @@ -52,7 +52,7 @@ }, "distribution": { "archive_verified": true, - "archive_sha256": "6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451", + "archive_sha256": "9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b", "package_verification": "reports/package_verification.json", "install_simulated": true, "install_simulation": "reports/install_simulation.json" @@ -76,7 +76,7 @@ "agent-skills-compatible" ], "package_metadata": "registry/packages/yao-meta-skill.json", - "package_sha256": "d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5" + "package_sha256": "af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068" } ] }, diff --git a/reports/registry_audit.md b/reports/registry_audit.md index 8f86738..d670717 100644 --- a/reports/registry_audit.md +++ b/reports/registry_audit.md @@ -6,8 +6,8 @@ - Maturity: `governed` - Owner: `Yao Team` - License: `MIT` -- Package SHA256: `d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5` -- Archive SHA256: `6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d451` +- Package SHA256: `af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068` +- Archive SHA256: `9208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b` - Install simulated: `True` ## Compatibility diff --git a/reports/review-studio.html b/reports/review-studio.html index 9bfc5df..b3d632a 100644 --- a/reports/review-studio.html +++ b/reports/review-studio.html @@ -235,8 +235,8 @@
Create, refactor, evaluate, and package agent skills from workflows, prompts, transcripts, docs, or notes. Use when asked to create a skill, turn a repeated process into a reusable skill, improve an existing skill, add evals, or package a skill for team reuse.
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 0; recorded 0; reviewed 0/5
initial load 987/1000; quality density 131.7
5 / 5 targets pass
0 secrets; 68 scripts; 2 network-capable scripts; 0 help smoke failures
3/3 permissions approved; gaps 0; required file_write, network, subprocess
3/3 targets probed; native 0; metadata fallback 3; residual risks 3
12 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 24 scoped non-actionable issues
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low
0 active waivers cover current warnings
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 3 adapters
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 0; recorded 0; reviewed 0/5; review pending 5
initial load 987/1000; quality density 131.7
5 / 5 targets pass
0 secrets; 68 scripts; 2 network-capable scripts; 0 help smoke failures
3/3 permissions approved; gaps 0; required file_write, network, subprocess
3/3 targets probed; native 0; metadata fallback 3; residual risks 3
12 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 24 scoped non-actionable issues
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low
0 active waivers; 1 warning gates still need reviewer decision
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 3 adapters
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
无。
无。
当前没有 blocker 或 warning。保持现有证据链即可。
补足 output eval 覆盖、execution evidence、blind A/B 和 reviewer adjudication。
没有输出质量和人工盲评证据时,Skill 只能证明会触发,不能证明输出真的更好且经得起审查。python3 scripts/run_output_execution.py对保留的 warning 写入 reviewer、理由、范围和到期时间,或修掉 warning。
warning 可以被接受,但必须可审计、会过期,并且不能掩盖 blocker。python3 scripts/render_review_waivers.py .d53de57f6593fca3dd937832f5923bd5381250e382631a6b6616e2405de748b5af6eab9b547a289ed5035eed9a9dd54cfdd01459acf1e2ba57d8ed3e2fbab068高风险 secret、远程 inline execution、缺失依赖策略或无法解释的脚本接口应阻断 governed release。
0 active waivers cover current warnings
0 active waivers; 1 warning gates still need reviewer decision
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 3 adapters
6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d4519208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
6972bb1d72746b6a16c7dbf1ec73f438252756dd850fe474bc3b75af8649d4519208791ccc286c01c0614f5de67ebcff69d90c5be8b09e19ad9617e4059d8d6b