5 targets in platform-neutral contract
5 targets in platform-neutral contract
target contracts compiled from Skill IR
5 cases; 1 file-backed
command 10; model 10; recorded 0
review pairs hide baseline vs with-skill labels
pending 5; answer key hidden
adjudication decisions; pending 5
4 blockers; local reproducible true
2.0 coverage; extensions partial 0, planned 0; evidence pending 4
target conformance pass rate
0 native; 4 installer-enforced
155 scripts scanned; secrets found
235 files scanned for Python 3.11
719 largest lines; 0 watchlist; 6 early; 71 CLI handlers; 18 in entrypoint
11 scanned skills; route collisions
1 metadata events; 0 missed triggers
proposal-review; approval 0; release lock false
curator-review; ready 1; top score 88
0 gates covered; human risk decisions
0 valid submissions; 0 invalid
194 public surfaces scanned
0 open blocker annotations
5 targets; MIT license
580 zip entries; package verification
4 adapters; 12 permissions enforced; 0 permission failures
declared minor; 0 breaking changes
5 targets in platform-neutral contract
target contracts compiled from Skill IR
5 cases; 1 file-backed
command 10; model 10; recorded 0
review pairs hide baseline vs with-skill labels
pending 5; answer key hidden
adjudication decisions; pending 5
4 blockers; local reproducible true
2.0 coverage; extensions partial 0, planned 0; evidence pending 4
target conformance pass rate
0 native; 4 installer-enforced
155 scripts scanned; secrets found
235 files scanned for Python 3.11
719 largest lines; 0 watchlist; 6 early; 71 CLI handlers; 18 in entrypoint
11 scanned skills; route collisions
1 metadata events; 0 missed triggers
proposal-review; approval 0; release lock false
curator-review; ready 1; top score 88
0 gates covered; human risk decisions
0 valid submissions; 0 invalid
201 public surfaces scanned
0 open blocker annotations
5 targets; MIT license
580 zip entries; package verification
4 adapters; 12 permissions enforced; 0 permission failures
declared minor; 0 breaking changes
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 10; recorded 0; reviewed 0/5; review pending 5
initial load 999/1000; deferred 531767/120000; top deferred scripts 469568; resource governance governed; quality density 140.1
5 / 5 targets pass
0 secrets; 155 scripts; 3 network-capable scripts; 0 help smoke failures
Python 3.11; 235 files; 0 compatibility issues; 0 syntax; 0 f-string 3.11 hazards
232 Python files; 0 hotspots; 0 watchlist files; 6 early watch files; 0 blockers; largest 719 lines; 71 CLI handlers; 18 in entrypoint
3/3 permissions approved; gaps 0; required file_write, network, subprocess
4/4 targets probed; native 0; metadata fallback 4; installer 4; residual risks 4
11 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 0 actionable drift; 22 scoped non-actionable issues
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock false; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock false
0 active waivers; 1 warning gates still need reviewer decision
4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 12/19 pass; 7 blocked; overclaim guard true
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 4 adapters; installer permissions 12 enforced / 0 failures
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 10; recorded 0; reviewed 0/5; review pending 5
initial load 987/1000; deferred 531767/120000; top deferred scripts 469568; resource governance governed; quality density 141.8
5 / 5 targets pass
0 secrets; 155 scripts; 3 network-capable scripts; 0 help smoke failures
Python 3.11; 235 files; 0 compatibility issues; 0 syntax; 0 f-string 3.11 hazards
232 Python files; 0 hotspots; 0 watchlist files; 6 early watch files; 0 blockers; largest 719 lines; 71 CLI handlers; 18 in entrypoint
3/3 permissions approved; gaps 0; required file_write, network, subprocess
4/4 targets probed; native 0; metadata fallback 4; installer 4; residual risks 4
11 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 0 actionable drift; 22 scoped non-actionable issues
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock false; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock false
0 active waivers; 1 warning gates still need reviewer decision
4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 12/19 pass; 7 blocked; overclaim guard true
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 4 adapters; installer permissions 12 enforced / 0 failures
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
补足 output eval 覆盖、execution evidence、blind A/B 和 reviewer adjudication。同步补足盲审声明。
没有输出质量和人工盲评证据时,Skill 只能证明会触发,不能证明输出真的更好且经得起审查。python3 scripts/adjudicate_output_review.py --write-template && python3 scripts/yao.py output-review{"id":"skill-package-contract","prompt":"Turn this repeated workflow into a reusable team skill package.","baseline_output":"I can write a prompt for that workflow and include a s…
# Output Quality Scorecard
# Output Execution Runs
# Output Blind A/B Review Pack
<title>Output Review Kit</title>
# Output Review Adjudication
对保留的 warning 写入 reviewer、理由、范围和到期时间,或修掉 warning。
warning 可以被接受,但必须可审计、会过期,并且不能掩盖 blocker。python3 scripts/render_review_waivers.py .# Review Waivers
# Review Waiver Method
补齐 provider、真人盲评、原生权限执行和真实客户端遥测证据,或明确本次发布不声明 world-class 完成。
世界级结论必须来自已接受的外部/人工证据;计划、metadata fallback、待评审和本地命令都不能替代完成证据。python3 scripts/yao.py world-class-runbook . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py review-studio .以下条目仍需真实外部或人工证据;提交文件、校验命令和阻断检查必须同时闭环。
model-executed 10; token-observed 10
evidence/world_class/submissions/provider-holdout.jsonevidence/world_class/templates/provider-holdout.intake.json暂无阻断检查。
owners: operator with provider credentialsevidence: provider-holdoutSet one provider API key in the operator shell, such as OPENAI_API_KEY or DEEPSEEK_API_KEY; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired env_any precheck is missing.Set one provider API key in the operator shell, such as OPENAI_API_KEY or DEEPSEEK_API_KEY; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key provider-holdout --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.0/5 decisions; pending 5
evidence/world_class/submissions/human-adjudication.jsonevidence/world_class/templates/human-adjudication.intake.jsonpending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.owners: human reviewerevidence: human-adjudicationAssign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: human reviewerevidence: human-adjudicationSet reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired human precheck is human-required.Assign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Set reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '==pair_count'.Every pair needs one valid human judgment.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 5 does not satisfy '==0'.Record a reviewer choice and reason for every pair.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Complete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Record reviewer and reviewed_at before adjudication can count.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key human-adjudication --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.native-enforced targets 0; installer-enforced targets 4
evidence/world_class/submissions/native-permission-enforcement.jsonevidence/world_class/templates/native-permission-enforcement.intake.jsonnative_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.owners: target client or installer integratorevidence: native-permission-enforcementAttach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: target client or installer integratorevidence: native-permission-enforcementCollect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Attach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Collect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-permission-enforcement --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.external source events 0; adoption samples 1
evidence/world_class/submissions/native-client-telemetry.jsonevidence/world_class/templates/native-client-telemetry.intake.jsonexternal_source_events: 0 / >0Import at least one metadata-only event from a real client.owners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryInstall a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryImport at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Install a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Import at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-client-telemetry --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.# World-Class Evidence Ledger
# World-Class Evidence Plan
# World-Class Evidence Intake
# World-Class Evidence Preflight
<title>World-Class Evidence Preflight</title>
# World-Class Submission Review
# World-Class Claim Guard
"title": "Yao World-Class Evidence Intake",
"evidence_key": "provider-holdout",
"evidence_key": "human-adjudication",
"evidence_key": "native-permission-enforcement",
"evidence_key": "native-client-telemetry",
# Skill OS 2.0 Audit
"winner_variant": "Use A or B after reading the blind review pack. Leave blank when pending.",
# Runtime Permission Probes
# Adoption And Drift Report
补足 output eval 覆盖、execution evidence、blind A/B 和 reviewer adjudication。同步补足盲审声明。
没有输出质量和人工盲评证据时,Skill 只能证明会触发,不能证明输出真的更好且经得起审查。python3 scripts/adjudicate_output_review.py --write-template && python3 scripts/yao.py output-review{"id":"skill-package-contract","prompt":"Turn this repeated workflow into a reusable team skill package.","baseline_output":"I can write a prompt for that workflow and include a s…
# Output Quality Scorecard
# Output Execution Runs
# Output Blind A/B Review Pack
<title>Output Review Kit</title>
# Output Review Adjudication
对保留的 warning 写入 reviewer、理由、范围和到期时间,或修掉 warning。
warning 可以被接受,但必须可审计、会过期,并且不能掩盖 blocker。python3 scripts/render_review_waivers.py .# Review Waivers
# Review Waiver Method
补齐 provider、真人盲评、原生权限执行和真实客户端遥测证据,或明确本次发布不声明 world-class 完成。
世界级结论必须来自已接受的外部/人工证据;计划、metadata fallback、待评审和本地命令都不能替代完成证据。python3 scripts/yao.py world-class-runbook . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py review-studio .以下条目仍需真实外部或人工证据;提交文件、校验命令和阻断检查必须同时闭环。
model-executed 10; token-observed 10
evidence/world_class/submissions/provider-holdout.jsonevidence/world_class/templates/provider-holdout.intake.json暂无阻断检查。
owners: operator with provider credentialsevidence: provider-holdoutSet one provider API key in the operator shell, such as OPENAI_API_KEY or DEEPSEEK_API_KEY; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired env_any precheck is missing.Set one provider API key in the operator shell, such as OPENAI_API_KEY or DEEPSEEK_API_KEY; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key provider-holdout --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.0/5 decisions; pending 5
evidence/world_class/submissions/human-adjudication.jsonevidence/world_class/templates/human-adjudication.intake.jsonpending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.owners: human reviewerevidence: human-adjudicationAssign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: human reviewerevidence: human-adjudicationSet reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired human precheck is human-required.Assign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Set reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '==pair_count'.Every pair needs one valid human judgment.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 5 does not satisfy '==0'.Record a reviewer choice and reason for every pair.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Complete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Record reviewer and reviewed_at before adjudication can count.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key human-adjudication --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.native-enforced targets 0; installer-enforced targets 4
evidence/world_class/submissions/native-permission-enforcement.jsonevidence/world_class/templates/native-permission-enforcement.intake.jsonnative_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.owners: target client or installer integratorevidence: native-permission-enforcementAttach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: target client or installer integratorevidence: native-permission-enforcementCollect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Attach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Collect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-permission-enforcement --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.external source events 0; adoption samples 1
evidence/world_class/submissions/native-client-telemetry.jsonevidence/world_class/templates/native-client-telemetry.intake.jsonexternal_source_events: 0 / >0Import at least one metadata-only event from a real client.owners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryInstall a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryImport at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Install a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Import at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-client-telemetry --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.# World-Class Evidence Ledger
# World-Class Evidence Plan
# World-Class Evidence Intake
# World-Class Evidence Preflight
<title>World-Class Evidence Preflight</title>
# World-Class Submission Review
# World-Class Claim Guard
"title": "Yao World-Class Evidence Intake",
"evidence_key": "provider-holdout",
"evidence_key": "human-adjudication",
"evidence_key": "native-permission-enforcement",
"evidence_key": "native-client-telemetry",
# Skill OS 2.0 Audit
"winner_variant": "Use A or B after reading the blind review pack. Leave blank when pending.",
# Runtime Permission Probes
# Adoption And Drift Report
initial load 999/1000; deferred 531767/120000; top deferred scripts 469568; resource governance governed; quality density 140.1
initial load 987/1000; deferred 531767/120000; top deferred scripts 469568; resource governance governed; quality density 141.8
Review reports/compiled_targets.md before packaging to inspect target adapter modes, generated files, preserved semantics, warnings, and unsupported features.
e0863533e0a0f0b5069bace6a24d7adfb43d615453f08cff7578076a3cb1a2a3db71157184dc649ee72dfcaf399204d44a366d942393d055571e0137d42c9f7c高风险 secret、远程 inline execution、缺失依赖策略或无法解释的脚本接口应阻断 governed release。
claim guard 扫描 README、docs 和 reports 中的完成态表述;ledger 未 ready 时,任何英文完成断言、true 状态声明或中文完成态都会阻断发布审查。
这里把每个指标压缩为状态、分数和第一条证据,先给出整体判断,再进入下方明细。This compresses each metric into status, score, and the first evidence item before the detailed cards below.
-| 类型Type | 证据Evidence | 建议Action |
|---|---|---|
| 强项Strength | 触发面保持精简,并锚定在 frontmatter description。The trigger surface stays lean and anchored in the frontmatter description. | 保留并复用Keep |
| 强项Strength | 已生成 Skill IR,核心语义可先于平台打包被审查和迁移。Skill IR is generated so core semantics can be reviewed and migrated before platform packaging. | 保留并复用Keep |
| 强项Strength | 已生成目标编译报告,可审查 IR 到 OpenAI、Claude、generic 等目标契约的映射。Target compilation evidence is generated to review how IR maps to OpenAI, Claude, generic, and other target contracts. | 保留并复用Keep |
| 缺口Gap | 上下文成本需要补强:入口约 342 个词/字,references 约 16934 个词/字。Context cost needs improvement: Entrypoint is about 342 words/characters; references are about 16934. | 纳入下一轮修复Fix next |
| 强项Strength | 触发面保持精简,并锚定在 frontmatter description。The trigger surface stays lean and anchored in the frontmatter description. | 保留并复用Keep |
| 强项Strength | 已生成 Skill IR,核心语义可先于平台打包被审查和迁移。Skill IR is generated so core semantics can be reviewed and migrated before platform packaging. | 保留并复用Keep |
| 强项Strength | 已生成目标编译报告,可审查 IR 到 OpenAI、Claude、generic 等目标契约的映射。Target compilation evidence is generated to review how IR maps to OpenAI, Claude, generic, and other target contracts. | 保留并复用Keep |
| 缺口Gap | 上下文成本需要补强:入口约 314 个词/字,references 约 16934 个词/字。Context cost needs improvement: Entrypoint is about 314 words/characters; references are about 16934. | 纳入下一轮修复Fix next |
| 风险Risk | 信号Signal | 应对Response |
|---|---|---|
| 误触发风险Trigger risk | frontmatter description 已存在,具备基础路由面。The frontmatter description exists, giving the skill a basic routing surface. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 输出漂移风险Output drift risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 证据不足风险Evidence gap risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 包体膨胀风险Package bloat risk | SKILL.md 约 342 个词/字。SKILL.md is about 342 words/characters. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 跨平台迁移风险Portability risk | agents/interface.yaml 已存在。agents/interface.yaml exists. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 误触发风险Trigger risk | frontmatter description 已存在,具备基础路由面。The frontmatter description exists, giving the skill a basic routing surface. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 输出漂移风险Output drift risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 证据不足风险Evidence gap risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 包体膨胀风险Package bloat risk | SKILL.md 约 314 个词/字。SKILL.md is about 314 words/characters. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 跨平台迁移风险Portability risk | agents/interface.yaml 已存在。agents/interface.yaml exists. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
这里把每个指标压缩为状态、分数和第一条证据,先给出整体判断,再进入下方明细。This compresses each metric into status, score, and the first evidence item before the detailed cards below.
-| 类型Type | 证据Evidence | 建议Action |
|---|---|---|
| 强项Strength | 触发面保持精简,并锚定在 frontmatter description。The trigger surface stays lean and anchored in the frontmatter description. | 保留并复用Keep |
| 强项Strength | 已生成 Skill IR,核心语义可先于平台打包被审查和迁移。Skill IR is generated so core semantics can be reviewed and migrated before platform packaging. | 保留并复用Keep |
| 强项Strength | 已生成目标编译报告,可审查 IR 到 OpenAI、Claude、generic 等目标契约的映射。Target compilation evidence is generated to review how IR maps to OpenAI, Claude, generic, and other target contracts. | 保留并复用Keep |
| 缺口Gap | 上下文成本需要补强:入口约 342 个词/字,references 约 16934 个词/字。Context cost needs improvement: Entrypoint is about 342 words/characters; references are about 16934. | 纳入下一轮修复Fix next |
| 强项Strength | 触发面保持精简,并锚定在 frontmatter description。The trigger surface stays lean and anchored in the frontmatter description. | 保留并复用Keep |
| 强项Strength | 已生成 Skill IR,核心语义可先于平台打包被审查和迁移。Skill IR is generated so core semantics can be reviewed and migrated before platform packaging. | 保留并复用Keep |
| 强项Strength | 已生成目标编译报告,可审查 IR 到 OpenAI、Claude、generic 等目标契约的映射。Target compilation evidence is generated to review how IR maps to OpenAI, Claude, generic, and other target contracts. | 保留并复用Keep |
| 缺口Gap | 上下文成本需要补强:入口约 314 个词/字,references 约 16934 个词/字。Context cost needs improvement: Entrypoint is about 314 words/characters; references are about 16934. | 纳入下一轮修复Fix next |
| 风险Risk | 信号Signal | 应对Response |
|---|---|---|
| 误触发风险Trigger risk | frontmatter description 已存在,具备基础路由面。The frontmatter description exists, giving the skill a basic routing surface. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 输出漂移风险Output drift risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 证据不足风险Evidence gap risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 包体膨胀风险Package bloat risk | SKILL.md 约 342 个词/字。SKILL.md is about 342 words/characters. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 跨平台迁移风险Portability risk | agents/interface.yaml 已存在。agents/interface.yaml exists. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 误触发风险Trigger risk | frontmatter description 已存在,具备基础路由面。The frontmatter description exists, giving the skill a basic routing surface. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 输出漂移风险Output drift risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 证据不足风险Evidence gap risk | 已生成 20 / 20 类报告证据。Generated 20 / 20 evidence report types. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 包体膨胀风险Package bloat risk | SKILL.md 约 314 个词/字。SKILL.md is about 314 words/characters. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
| 跨平台迁移风险Portability risk | agents/interface.yaml 已存在。agents/interface.yaml exists. | 先补证据和边界,再增加包体复杂度。Improve evidence and boundaries before adding package complexity. |
Supporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.
copy to artifact_refs:false
Supporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.
copy to artifact_refs:false