意图画布
intent confidence 100/100; Intent is clear enough to package the first routeable version.
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 0; recorded 0; reviewed 0/5; review pending 5
initial load 990/1000; deferred 512782/120000; top deferred scripts 451479; resource governance governed; quality density 131.3
5 / 5 targets pass
0 secrets; 152 scripts; 3 network-capable scripts; 0 help smoke failures
Python 3.11; 231 files; 0 compatibility issues; 0 syntax; 0 f-string 3.11 hazards
228 Python files; 0 hotspots; 0 watchlist files; 3 early watch files; 0 blockers; largest 696 lines; 68 CLI handlers; 18 in entrypoint
3/3 permissions approved; gaps 0; required file_write, network, subprocess
4/4 targets probed; native 0; metadata fallback 4; installer 4; residual risks 4
12 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 0 actionable drift; 24 scoped non-actionable issues
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock true; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock true
0 active waivers; 1 warning gates still need reviewer decision
4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 10/19 pass; 9 blocked; overclaim guard true
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 4 adapters; installer permissions 12 enforced / 0 failures
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
intent confidence 100/100; Intent is clear enough to package the first routeable version.
13 trigger cases; 0 misroutes; 0 ambiguous
5/5 cases; with-skill 100.0; baseline 0.0; file-backed 1; near-neighbor 1; blind A/B 5; exec 10; command 10; model 0; recorded 0; reviewed 0/5; review pending 5
initial load 990/1000; deferred 512782/120000; top deferred scripts 451479; resource governance governed; quality density 131.3
5 / 5 targets pass
0 secrets; 152 scripts; 3 network-capable scripts; 0 help smoke failures
Python 3.11; 231 files; 0 compatibility issues; 0 syntax; 0 f-string 3.11 hazards
228 Python files; 0 hotspots; 0 watchlist files; 3 early watch files; 0 blockers; largest 696 lines; 68 CLI handlers; 18 in entrypoint
3/3 permissions approved; gaps 0; required file_write, network, subprocess
4/4 targets probed; native 0; metadata fallback 4; installer 4; residual risks 4
12 skills, 1 actionable; 0 actionable route collisions; 0 actionable owner gaps; 0 actionable stale; 0 actionable drift; 24 scoped non-actionable issues
1 metadata events; adoption 0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock true; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock true
0 active waivers; 1 warning gates still need reviewer decision
4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 9/19 pass; 10 blocked; overclaim guard true
yao-meta-skill 1.1.0; 6/6 compatibility entries pass; install pass with 4 adapters; installer permissions 12 enforced / 0 failures
0 promote; 3 keep current; 0 blocked; upgrade minor declared / minor recommended
无。
补足 output eval 覆盖、execution evidence、blind A/B 和 reviewer adjudication。同步补足盲审声明。
没有输出质量和人工盲评证据时,Skill 只能证明会触发,不能证明输出真的更好且经得起审查。python3 scripts/adjudicate_output_review.py --write-template && python3 scripts/yao.py output-review{"id":"skill-package-contract","prompt":"Turn this repeated workflow into a reusable team skill package.","baseline_output":"I can write a prompt for that workflow and include a s…
# Output Quality Scorecard
# Output Execution Runs
# Output Blind A/B Review Pack
<title>Output Review Kit</title>
# Output Review Adjudication
对保留的 warning 写入 reviewer、理由、范围和到期时间,或修掉 warning。
warning 可以被接受,但必须可审计、会过期,并且不能掩盖 blocker。python3 scripts/render_review_waivers.py .# Review Waivers
# Review Waiver Method
补齐 provider、真人盲评、原生权限执行和真实客户端遥测证据,或明确本次发布不声明 world-class 完成。
世界级结论必须来自已接受的外部/人工证据;计划、metadata fallback、待评审和本地命令都不能替代完成证据。python3 scripts/yao.py world-class-runbook . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py review-studio .以下条目仍需真实外部或人工证据;提交文件、校验命令和阻断检查必须同时闭环。
model-executed 0; token-observed 0
evidence/world_class/submissions/provider-holdout.jsonevidence/world_class/templates/provider-holdout.intake.jsonmodel_executed_count: 0 / >0Run provider-backed output-exec with real credentials.token_observed_count: 0 / >0Provider execution should return non-estimated token usage.owners: operator with provider credentialsevidence: provider-holdoutSet OPENAI_API_KEY in the operator shell; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: operator with provider credentialsevidence: provider-holdoutRun provider-backed output-exec with real credentials.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired env precheck is missing.Set OPENAI_API_KEY in the operator shell; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Run provider-backed output-exec with real credentials.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Provider execution should return non-estimated token usage.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key provider-holdout --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.0/5 decisions; pending 5
evidence/world_class/submissions/human-adjudication.jsonevidence/world_class/templates/human-adjudication.intake.jsonpending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.owners: human reviewerevidence: human-adjudicationAssign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: human reviewerevidence: human-adjudicationSet reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired human precheck is human-required.Assign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Set reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '==pair_count'.Every pair needs one valid human judgment.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 5 does not satisfy '==0'.Record a reviewer choice and reason for every pair.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Complete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Record reviewer and reviewed_at before adjudication can count.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key human-adjudication --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.native-enforced targets 0; installer-enforced targets 4
evidence/world_class/submissions/native-permission-enforcement.jsonevidence/world_class/templates/native-permission-enforcement.intake.jsonnative_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.owners: target client or installer integratorevidence: native-permission-enforcementAttach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: target client or installer integratorevidence: native-permission-enforcementCollect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Attach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Collect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-permission-enforcement --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.external source events 0; adoption samples 1
evidence/world_class/submissions/native-client-telemetry.jsonevidence/world_class/templates/native-client-telemetry.intake.jsonexternal_source_events: 0 / >0Import at least one metadata-only event from a real client.owners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryInstall a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryImport at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Install a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Import at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-client-telemetry --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.# World-Class Evidence Ledger
# World-Class Evidence Plan
# World-Class Evidence Intake
# World-Class Evidence Preflight
<title>World-Class Evidence Preflight</title>
# World-Class Submission Review
# World-Class Claim Guard
"title": "Yao World-Class Evidence Intake",
"evidence_key": "provider-holdout",
"evidence_key": "human-adjudication",
"evidence_key": "native-permission-enforcement",
"evidence_key": "native-client-telemetry",
# Skill OS 2.0 Audit
"winner_variant": "Use A or B after reading the blind review pack. Leave blank when pending.",
# Runtime Permission Probes
# Adoption And Drift Report
补足 output eval 覆盖、execution evidence、blind A/B 和 reviewer adjudication。同步补足盲审声明。
没有输出质量和人工盲评证据时,Skill 只能证明会触发,不能证明输出真的更好且经得起审查。python3 scripts/adjudicate_output_review.py --write-template && python3 scripts/yao.py output-review{"id":"skill-package-contract","prompt":"Turn this repeated workflow into a reusable team skill package.","baseline_output":"I can write a prompt for that workflow and include a s…
# Output Quality Scorecard
# Output Execution Runs
# Output Blind A/B Review Pack
<title>Output Review Kit</title>
# Output Review Adjudication
对保留的 warning 写入 reviewer、理由、范围和到期时间,或修掉 warning。
warning 可以被接受,但必须可审计、会过期,并且不能掩盖 blocker。python3 scripts/render_review_waivers.py .# Review Waivers
# Review Waiver Method
补齐 provider、真人盲评、原生权限执行和真实客户端遥测证据,或明确本次发布不声明 world-class 完成。
世界级结论必须来自已接受的外部/人工证据;计划、metadata fallback、待评审和本地命令都不能替代完成证据。python3 scripts/yao.py world-class-runbook . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissions && python3 scripts/yao.py review-studio .以下条目仍需真实外部或人工证据;提交文件、校验命令和阻断检查必须同时闭环。
model-executed 0; token-observed 0
evidence/world_class/submissions/provider-holdout.jsonevidence/world_class/templates/provider-holdout.intake.jsonmodel_executed_count: 0 / >0Run provider-backed output-exec with real credentials.token_observed_count: 0 / >0Provider execution should return non-estimated token usage.owners: operator with provider credentialsevidence: provider-holdoutSet OPENAI_API_KEY in the operator shell; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: operator with provider credentialsevidence: provider-holdoutRun provider-backed output-exec with real credentials.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired env precheck is missing.Set OPENAI_API_KEY in the operator shell; never commit or print the value.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Run provider-backed output-exec with real credentials.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Provider execution should return non-estimated token usage.python3 scripts/yao.py output-exec --provider-runner openai --timeout-seconds 60 && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key provider-holdout --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.0/5 decisions; pending 5
evidence/world_class/submissions/human-adjudication.jsonevidence/world_class/templates/human-adjudication.intake.jsonpending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.owners: human reviewerevidence: human-adjudicationAssign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: human reviewerevidence: human-adjudicationSet reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired human precheck is human-required.Assign a real reviewer identity before claiming human adjudication.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Set reviewer_attestation only after choices are completed before opening the answer key.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '==pair_count'.Every pair needs one valid human judgment.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 5 does not satisfy '==0'.Record a reviewer choice and reason for every pair.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Complete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value False does not satisfy 'true'.Record reviewer and reviewed_at before adjudication can count.python3 scripts/yao.py output-review && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key human-adjudication --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.native-enforced targets 0; installer-enforced targets 4
evidence/world_class/submissions/native-permission-enforcement.jsonevidence/world_class/templates/native-permission-enforcement.intake.jsonnative_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.owners: target client or installer integratorevidence: native-permission-enforcementAttach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: target client or installer integratorevidence: native-permission-enforcementCollect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Attach a real target-client or external installer runtime guard; metadata fallback is not enough.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Collect real target-client or external runtime guard proof.python3 scripts/yao.py runtime-permissions . --package-dir dist && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-permission-enforcement --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.external source events 0; adoption samples 0
evidence/world_class/submissions/native-client-telemetry.jsonevidence/world_class/templates/native-client-telemetry.intake.jsonexternal_source_events: 0 / >0Import at least one metadata-only event from a real client.adoption_sample_count: 0 / >0Telemetry must include adoption outcome evidence.owners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryInstall a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsowners: Browser/Chrome/IDE/provider client integratorevidence: native-client-telemetryTelemetry must include adoption outcome evidence.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsRequired external precheck is external-required.Install a real Browser, Chrome, IDE, or provider client that emits metadata-only events.python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Telemetry must include adoption outcome evidence.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsCurrent value 0 does not satisfy '>0'.Import at least one metadata-only event from a real client.python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-kit . --evidence-key native-client-telemetry --output-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-intake . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-submission-review . --submissions-dir evidence/world_class/submissionspython3 scripts/yao.py world-class-ledger . --submissions-dir evidence/world_class/submissionssource: world-class-submission-kit; counts as evidence: false; prefill counts as evidence: false
artifact_refs: trueRows marked submission-ref are the aggregate paths expected in artifact_refs.artifact_refs: falseSupporting-evidence rows help reviewers audit the packet but do not all need to be copied into artifact_refs.# World-Class Evidence Ledger
# World-Class Evidence Plan
# World-Class Evidence Intake
# World-Class Evidence Preflight
<title>World-Class Evidence Preflight</title>
# World-Class Submission Review
# World-Class Claim Guard
"title": "Yao World-Class Evidence Intake",
"evidence_key": "provider-holdout",
"evidence_key": "human-adjudication",
"evidence_key": "native-permission-enforcement",
"evidence_key": "native-client-telemetry",
# Skill OS 2.0 Audit
"winner_variant": "Use A or B after reading the blind review pack. Leave blank when pending.",
# Runtime Permission Probes
# Adoption And Drift Report
1 metadata events; adoption 100.0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock true; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock true
1 metadata events; adoption 0; missed 0; bad-output 0; risk low; daily proposals 5; daily decision proposal-review; daily release lock true; weekly queue 5 unique; weekly ready 1; weekly top 88; weekly release lock true
这里列出每个 world-class 证据项的当前状态、完成定义、证据来源、隐私约束和下一步;计划、metadata fallback、待评审和本地命令不会被当成完成证据。
-Collect at least one provider-backed output-eval holdout run with model, timing, and token metadata.
model_executed_count: 0 / >0Run provider-backed output-exec with real credentials.timing_observed_count: 10 / >0Provider execution should record timing metadata.token_observed_count: 0 / >0Provider execution should return non-estimated token usage.Record real blind A/B reviewer decisions before claiming human output review completion.
pair_count: 5 / >0Generate the blind A/B review pack.pending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.invalid_decision_count: 0 / ==0Fix malformed winner/confidence entries.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.reason_required: True / trueKeep reason mandatory for every imported or direct reviewer decision.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.raw_content_excluded_attested: True / trueAttest that reviewer decisions exclude raw prompts, outputs, transcripts, messages, and private user content.raw_content_allowed: False / falseAdjudication evidence should store prompt hashes and reviewer metadata, not raw prompts, outputs, transcripts, or messages.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.Prove at least one real target client or external installer runtime guard enforces approved high-permission capabilities.
native_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.failure_count: 0 / ==0Runtime permission probes must stay clean.installer_enforcement_ready: True / trueInstaller enforcement is supporting evidence, not native proof.Import production metadata-only events from a real external client into the local drift loop.
external_source_events: 0 / >0Import at least one metadata-only event from a real client.adoption_sample_count: 1 / >0Telemetry must include adoption outcome evidence.raw_content_allowed: False / falseTelemetry must stay metadata-only.Collect at least one provider-backed output-eval holdout run with model, timing, and token metadata.
model_executed_count: 0 / >0Run provider-backed output-exec with real credentials.timing_observed_count: 10 / >0Provider execution should record timing metadata.token_observed_count: 0 / >0Provider execution should return non-estimated token usage.Record real blind A/B reviewer decisions before claiming human output review completion.
pair_count: 5 / >0Generate the blind A/B review pack.pending_count: 5 / ==0Record a reviewer choice and reason for every pair.judgment_count: 0 / ==pair_countEvery pair needs one valid human judgment.invalid_decision_count: 0 / ==0Fix malformed winner/confidence entries.reviewer_metadata_present: False / trueRecord reviewer and reviewed_at before adjudication can count.reason_required: True / trueKeep reason mandatory for every imported or direct reviewer decision.blind_review_attested: False / trueSet reviewer_attestation only after choices are completed before opening the answer key.raw_content_excluded_attested: True / trueAttest that reviewer decisions exclude raw prompts, outputs, transcripts, messages, and private user content.raw_content_allowed: False / falseAdjudication evidence should store prompt hashes and reviewer metadata, not raw prompts, outputs, transcripts, or messages.ready_for_human_evidence: False / trueComplete all reviewer decisions with metadata and rationale, plus blind-review attestation and integrity fingerprints.Prove at least one real target client or external installer runtime guard enforces approved high-permission capabilities.
native_enforcement_count: 0 / >0Collect real target-client or external runtime guard proof.failure_count: 0 / ==0Runtime permission probes must stay clean.installer_enforcement_ready: True / trueInstaller enforcement is supporting evidence, not native proof.Import production metadata-only events from a real external client into the local drift loop.
external_source_events: 0 / >0Import at least one metadata-only event from a real client.adoption_sample_count: 0 / >0Telemetry must include adoption outcome evidence.raw_content_allowed: False / falseTelemetry must stay metadata-only.4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 10/19 pass; 9 blocked; overclaim guard true
4 pending world-class evidence entries; 1 human pending; 3 external pending; source checks 9/19 pass; 10 blocked; overclaim guard true
世界级证据尚未完成:4 项待补,0 项已接受。World-class evidence is not complete: 4 pending, 0 accepted.
缺少真实 provider 模型运行和 token metadata。Missing a real provider model run and token metadata.
盲评 pair 仍待真实 reviewer 决策。Blind-review pairs still need real reviewer decisions.
原生 runtime enforcement 仍待目标客户端或外部安装器证明。Native runtime enforcement still needs target-client or external-installer proof.
真实外部客户端 metadata-only 事件仍未导入。Real external-client metadata-only events have not been imported yet.
世界级证据尚未完成:4 项待补,0 项已接受。World-class evidence is not complete: 4 pending, 0 accepted.
缺少真实 provider 模型运行和 token metadata。Missing a real provider model run and token metadata.
盲评 pair 仍待真实 reviewer 决策。Blind-review pairs still need real reviewer decisions.
原生 runtime enforcement 仍待目标客户端或外部安装器证明。Native runtime enforcement still needs target-client or external-installer proof.
真实外部客户端 metadata-only 事件仍未导入。Real external-client metadata-only events have not been imported yet.
世界级证据尚未完成:4 项待补,0 项已接受。World-class evidence is not complete: 4 pending, 0 accepted.
缺少真实 provider 模型运行和 token metadata。Missing a real provider model run and token metadata.
盲评 pair 仍待真实 reviewer 决策。Blind-review pairs still need real reviewer decisions.
原生 runtime enforcement 仍待目标客户端或外部安装器证明。Native runtime enforcement still needs target-client or external-installer proof.
真实外部客户端 metadata-only 事件仍未导入。Real external-client metadata-only events have not been imported yet.
世界级证据尚未完成:4 项待补,0 项已接受。World-class evidence is not complete: 4 pending, 0 accepted.
缺少真实 provider 模型运行和 token metadata。Missing a real provider model run and token metadata.
盲评 pair 仍待真实 reviewer 决策。Blind-review pairs still need real reviewer decisions.
原生 runtime enforcement 仍待目标客户端或外部安装器证明。Native runtime enforcement still needs target-client or external-installer proof.
真实外部客户端 metadata-only 事件仍未导入。Real external-client metadata-only events have not been imported yet.
This operator view shows which external and human evidence is still blocked, which commands prepare editable submission drafts, and why preflight never counts as accepted evidence.
-collect-source10>0python3 scripts/yao.py telemetry-import . --input-jsonl .yao/telemetry_spool/external_events.jsonl && python3 scripts/yao.py world-class-preflight . --submissions-dir evidence/world_class/submissionsA single operating page for collecting the remaining human and external evidence. It coordinates action, but does not accept evidence or change the ledger.
-pending12evidence/world_class/submissions/native-client-telemetry.jsonexternal_source_events: 0 / >0Import at least one metadata-only event from a real client.adoption_sample_count: 1 / >0Telemetry must include adoption outcome evidence.raw_content_allowed: False / falseTelemetry must stay metadata-only.external_source_events: 0 / >0Import at least one metadata-only event from a real client.adoption_sample_count: 0 / >0Telemetry must include adoption outcome evidence.raw_content_allowed: False / falseTelemetry must stay metadata-only.