From 9681707ad905f95879b6dec5e8c846037ccc57b7 Mon Sep 17 00:00:00 2001 From: Coden Date: Wed, 2 Sep 2026 09:26:12 +0900 Subject: [PATCH] =?UTF-8?q?docs:=20=EC=99=B8=EB=B6=80=20=EB=A6=AC=EB=B7=B0?= =?UTF-8?q?=EC=96=B4=EC=99=80=EC=9D=98=20=EB=B0=A9=ED=96=A5=20=EC=97=B0?= =?UTF-8?q?=EA=B5=AC=20=EB=9D=BC=EC=9A=B4=EB=93=9C=EB=A5=BC=20=ED=95=B8?= =?UTF-8?q?=EB=93=9C=EC=98=A4=ED=94=84=EC=97=90=20=EA=B8=B0=EB=A1=9D?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 스크러빙된 GLM 패킷 두 건으로 `wclass` 와 advisory 동반 도구의 다음 방향을 검토했다. 코드/설정 변경 없음, 벤더 라우트 미실행, 캠페인·사용량·자격증명 미접근. 공급자 출력은 신뢰하지 않는 입력으로 다루고 전부 소스와 대조했다. 검증을 통과한 것: 휠 소속 판단 규칙, 60/12 게이트는 재색인이 아니라 폐기해야 하는 이유, 항목 D 의 가장 강한 형태가 주입이 아니라 수확이라는 점, check_test_vacuity.py 가 추정량이 아니라 엔진이라는 구분, list-don't-judge 의 분할, 잎 단위 위로 집계할 때 가려지는 것. 기각한 것과 근거: - 항목 B 의 심볼 수가 한 릴리스 낡았다. 0.30.0 의 declares_stream_json_input 이 빠져 일곱이 아니라 여덟이며, 같은 파일을 import 하는 모듈이 넷 더 있어 ask 경로만 떼어내도 5,526 줄 파일은 휠에 남는다. - 관측 세 건짜리 표본에 증거 라벨이 붙었다. 항목 5 가 이미 p 추정에 턱없이 부족하다고 못 박은 수치다. - 제안된 결정 규칙이 z=1.96 을 쓴다. speculative_report.py:798-843 의 t 분위수 규칙을 적용하면 15/25 의 하한이 0.4074 에서 0.3983 으로 내려가 기각에서 미달로 분기가 바뀐다. - 파일럿 kill gate 가 4% 수율에 걸려 있다. 원 발견의 기저율은 14 과제 중 6 건 이라 한 자릿수 낮고, 재현과 열 배 약한 효과를 구분하지 못한다. 항목 A-D 는 여전히 미승인이며, 이 기록은 항목 D 를 승격시키지 않고 좁힌다. ./.weightclass/verify 는 1,599 테스트 35 skip 으로 통과했다. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01P3dXGuZjrBWD669SgUuk65 --- HANDOFF.md | 131 ++++++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 129 insertions(+), 2 deletions(-) diff --git a/HANDOFF.md b/HANDOFF.md index 2f9c8b6..e725460 100644 --- a/HANDOFF.md +++ b/HANDOFF.md @@ -1,9 +1,135 @@ # Handoff -_Last updated: 2026-09-02 KST by Claude (agent guidance split)_ +_Last updated: 2026-09-02 KST by Claude (direction research, no code change)_ _Flexible advisory vendor support follow-up: 2026-08-23 KST._ +## Direction research with an external reviewer (2026-09-02, no code change) + +Two scrubbed GLM packets explored where `wclass` and the advisory companion go next. No code, +document, or configuration outside this file changed; no vendor route ran; no campaign, usage +record, or credential was touched. This section records what the round produced and, equally, +which of its claims did not survive checking against the code. + +Both packets carried only files that are already public, and the scrubber reported zero +redactions in each: + +- Strategy packet: root, core, and advisory `AGENTS.md`, `README.md`, and + `docs/advisory-product-roadmap-v2.md`, plus an inline brief assembled from this file. + 135,719 bytes, SHA-256 prefix `f4b4be0e2468`, 565 s. +- Harness-design packet: `AGENTS.md`, `docs/speculative-cheap-route-design.md`, + `tools/check_test_vacuity.py`, `docs/policy4-fresh-blind-evaluation.md`, + `docs/phase4-go-no-go-template.md`, and `docs/measuring-p-at-work.md`. + 77,342 bytes, SHA-256 prefix `cbff0490e80f`, 599 s. + +Provider output was treated as untrusted throughout and verified against the source before +being recorded here. + +### What survived checking + +**A membership rule for the shipped wheel.** A surface stays in the wheel only if a first-time +user can obtain its entire honest value in a single invocation — no init, no sealed contract, +no price table, no accumulated population — and every claim it makes is enforced by code inside +the wheel. This is the advisory `AGENTS.md` definition of `ask` promoted to a project-wide rule. +It is durable because it is indexed to the user's job rather than to code quality, and the +campaign apparatus fails it by construction: its maximum payoff is `eligible_for_human_review`. + +**Abandon the 60-task/12-advised-failure gate rather than re-index it.** Next Steps item C is +right that the gate counts the wrong event, but re-indexing means changing a pre-registered +counting rule to make completion easier. That is the move this project refused at Phase 2, when +it reported the 9/36 shortfall instead of lowering the floor. The alternative is to close the +natural-population study as under-powered by design mismatch, publish the descriptive record +(`s = 0/2`, retry 1 passed / 4 same / 3 degraded, 14/60 tasks, 9/12 advised failures), and +replace the instrument rather than feed it. + +**The verification hypothesis reduces to a three-to-four day pilot, not a project.** The +strongest form of Next Steps item D is not injection. Running the cheap model on synthetic +tasks under the production setup and harvesting the defects it produces naturally makes the +tamper-artifact confound structurally impossible — nothing was injected, so no authorship signal +can exist — and the harvested process is the same one that produced the original finding. +Injection survives only as a supplement for classes the harvest underproduces. + +**`tools/check_test_vacuity.py` is the engine, not the estimand.** Identity mutation asks +whether a suite's verdict depends on the component at all; that is a necessary condition and the +easiest mutant to catch. Semantic mutation asks whether the verdict separates correct from +plausibly wrong. A suite can pass the vacuity audit on every leaf and still miss a `p21`-class +defect entirely, because its assertions check that output is produced rather than that a +boundary rejects `True`. The reusable part is therefore the temp-copy isolation, the fail-closed +exactly-once `neutralize_source` check, `LeafRecorder`, the unrepresented-method-to-NG move, and +byte-stable output — verified present at `tools/check_test_vacuity.py:79`, `:148-155`, and +`:182-183`. Its natural role is the calibration arm: the identity operator run against each +task's reference solution is the pre-registered proof that the corpus's suites are not vacuous. + +**Split `list-don't-judge` rather than transferring it.** Mechanical outcomes — suite exit code, +compile failure, diff size — are property checks and must be scored mechanically. Materiality is +the one predicate no probe covers, and it is where the judgment lives: three of the original nine +routed-tier wins were test organisation and line wrapping. The binding budget is therefore rater +hours, not model calls. + +**What aggregating above the leaf hides.** Aggregating to the task hides which classes escape, so +a suite that catches five of six reads as checked. Aggregating to the class hides suite-weakness +concentration, so one class's miss rate may be three suite-poor tasks wearing a class's face. + +**`p21` is a property, not a shape.** In Python `True == 1`, so the natural `if version != 1` +already admits `True`; the reference must deliberately exclude `bool`, and the defect is dropping +that exclusion. Two injections written differently satisfy the identical property and the probe +cannot separate them, which is what makes the class machine-checkable under the repository rule +that a format is judged on its own properties. + +**Arithmetic that was checked and is correct.** Wilson 95% for 6/6 is `[0.610, 1.000]`; the +per-class half-width at `p = 0.6` is 24.4 points at `n = 12` and 17.9 points at `n = 25`; the +two-sided 5% critical proportion for 180 forced-choice trials is 57.3%. A leave-one-task-out +jackknife was proposed as the clustered estimator and matches the delete-one convention already +implemented at `src/weightclass/advisory/speculative_report.py:865`. + +### What did not survive checking + +**Item B's symbol count is one release stale.** `ask` uses eight symbols from +`speculative_run.py`, not seven: `declares_stream_json_input` was added in 0.30.0 and appears at +`src/weightclass/advisory/advisory_quick.py:1069`. The low-coupling conclusion is unaffected. +More consequentially, four other modules import that file — `managed_advisory.py` (13 references), +`advisory_consult.py` (8), `wclass_advisory.py` (3), `managed_cli.py` (1) — so extracting the +shared runtime decouples the `ask` path but does not by itself remove the 5,526-line file from +the wheel. Item B is cheaper than the rest of the move, not cheap in absolute terms. + +**A small-`n` observation was labelled as evidence.** The review presented cheap acceptance 2/3 +as `p ~ 0.33 < 0.69` and concluded the speculative route's economics are fine. Next Steps item 5 +already states that three observations are far too few to estimate `p`, and that moving the V1 +boundary needs a larger sample landing under 20%. The reviewer's operational conclusion — freeze +that work until verifier recall is known — is unaffected, but the evidence label was wrong and is +rejected. + +**The proposed decision rule uses the wrong quantile and flips branch under this project's own +convention.** The review fixed `z = 1.96` throughout and concluded that at `n = 25` a point +estimate of 0.60 rejects `m <= 0.40`. `speculative_report.py:798-843` deliberately uses a +`t` quantile at small `n`, states that a narrow interval can flip a break-even judgment, and +requires rounding to the conservative table entry. Recomputed: + +- 15/25 at `z = 1.96` gives a lower bound of 0.4074, which rejects. +- 15/25 at `t(df=24) = 2.060` gives 0.3983, which does not, and lands in the shortfall branch. +- Clearing 0.40 under the project's rule needs 16/25, a point estimate of 0.64. + +One reclassified instance changes the outcome. The review's own table reports the 18-point width +at `n = 25` and the decision rule then ignores it. + +**The proposed pilot kill gate is indexed to the wrong quantity.** It would declare the +phenomenon absent at fewer than 8 confirmed material defects in 200 runs, a 4% yield. The +original finding's own rate is six material defects across fourteen cheap-arm tasks, roughly 43%. +A floor an order of magnitude below the effect it is meant to detect cannot distinguish +reproduction from a tenfold weaker effect, and it is the same failure item C identifies in the +campaign gate. Any pilot must set this near the fixture's observed rate before it is run. + +### Status + +Nothing here is authorized. Next Steps items A-D remain the reviewed direction rather than a plan +of record, and this section does not promote them. It narrows item D: its strongest form is a +harvest pilot of roughly three to four days that writes almost no new code, with the two defects +above corrected first, and with the taxonomy document as the standalone deliverable that survives +either result. The reviewer's own recommendation was to build the pilot and nothing else until it +reports, and to keep the result out of the shipped tool in every branch. + +The next safe action is unchanged: ordinary explicit use of `$advisory` or `wclass-advisory ask`. + ## Scoped agent guidance (merged, unreleased) `AGENTS.md` is no longer one file. PR #181 merged at `57ae7ae` split it into a root @@ -2032,7 +2158,8 @@ councils, session consent, readable preview). The fifth is not: 5,463-line `speculative_run.py` (`run_child`, `extract_evidence_result`, `default_child_env`, `AGENT_SCAFFOLDING`, `CHILD_TIMEOUT`, `RunFailure`, `MAX_TASK_FILE_BYTES`). Extracting those into a small shared runtime module decouples the shipped path from the research runner and - makes the "vacuity anchors cannot move" constraint irrelevant to `ask`. + makes the "vacuity anchors cannot move" constraint irrelevant to `ask`. That symbol + count is one release stale; see the direction-research section above. - **C. The 60-task / 12-advised-failure gate is indexed to the wrong quantity.** The event it counts is an advised failure, so the cheap route succeeding — good product news — starves the study. That is why the population sits at 14/60 tasks while already at 9/12 advised failures.