Add primer-guided deck improvements and fix commander/import reliability - #85
Add primer-guided deck improvements and fix commander/import reliability#85LlamaAdam wants to merge 10 commits into
Conversation
40 community primers (21 Moxfield + 19 Archidekt, harvested 2026-08-29) distilled to per-deck structured records, shipped as package data with a fail-quiet, immutable, offline loader. Profiles-not-one-truth (§13): profiles_for_commander returns every build of a commander; win-line card names carry per-card mainboard-verification flags; budget swaps flatten into a §10 function-preserving swap corpus; prompt_block_for_commander renders clipped LLM context via primer.clip_for_prompt.
CONSISTENCY_TARGETS names the section-1 numbers that converged across independent primer authors (85% third land drop, 85% commander on curve, 90% card advantage by T5, the 13-enabler free-mulligan rule, at most 2 unconditionally-tapped fetchables) and evaluate_consistency_targets grades a deck against them, reusing the Monte-Carlo projection deck_health already computes plus closed-form hypergeometrics. Conditional floors (t1_enabler_plan / proactive_t2_plan) report their value but stay unjudged until a caller declares the plan. Wired as an additive-only deck_health tile under the standard outage contract; the health grade keeps ignoring it.
Renders met/evaluated with per-check tooltip, mirroring the consistency tile's three-state null handling. Display-only; the letter grade ignores it. Key-set pins in the audit tests extended for the additive consistency_targets key.
…e classifier contextual_role_targets adjusts ROLE_TARGETS by archetype, commander role, average mana value, and bracket, per the primer synthesis (Edgar aggro spec, Gishath resolve-engine deltas, Winota trigger-density trade, bracket-scaled interaction floors). One documented delta table is the tuning surface; unknown labels degrade to the flat table. role_target_report gains a keyword-only context= that is byte-identical to the old report when absent. infer_commander_role classifies only the regex-trustworthy roles (cost cheater, trigger multiplier, resolve engine by mana value) and returns None when unsure.
Four narrow, oracle-signature modifiers from the primer synthesis: commander_dependence (-8, effect gated on controlling the commander), tempo_fail (-6, value a turn-cycle late under an aggro plan), capped_engine (-4, printed once-each-turn limiter), tutor_top_delta (-5, Vampiric-class top-of-library tutors outside combo shells). Patterns are deliberately conservative — a miss costs nothing, a false positive would down-rank a fine card on every audit. Same flag and validation gate as the rest of FP-015: these refine the ranking prior, they do not enable it.
nonbo_lint encodes the primer synthesis section-14 table as data-driven pairwise checks (Skullclamp vs own anthems, Cursed Totem vs own dorks, forced draw vs Thassa's Oracle, Heartless Summoning vs one-toughness creatures, mode-exclusive Asceticism/Everlasting Torment, own shroud vs own targeting, symmetric-effect notes) — the first anti-synergy detector in the tree. Selectors match by exact name or oracle/type predicates; a card never pairs with itself; unresolvable cards match nothing. Wired as the additive-only 'nonbos' deck_health tile with the standard degrade contract.
…, budget swaps build_judge_prompt gains a COMMUNITY PRIMER CONTEXT block (clipped, attention-steering only, identical across A/B orderings) when the bundled KB covers the deck's commander. The Claude advisor payload carries community_primer_consensus under the same omit-when-absent policy as bracket peers, and its system prompt gains the distilled card-evaluation principles. Budget mode surfaces author-documented function-preserving swaps for the deck's own commander as paired add/cut recommendations (supplemental source, same post-processing as lift picks, fail-quiet).
LlamaAdam
left a comment
There was a problem hiding this comment.
Round-3 adversarial review of this PR (maximum-effort pass; every item below was verified against source with file:line and executed reproductions, then cross-examined from the defense side — the cross-examiner's corrected version is what's stated here). Full report with evidence: docs/ollama-analysis/NEGATIVE_MODE_ROUND3.md on LlamaAdam/mtga-advisor#5, §1a and §6.
Owner decisions this PR needs before merge (recorded as R3-D1/R3-D2 in DECISIONS_FOR_REVIEW.md on that PR):
- Auto-Protect reversal (PR-06, major). This PR removes FP-018.3's primer card-link auto-protection and inverts the test that pinned it (
tests/test_adopt.py:239-250). The rationale in the commit is sound engineering — which is exactly why it's a product trade for the owner, not a bug fix; no decision record or review exists for it. - Knowledge-base provenance (PR-05 B3).
data/primer_kb.json,consistency_targets.py,staples.py,card_score.py,nonbo_lint.pyand_advisor_claude.pyciteprimer_harvest/deckbuilding_heuristics.md§1–§16 and two harvest JSONs that exist in no tree, branch or commit; 35 of the KB's 40 decks have no list or prose anywhere in the repo. Not a code defect — a question of whether an unreproducible data asset ships. - Duplicate
/api/deck_commander(PR-08, major). #84 adds the same endpoint with a different contract;git merge-treeof the two conflicts in 7 files / 18 hunks, and both declare the same top-levellets inapp.js(a SyntaxError for the bundle if both land). Each PR ismergeable_state: cleanalone, so GitHub warns nobody.
Fixes for this PR, in priority order:
- PR-02 (major) — the import "hardening" now 400s whole pastes master accepted:
Maybeboard/Commanders/Tokensheadings and Archidekt's default[Commander{top}]tail. Recognise them innormalize_card_line/deck_text_opsrather than rejecting. - PR-03 (major) —
quoted_win_linestreats any paragraph mentioning "maybeboard" / "Cons:" / "Updates" as a section heading and silences every later paragraph; one in-sentence mention blanks a primer's win lines. Heading heuristic in_primer_tokensmust not fire on in-sentence words. - PR-05(A) (major) — the KB budget-swap table emits non-card strings as advisor adds (14/36 rows, e.g.
"Kamahl, Heart of Krosa / End-Raze Forerunners","$200-tier builds");_safe_ci_lookupskips rather than drops them, so they reach output. Validate rows at load; pin every shipped row as one resolvable card name. Also PR-17key_cardsparses to empty for all 40 profiles (data shape ≠ loader), PR-18Bruce Banner // The Hulk (gift build)commander key, PR-S2 Ur-Dragon "d10 tokens" (d20). - PR-07 (major) —
_primer_kb_blockputs community-consensus card names into the judge prompt, with noprompt_versiononJudgeReport, so the Phase-1 agreement study now pools two prompts with no column to split them — the exact G3 confound the scope doc pre-registered against. Stamp a prompt version; render the block without card names or gate it off by default. - PR-09 (major) — partner decks structurally never reach
legalon the dashboard (only the hero commander gets the guarded lookup); fetch the ≤2 command-zone cards the same way. - PR-04 (major) — three of the six FP-019 slices (context-aware roles/quotas, plan-conditional floors, nonbo advisor pre-filter) have zero production callers; either wire them or reword the PR body/CHANGELOG to the
future-plans.mdSTATUS line. - PR-12 / PR-21 — the commander editor and
import_deckaccept"2 Krenko, Mob Boss"as a card name (executed: the import route writes1 1 Krenko, Mob Bossand returns 200); the editor also demotes aProtect=-locked card without touching the lock. Reject^\d+\snames; warn on a locked choice. - PR-01 (minor after cross-exam) — "tests isolate Forge paths" is true for the corpus-reading paths but 13 modules bind
VENDOR_FORGE-derived constants at import time, so the autouse fixture cannot reach them; resolveDECK_DIR-style constants at call time and softenCHANGELOG.md:32. PR-13 reportnormalized_lines; PR-20main_countchanged meaning silently; PR-S1 the Asceticism/Everlasting Torment nonbo "why" states the wrong mechanism; PR-14 five modules now exceed the documented 800-line ceiling; PR-19 the "independent backend/Python/frontend reviews", wheel-build and manual-smoke claims have no artifact (GitHub shows zero reviews) — label or drop. - Mirrors of master findings this PR carries: F-02
adopt.py:165-166compares card-link names (Front // Back) against front-face.dcknames, so every linked DFC is reported "NOT in the list"; W-03 a UTF-8 BOM should be accepted, not 400'd; W-07 thecreate_appconsumer of an unvalidateddeck_dir; W-10 PUT line endings.
What held on Linux: the "4,590 passed / 26 Playwright" numbers are substantiated by this PR's own CI; the --run-live gate resists -m/-k/node-id selection; the deck-dir precedence table (8×3 combinations) matches the description; the desktop-lock test isolation is a real fix.
Generated by Claude Code
Summary
--run-liveconsent for real-service tests, independent of slow-test selection.Verification
--run-slow): 4,590 passed, 1 skipped in 636.85 seconds. The sole skip is the explicit live-service opt-in; all offline slow tests ran. Existing positional-maxsplitdeprecation warnings remain.Limits
The manual/UI tests use stub card data and simulation reports. No live Forge gameplay or successful model-service request is claimed. The initial full run exposed a legacy live-Claude test missing its optional SDK (4,582 other tests passed); the live lane now requires an explicit opt-in and configured optional dependencies. User deck libraries and raw primer harvests were not modified.
See
docs/negative-mode-audit-2026-09-01.mdfor the original findings and resolution details. This PR includes the nine existing FP-019 commits; it does not merge or deploy them.