Add Medvedev-De Minaur break-point ledger task - #116
Conversation
| return previous[-1] | ||
|
|
||
|
|
||
| def layer_atoms(event, reference): |
There was a problem hiding this comment.
The first token’s stroke is contractually always serve, but layer_atoms currently awards it ordinary field credit. This is material: the scoreboard-only run’s 16 correct shot fields are exactly these 16
schema-guaranteed atoms, which open the bottleneck despite zero correct serve-direction or terminal-detail fields. The fixed-prior’s 23 shot hits are the same 16 constants plus seven guessed directions.Excluding only this constant atom regrades oracle/full/scoreboard/fixed-prior to approximately 1.0 / 0.2989 / 0.0 / 0.0449. Please score the first token on direction only, or otherwise exclude its constant
stroke field, and add the scoreboard-only and fixed-prior artifacts as regressions.
|
Closes #93
WIP status
This PR is intentionally submitted as a draft calibration candidate, not as a
conditions-satisfied task. The current official hierarchical score for GPT-5.6 Sol
is
0.3305, above the preferred<0.10gate. The oracle-assisted fixed-priordiagnostic is
0.1833, above the usual0.15ablation ceiling. Both are disclosedso the task can be reviewed before another hardening/calibration cycle.
Summary
Adds one new
agentic_vbench_understandingtask:tasks/agentic_vbench_understanding/medvedev-de-minaur-2023-us-open-break-point-ledger/It does not modify shared benchmark infrastructure or existing tasks.
Given the complete silent 2023 US Open fourth-round match between Daniil Medvedev
and Alex De Minaur, the agent reconstructs every break-point opportunity in
chronological order. Each event includes:
stroke/directionshot sequence.The provisional reference contains 16 break-point opportunities and 112 shot
tokens. Repeated break points in one service game remain separate events.
Media
The Dockerfile follows the direct pinned-release pattern used by the existing
GSW–Cleveland understanding task.
d78c9246d5dd36b812c71b5f39bfa43ab86d1a7adf711fb7bbc6ff1d66d618b2;d61bee17596a28dbc8f8b607e4fc0dd6542885dbe9cdad18e1953e89363b0860;685013111;214363 decoded frames, zero audio streams.
The hosted source contains commentary. The Docker build verifies its checksum and
copies only video stream
0:v:0, so the agent receives a silent artifact. The videobinary is not committed.
Ground truth and provenance
The official aggregate statistics fix 16 opportunities:
The current provisional identity, serve, rally, terminal, and shot labels are
reproducibly decoded from the pinned Match Charting Project record
20230904-M-US_Open-R16-Daniil_Medvedev-Alex_De_Minaurat commit2c59eef194967e688b69e73df344184a06322cd8.The instruction now defines observable criteria for
forced_error,unforced_error,error_unknown, ace/unreturnable, terminal position/error, servedirection, rally counting, and shot direction. MCP is treated as attributed
third-party annotation, not machine truth.
The target
human-verifiedtier is still pending. The PR includes a complete blindtwo-annotator packet/freeze/union/adjudication protocol and does not claim that the
two independent full-video ledgers already exist.
Official hierarchical scoring
The verifier is deterministic Python with no VLM/LLM judge. It implements
hierarchical-bottleneck-v1:stroke/directionatoms in order and divide by the larger predictedor reference shot-atom count.
The bottleneck requires both detail layers and prevents summary-only or shot-only
submissions from receiving event credit. It intentionally has plateaus: improving
the stronger layer does not change reward until it crosses the weaker layer. False,
duplicate, malformed, and extra events stay in the prediction denominator.
Chronological exact-event F1 remains a strict diagnostic only. Oracle reward is
1.0; empty reward is0.0.Reviewer feedback addressed
field-level scoring.
forced/unforced errors.
silent artifact without host-only media-analysis dependencies.
generation-time validation, current verifier details, and SHA manifests.
and
NO_SCOREharness failures following PR Add Dota 2 pre-death minimap trajectory task #84's artifact layout.Measured calibration
The full run aligned all 16 identities from 17 predictions, recovered 79/144
summary atoms and 136 ordered shot-field atoms, and had zero exact events. Its
hierarchical true-positive credit is
5.4539, precision0.3208, recall0.3409,and F1
0.3305.The original
.reward.json/.validation.jsonfiles remain unalteredgeneration-time exact-judge records. The new
*.hierarchical-verifier-details.jsonfiles are deterministic regrades of thefrozen submitted solutions. This distinction preserves provenance instead of
rewriting old validations as if they had used the new metric.
Anti-shortcut results and limits
0.0000.0.0000.0.1175.0.0222.0.1833(non-agent diagnostic,not passing).
The contact-sheet record proves zero calls observed in the retained session, not an
auditable backend
tools=[]/tool_choice=nonesetting. CLI image preprocessing anddownsampling also remain explicit limitations.
Cross-harness status
No numeric score is assigned to an invalid or unstarted trajectory:
inference.
terminal trajectory classification and are
NO_SCORE.Compact privacy-safe failure records contain only versions, failure stage/reason,
and sanitized-package hashes. Raw failure packages, machine-local paths, credentials,
and terminal reports are excluded.
Validation
1.0.0.0.skipped.
git diff --check: pass.sha256:390fb56051fafb49b5b4b797cb15704469294f816255994e9ee0fd21fe2da06b.strong-agent reward
0.3305is not<0.10.Open acceptance items
reward is below the required gate.
0.1833oracle-assisted fixed-prior shortcut diagnostic.claiming
human-verifiedground truth.replace it with auditable backend-disabled tools.