Skip to content

fix: require n>=3 for stable benchmark rankings - #5

Merged
Hubujiu merged 5 commits into
mainfrom
fix/delivery-stable-ranking-gate
Aug 24, 2026
Merged

fix: require n>=3 for stable benchmark rankings#5
Hubujiu merged 5 commits into
mainfrom
fix/delivery-stable-ranking-gate

Conversation

@Hubujiu

@Hubujiu Hubujiu commented Aug 24, 2026

Copy link
Copy Markdown
Owner

Summary

  • add a machine-readable stability gate for benchmark artifacts
  • mark n=1/incomplete/infrastructure-failed evidence as provisional
  • add -RequireStableRanking to benchmarks/run.ps1
  • require production build evidence for stable Delivery rankings
  • add stability-gate regression tests and run benchmark harness tests in CI
  • document the exact n=3 Delivery rerun path

Why

The published v1.11 Delivery calibration currently has six differentiating cases at n=1: Practical v1.11, frozen v1.10, and Ponytail each build-pass 5/6. That is useful smoke evidence but insufficient for a stable ranking. The runner already defaults standard/full profiles to three repetitions, but nothing previously prevented an n=1 artifact from being presented as stable.

This change makes that evidence boundary executable without treating real behavioral/build failures as infrastructure failures.

@Hubujiu
Hubujiu merged commit 162175a into main Aug 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant