## Proposal Add **REFUTE** to related evaluation / scientific AI tooling docs if relevant. REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE). - Product: https://bgpt.pro/refute - Inspect task: https://github.com/connerlambden/refute-inspect - Interim preprint release: https://github.com/connerlambden/refute-inspect/releases/tag/v3.0.0-preprint Happy to adjust wording / PR if preferred.
Proposal
Add REFUTE to related evaluation / scientific AI tooling docs if relevant.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.