Skip to content

Judge calibration: 40 human labels, weighted kappa 0.53 adequacy / 0.27 fluency / 0.21 localization#7

Merged
0Smallcat0 merged 1 commit into
mainfrom
feat/judge-calibration
Jul 10, 2026
Merged

Judge calibration: 40 human labels, weighted kappa 0.53 adequacy / 0.27 fluency / 0.21 localization#7
0Smallcat0 merged 1 commit into
mainfrom
feat/judge-calibration

Conversation

@0Smallcat0

Copy link
Copy Markdown
Owner

Closes the calibration loop from #6: 40 blind human ratings over the seeded stratified sample (eval/results/labeling.html), computed by pnpm bench:agreement.

Axis Raw agreement Cohen k Weighted k (quadratic)
adequacy 62.5% 0.339 0.526
fluency 60.0% 0.191 0.267
localization 47.5% 0.104 0.213

Interpretation shipped into docs/BENCHMARK.md / README: the qwen3.5 judge is moderately trustworthy on adequacy, weak on fluency and Taiwan-localization (systematically more lenient than a Taiwanese reader), so the report leans on chrF + adequacy and reads the judge localization column as an upper bound. Calibration told us which parts of the judge to trust - that is the point of the exercise.

eval/dataset/human-labels.json committed for reproducibility; eval/AGREEMENT.md generated.

🤖 Generated with Claude Code

…3/0.27/0.21)

40 blind human labels over the seeded stratified sample. Quadratic-weighted
Cohen's kappa vs the qwen3.5 judge: adequacy 0.526 (moderate - usable),
fluency 0.267 and localization 0.213 (weak - judge systematically more
lenient than a Taiwanese reader on terminology). Docs updated so quality
claims lean on chrF + adequacy and treat judge localization as upper bound.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@0Smallcat0
0Smallcat0 merged commit 9cefc74 into main Jul 10, 2026
2 checks passed
@0Smallcat0
0Smallcat0 deleted the feat/judge-calibration branch July 10, 2026 03:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant