Skip to content

Qwen35 397b GRPO ruleset and rcps - #468

Merged
ShriyaRishab merged 4 commits into
mlcommons:masterfrom
jepio:qwen35_397b_grpo_ruleset
Jul 27, 2026
Merged

Qwen35 397b GRPO ruleset and rcps#468
ShriyaRishab merged 4 commits into
mlcommons:masterfrom
jepio:qwen35_397b_grpo_ruleset

Conversation

@jepio

@jepio jepio commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

The ruleset requires evaluation to start at a fixed step and happen every step after that until target accuracy is reached.

GBS val start step val start samples
256 18 4608
512 10 5120
1024 7 7168

jepio added 3 commits July 24, 2026 17:41
Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
@jepio
jepio requested review from a team as code owners July 24, 2026 15:47
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

Comment thread mlperf_logging/compliance_checker/training_6.1.0/closed_qwen35_397b_grpo.yaml Outdated
Comment thread mlperf_logging/compliance_checker/training_6.1.0/closed_qwen35_397b_grpo.yaml Outdated
@ShriyaRishab
ShriyaRishab merged commit d2afcac into mlcommons:master Jul 27, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants