AIWG training-complete framework — corpus-to-dataset pipeline with SKILL.md agentic surface and optional Python runtime backend. Marketplace plugin for AIWG.
-
Updated
Apr 16, 2026 - Python
AIWG training-complete framework — corpus-to-dataset pipeline with SKILL.md agentic surface and optional Python runtime backend. Marketplace plugin for AIWG.
Reproducible code and metrics for measuring exploitability and contamination in LLM evaluation harnesses.
Do quality filters pull benchmark questions into pretraining corpora? Injects MMLU, GSM8K, GPQA, ARC, HellaSwag, PIQA and TruthfulQA items into a corpus and ranks them with six quality classifiers, including DCLM, FineWeb-Edu and Gaperon.
End-to-end Python research pipeline replicating "Beyond Benchmark Rankings": treats LLM selection as portfolio construction rather than leaderboard ranking. Estimates marginal utility, informational novelty, and correlated-failure risk per model, then runs budget-feasible greedy optimization to select LLM ensembles.
Detecting benchmark contamination by probing a language model's residual stream, with the controls that make the result mean something. Paper: arXiv:2608.12652
Add a description, image, and links to the benchmark-contamination topic page so that developers can more easily learn about it.
To associate your repository with the benchmark-contamination topic, visit your repo's landing page and select "manage topics."