Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
-
Updated
Aug 22, 2026 - Python
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
Prompt-only boundary prediction for IFEval-style instruction-checker pass/fail behavior.
Task-dependent benchmark gap between /v1/completions and /v1/chat/completions on instruction-tuned LLMs -- two case studies on Qwen3.x GGUFs, reproduction recipe, and probes.
🎯 现代化大语言模型标准自动化评测平台 (LLMEval) | 支持开源/商业模型 · 五维能力雷达图 · Bad Case 逐题审查 · SQLite 历史持久化 · Gradio Web & CLI
Parameter-efficient instruction tuning of TinyLlama-1.1B with LoRA, evaluated on IFEval.
LoRA-based parameter-efficient fine-tuning for TinyLlama instruction following and ALBERT IMDb sentiment classification.
Add a description, image, and links to the ifeval topic page so that developers can more easily learn about it.
To associate your repository with the ifeval topic, visit your repo's landing page and select "manage topics."