Headless LLM fine-tuning in 3 lines — smart defaults, VRAM-aware batch sizing, multi-run SLAO, GGUF export for Ollama.
-
Updated
Jul 17, 2026 - Python
Headless LLM fine-tuning in 3 lines — smart defaults, VRAM-aware batch sizing, multi-run SLAO, GGUF export for Ollama.
End-to-end RLHF pipeline with reward debiasing, DPO vs SimPO comparison, and statistical significance testing on Anthropic hh-rlhf dataset.
Domain-specific benchmark for B2B sales agents — 250 tasks, SimPO judge model, published on HuggingFace.
Full post-training pipeline for Qwen2.5-1.5B — SFT → SimPO → GRPO on free T4/P100 GPUs. GSM8K accuracy jumps from 23% (base) to 61% (GRPO) using Unsloth 4-bit LoRA, TRL, and HuggingFace Hub checkpointing.
Add a description, image, and links to the simpo topic page so that developers can more easily learn about it.
To associate your repository with the simpo topic, visit your repo's landing page and select "manage topics."