Experiments with self-improving ML pipelines — active-learning curation and incremental LoRA fine-tuning.
lora active-learning ewc fine-tuning continual-learning mlops llm rlhf llmops self-improving-ai data-flywheel proxy-reward-model
-
Updated
May 28, 2026 - Python