Phase-by-phase benchmark of LLM instance cold start on RTX 2070: weight loading, PCIe transfer, CUDA init, KV cache allocation, and JIT warmup — with pre-warming savings and extrapolation to 7B+ models.
benchmarking cuda inference pytorch cold-start autoscaling pcie kv-cache llm-serving warm-pool mlsystems startup-latency weight-loading
-
Updated
Jul 19, 2026 - Python