Serverless-GPU LLM serving: scale-to-zero with fast GPU snapshot/restore (cuda-checkpoint), multi-tenant packing, and an OpenAI-compatible API — built on vLLM.
serverless gpu cuda lora autoscaling scale-to-zero openai-api llm-serving vllm llm-inference snapshot-restore cuda-checkpoint
-
Updated
Jul 5, 2026 - Python