Lattice and trellis-coded quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
-
Updated
Aug 19, 2026 - Python
Lattice and trellis-coded quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
Add a description, image, and links to the tcq topic page so that developers can more easily learn about it.
To associate your repository with the tcq topic, visit your repo's landing page and select "manage topics."