-
Updated
Jul 4, 2026 - C++
#
flash-decoding
Here are 2 public repositories matching this topic...
amd hip gpu-computing rocm cpp20 bitnet llm-inference flash-decoding strix-halo gfx1151 1-58-bit ternary-llm
Benchmark comparing Flash-Decoding-style chunked KV-parallel attention against standard SDPA at decode time (q_len=1) on RTX 2070, measuring the crossover point where KV-parallelism outperforms fused attention kernels.
benchmarking cuda pytorch transformer decode attention kv-cache long-context llm-inference flash-attention flash-decoding mlsystems
-
Updated
Jul 19, 2026 - Python
Improve this page
Add a description, image, and links to the flash-decoding topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the flash-decoding topic, visit your repo's landing page and select "manage topics."