Run Qwen3.6-35B-A3B full native 258K context on 12GB VRAM (llama.cpp): ncmoe cliff rule, q4_0 KV prefill fix, 4 tuned profiles with scripts
-
Updated
Aug 11, 2026 - PowerShell
Run Qwen3.6-35B-A3B full native 258K context on 12GB VRAM (llama.cpp): ncmoe cliff rule, q4_0 KV prefill fix, 4 tuned profiles with scripts
Run Qwen3.6-35B LLM + ComfyUI SDXL/Pony image gen simultaneously on 12GB VRAM — measured config (36-37 tok/s during render), scripts included
Add a description, image, and links to the rtx4070-super topic page so that developers can more easily learn about it.
To associate your repository with the rtx4070-super topic, visit your repo's landing page and select "manage topics."