Reproducible refusal-subspace editing and TP=2 deployment of DeepSeek V4 Flash 0731 on two DGX Sparks
-
Updated
Aug 11, 2026 - Python
Reproducible refusal-subspace editing and TP=2 deployment of DeepSeek V4 Flash 0731 on two DGX Sparks
Production-oriented Qwen3.6-35B-A3B-NVFP4-Fast vLLM deployment for NVIDIA DGX Spark / GB10
DeepSeek-V4-Flash + DSpark speculative decoding on a pair of NVIDIA DGX Sparks (vLLM TP=2 over RoCE) — tuned recipe, overlays that halve multi-turn TTFT, contamination-guarded benchmarks, ops runbook
Measured Muse Glimmer 30B recipe for NVIDIA DGX Spark, with llama.cpp, DFlash parity checks, vision, tools, and GB10 benchmarks.
Benchmark for GB10 - Nvidia DGX Spark
Reproducible high-speed Qwen3.6-35B-A3B inference on NVIDIA GB10
Production-ready local AI agent stack for NVIDIA DGX Spark / ASUS GB10. Gemma 4 31B NVFP4 + bge-m3 embeddings + OpenClaw gateway, one docker compose up.
Auditable DeepSeek V4 Flash inference evidence on two NVIDIA GB10 systems
Auditable DeepSeek V4 Flash inference evidence on two NVIDIA GB10 systems
Empirical NVIDIA GB10 and SM121 microarchitecture characterization
Add a description, image, and links to the nvidia-gb10 topic page so that developers can more easily learn about it.
To associate your repository with the nvidia-gb10 topic, visit your repo's landing page and select "manage topics."