-
Notifications
You must be signed in to change notification settings - Fork 0
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
docs(readme): refresh the benchmark table — accumulated drift (stale numbers + descriptions) found during #440
maintenanceRecurring upkeep / sync with upstreamRecurring upkeep / sync with upstreamStatus: Open.#442 In pekkah/SharpInference;perf(cuda): DSpark GPU draft round runs ~3× the bandwidth floor — small-M (B=7) projection GEMMs → weight-stationary matvec (#428 lever 1)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#441 In pekkah/SharpInference;- Status: Open.#437 In pekkah/SharpInference;
GPU Lloyd-Max TqAttention writes compressed-region V aggregate in the rotated basis (no inverse WHT)
Status: Open.#435 In pekkah/SharpInference;KVarN follow-ups: int8-TC prefill GEMM, split-tile decode kernel, minors (#180 successor)
enhancementNew feature or requestNew feature or requestStatus: Open.#433 In pekkah/SharpInference;perf(cuda): CUDA-graph-captured k-token spec verify — the ~10.6 ms fixed verify overhead is now DSpark's whole gap (#428 lever 2)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#430 In pekkah/SharpInference;- Status: Open.#427 In pekkah/SharpInference;
perf(cuda): dense int8-MMQ prefill ~2–2.6× behind llama.cpp (cp.async MMQ + flash attn) — Qwen3-8B / Gemma4
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#409 In pekkah/SharpInference;perf(cuda): full-offload MoE decode ~3.4× behind llama.cpp — OLMoE 126 vs 426 t/s (on-GPU experts, kernel-bound)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#408 In pekkah/SharpInference;perf(cuda): batched int8-MMQ trunk prefill gated off for dense NORM-RoPE & QKV-bias models — SmolLM2 105×, VibeThinker/Qwen2 159×
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#407 In pekkah/SharpInference;perf(cuda): batched int8-MMQ trunk prefill gated off for MoE — OLMoE prefill 97 vs llama.cpp 17,137 t/s (176×)
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#406 In pekkah/SharpInference;perf(cuda): README vs llama.cpp CUDA benchmark gaps — full results + tracking
perfPerformance optimization opportunityPerformance optimization opportunityStatus: Open.#405 In pekkah/SharpInference;