-
Notifications
You must be signed in to change notification settings - Fork 74
Pull requests: intel/llm-scaler
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Omni: add Windows Torch 2.13 CUTE and SolAttn
#659
opened Sep 1, 2026 by
xiangyuT
Contributor
Loading…
omni: INT4 per-block zero-point GEMM + calling adapter (torchao asymmetric INT4)
#658
opened Aug 31, 2026 by
JWLHS
Loading…
docker: upgrade production vLLM image to 0.26.0
#655
opened Aug 31, 2026 by
gc-fu
Contributor
Loading…
omni: W4A8 s8u4 oneDNN GEMM + s8 activation quantizer (per-block zero point)
#629
opened Aug 19, 2026 by
JWLHS
Loading…
docs(vllm): Qwen3.8-27B DFlash drafter recipe — 72.2 tok/s on Arc Pro B70
#620
opened Aug 16, 2026 by
rmacy
Loading…
docs(omni): add HiDream-O1, Krea2 Turbo, Boogu-Image, Mage-Flow and LTX-2.3 to supported models table
#592
opened Aug 4, 2026 by
KristianZeng
Contributor
Loading…
[XPU] Harden fused GDN recurrent state updates
#570
opened Jul 27, 2026 by
gc-fu
Contributor
Loading…
Fix #533: guard out_proj.weight access in GDN out-projection ESIMD probes
#557
opened Jul 22, 2026 by
joaovgaraujo
Loading…
refactor(omni): multi-stage Docker build with decoupled builder/runtime
#549
opened Jul 17, 2026 by
KristianZeng
Contributor
•
Draft
page_attn: make global-max reduction order-invariant
#535
opened Jul 12, 2026 by
taste-software
Loading…
Fix Qwen3.5/3.6 load_weights stacked-mapping name mutation (gate_gate_up_proj / qkqkv_proj)
#475
opened Jun 12, 2026 by
bongmiin
Loading…
vllm: add MiniCPM-V 4.6 support (MiniCPMV4_6ForConditionalGeneration)
#472
opened Jun 12, 2026 by
Zjq9409
Loading…
Add Lunar Lake Xe2 iGPU compatibility report and benchmarks
#342
opened Apr 1, 2026 by
MegaStood
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.