Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
-
Updated
Aug 12, 2026 - Python
Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
Galactus executes 744B and 235B MoE models on undersized Macs with bit‑perfect llama.cpp parity, RAM-as-cache execution, and a full local app featuring an agent, permission gate, code editor, authenticated server mode, scheduled unattended runs, and fully published measurements.
Workload-aware persistent expert-tile scheduling for irregular MoE inference kernels on NVIDIA Blackwell GPUs
Running Kimi K3 2.8T on a 64GB Mac mini M4 Pro with Deltafin, Metal/MPS, and full-local SSD-backed MoE inference.
Simulates MoE expert placement across GPU memory, RAM, and NVMe.
Add a description, image, and links to the moe-inference topic page so that developers can more easily learn about it.
To associate your repository with the moe-inference topic, visit your repo's landing page and select "manage topics."