Popular repositories Loading
-
apus-deepseek-v4-flash
apus-deepseek-v4-flash PublicLocal inference engine for DeepSeek-V4-Flash (284B MoE, MXFP4 experts streamed from disk) on consumer hardware — C11, zero deps, bit-exactness gated. macOS / Linux / Windows.
-
-
NVIDIA-locateanything
NVIDIA-locateanything PublicRun NVIDIA LocateAnything-3B natively on Apple Silicon with MLX — local vision-language object detection with bounding boxes, single-image and batch modes. ~33 tok/s on M1, 16 GB RAM, zero cloud. N…
Python
-
apus-qwen3.6-35B-A3B
apus-qwen3.6-35B-A3B PublicLocal inference engine for Qwen3.6-35B-A3B (35B-total / 3B-active hybrid-linear MoE, experts streamed from NVMe) on consumer hardware — C11, zero deps, bit-exactness gated. Runs on 16 GB RAM. macOS…
C
-
apus-glm5.3-flash
apus-glm5.3-flash PublicLocal inference engine for GLM-5.3-Flash (320B MoE, glm5_next) on consumer hardware — runs the full 306 GiB FP8 model on 32 GB RAM by streaming experts from NVMe. Bitwise-exact C11 engine (NEON/AVX…
C
If the problem persists, check the GitHub status page or contact support.