Skip to content
View bricesommers's full-sized avatar

Block or report bricesommers

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. apus-deepseek-v4-flash apus-deepseek-v4-flash Public

    Local inference engine for DeepSeek-V4-Flash (284B MoE, MXFP4 experts streamed from disk) on consumer hardware — C11, zero deps, bit-exactness gated. macOS / Linux / Windows.

    C 3 2

  2. nanoGPT2 nanoGPT2 Public

    The simplest repository for fine-tuning GPT-2 into a chatbot

    Python 1 1

  3. docs docs Public

  4. NVIDIA-locateanything NVIDIA-locateanything Public

    Run NVIDIA LocateAnything-3B natively on Apple Silicon with MLX — local vision-language object detection with bounding boxes, single-image and batch modes. ~33 tok/s on M1, 16 GB RAM, zero cloud. N…

    Python

  5. apus-qwen3.6-35B-A3B apus-qwen3.6-35B-A3B Public

    Local inference engine for Qwen3.6-35B-A3B (35B-total / 3B-active hybrid-linear MoE, experts streamed from NVMe) on consumer hardware — C11, zero deps, bit-exactness gated. Runs on 16 GB RAM. macOS…

    C

  6. apus-glm5.3-flash apus-glm5.3-flash Public

    Local inference engine for GLM-5.3-Flash (320B MoE, glm5_next) on consumer hardware — runs the full 306 GiB FP8 model on 32 GB RAM by streaming experts from NVMe. Bitwise-exact C11 engine (NEON/AVX…

    C