One Mac process. Many models. Real speed. Multi-model LLM serving with prefix reuse, MTP acceleration, and OpenAI APIs — built for Apple Silicon, measured against mlx-lm and llama.cpp.
macos rust ocr metal embeddings multi-model gemma mlx model-serving multimodal on-device-ai apple-silicon openai-api llm local-llm qwen speculative-decoding nemotron holo3 orinth
-
Updated
Sep 10, 2026 - Rust