AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
-
Updated
Sep 9, 2026 - Rust
AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
Prefix-Aware Attention for LLM Decoding
Proxima lets existing GPUs serve 4x more concurrent requests
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
A curated list of plugins built on top of vLLM
A honest port of vLLM-ROCm for windows.
Spark-plugin for Spark2.5 model, enables seamless loading and serving Spar2.5 large language model within vLLM framework. Implements custom model plugin to support weights loading, tokenizer and inference runtime without modifying original vLLM source code.
独立、可单独安装的 vLLM KV-cache 池 + 空闲队列可视化插件。 与具体缓存方案解耦,自动适配: 三区 (ThreePhaseBlockQueue, vllm-kv-cache-plugin) — Cold / Warm / Hot 双区 (TwoPhaseBlockQueue, kv_cache_affinity) — Aged / Fresh 原生 vLLM 队列兜底 — 单 Free 区
A manager to load vllm plugins without rebuilding image for each new plugin.
vLLM Plugins for additional features like decoding strategies, monitoring, models etc
A focused Windows runtime and build setup for running vLLM on an AMD Radeon RX 9070 XT with native ROCm/HIP. The project uses the OpenAI-compatible vLLM server and does not require WSL, Docker, Linux, or a virtual machine.
Exact dynamic output budgeting for Hermes Agent
Declarative K3s stack serving vLLM on AMD ROCm with a Streamlit frontend
To associate your repository with the vllm-plugins topic, visit your repo's landing page and select "manage topics."