VibeHPC builds efficient AI systems through customized models,
high-performance GPU kernels, and end-to-end inference optimization.
- Models - customized LLMs and AI agents designed for real workloads.
- Kernels - high-performance GPU kernels built through low-level optimization.
- Systems - end-to-end inference systems that turn hardware capability into application performance.
End-to-end optimization for MLPerf Inference: Datacenter v6.1, with a focus on higher throughput, lower latency, and efficient serving.
An open leaderboard for high-performance GPU matrix multiplication, comparing kernels across hardware, numerical precisions, and matrix shapes.
Research into how key information affects large language models for factual inconsistency detection.
View the repository | Read the research blog
We study AI performance across the full stack, from model behavior and GPU kernels to system-level inference optimization. Our work combines empirical research with practical engineering.
We welcome conversations about research collaboration, high-performance AI engineering, benchmarking, and the VibeHPC projects.
Email: office@vibehpc.com