Skip to content
View VibeHPC's full-sized avatar

Block or report VibeHPC

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
VibeHPC/README.md
VibeHPC

AI Systems, Engineered for High Performance

VibeHPC builds efficient AI systems through customized models,
high-performance GPU kernels, and end-to-end inference optimization.

Website | Research | Contact


What we work on

  • Models - customized LLMs and AI agents designed for real workloads.
  • Kernels - high-performance GPU kernels built through low-level optimization.
  • Systems - end-to-end inference systems that turn hardware capability into application performance.

Projects

LLM Inference

End-to-end optimization for MLPerf Inference: Datacenter v6.1, with a focus on higher throughput, lower latency, and efficient serving.

Explore LLM Inference

VibeGEMM

An open leaderboard for high-performance GPU matrix multiplication, comparing kernels across hardware, numerical precisions, and matrix shapes.

Explore VibeGEMM

Attend to Fragments

Research into how key information affects large language models for factual inconsistency detection.

View the repository | Read the research blog

Research

We study AI performance across the full stack, from model behavior and GPU kernels to system-level inference optimization. Our work combines empirical research with practical engineering.

Explore VibeHPC Research

Work with us

We welcome conversations about research collaboration, high-performance AI engineering, benchmarking, and the VibeHPC projects.

Email: office@vibehpc.com

Popular repositories Loading

  1. VibeGEMM VibeGEMM Public

    The open-source leaderboard for high-performance GPU GEMM.

    Cuda 3 1

  2. attend-to-fragments attend-to-fragments Public

    Attend to Fragments: How Key Information Affects Large Language Models for Factual Inconsistency Detection

    Python 2

  3. VibeHPC VibeHPC Public

    Engineering high-performance AI systems across models, GPU kernels, and end-to-end LLM inference.