CS + Math at UT Austin. I build performance-critical ML infrastructure across inference serving, compilers, distributed systems, and online learning.
|
Inference deployment compiler + adaptive runtime Profiles deployment candidates, searches constrained configurations, adapts online, and validates behavior through replay, fault injection, and machine-readable evidence.
|
Outcome-decoupled speculative LLM serving Per-request scheduling, fission and re-coalescing, an exact sampling oracle, transactional KV state, and reproducible CPU experiments.
|
|
End-to-end int8 inference accelerator stack A custom ISA, cycle-approximate simulator, MLIR compiler, post-training quantization frontend, and reproducible model execution.
|
Correctness infrastructure for online-RL serving Failure-atomic publication, causal attestation, compact proofs, crash recovery, and formal TLA+ models for continuously updated serving.
|
|
Hybrid inference control plane Typed routing, admission, fallback, canary, and evidence semantics across local and remote model backends, with deterministic replay.
|
Correctness-first KV-state transport Content-addressed objects, Python/Rust validation, LMCache remote storage, fault injection, and replayable benchmarks.
|
LLM serving · ML systems · compilers · distributed systems · performance engineering
Currently working across production infrastructure and ML systems research.


