AI researcher working at the intersection of LLM post-training and inference systems — training agents with RL, then making them run fast.
- Agentic RL & post-training — sync vs. fully-async multi-turn tool-calling GRPO on verl; diffusion RL post-training (Flow-GRPO on Qwen-Image) with a CPU OCR reward at 6.9× headroom
- Inference systems — hand-built MoE expert parallelism (fused grouped-GEMM Triton kernels, 1.28×), INT8 KV-cache quantization (0.50× memory, near-lossless PPL), CUDA pipeline-bubble profiling on sm_120
- Open source — contributing to SGLang and vLLM
- Research — conditional injection in diffusion transformers for compositional 3D scene generation
Seeking PhD / RA opportunities in LLM alignment & systems (Fall 2027).
Training: PyTorch · verl · HuggingFace Transformers · PEFT/LoRA Systems: CUDA · Triton · vLLM · SGLang · Nsight Systems/Compute Dev: Python · Docker · Next.js · PostgreSQL
- M.S. Engineering Sciences & Applied Mathematics — Northwestern University (2025)
- B.S. Applied Mathematics — University of California, San Diego (2024)
✉️ samuelwang997@gmail.com · 🌐 chaoyuwang.vercel.app · 🤗 huggingface.co/SamWang0405
