I am a Site Reliability Engineer, working on rack-scale GPU fleet reliability, AI infrastructure and Kubernetes-based AI model serving.
I write at ezgitastan.systems
GPU & AI infrastructure GB300 NVL72 · B300 / B200 · H200 / H100 · NVLink / NVSwitch · InfiniBand · NVIDIA GPU Operator · MIG slicing · NFD · DCGM · NCCL · vLLM · CUDA · Redfish / IPMI
Orchestration & platform Kubernetes · OpenShift (ROSA) · Helm · Kustomize · ArgoCD · Slurm
Observability & reliability Prometheus · VictoriaMetrics · Grafana · eBPF / bpftime · Datadog · Sentry · Langfuse · K6 · PagerDuty
Cloud, IaC & automation AWS · GCP · Terraform · Terragrunt · SaltStack · Packer · Vagrant · Go · Bash
Datacenter & storage Ceph · MAAS · NetBox · libvirt / KVM
- Performance-regression detection for GPU fleets
- NVLink and NVSwitch fault isolation
- eBPF for GPU observability
- vllm-project/production-stack — least-privilege RBAC for Kubernetes secrets
- kubernetes-sigs/gateway-api-inference-extension — release-process fix
- eunomia-bpf/bpftime — config rename across 36 files of the userspace eBPF runtime
- dstackai/gpuhunt — Azure H100 NVL + H200 support
- aws-actions/amazon-ecs-deploy-task-definition — bounded exponential backoff in deployment waiters
- Also agentgateway, llm-d, bifrost, ratty, rustypaste → see all 13 merged
- Profiling Voice AI on H100 with eBPF — traced STT and TTS containers with bpftime: the GPU sat at 0% while the CPU ran 169K malloc/s
- Optimizing LLM Inference with Kubernetes and vLLM



