Skip to content
View EzgiTastan's full-sized avatar
◻️
Learning
◻️
Learning

Block or report EzgiTastan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
EzgiTastan/README.md

Hi, I'm Ezgi Taştan

I am a Site Reliability Engineer, working on rack-scale GPU fleet reliability, AI infrastructure and Kubernetes-based AI model serving.

I write at ezgitastan.systems

Skills

GPU & AI infrastructure GB300 NVL72 · B300 / B200 · H200 / H100 · NVLink / NVSwitch · InfiniBand · NVIDIA GPU Operator · MIG slicing · NFD · DCGM · NCCL · vLLM · CUDA · Redfish / IPMI

Orchestration & platform Kubernetes · OpenShift (ROSA) · Helm · Kustomize · ArgoCD · Slurm

Observability & reliability Prometheus · VictoriaMetrics · Grafana · eBPF / bpftime · Datadog · Sentry · Langfuse · K6 · PagerDuty

Cloud, IaC & automation AWS · GCP · Terraform · Terragrunt · SaltStack · Packer · Vagrant · Go · Bash

Datacenter & storage Ceph · MAAS · NetBox · libvirt / KVM

What I'm working on

  • Performance-regression detection for GPU fleets
  • NVLink and NVSwitch fault isolation
  • eBPF for GPU observability

Merged upstream

Writing

Pinned Loading

  1. gateway-api-inference-extension gateway-api-inference-extension Public

    Forked from kubernetes-sigs/gateway-api-inference-extension

    Gateway API Inference Extension

    Go

  2. llm-d llm-d Public

    Forked from llm-d/llm-d

    Achieve state of the art inference performance with modern accelerators on Kubernetes

    Shell