DevOps & Cloud Engineer | Kubernetes, GPU & LLM Infrastructure
I design, deploy, and operate scalable cloud-native systems with a focus on Kubernetes platforms, GPU workloads, and production LLM inference. My work spans architecture, infrastructure as code, CI/CD, observability, and reliability across Azure, AWS, and hybrid environments.
My current focus is building practical infrastructure for distributed AI workloads: running Ray and vLLM on Kubernetes, managing GPU capacity and MIG-based workloads, improving inference observability, and automating secure, repeatable delivery pipelines.
- Kubernetes platform engineering across AKS, RKE2, on-premises, and hybrid environments
- GPU infrastructure for distributed AI and LLM inference workloads
- Ray, KubeRay, Ray Serve, and vLLM deployment and operations
- Azure cloud architecture, Azure DevOps, and self-hosted CI/CD agents
- Infrastructure as code with Terraform and Helm-based application delivery
- Production observability with Prometheus, Grafana, and NVIDIA DCGM Exporter
- Secure, scalable storage and API platforms for AI and data workflows
| Project | Description | Technologies |
|---|---|---|
| s3BEAR | A secure, self-hosted S3 gateway and web console for AWS S3, MinIO, and other S3-compatible backends. It includes Microsoft Entra SSO, group-based access control, audit logging, presigned URLs, multipart uploads, scheduled cleanup, and secure image/file serving for AI and LLM workflows. | FastAPI, async SQLAlchemy, PostgreSQL, React, Vite, Ant Design, boto3, Docker, Helm, S3, MinIO |
| Ray-LLM-Dashboard-Grafana | An LLM observability pack for Ray Serve and vLLM workloads. It brings together TTFT, TPOT, token throughput, request queues, KV-cache usage, replica health, autoscaling, and GPU capacity metrics in Grafana. | Ray Serve, vLLM, Kubernetes, Prometheus, Grafana, NVIDIA DCGM Exporter |
| Kubernetes_as_Azure_Agent | A KEDA-based platform for running self-hosted Azure DevOps agents on AKS. Agent pods scale from zero according to pipeline demand and are deployed through Helm, Kubernetes manifests, and container images. | AKS, Kubernetes, KEDA, Azure DevOps, Helm, Docker, ACR, Go, Node.js, Python |
| MLOPS-Terraform | Infrastructure automation for an Azure Machine Learning workspace, including a blue-green deployment approach. | Azure Machine Learning, Terraform |
| Area | Technologies |
|---|---|
| Cloud & Platform | Kubernetes, AKS, RKE2, Azure, AWS, Terraform, Helm, KEDA, Azure Container Registry |
| AI & LLM Infrastructure | Ray, KubeRay, Ray Serve, vLLM, NVIDIA GPUs, MIG, NVIDIA Device Plugin, NVIDIA Container Toolkit |
| Observability | Prometheus, Grafana, kube-prometheus-stack, NVIDIA DCGM Exporter, PodMonitor, ServiceMonitor |
| Backend & APIs | Python, FastAPI, asynchronous services, REST APIs, gRPC, React, NGINX, APISIX |
| Data & Storage | PostgreSQL, S3-compatible storage, AWS S3, MinIO, Azure Blob Storage, Azure Data Lake |
| Delivery & Automation | Azure DevOps, GitHub Actions, Docker, containerd, CI/CD pipelines, shell scripting |
I write about real production challenges rather than purely theoretical examples. My recent topics include operating Ray and KubeRay on Kubernetes, building Grafana and Prometheus observability for LLM inference, deploying vLLM on NVIDIA GPUs, configuring MIG on H200 NVL hardware, running GPU workloads on bare-metal RKE2, migrating application storage to S3, and scaling Azure DevOps agents with KEDA.
The certifications currently visible on my public LinkedIn profile include CKAD: Certified Kubernetes Application Developer, GitHub Actions, and GitHub Administration. My certification history also includes Microsoft Certified: Azure Fundamentals, Microsoft Certified: Azure Administrator Associate, Microsoft Certified: Azure Developer Associate, and Microsoft Certified: DevOps Engineer Expert. Certification validity and renewal status are maintained on LinkedIn.