I'm an SRE / Cloud Platform Engineer focused on LLM inference platforms, GPU-enabled Kubernetes, API gateways, observability, data systems and reliable multi-cloud infrastructure.
My current work is centered on production-grade AI platforms: deploying vLLM behind OpenAI-compatible APIs, operating NVIDIA GPU workloads on Kubernetes, designing gateway patterns for inference traffic, and making the platform observable enough to debug latency, throughput, saturation and cost.
I also bring a strong background in cloud, DevOps, cybersecurity, data, and operations across Azure, Google Cloud and AWS.
- LLM serving platforms with vLLM, OpenAI-compatible APIs and model routing.
- GPU Kubernetes platforms with NVIDIA GPU Operator, DCGM metrics and dedicated node pools.
- API gateway patterns for authentication, rate limits, tenant isolation, model versioning and traffic control.
- Observability systems with Grafana, Prometheus/Mimir, Loki, Tempo, OpenTelemetry and Grafana Alloy.
- Cloud platforms using Terraform, Helm, Argo CD, GitOps and automation.
- Data and AI infrastructure for APIs, pipelines, retrieval patterns and operational analytics.
Public engineering blueprint for production-grade LLM deployment patterns on Kubernetes.
- vLLM deployment manifests for OpenAI-compatible model serving.
- NVIDIA GPU Operator values, GPU node scheduling and validation scripts.
- Gateway API routing for inference endpoints and long-running generation requests.
- Prometheus rules for queue depth, request latency, time to first token, token throughput and KV cache pressure.
- Runbooks for GPU readiness, pod scheduling, model health and gateway troubleshooting.
- Multi-cloud planning for AKS, GKE and EKS GPU platforms.
Repo: vllm-gpu-kubernetes-blueprint
Public cloud-native platform lab for Azure Kubernetes Service, GitOps delivery, separated app and observability stacks, and Grafana Cloud telemetry.
- Terraform-managed Azure infrastructure.
- Argo CD app-of-apps delivery model.
- Helm-based platform components.
- Grafana Alloy pipelines for metrics, logs, events and traces.
- Prometheus/Mimir, Loki, Tempo and OpenTelemetry integrations.
- Dashboards, scripts, runbooks and troubleshooting workflows.
Repo: aks-argocd-alloy
Public portfolio website focused on LLM inference platform engineering.
- Full static GitHub Pages site.
- Hero visual for AI/GPU platform operations.
- Sections for vLLM, GPU Kubernetes, gateways, observability, cloud, data and security.
Site: alecscodehub.github.io
Repo: alecscodehub.github.io
LLM serving: vLLM, OpenAI-compatible APIs, inference gateways, model routing, batching, latency and token-throughput analysis.
GPU platforms: NVIDIA GPU Operator, Kubernetes device plugin, DCGM Exporter, CUDA runtime, GPU scheduling, node pools and capacity planning.
Kubernetes and delivery: AKS, GKE, EKS, Helm, Argo CD, GitOps, ingress, Gateway API, autoscaling and release workflows.
Cloud and infrastructure: Azure, Google Cloud, AWS, Terraform, Bash, identity, networking, secret management and automation.
Observability: Grafana, Prometheus/Mimir, Loki, Tempo, OpenTelemetry, Grafana Alloy, dashboards, alerts and runbooks.
Data and security: SQL, Python, APIs, data pipelines, operational analytics, cloud security, secrets and platform hardening.
- Building portfolio-grade LLM platform projects that show real infrastructure depth.
- Making vLLM deployments operable on GPU Kubernetes clusters.
- Connecting API gateways, observability and data platforms around inference workloads.
- Improving public engineering evidence through focused repos, runbooks and deployment documentation.
Verified public Credly credential catalog: 117 badges.
| Type | Count |
|---|---|
| Certification | 18 |
| Learning | 65 |
| Validation | 21 |
| Other / uncategorized | 13 |
View all 117 Credly credentials
Full verification profile: credly.com/users/alexcuriman/badges
- Portfolio: alecscodehub.github.io
- Credly: alexcuriman verified badges
- LinkedIn: alexcuriman
- GitHub: alecscodehub

