Production ML Platform on Kubernetes

An autoscaling Kubernetes platform with Ray, Flyte, observability, experiment tracking, and model versioning gave cross-functional teams an on-demand path to production ML workloads.

Annual cloud-cost reduction
$500K+
Time to insight
Days → minutes
How the production ML platform on Kubernetes serves model users Model users launch on-demand runs from internal apps; Flyte orchestrates autoscaling Ray workers on Kubernetes, with experiment tracking, model versioning, and alerting built into the platform. Model users (Cross-functional) flows to Internal apps (On-demand runs) [requests]; Internal apps (On-demand runs) flows to Flyte workflows (Orchestration) [workflows]; Flyte workflows (Orchestration) flows to Ray on Kubernetes (Autoscaling) [tasks]; Ray on Kubernetes (Autoscaling) flows to MLflow (Runs · versions) [runs]; Ray on Kubernetes (Autoscaling) drives Prometheus (Grafana · alerts) [metrics]. Owned by Sudhanva: Kubernetes ML platform. Outcomes: Annual cloud spend: Idle capacity to $500K+ saved; Time to insight: Days to Minutes. KUBERNETES ML PLATFORM CROSS-FUNCTIONAL Model users ON-DEMAND RUNS Internal apps ORCHESTRATION Flyte workflows AUTOSCALING Ray on Kubernetes RUNS · VERSIONS MLflow GRAFANA · ALERTS Prometheus requests workflows tasks runs metrics ANNUAL CLOUD SPEND Idle capacity $500K+ saved TIME TO INSIGHT Days Minutes Owned by Sudhanva Data Control
How the production ML platform on Kubernetes serves model users KUBERNETES ML PLATFORM CROSS-FUNCTIONAL Model users ON-DEMAND RUNS Internal apps ORCHESTRATION Flyte workflows AUTOSCALING Ray on Kubernetes RUNS · VERSIONS MLflow GRAFANA · ALERTS Prometheus requests workflows tasks runs metrics ANNUAL CLOUD SPEND Idle capacity $500K+ saved TIME TO INSIGHT Days Minutes Owned by Sudhanva Data Control
Model users launch on-demand runs from internal apps; Flyte orchestrates autoscaling Ray workers on Kubernetes, with experiment tracking, model versioning, and alerting built into the platform.

Confidentiality note. Employer-specific implementation details, internal names, and proprietary scientific context have been intentionally generalized. The scale, responsibilities, technologies, and outcomes below are limited to facts already published in Sudhanva’s résumé and professional profile.

Problem

Research and engineering teams needed a dependable shared path for running models, tracking experiments, and moving versioned workloads into production without rebuilding infrastructure for each project.

Scale and constraints

Scale

  • Shared production infrastructure supported cross-functional model users.
  • The platform hosted large batch prediction and multi-GPU workflows.
  • Resource optimization reduced annual cloud spend by more than $500,000.

Constraints

  • Support different workload shapes without turning every request into platform-engineering work.
  • Keep production operations observable and recoverable.
  • Control cloud cost while allowing workloads to scale on demand.

Sudhanva’s ownership

Sudhanva owned the end-to-end ML infrastructure and built internal applications on top of it.

  • Autoscaling Kubernetes infrastructure and workload orchestration.
  • Ray and Flyte platform integration.
  • Prometheus, Grafana, alerting, experiment tracking, and model versioning.
  • Documentation and training for cross-functional users.

Architecture at a safe level

This public architecture describes system responsibilities and leaves out internal service names and proprietary topology. The relevant boundaries were workload execution, orchestration, data movement, observability, and delivery.

  • Built a shared platform layer around orchestration, execution, observability, and lifecycle concerns.
  • Exposed common workflows through internal applications so users could run models on demand.
  • Used measured resource behavior to tune capacity and eliminate avoidable cloud spend.

Engineering decisions and tradeoffs

  • Shared abstractions had to remain flexible enough for research while preserving production guardrails.
  • Autoscaling reduced idle cost but required careful observability and workload-aware limits.

Results

  • Reduced cloud infrastructure costs by more than $500,000 annually.
  • Reduced time to insight for internal model users from days to minutes.
  • Enabled on-demand model runs with experiment tracking and model versioning.

Reliability and operations

  • Prometheus metrics, Grafana dashboards, and alerting made failures and capacity constraints visible.
  • Documented, versioned workflows replaced ad hoc execution paths.

Transferable lessons

  • A production ML platform succeeds when it removes repeated operational work for model users.
  • Build cost management and observability into the platform from the start.

Technologies

  • Kubernetes
  • Ray
  • Flyte
  • Prometheus
  • Grafana
  • MLflow
  • Terraform

Continue exploring

From the blog

Next case study · 2× release velocity Multi-GPU Delivery Platform