Sudhanva Narayana · San Francisco Bay Area

Senior Machine Learning Engineer.

Production ML Systems · Infrastructure · Inference

I build production ML systems for high-throughput inference, reliable model serving, and scalable AI platforms.

Selected impact

Production outcomes in context

1B+ rows
Batch prediction across TensorFlow workloads

Processed end to end in under three hours for scientific model evaluation.

$50K+ annually
Cloud infrastructure cost reduction

Removed through resource optimization across production ML infrastructure.

50% faster
Model build and deployment

A multi-GPU build system and automated tuning also saved about ten hours each week.

1M+ daily
Interactions analyzed for next-click prediction

A high-throughput transformer served on a real-time event pipeline at Autodesk.

Professional evidence

Featured production ML case studies

01 1B+ rows

Billion-Row Batch Inference

A production TensorFlow prediction pipeline processed more than one billion rows end to end in under three hours, helping scientific teams evaluate candidates and iterate faster.

  • TensorFlow
  • Kubernetes
  • Ray
  • Flyte
Read the case study →
02 $50K+ annual cloud-cost reduction

Kubernetes ML Platform

An autoscaling Kubernetes platform with Ray, Flyte, observability, experiment tracking, and model versioning gave cross-functional teams an on-demand path to production ML workloads.

  • Kubernetes
  • Ray
  • Flyte
  • Prometheus
Read the case study →
03 50% shorter deployment time

Multi-GPU Deployment System

A multi-GPU build system, automated tuning, and CI/CD reduced model deployment time by 50% and saved about ten engineering hours each week.

  • Multi-GPU systems
  • TensorFlow
  • PyTorch
  • CI/CD
Read the case study →

Core engineering capabilities

Systems first, tools in service of outcomes

Production ML Platforms
Kubernetes, Ray, Flyte, Airflow, infrastructure as code, environment management, and observability that give models a repeatable path to production.
Inference and Serving
Batch inference, real-time prediction, multi-GPU execution, model serving, input-pipeline optimization, performance tuning, and deployment automation.
Reliability and Operations
Monitoring, alerting, safe rollouts, recovery, data validation, resource management, and cost control for long-running ML systems.
Scientific and Data-Intensive ML
Large-scale prediction, geospatial analytics, vector search, scientific workloads, data pipelines, experiment tracking, and model evaluation.
Explore the production ML systems approach →

Engineering notes

Selected production-focused writing

Open-source proof

Infrastructure labs and public repositories

Personal projects make the implementation inspectable without confusing lab work with employer production systems.

01 Docker 25

HTTPS Caddy + Tailscale

Zero-config HTTPS for self-hosted services

Docker Compose setup that pairs Tailscale with Caddy for automatic HTTPS on any service, even behind CGNAT.

caddy docker docker-compose tailscale
GitHub Project page →
02 HCL 8

K3s on Oracle Always Free

A real Kubernetes cluster on $0/month

Three-node K3s cluster on Oracle Cloud Always Free ARM64 with Argo CD, Gateway API, Envoy, and automatic HTTPS.

kubernetes k3s oracle-cloud terraform
GitHub Live Project page →
03 Shell 5

Bare-Metal Kubernetes Homelab

GitOps on physical hardware

Bare-metal Kubernetes on Ubuntu 24.04 with Ansible provisioning, ArgoCD GitOps, Longhorn storage, and Tailscale operator.

ansible argocd gitops homelab
GitHub Live Project page →

Frequently asked

Is Sudhanva open to new opportunities?

Open to conversations about thoughtful production ML, inference, and AI-platform engineering work. Reach out via email or LinkedIn.

What does Sudhanva specialize in?

Sudhanva specializes in production ML systems, infrastructure, and inference: large-scale batch prediction, model serving, multi-GPU execution, Kubernetes platforms, observability, reliability, and cloud cost management.

What production-scale systems has he built?

Published examples include a TensorFlow pipeline processing more than one billion rows in under three hours, an autoscaling Kubernetes ML platform, a multi-GPU model build system, a real-time transformer pipeline serving more than one million daily interactions, and multi-regional geospatial ML infrastructure.

Where is Sudhanva based?

Sudhanva is based in the San Francisco Bay Area and works remotely as a Senior Machine Learning Engineer at Montai Therapeutics.

What technologies does he use?

Core tools include Kubernetes, Ray Data and Serve, Flyte, TensorFlow, PyTorch, Terraform, Argo CD, Prometheus, Grafana, Python, SQL, AWS, GCP, and Azure. The technology choice follows the workload and operational requirements.

Where can someone see his work?

The Work page leads with production ML case studies and separates them from personal open-source projects. The Blog contains technical notes, while the About and HTML résumé pages provide the full professional timeline.

Search

Find public work and writing

Search case studies, production ML expertise, and technical notes. This server-rendered form is also exposed as a declarative browser tool for compatible agents.

Sudhanva developer resources for agents →

Contact

Open to conversations about thoughtful production ML, inference, and AI-platform engineering work.