Billion-Row Batch Inference
A production TensorFlow prediction pipeline processed more than one billion rows end to end in under three hours, helping scientific teams evaluate candidates and iterate faster.
- TensorFlow
- Kubernetes
- Ray
- Flyte
Sudhanva Narayana · San Francisco Bay Area
Production ML Systems · Infrastructure · Inference
I build production ML systems for high-throughput inference, reliable model serving, and scalable AI platforms.
Selected impact
Processed end to end in under three hours for scientific model evaluation.
Removed through resource optimization across production ML infrastructure.
A multi-GPU build system and automated tuning also saved about ten hours each week.
A high-throughput transformer served on a real-time event pipeline at Autodesk.
Professional evidence
A production TensorFlow prediction pipeline processed more than one billion rows end to end in under three hours, helping scientific teams evaluate candidates and iterate faster.
An autoscaling Kubernetes platform with Ray, Flyte, observability, experiment tracking, and model versioning gave cross-functional teams an on-demand path to production ML workloads.
A multi-GPU build system, automated tuning, and CI/CD reduced model deployment time by 50% and saved about ten engineering hours each week.
Core engineering capabilities
Engineering notes
A production-safe walkthrough for migrating a standalone Milvus deployment from PVCs spread across availability zones to one dedicated Kubernetes node, while preserving vector data.
Machine Learning best practices and guidelines, tools to be used, while you're a developer. It applies to Data Analysts, Data Engineers, Machine Learning Engineers, Data Scientists or any research team in general
CoGeoTIFF Research on Space Tech SaaS Platform at Pixxel
Open-source proof
Personal projects make the implementation inspectable without confusing lab work with employer production systems.
Zero-config HTTPS for self-hosted services
Docker Compose setup that pairs Tailscale with Caddy for automatic HTTPS on any service, even behind CGNAT.
Sudhanva specializes in production ML systems, infrastructure, and inference: large-scale batch prediction, model serving, multi-GPU execution, Kubernetes platforms, observability, reliability, and cloud cost management.
Published examples include a TensorFlow pipeline processing more than one billion rows in under three hours, an autoscaling Kubernetes ML platform, a multi-GPU model build system, a real-time transformer pipeline serving more than one million daily interactions, and multi-regional geospatial ML infrastructure.
Sudhanva is based in the San Francisco Bay Area and works remotely as a Senior Machine Learning Engineer at Montai Therapeutics.
Core tools include Kubernetes, Ray Data and Serve, Flyte, TensorFlow, PyTorch, Terraform, Argo CD, Prometheus, Grafana, Python, SQL, AWS, GCP, and Azure. The technology choice follows the workload and operational requirements.
The Work page leads with production ML case studies and separates them from personal open-source projects. The Blog contains technical notes, while the About and HTML résumé pages provide the full professional timeline.
Search
Search case studies, production ML expertise, and technical notes. This server-rendered form is also exposed as a declarative browser tool for compatible agents.
Sudhanva developer resources for agents →Contact