Billion-Row Batch Inference
A production TensorFlow prediction pipeline processed more than one billion rows end to end in under three hours, helping scientific teams evaluate candidates and iterate faster.
- TensorFlow
- Kubernetes
- Ray
- Flyte
Expertise
Models become products only when inference, infrastructure, observability, and operations work together. Sudhanva Narayana builds that production path for data-intensive and scientific ML workloads.
Production ML is the engineering system around a model: how data reaches it, how it runs at the required throughput, how versions are delivered, how failures are detected, and how teams reproduce an output. The system boundary includes data pipelines, model execution, orchestration, deployment, monitoring, and cost—not only training code.
Batch inference emphasizes input throughput, accelerator utilization, repeatability, and end-to-end completion time. Online inference and model serving add latency, availability, rollout, and failure-isolation constraints. Both depend on the same operational foundation: versioned workloads, observable infrastructure, safe resource limits, and a clear recovery path.
Kubernetes, Ray, and Flyte can provide a useful platform substrate, while TensorFlow and PyTorch provide model execution. Those technologies are means rather than the outcome. The outcome is a reliable service or pipeline that moves scientific and product teams from an experiment to a measured, repeatable production workflow.
Evidence
A production TensorFlow prediction pipeline processed more than one billion rows end to end in under three hours, helping scientific teams evaluate candidates and iterate faster.
An autoscaling Kubernetes platform with Ray, Flyte, observability, experiment tracking, and model versioning gave cross-functional teams an on-demand path to production ML workloads.
A multi-GPU build system, automated tuning, and CI/CD reduced model deployment time by 50% and saved about ten engineering hours each week.
At Autodesk, a high-throughput transformer ran on a real-time event pipeline that analyzed more than one million daily user interactions for next-click prediction.
At Pixxel, multi-regional ML infrastructure improved inference efficiency by 75%, reduced map-rendering latency by 50%, and supported ETL and ML pipelines processing more than 1 TB each day.
A production-safe walkthrough for migrating a standalone Milvus deployment from PVCs spread across availability zones to one dedicated Kubernetes node, while preserving vector data.
Machine Learning best practices and guidelines, tools to be used, while you're a developer. It applies to Data Analysts, Data Engineers, Machine Learning Engineers, Data Scientists or any research team in general
Install JupyterHub on AWS Elastic Kubernetes Service (EKS)
Review the professional timeline and résumé for context, or reach out about production inference, ML infrastructure, and AI-platform engineering work.