Blog #machine-learning #ai #portfolio
Real-Time Transformer Prediction at Autodesk
Notes from my 2022 machine learning internship at Autodesk, where I deployed a high-throughput transformer on a real-time event pipeline that analyzed more than one million user interactions a day for next-click prediction.
- Published
- Reading
- 1 min
- Author
- Sudhanva Narayana
In the summer of 2022, between the two years of my master’s at Northeastern, I worked as a Machine Learning Engineer Intern at Autodesk in San Francisco. My project was next-click prediction. I deployed a high-throughput transformer on a real-time event pipeline that analyzed more than one million user interactions every day.
The problem
Next-click prediction puts the model in the path of a high-volume event stream. Every interaction becomes an event, and the prediction has to come out of the same pipeline that carries those events.
The pipeline already existed and already served a product, which gave me two constraints. The prediction path had to stay compatible with the real-time event pipeline as it was, and the transformer’s throughput had to fit the operational needs of an existing product system.
What I did
I treated serving and event integration as part of delivering the model from the start. Most of the optimization work went into throughput at that volume. The transformer shipped on the production event path.
Real-time serving puts tighter latency and failure constraints on model execution than offline analysis does. A slow batch job finishes late, and a failed one can be rerun. On the event path the model runs as the events arrive, so it has to keep up with the stream.
What I took away
Model serving begins with the latency, throughput, and failure contract of the product path, and the model has to fit inside that contract. For real-time ML, production integration work matters as much as model selection.
The batch side of the same problem, billions of rows on a schedule rather than a live stream, is in the billion-row batch inference case study.
Related work
Production ML context
See how this topic connects to production ML systems, infrastructure, and inference.