High-performance Model Serving in Python, using BentoML
Model serving and deployment is the last mile to success in any machine learning project. However building the entire MLOps workflow for serving is challenging. ML teams often end up with complex and inefficient solutions, using tools that are not designed for this problem. BentoML is the open source framework for model serving, offering a simple workflow towards high performance model serving. In this talk, we will explore the key challenges in building a modern ML model serving stack and how BentoML could help.
