ML OPs & Platforms

Scalable and Sustainable Feature Engineering with Hamilton

In this talk we present Hamilton, a novel open-source framework for developing and maintaining scalable feature engineering dataflows. Hamilton was initially built to solve the problem of managing a codebase of transforms on pandas dataframes, enabling a data science team to scale the capabilities they offer with the complexity of their business. Since then, it has grown to be an end-to-end tool for writing, managing, and iterating on Machine Learning pipelines. We introduce the framework, discuss its motivations and initial successes at Stitch Fix, showcase its lightweight data lineage and catalog abilities, and share recent extensions that seamlessly integrate it with distributed compute offerings, such as Dask, Ray, and Spark.

Speakers