Continuous Data Pipeline for Real-time Benchmarking & Data Set Augmentation
Building and curating representative datasets during the model research and development is a critical component of getting a ML system with accuracy meeting the project's requirements. After the deployment of said model, monitoring accuracy and other statistical metrics in order to improve and adjust the model is a natural workflow. Models working with unstructured language data might experience data shift resulting in unpredictable and non-representative inference. With the help of open-source APIs and commercial or open-source annotation tools the building of annotations can be operationalized and the analyst workload reduced. In this talk I will cover the process of generating datasets and using them for real time precision/recall splits with the goal of detecting data shifts away from the in-sample space to prioritize future data collection and model retraining.
