Scaling Data Products Under Startup Constraints: A Case Study of ML Bias Testing
Modern machine learning can be incredibly powerful, but the reality is that even the best systems are far from perfect. Unlike traditional software with deterministic outcomes, ML services operate in probabilities and can return different results over time. To demonstrate this, I will showcase the approach we took to testing 4 commercial ML systems for gender bias at our small startup. By startup necessity, we had to scale our blackbox testing with practical automation. By building tools that allowed us to rapidly generate new data sets, perform queries, and version control results, we were able to find large categories of images with gender labelling errors.
