Case Studies from a Methodologist on an Experimentation Platform
Microsoft's experimentation platform team aims to enable users to make trustworthy decisions in A/B tests. In case studies, I’ll describe how we use statistical evaluation and simulation frameworks in meeting this trustworthiness promise and making the right platform decisions for our user needs. How should we choose an analysis window and inclusion criteria for a single-step vs. ramp-up policy for controlled feature rollout, and is there one “right” choice? How did we ground the performance of our platform’s variance reduction (VR) estimator and understand cases of poor efficacy? How did we evaluate the complexity vs. efficacy tradeoffs of ML-assisted VR techniques head-to-head with simpler approaches?
