Reproducibility in Data Science
The abundance of data, coupled with cheap and widely-available computing and storage, has revolutionized science, industry and government alike. Now, to a large extent, the bottleneck to extracting actionable insights lies with people. Complex computational pipelines are required to ingest, clean, analyze, visualize and create models from data. But the process to assemble these is inherently iterative and time consuming. In addition, after a series of steps, there are many ways in which the computations, the data, and the analyst could have been wrong. Thus, when results are derived, an important question is whether you can trust them. In this talk, I will discuss the importance of computational provenance for data science and how it enables reproducibility, transparency, and helps build trust in results obtained from data-driven exploration. I will also present techniques and tools that support automatic provenance capture and simplify the reproducibility of computations.