The Trade-off between Strict Validation and Accepting Anything

When building a data pipeline, we need to decide if we should strictly validate incoming data, and discard anything that we don't support, or if we should be flexible, and accept anything so we can analyze it later. In this talk, I'll discuss how the compromise we've reached at Bluecore, where we both record the "raw" data to recover from bugs or mistakes, as well as strictly validated data. I'll talk about why we think that validating up front is the better choice when building data intensive applications.