Real-time Schema Discovery

Nearly every data-engineer has had to deal with schema related issues. Product makes a change, backend adds some fields, data envelopes change and now your pipelines need to be updated. This is a painful reality that most data engineers deal with on a constant basis and is a significant time-waste in every data engineering org. In this talk, I will show you how we developed a schema discovery process that is able to automatically evolve schemas in a complex distributed system that is processing upwards of a 100,000 messages per second. I will dive deep into the details of schema versioning, detecting schema conflicts, compatibility and normalization, all without the use of any batching processes. In this talk, I will show you how we developed a schema discovery process that is able to automatically evolve schemas in a complex distributed system that is processing upwards of a 100,000 messages per second. I will dive deep into the details of how to detect schema drift, how to determine compatibility and ultimately how to do all of this, without having to involve batching.