Split Learning: A Resource Efficient Distributed Deep Learning Method without Sensitive Data Sharing
Collaboration in health is heavily impeded by lack of trust, data sharing regulations and limited consent of patients. In settings where different institutions hold different modalities of patient data in the form of electronic health records (EHR), picture archiving and communication systems (PACS) for radiology and other imaging data, pathology test results, or other sensitive data such as genetic markers for disease, collaborative training of distributed machine learning models without any data sharing or leakage of patterns about raw data is desired. In addition the solution needs to be resource efficient in terms of communication bandwidth, computations and memory. This talk is primarily about a recently developed, highly resource efficient method called 'Split Learning' for this very purpose by allowing to perform distributed deep learning under these constraints.
