Praneeth Vepakomma

Researcher, Massachusetts Institute of Technology

Praneeth Vepakomma is currently a grad student and researcher at MIT in Camera Culture research group where his focus is on developing algorithms to support distributed and collaborative machine learning. He was previously a scientist at Amazon, Motorola Solutions and at various startups all of which were successfully acquired.

Praneeth Vepakomma

Sessions / 2019 / 2 talks

  • Collaboration in health is heavily impeded by lack of trust, data sharing regulations and limited consent of patients. In settings where different institutions hold different modalities of patient data in the form of electronic health records (EHR), picture archiving and communication systems (PACS) for radiology and other imaging data, pathology test results, or other sensitive data such as genetic markers for disease, collaborative training of distributed machine learning models without any data sharing or leakage of patterns about raw data is desired. In addition the solution needs to be resource efficient in terms of communication bandwidth, computations and memory. This talk is primarily about a recently developed, highly resource efficient method called 'Split Learning' for this very purpose by allowing to perform distributed deep learning under these constraints.

  • Personally Identifiable Information (PII) is piling up in databases and on filesystems across the globe. Smart companies are hard at work generating insights from this data, while World-dominating companies are intentionally generating it, mining it and in various ways obtaining clear value from it. GDPR and HIPPA are game-changing government regulations affecting data storage and transmission, while, in the meantime, advances in machine and deep Learning are powering huge leaps in analytical insights and business innovation. In addition, an unlevel playing field exists between the sheer size of the data accumulated at the biggest tech cos vs. the nimbleness and inspiration of the smallest startups. Yes, companies of all sizes are competing with each other in an attempt to add significant value to their users. At best, large datasets represent the bedrock for meaningful consumer insights; value-added customer features, services, and products; and massive amounts of rich training data to increase model efficiency. At worst, new systems, algorithms and data architectures represent a plethora of nefarious new opportunities to de-anonymize, leak or blatantly distribute data that was previously secret and/or obfuscated. So in this brave new world of data and algos and regulations what are the privacy concerns surrounding data access and security? Our panelists will explore these issues, from the hands-on perspective of building some of the most sophisticated data mining systems in the world. They are all hands-on technologists - data scientists, engineers, researchers and technical founders - and will share from their deep experience in building massively scalable data systems. They will also help us contemplate the thorny issues of technical and ethical responsibility - issues essential to consider as we all work together to build the data-driven systems of the present, and the future.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker