Vinoth Chandar

CEO & Founder, Onehouse

Vinoth Chandar is the original creator & VP of the Apache Hudi project, which pioneered transactional data lakes as we know it today, during his time as Uber's data architect. Vinoth has unique perspectives and deep experience with databases, distributed-systems and data systems at planet-scale, through his work at Oracle, Linkedin, Uber & Confluent, on systems like Oracle Streams, Voldemort, Apache Kafka/Streams, ksqlDB.

Vinoth Chandar

Sessions / 2022 / 1 talk

  • Data lakes and Data warehouses have co-existed for over a decade, but their paths appear to converge in the cloud. But, do they really? In this talk, you will learn how to build a future-proof data architecture, that leverages both. We discuss how the modern data architecture has evolved over the last decade, from the front row seats we have had as creators of Apache Hudi. We walk through the forces behind the move from on-prem data warehouses to Hadoop data lakes and their incarnations on the cloud, rise of cloud data warehouses and most recently the emergence of the Lakehouse technologies. We then present major lakehouse and cloud warehouse technology stacks, discuss their pros/cons, features and cost/performance tradeoffs. Finally, we present a practical approach to using an interoperable lakehouse as the bedrock of your data architecture, showing how it unlocks cost effectiveness, AI/ML ecosystem and massive scale, while still retaining your cloud warehouse for traditional BI/Analytics workloads.

Sessions / 2017 / 1 talk

  • Even after a decade, the name “Hadoop" remains synonymous with "big data”, even as new options for processing/querying (stream processing, in-memory analytics, interactive sql) and storage services (S3/Google Cloud/Azure) have emerged & unlocked new possibilities. However, the overall data architecture has become more complex with more moving parts and specialized systems, leading to duplication of data and strain on usability . In this talk, we argue that by adding some missing blocks to existing Hadoop stack, we are able to a provide similar capabilities right on top of Hadoop, at reduced cost and increased efficiency, greatly simplifying the overall architecture as well in the process. We will discuss the need for incremental processing primitives on Hadoop, motivating them with some real world problems from Uber. We will then introduce “Hoodie”, an open source spark library built at Uber, to enable faster data for petabyte scale data analytics and solve these problems. We will deep dive into the design & implementation of the system and discuss the core concepts around timeline consistency, tradeoffs between ingest speed & query performance. We contrast Hoodie with similar systems in the space, discuss how its deployed across Hadoop ecosystem at Uber and finally also share the technical direction ahead for the project.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker