Julien Le Dem

Principal Engineer, Datadog

Julien Le Dem is a Principal Engineer at Datadog, serves as an officer of the ASF and is a member of the LFAI&Data Technical Advisory Council. He co-created the Parquet, Arrow and OpenLineage open source projects and is involved in several others. His career leadership began in Data Platforms at Yahoo! - where he received his Hadoop initiation - then continued at Twitter, Dremio and WeWork. He then co-founded Datakin (acquired by Astronomer) to solve Data Observability. His French accent makes his talks particularly attractive.

Julien Le Dem

Sessions / 2026 / 1 talk

  • Datadog has grown from a startup focused on infrastructure monitoring into a platform processing over a hundred trillion events daily. Over the years, we expanded beyond metrics to include traces, logs, profiling, real user monitoring, and security. Our user base has broadened from operations to developers, analysts, and business users. More recently, automated agents have also become key consumers of our platform. Our bottom-up culture encourages teams to take initiative. Consequently, we developed multiple specialized ingestion pipelines and query engines. These were built to satisfy strict real-time requirements and interactive experience, providing users with the insights necessary for success. This focus on efficiency led to custom-built, proprietary solutions designed for our unique constraints. Today, however, the evolving landscape allows us to reconcile these specialized engines with open standards, blending the versatility of the ecosystem with our purpose-built designs. In recent years, we have refactored the interfaces of these query engines to create a composable data system. This allows us to better leverage shared capabilities, enabling cross-dataset querying, advanced analytics, and more versatile access patterns. Our goal is to scale our bottom up culture. By defining clear contracts and high-level components, we enable decentralized decision-making. This improves performance, efficiency, and flexibility across the platform, while reducing silos. By adopting a deconstructed stack, we combine the efficiency of the open-source ecosystem with our internal capabilities to build a truly composable system. This architecture provides the flexibility to adapt to immediate and future demands, specifically addressing requirements for scale, velocity, and operational resilience, while ensuring readiness for growing challenges such as data intensive operations like AI. In this talk, we will discuss how we rely on and contribute to key projects in the data ecosystem: Arrow for data interchange, Substrait for plans, Calcite as an optimizer, DataFusion as an execution core, and Parquet for columnar storage.

Sessions / 2025 / 1 talk

  • In 2018, Julien Le Dem described how the components of databases, distributed or not, were being commoditized as individual parts that anyone could recombine into use-case-specific engines. Given one's constraints, they could leverage those components to build a query engine that solves a specific problem much faster than building everything from the ground up. He called this idea "the Deconstructed Database" and spoke about it at a previous edition of Data Council. Fast forward to today, the big data ecosystem has matured and evolved from a melting pot of competing projects into a more composable ecosystem organized around a few open source standards. It's been incredible to see the vision he outlined in his talk crystallize with the adoption of key components like Parquet, Arrow, Iceberg, Calcite, Substrait and OpenLineage. These tools, and others like them, provide an interoperability layer that enables harnessing data for many purposes without creating silos. In this talk, Julien will discuss the impact of the cloud and the advent of the open data lake, breaking silos to form the foundation of this ecosystem. As compute and storage can be efficiently decoupled, a common storage layer enables a vibrant ecosystem of on-demand tools specialized to specific use cases that avoid vendor lock-in. He'll go over the core components, how they work together and more importantly, the contracts that keep them decoupled and composable.

Sessions / 2024 / 1 talk

  • Explore the journey of building open-source standards with insights from Julien Le Dem, covering the birth of Parquet, transition to low-latency systems and the development of Arrow. Learn how community-driven collaboration paved the way for groundbreaking innovations in data processing. Tune in now to delve into the dynamic landscape of data engineering over the past decade.

Sessions / 2023 / 1 talk

  • Over the last decade I have been lucky enough to contribute a few successful open source projects to the data ecosystem. In this talk I’ll share the story of how these projects came to be and what made their success possible. I’ll describe the ideation process and early growth of the Apache Parquet columnar format and show how this led to the creation of its in-memory alter-ego Apache Arrow. I’ll end with showing how this experience enabled the success of OpenLineage, an LFAI & Data project that brings observability to the data ecosystem. Along the way, I’ll talk about the key elements that catalyzed their growth, from project focus to governance to community.

Sessions / 2022 / 1 talk

  • As workflows increase in complexity, companies have come to depend on Airflow to manage inter-DAG dependencies. Airflow has quickly become an important component of the Modern Data Stack powering analytical reports, business metrics, and dashboards. But what effects (if any) would upstream DAGs have on downstream DAGs if dataset consumption were delayed? What alerting rules should be in place to notify downstream DAGs of possible upstream processing issues or failures? How can we use data lineage to achieve the data observability we need to answer these questions? In this talk, OpenLineage will be introduced, an open standard for collecting lineage metadata for jobs under execution, and how it works with Airflow. The presentation will walk through a practical example using Marquez, the reference implementation of OpenLineage. It will be explained how OpenLineage can help data teams maintain inter-DAG dependencies within their Airflow instance, capture metadata on historical DAG runs, and minimize data quality issues.

Sessions / 2018 / 1 talk

  • Over the past ten years the Big Data infrastructure has evolved from flat files lying down in a distributed file system to a more efficient ecosystem and is turning into a fully deconstructed database. With Hadoop, we started from a system that was good at looking for a needle in a haystack using snowplows. We had a lot of horsepower and scalability but lacked the subtlety and efficiency of relational databases. Since Hadoop provided the ultimate flexibility compared to the more constrained and rigid RDBMSes we didn’t mind and plowed through with confidence. Machine Learning, Recommendations, Matching, Abuse detection and in general data driven products require a more flexible infrastructure. Over time we started applying everything that had been known to the Database world for decades to this new environment. They told us loud enough how Hadoop was a huge step backwards. And they were right in some way. The key difference was the flexibility of it all. There are many highly integrated components in a relational database and decoupling them took some time. Today we see the emergence of key components (optimizer, columnar storage, in-memory representation, table abstraction, batch and streaming execution) as standards that provide the glue between the options available to process, analyze and learn from our data. We’ve been deconstructing the tightly integrated Relational database into flexible reusable open source components. Storage, compute, multi-tenancy, batch or streaming execution are all decoupled and can be modified independently to fit every use case. This talk will go over key open source components of the Big Data ecosystem (including Apache Calcite, Parquet, Arrow, Avro, Kafka, Batch and Streaming systems) and will describe how they all relate to each other and make our Big Data ecosystem more of a database and less of a file system. Parquet is the columnar data layout to optimize data at rest for querying. Arrow is the in-memory representation for maximum throughput execution and overhead-free data exchange. Calcite is the optimizer to make the most of our infrastructure capabilities. We’ll discuss the emerging components that are still missing or haven’t become standard yet to fully materialize the transformation to an extremely flexible database that lets you innovate with your data.

Sessions / 2016 / 1 talk

  • In pursuit of speed and efficiency, big data processing is continuing its logical evolution toward columnar execution. Julien Le Dem offers a glimpse into the future of column-oriented data processing with Arrow and Parquet. A number of key big data technologies have or will soon have in-memory columnar capabilities. This includes Kudu, Ibis, Drill and many others. Modern CPUs will achieve higher throughput using SIMD instructions and vectorization on Apache Arrow’s columnar in-memory representation. Similarly Apache Parquet will provide storage and I/O optimized columnar data access using statistics and appropriate encodings. For interoperability, row-based encodings (CSV, Thrift, Avro) combined with general-purpose compression algorithms (GZip, LZO, Snappy) are common but inefficient. Julien explains why the Arrow and Parquet Apache projects define standard columnar representations that allow interoperability without the usual cost of serialization. This solid foundation for a shared columnar representation across the big data ecosystem promises great things for the future. Julien discusses the future of columnar data processing and the hardware trends it can take advantage of. Arrow-based interconnection between the various big data tools (SQL, UDFs, machine learning, big data frameworks, etc.) will allow using them together seamlessly and efficiently without overhead. When collocated on the same processing node, read-only shared memory and IPC avoid communication overhead; when remote, scatter-gather I/O sends the memory representation directly to the socket, avoiding serialization costs; and soon RDMA will allow exposing data remotely.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker