Mark Grover

CEO & Co-Founder, Stemma

Mark is the co-founder/CEO of Stemma - a modern data catalog for building self-serve data culture used by Grafana, iRobot, SoFi, Convoy and many others. He is the co-creator of the leading open-source data catalog, Amundsen, used by Lyft, Instacart, Square, ING, Snap and many more! ​Mark was previously a developer on Apache Spark at Cloudera and is a committer and PMC member on a few open-source Apache project. He is a co-author of Hadoop Application Architectures.

Mark Grover

Sessions / 2023 / 1 talk

  • At Lyft, Mark built the Amundsen data catalog so data scientists could navigate hundreds of thousands of tables to distinguish trustworthy data from sandboxed, out-of-date data. When he took Amundsen open source, he helped dozens of data teams support a variety of demands to make data discoverable and self-serve. Time and again, Mark sees processes that seem “good enough” come back to bite data teams. For example, because it doesn’t directly involve creating the canonical data set, sending a Slack blast about an upcoming change feels like an adequate effort to warn users. But, key stakeholders frequently miss the message, the change still causes pain, and angry fingers still point back at the data team. Fortunately, there is a better way. With a little bit of inside knowledge, you can see who is actually using your data so you can better serve them as a customer. So roll up your sleeves because Mark is going to take you deep into query logs and APIs to see where all of that metadata lives, and he'll show you how to use it so you don’t lose any fingers during your next data change.

Sessions / 2020 / 1 talk

  • At Lyft, like many other organizations, analysts and data scientists were spending more than 1/3rd of their time discovering and establishing trust in the data they use. Lyft has made its analysts and data scientists over 20% more productive by creating an open source data discovery and metadata engine, Amundsen. In this talk, we will deep dive into the product and architecture of Amundsen and discuss how Amundsen leverages centralized metadata, page rank, and a comprehensive data graph to achieve its goal. Square has been leveraging Amundsen for a different use case. We will share the product Square has built on top of Amundsen to power more granular column-level access control for its data lakes, including Snowflake and BigQuery. We will provide an overview of how Square is tagging additional metadata to understand the data subjects, data storage security, and PII semantic types associated with columns, and uses this enriched metadata to drive purpose-driven access control for its data users. The talk will end with an insight into current challenges and how we may solve them in the future.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker