AI Council 2018

A look back on 2018

The talks that shaped AI Council 2018.

Data Council Barcelona 2018 — Talks

All 2018 Talks

Every session from 2018 — filter by topic, speaker, or company.

· Talk

Yara is the world’s leading fertilizer company, is headquartered in Oslo, Norway, and has more than 16,000 employees worldwide. Our mission is to responsibly feed the world while respecting the planet. To this end, the digital transformation in agriculture will allow a drastically more precise use of crop nutrition products, conceivably down to the single plant. Yara Digital Labs is shaping these future tools to help farmers achieve their most ambitious goals reliably and with ease. Tapping into Yara’s extensive agronomic knowledge, we collect, transform and integrate current and historical datasets while building new tools from them. Our vision is a world free from hunger – a difficult task in the face of steady population growth and changing environmental conditions. Saving water - especially in farming - will be one of the huge challenges to overcome in the near future: In areas which have struggled with this in the past years, the conditions have become increasingly severe, and even in areas where irrigation has not been of such concern in the past, the farmers are often not well prepared for dealing with extensive droughts. In his talk, Richard will present the Yara Water Solution, a hardware based system for irrigation monitoring and -optimization in tree orchards, as an example for operating an analytics back end based on R. On the development side, the focus is on close collaboration with agronomists and maintaining continuously high data quality. On the operations side, the main emphasis is on compatibility and stability, as well as using the new resources within the context of Yara Digital Labs.

Using R in a Mid-Sized Data Analysis Scenario
· Talk

Making data available to each employee is important for companies to speed up their decision making. However, there are some challenges to achieve it. Controlling each employee's data access strictly is the most important thing to make data public within a company securely. In addition, a data analysis platform must be scalable and stable so that many employees can access and analyze data simultaneously and smoothly. LINE Corporation, which is a communication platform provider based in Tokyo, Japan, has succeeded in overcoming these challenges by creating a brand-new web-based data analysis platform named "OASIS" since this year. In the talk, 1) motivation for creating OASIS, 2) its features and system architecture, and 3) its use cases at LINE Corporation are going to be described in detail.

OASIS – Data Analysis Platform for Enterprise
· Talk

GDPR is officially here, and technical organizations have made significant adjustments to support it. While there are still many legal and policy questions & concerns surrounding its use, technical managers have been forced to make sense of the regulation and create their own specific implementations. Come and hear from several EU CTOs who have grappled with early implementations at their own companies, and discover the main technical challenges and considerations they they faced. GDPR "Bill of Rights": The rights of the user/client (referred to as “data subject” in the regulation) that are relevant for developers are: The right to erasure (the right to be forgotten/deleted from the system) Right to restriction of processing (you still keep the data, but mark it as “restricted” and don’t touch it without further consent by the user) The right to data portability (the ability to export one’s data in a machine-readable format) The right to rectification (the ability to get personal data fixed) The right to be informed (getting human-readable information, rather than long terms and conditions) The right of access (the user should be able to see all the data you have about them).

GDPR: Discover The Main Challenges & Considerations
· Talk

Traditional data architectures are not enough to handle the huge amounts of data generated from millions of users. In addition, the diversity of data sources are increasing every day: Distributed file systems, relational, columnar-oriented, document-oriented or graph databases. Letgo has been growing quickly during the last years. Because of this, we needed to improve the scalability of our data platform and endow it further capabilities, like “dynamic infrastructure elasticity”, real-time processing or real-time complex event processing. In this talk, we are going to dive deeper into our journey. We started from a traditional data architecture with ETL and Redshift, till nowadays where we successfully have made an event oriented and horizontally scalable data architecture. We will explain in detail from the event ingestion with Kafka / Kafka Connect to its processing in streaming and batch with Spark. On top of that, we will discuss how we have used Spark Thrift Server / Hive Metastore as glue to exploit all our data sources: HDFS, S3, Cassandra, Redshift, MariaDB ... in a unified way from any point of our ecosystem, using technologies like: Jupyter, Zeppelin, Superset … We will also describe how to made ETL only with pure Spark SQL using Airflow for orchestration. Along the way, we will highlight the challenges that we found and how we solved them. We will share a lot of useful tips for the ones that also want to start this journey in their own companies.

Event-Driven Data Architecture at Letgo
· Talk

SQL is the lingua franca of data processing, and everybody working with data knows SQL. Apache Flink provides SQL support for querying and processing batch and streaming data. Flink's SQL support powers large-scale production systems at Alibaba, Huawei, and Uber. Based on Flink SQL, these companies have built systems for their internal users as well as publicly offered services for paying customers. In my talk I will show how to leverage the simplicity and power of SQL on Flink. I’ll explain why unified batch and stream processing is important and what it means to run SQL queries on streams of data. Once we’ve covered the basics, I will spend the remainder of the talk demonstrating the capabilities of Flink SQL. We will explore different use cases that Flink SQL was designed for by running queries on Flink’s SQL shell. In particular, I will demonstrate the unified batch and streaming engine by running the same query on batch and streaming data and show how to build a real-time dashboard that is powered by a streaming SQL query, which continuously updates an external result table.

Flink SQL in Action
· Talk

In Schibsted we have billions of events stored in their raw format on S3 buckets every day. Our analyst and data scientist have been fighting to get this data and start using it to: get insights, make analysis and build models. Exploring this data is complicated because of the evolving schema, the size and the lack of supporting tooling. We have worked on democratizing access to data by providing tooling to reduce time to data, and time to insights. We started with Jupyter, providing a serverless solution with some extra features and 0 infrastructure work. as Easy as clicking a button on your SSO dashboard. But this wasn't enough, and later on, we started offering an alternative driven by the use of SQL and JDBC connectivity. After a Beta version with Athena and few data, we have moved to Presto with our own patched solution. We are promoting some of these features to the OpenSource community and exploring ways to offer the others (like per-user data access authorization) to our DataENgineer colleagues outside Schibsted. We will speak about this journey and get deeper into the Presto chapter. How we have achieved a Continuous Delivery Pipeline using mixing Travis, spinnaker, cloudformation and AWS. What are the downsides of maintaining your patched presto version, the cost of maintaining it up to date and what you should take into account before choosing a query engine solution for your company.

Easy Access to Data with Presto
· Talk

This talk will focus on a recent prototype we develop of a DeepNet that can accurately predict ETAs of trips without having ever seen a map of the city, just learning from past trajectories of our assets by leveraging embeddings coming from Natural Language Processing (we found that trips are remarkably similar to sentences xD). I would like to frame it with the experiments we ran in real-life, proving how such system can significantly improve the experience of our riders and drivers (mostly by doing smarter assignments and fairer pricing), and thus helps us build smarter and overall better cities. I will touch also on the technical challenge of bringing such system to production considering the scale Cabify has already.

Driving in Dataland
· Talk

Apache Beam is a unified batch and streaming programming model for distributed data processing. Unlike other systems, Beam supports a range of different execution engines, e.g. Apache Flink, Apache Spark, or Google Cloud Dataflow. But it doesn't stop there. The Beam API is not only available in Java, but you can also write your data processing jobs in Python or Go. This gives you much greater flexibility compared to other data processing APIs. You can finally leverage the features and libraries of your favorite programming language. In this talk, I would like to give an introduction to the Beam programming model and explain how Beam achieves portability for different languages and execution engines.

Data Processing with Apache Beam: Towards Portability and Beyond
· Talk

Fishtown Analytics works with companies like Casper, Invision, Away Travel, and many more to help them build out effective analytics practices. These companies have complex data sets that are best understood by their analysts and business users (not their engineers!). To empower these users, the team at Fishtown Analytics has built dbt, an open source data transformation and democratization tool. It allows analysts and other non-engineers to write data transformations, while giving data engineers the ability to govern the process and ensure data quality. In this talk, we'll explore the key practices that make this setup work, including continuous integration, data lineage, quality testing, and documentation.

DBT: Powerful, Open Source Data Transformations
· Talk

Uber's mission is to provide transportation as reliable as running water, everywhere, and for everyone. To fulfill this mission, Uber relies heavily on making data-driven decisions at every level. Thus, we need to store an ever-increasing amount of data as the business grows in addition to providing faster, more-reliable, and more-performant access to our analytical data. In practice, this has resulted in 100+ PetaBytes of analytical data with minute-level data latency. On the other hand, due to Uber's global presence, regional regulations (such as GDPR) require additional and potentially more complex operations to be supported on the stored analytical data. These additional operations are usually unknown ahead of time, in many cases contradict the way data lakes are traditionally built/stored, and may require fundamental changes to the underlying assumptions/architecture. A good example is the GDPR regulation need to support update/delete operation on all historical Hadoop data that is traditionally considered append-only and stored in a read-only columnar file format within the analytical data lake. In this talk, we'll dive into how to build a generic big data platform that is flexible enough to support many of these unknown additional regulations/requirements out of the box and with minimum effort. This is not a review of how Uber tackled all the requirements of the GDPR but a deep dive into how Uber's big data platform came up with the fundamental primitives that enabled all other teams across the company to build their solution on top of Hadoop. We'll look into what technologies we were able to use from the open-source community (e.g. Hadoop, Spark, Hive, Presto, Kafka, Avro, and Vertica) and what solutions we had to build in-house (and open-source) to make this happen. You'll leave the talk with greater insight into how things work at Uber and will be inspired to re-envision your own data platform to make it more generic and flexible for future new requirements.

Creating an Extensible Big Data Platform

Showing 10 of 18

Data Council Barcelona 2018 — Speakers

The voices that shaped 2018

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Wes McKinney, Principal Architect, Posit

Principal Architect, Posit

Willem Pienaar, Co-Founder & CTO, Cleric

Co-Founder & CTO, Cleric

Albert Franzi Cros, Data Engineer, Alpha Health

Data Engineer, Alpha Health

Alberto Betella, CTO, Badi

CTO, Badi

Carlos Herrera, Head of Data Science & Research  Product Team, Cabify

Head of Data Science & Research Product Team, Cabify

Connor McArthur, Co-founder & Data Engineer, Fishtown Analytics / DBT

Co-founder & Data Engineer, Fishtown Analytics / DBT

Dave Garcia, SVP of Engineering, TravelPerk

SVP of Engineering, TravelPerk

Iker Martinez de Apellaniz, Data Engineer, Schibsted Classified Media

Data Engineer, Schibsted Classified Media

Data Council Barcelona 2018 — Sponsors

Supported by leaders in AI infrastructure

Snowflake
TextQL
HEX
Databricks
Braintrust
ClickHouse
Snorkel
Datalinks
Airbyte
Render
Turbopuffer
DigitalOcean
CockroachDB
bem
Preset
LanceDB
Chalk
Unstructured
MotherDuck
Crux
TOPK
Data Council Barcelona 2018 — Testimonials

Voices from 2018

AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck
Priya Nair at the panel discussion