Albert Franzi Cros

Data Engineer, Alpha Health

Albert Franzi is a Software Engineer who fell so in love with data that ended up as Data Engineer for Alpha Health. He believes in a world where distributed organizations can work together to build common and reusable tools to empower their users and projects. Albert cares deeply about unified data and models as well as data quality and enrichment. He also has a secret plan to conquer the world with data, insights, and penguins.

Albert Franzi Cros

Sessions / 2019 / 1 talk

  • At Alpha Health we are developing genuine data products and hence we regard data as one of our most valuable assets. For such products that are in constant development, data also evolves continuously to meet the requirement. Hence, from the product version to version, some data entities are expected to change and even to look completely different. To keep track of the changes over time while improving the understanding of our data, we need to accurately define our data entities using schemas and how these schemas evolve from version to version. This talk aims to cover the integration of JSON-schemas in our data flow in real-time, the benefits of using VEA (Validating, Evolving & Anonymizing) and how this approach can empower the whole company to bring data to the next level. We will focus on how to detect and put invalid data into quarantine, how we evolve data into its latest schema version in a streamlined manner and how we generate de-identified, GDPR-compliant data. We will go over the multiple benefits and challenges we found during the implementation from managing data models, to iterating the infrastructure.

Sessions / 2018 / 1 talk

  • In Schibsted we have billions of events stored in their raw format on S3 buckets every day. Our analyst and data scientist have been fighting to get this data and start using it to: get insights, make analysis and build models. Exploring this data is complicated because of the evolving schema, the size and the lack of supporting tooling. We have worked on democratizing access to data by providing tooling to reduce time to data, and time to insights. We started with Jupyter, providing a serverless solution with some extra features and 0 infrastructure work. as Easy as clicking a button on your SSO dashboard. But this wasn't enough, and later on, we started offering an alternative driven by the use of SQL and JDBC connectivity. After a Beta version with Athena and few data, we have moved to Presto with our own patched solution. We are promoting some of these features to the OpenSource community and exploring ways to offer the others (like per-user data access authorization) to our DataENgineer colleagues outside Schibsted. We will speak about this journey and get deeper into the Presto chapter. How we have achieved a Continuous Delivery Pipeline using mixing Travis, spinnaker, cloudformation and AWS. What are the downsides of maintaining your patched presto version, the cost of maintaining it up to date and what you should take into account before choosing a query engine solution for your company.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker