Data Processing with Apache Beam: Towards Portability and Beyond
Apache Beam is a unified batch and streaming programming model for distributed data processing. Unlike other systems, Beam supports a range of different execution engines, e.g. Apache Flink, Apache Spark, or Google Cloud Dataflow. But it doesn't stop there. The Beam API is not only available in Java, but you can also write your data processing jobs in Python or Go. This gives you much greater flexibility compared to other data processing APIs. You can finally leverage the features and libraries of your favorite programming language. In this talk, I would like to give an introduction to the Beam programming model and explain how Beam achieves portability for different languages and execution engines.
Speakers