
1. Abstract
Despite its limitations, Apache Spark is still the go-to choice for big data workloads across organizations in the industry. However, organizations around the world waste money and productive time by running inefficient Spark jobs. The difference between an efficient Spark pipeline and an inefficient one could be an order of magnitude greater in terms of both compute cost and wall-clock time, and investing in an efficient pipeline could yield more than 75% savings in money and time.
In this workshop, Ammar Chalifah will cover best practices for optimizing a Spark job, from reading the physical plan, minimizing shuffle and skew, avoiding UDFs, choosing the right storage format and storage layout, and right-sizing the cluster.
2. Agenda
3. Objectives
Attendees understand the biggest bottlenecks in Spark pipelines, know how to identify them, are able to implement an optimization technique, and are aware of production best practices.
4. Target audience and Prerequisites