Ammar Chalifah

Senior Data Engineer
Modash
Estonia

About

Ammar is a Data Engineer with strong focus in building scalable and optimized data pipelines, with experience in saving cost and wall-clock time through compute & storage tuning. Throughout his career, Ammar has generated €1M+ in compute savings for the organizations he worked for. Currently Ammar is working on building influencer marketing platform at Modash.
Workshop

Ammar Chalifah | Spark Pipeline Optimization

Spark, Iceberg, Optimization
1.Abstract Despite its limitations, Apache Spark is still the go-to choice for big data workloads across organizations in the industry. However, organizations around the world waste money and productive time by running inefficient Spark jobs. The difference between an efficient Spark pipeline and an inefficient one could be an order of magnitude greater in terms of both compute cost and wall-clock time, and investing in an efficient pipeline could yield more than 75% savings in money and time. In this workshop, Ammar Chalifah will cover best practices for optimizing a Spark job, from reading the physical plan, minimizing shuffle and skew, avoiding UDFs, choosing the right storage format and storage layout, and right-sizing the cluster. 2.Agenda-Brief introduction to the topic: problems around Spark pipelines (10 minutes) -Setting up repository for attendees (10 minutes) -Reading Spark UI (10 minutes) -Optimization case 1: shuffle. Demonstration + practice (15 minutes) -Optimization case 2: shuffle, StoragePartitionedJoin (10 minutes) -Optimization case 3: lazy execution, solving it through cache/checkpoint/materialization (15 minutes) -Optimization case 4: UDF vs native Spark (10 minutes) -Optimization case 5: native acceleration, Apache Gluten (15 minutes) -Closing, questions (5 minutes)3.Objectives Attendees understand the biggest bottlenecks in Spark pipelines, know how to identify them, are able to implement an optimization technique, and are aware of production best practices.4.Target audience and Prerequisites -Data Engineer in early-to-mid career -Engineers looking to deepening expertise in Spark
2026-11-24
09:00
17:00