
This talk dives into Apache Spark’s quirks in real-world production workloads involving terabytes of data. Marcin Szymaniuk begins with a quick recap of last year’s lessons and then explores practical corner cases in which Spark still struggles today. Topics include user-defined function (UDF) bottlenecks, driver-heavy tasks, large directed acyclic graphs (DAGs), and challenges with non-standard data sources. Attendees can expect a no-fluff, hands-on walkthrough of edge cases Marcin has encountered in production, offering insights to help their Spark jobs run more smoothly.