This intensive three-day workshop is dedicated to constructing and refining high-efficiency data-processing workloads that leverage PySpark, Pandas, and Polars within Kubernetes-based environments.
Learners will gain a practical insight into the execution mechanics of Spark applications on Kubernetes, exploring how application-level configuration choices directly impact performance, scalability, resource consumption, and overall cost. Key optimisation topics include executor sizing, memory allocation, dynamic allocation, partitioning strategies, shuffle mechanics, the small-file problem, and the efficient handling of Parquet data.
The curriculum also tackles frequent challenges associated with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance alternative for specific data-processing tasks. Through interactive exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and implement optimisation techniques in realistic ETL and machine learning contexts.
The core focus of the course remains on practical decision-making: mastering the ability to identify performance bottlenecks, choose the right tools, configure Spark effectively, and strike a balance between processing speed and infrastructure resource expenditure.
Read more...