This hands-on, three-day programme is dedicated to constructing and fine-tuning high-performance data-processing workloads by leveraging PySpark, Pandas and Polars within Kubernetes-based ecosystems.
Attendees will cultivate a practical grasp of how Spark applications are executed on Kubernetes, gaining insight into how application-level configuration choices directly impact performance, scalability, resource utilisation and overall costs. Key optimisation topics covered include executor sizing, memory allocation, dynamic allocation, partitioning strategies, shuffle mechanics, the challenges of small files, and efficient Parquet processing.
The curriculum also tackles common pain points associated with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance alternative for specific data-processing tasks. Through practical exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and apply optimisation techniques to realistic ETL and machine learning scenarios.
The core focus of the course is on practical decision-making: mastering the ability to identify performance bottlenecks, select the most appropriate tools, configure Spark efficiently, and strike a balance between performance gains and infrastructure resource consumption and cost.
Read more...