Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Getting Started:
- Positioning Apache Spark within the Hadoop ecosystem
- Brief overview of Python and Scala
Fundamental Concepts (Theory):
- System architecture
- Resilient Distributed Datasets (RDDs)
- Transformations versus Actions
- Stages, tasks, and dependencies
Practical Application in Databricks (Hands-on Workshop):
- Exercises utilizing the RDD API
- Core action and transformation functions
- Working with PairRDDs
- Implementing Joins
- Caching strategies
- Exercises utilizing the DataFrame API
- SparkSQL integration
- DataFrame operations: select, filter, group, and sort
- User-Defined Functions (UDFs)
- Introduction to the Dataset API
- Streaming concepts
Deployment Strategies in AWS (Hands-on Workshop):
- Foundations of AWS Glue
- Distinguishing between AWS EMR and AWS Glue
- Executing sample jobs across both environments
- Evaluating advantages and limitations
Supplementary Topics:
- Overview of Apache Airflow for orchestration
Requirements
Programming proficiency (Python and Scala preferred)
Foundational SQL knowledge
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift