Get in Touch
 Duration 21 hours

Course Outline

Getting Started:

  • Positioning Apache Spark within the Hadoop ecosystem
  • Brief overview of Python and Scala

Fundamental Concepts (Theory):

  • System architecture
  • Resilient Distributed Datasets (RDDs)
  • Transformations versus Actions
  • Stages, tasks, and dependencies

Practical Application in Databricks (Hands-on Workshop):

  • Exercises utilizing the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Implementing Joins
  • Caching strategies
  • Exercises utilizing the DataFrame API
  • SparkSQL integration
  • DataFrame operations: select, filter, group, and sort
  • User-Defined Functions (UDFs)
  • Introduction to the Dataset API
  • Streaming concepts

Deployment Strategies in AWS (Hands-on Workshop):

  • Foundations of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Executing sample jobs across both environments
  • Evaluating advantages and limitations

Supplementary Topics:

  • Overview of Apache Airflow for orchestration

Requirements

Programming proficiency (Python and Scala preferred)

Foundational SQL knowledge

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories