Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Course goals, alignment with participant profiles, and success criteria
- High-level migration approaches and associated risk considerations
- Setup of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Distinctions between SMP and MPP and their impact on migration
- Medallion (Bronze to Silver to Gold) design principles and an overview of Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure to a notebook
- Converting temp tables and cursors into DataFrame transformations
- Validation and comparison against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning techniques
Day 2 Lab — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM; verification of results
- Evaluation of read and write performance improvements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, JSON and array handling
- Interpreting the Spark UI, DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Query tuning strategies: broadcast joins, hints, caching, and reducing spills
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Leveraging Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking performance before and after changes and documenting tuning steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs and pandas UDFs, widgets, and building reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Transforming a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub or Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze to Silver to Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage, and access controls
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and creating runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommendations for follow-up activities, and handover of training materials
- References, further learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse or SQL Server)
- Knowledge of ETL orchestration concepts (ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers shifting procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption