Get in Touch
 Duration 35 hours

Course Outline

Introduction, Objectives, and Migration Strategy

  • Course aims, alignment with participant profiles, and definitions of success criteria
  • High-level migration methodologies and associated risk assessments
  • Configuration of workspaces, repositories, and laboratory datasets

Day 1 — Migration Fundamentals and Architecture

  • Lakehouse concepts, Delta Lake overview, and Databricks architectural components
  • Differences between SMP and MPP and their impact on migration strategies
  • Medallion (Bronze to Silver to Gold) design principles and Unity Catalog overview

Day 1 Laboratory — Translating a Stored Procedure

  • Practical migration of a sample stored procedure into a notebook format
  • Mapping temporary tables and cursors to DataFrame transformations
  • Validation and comparison against original output results

Day 2 — Advanced Delta Lake & Incremental Loading

  • ACID transactions, commit logs, versioning, and time travel capabilities
  • Auto Loader, MERGE INTO patterns, upserts, and schema evolution techniques
  • OPTIMIZE, VACUUM, Z-ORDER, partitioning strategies, and storage optimisation

Day 2 Laboratory — Incremental Ingestion & Optimization

  • Implementation of Auto Loader ingestion and MERGE workflows
  • Application of OPTIMIZE, Z-ORDER, and VACUUM; verification of outcomes
  • Assessment of read/write performance enhancements

Day 3 — SQL in Databricks, Performance & Debugging

  • Analytical SQL features: window functions, higher-order functions, and JSON/array management
  • Interpreting Spark UI, DAGs, shuffles, stages, tasks, and identifying bottlenecks
  • Query tuning patterns: broadcast joins, hints, caching, and spill minimisation

Day 3 Laboratory — SQL Refactoring & Performance Tuning

  • Refactoring a complex SQL process into optimized Spark SQL
  • Utilising Spark UI traces to identify and resolve skew and shuffle issues
  • Benchmarking performance before and after, and documenting tuning procedures

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Spark execution model: driver, executors, lazy evaluation, and partitioning tactics
  • Converting loops and cursors into vectorized DataFrame operations
  • Modularization, UDFs/pandas UDFs, widgets, and creation of reusable libraries

Day 4 Laboratory — Refactoring Procedural Scripts

  • Refactoring a procedural ETL script into modular PySpark notebooks
  • Introducing parametrization, unit-style tests, and reusable functions
  • Code review and application of best-practice checklists

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error management
  • Designing incremental Medallion pipelines with quality rules and schema validation
  • Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic

Day 5 Laboratory — Build a Complete End-to-End Pipeline

  • Assembly of a Bronze to Silver to Gold pipeline orchestrated via Workflows
  • Implementation of logging, auditing, retries, and automated validations
  • Execution of the full pipeline, verification of outputs, and preparation of deployment notes

Operationalization, Governance, and Production Readiness

  • Unity Catalog governance, lineage tracking, and access control best practices
  • Cost management, cluster sizing, autoscaling, and job concurrency patterns
  • Deployment checklists, rollback strategies, and runbook creation

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations on migration work and key lessons learned
  • Gap analysis, suggested follow-up activities, and handover of training materials
  • Reference materials, further learning paths, and support options

Requirements

  • A solid grasp of data engineering concepts
  • Practical experience with SQL and stored procedures (e.g., Synapse or SQL Server)
  • Knowledge of ETL orchestration concepts (e.g., ADF or similar tools)

Target Audience

  • Technology managers with a background in data engineering
  • Data engineers shifting procedural OLAP logic towards Lakehouse patterns
  • Platform engineers tasked with overseeing Databricks adoption

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 6500 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories