Course Outline
Introduction, Objectives, and Migration Strategy
- Course aims, alignment with participant profiles, and definitions of success criteria
- High-level migration methodologies and associated risk assessments
- Configuration of workspaces, repositories, and laboratory datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake overview, and Databricks architectural components
- Differences between SMP and MPP and their impact on migration strategies
- Medallion (Bronze to Silver to Gold) design principles and Unity Catalog overview
Day 1 Laboratory — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook format
- Mapping temporary tables and cursors to DataFrame transformations
- Validation and comparison against original output results
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel capabilities
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution techniques
- OPTIMIZE, VACUUM, Z-ORDER, partitioning strategies, and storage optimisation
Day 2 Laboratory — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM; verification of outcomes
- Assessment of read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL features: window functions, higher-order functions, and JSON/array management
- Interpreting Spark UI, DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Query tuning patterns: broadcast joins, hints, caching, and spill minimisation
Day 3 Laboratory — SQL Refactoring & Performance Tuning
- Refactoring a complex SQL process into optimized Spark SQL
- Utilising Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking performance before and after, and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning tactics
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs/pandas UDFs, widgets, and creation of reusable libraries
Day 4 Laboratory — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Introducing parametrization, unit-style tests, and reusable functions
- Code review and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Laboratory — Build a Complete End-to-End Pipeline
- Assembly of a Bronze to Silver to Gold pipeline orchestrated via Workflows
- Implementation of logging, auditing, retries, and automated validations
- Execution of the full pipeline, verification of outputs, and preparation of deployment notes
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, lineage tracking, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook creation
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key lessons learned
- Gap analysis, suggested follow-up activities, and handover of training materials
- Reference materials, further learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Practical experience with SQL and stored procedures (e.g., Synapse or SQL Server)
- Knowledge of ETL orchestration concepts (e.g., ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers shifting procedural OLAP logic towards Lakehouse patterns
- Platform engineers tasked with overseeing Databricks adoption
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 6500 € + VAT*
Contact us for an exact quote and to hear our latest promotions