The Data Engineering with Microsoft Fabric, Data Factory, DAX, Python & SQL Training Course is a comprehensive program designed to develop practical skills for building, integrating, transforming, managing, and analyzing enterprise data using modern Microsoft data technologies.
The course covers the end-to-end data engineering lifecycle, beginning with data engineering fundamentals, relational data concepts, and SQL, then progressing to Python-based data processing, Microsoft Fabric, OneLake, Lakehouse architecture, Data Factory, data pipelines, notebooks, Delta tables, semantic models, and DAX (Data Analysis Expressions).
Participants will learn how these technologies work together to ingest data from multiple sources, perform transformations, build scalable data pipelines, organize data within lakehouse and warehouse architectures, prepare analytics-ready datasets, and support reporting and business intelligence requirements.
The program is suitable for organizations developing modern data platforms using the Microsoft ecosystem and for technical professionals who need integrated skills across SQL, Python, Microsoft Fabric, Data Factory, and Power BI-related analytical technologies.
Duration 10 Days – 70 hrs.
Objectives
- Understand modern data engineering concepts, architectures, and workflows.
- Explain the role of data engineers within modern analytics environments.
- Understand structured, semi-structured, and unstructured data.
- Use SQL to query, filter, join, aggregate, and transform data.
- Apply intermediate and advanced SQL techniques to data engineering tasks.
- Use Python for data manipulation, transformation, and automation.
- Work with common Python data-processing libraries and data structures.
- Understand the Microsoft Fabric architecture and its major workloads.
- Navigate Microsoft Fabric workspaces and understand capacity concepts.
- Understand OneLake and its role in Microsoft Fabric.
- Build and manage Fabric Lakehouses and Warehouses.
- Work with tables, files, Delta tables, and related storage structures.
- Use Microsoft Fabric Data Factory for data ingestion and orchestration.
- Create and configure data pipelines.
- Perform data transformations using Dataflows Gen2 and related Fabric capabilities.
- Use notebooks for data engineering and transformation.
- Integrate SQL and Python within data engineering workflows.
- Understand medallion architecture and Bronze, Silver, and Gold data layers.
- Prepare curated datasets for analytics and reporting.
- Understand semantic models and their relationship to data engineering.
- Create calculated columns, measures, and analytical calculations using DAX.
- Apply DAX filter and evaluation concepts.
- Implement data quality, monitoring, security, and governance considerations.
- Design an integrated end-to-end data engineering solution using Microsoft technologies.
Target Audience
- Data Engineers
- Junior Data Engineers
- Data Analysts
- BI Developers
- Business Intelligence Analysts
- Database Developers
- SQL Developers
- ETL/ELT Developers
- Data Integration Specialists
- Analytics Engineers
- Power BI Developers
- Python Developers transitioning into data engineering
- Database Administrators
- Cloud Data Professionals
- Application Developers working with data platforms
- IT Professionals supporting analytics environments
- Technical professionals preparing to work with Microsoft Fabric
- Professionals transitioning into data engineering roles
Prerequisites
- Basic understanding of databases and data concepts.
- Basic familiarity with spreadsheets, reporting, or data analysis.
- Basic knowledge of SQL is beneficial but not mandatory.
- Basic programming knowledge is helpful for the Python modules.
- General understanding of cloud computing concepts is beneficial.
- Familiarity with Microsoft Power BI is helpful for DAX and semantic modeling topics but is not mandatory.
- Basic understanding of data tables, columns, relationships, and data types.
- Access to the appropriate Microsoft Fabric environment and development tools for practical exercises.
Course Outline
Day 1 – Data Engineering Fundamentals and Modern Data Architecture
Module 1: Introduction to Data Engineering
- What is data engineering?
- Data engineering vs. data analytics vs. data science
- Roles and responsibilities of a data engineer
- Modern enterprise data ecosystems
- Data engineering lifecycle
- Batch and streaming data concepts
- Structured, semi-structured, and unstructured data
- Common data sources and destinations
Module 2: Modern Data Architecture
- Traditional data warehouses
- Data lakes
- Lakehouse architecture
- Modern cloud data platforms
- ETL vs. ELT
- Data ingestion and transformation patterns
- Analytical vs. transactional workloads
- Introduction to medallion architecture
- Bronze, Silver, and Gold layers
- Designing analytics-ready datasets
Module 3: Microsoft Data Engineering Ecosystem
- Overview of the Microsoft data platform
- Microsoft Fabric
- OneLake
- Data Factory
- Lakehouse
- Data Warehouse
- Power BI
- SQL and Python
- How the technologies work together
Day 2 – SQL Fundamentals for Data Engineering
Module 4: Relational Database and SQL Foundations
- Relational database concepts
- Tables, rows, and columns
- Primary and foreign keys
- Relationships
- Data types
- SQL syntax fundamentals
- SELECT statements
- Filtering with WHERE
- Sorting with ORDER BY
- DISTINCT
- Aliases
- NULL handling
Module 5: SQL Data Manipulation
- INSERT
- UPDATE
- DELETE
- Creating and modifying tables
- Data type considerations
- Built-in SQL functions
- String functions
- Date and time functions
- Numeric functions
- Conditional logic with CASE
Module 6: Aggregating Data with SQL
- Aggregate functions
- COUNT, SUM, AVG, MIN, and MAX
- GROUP BY
- HAVING
- Creating analytical summaries
- Working with business datasets
Day 3 – Intermediate and Advanced SQL for Data Engineering
Module 7: SQL Joins and Data Integration
- INNER JOIN
- LEFT and RIGHT JOIN
- FULL OUTER JOIN
- CROSS JOIN
- Self joins
- Combining datasets
- UNION and UNION ALL
- Data integration using SQL
Module 8: Advanced SQL Queries
- Subqueries
- Common Table Expressions
- Recursive concepts
- Window functions
- ROW_NUMBER
- RANK and DENSE_RANK
- LAG and LEAD
- Partitioning analytical calculations
- Advanced aggregation techniques
Module 9: SQL for Data Engineering
- Data cleansing with SQL
- Data transformation
- Duplicate detection
- Missing-value handling
- Data validation queries
- Views
- Stored procedures concepts
- Query optimization fundamentals
- SQL in ETL/ELT workflows
Day 4 – Python for Data Engineering
Module 10: Python Fundamentals for Data Professionals
- Python data engineering use cases
- Variables and data types
- Lists, tuples, sets, and dictionaries
- Conditional statements
- Loops
- Functions
- Exception handling
- Working with modules and packages
Module 11: Data Processing with Python
- Introduction to NumPy concepts
- Introduction to pandas
- Series and DataFrames
- Reading data from files
- Selecting and filtering data
- Sorting data
- Handling missing values
- Removing duplicates
- Data type conversion
- Creating calculated fields
Module 12: Data Transformation with Python
- Combining datasets
- Merge and join operations
- Grouping and aggregation
- Reshaping data
- Date and time processing
- String transformations
- Data validation
- Exporting transformed datasets
Day 5 – Microsoft Fabric and OneLake Fundamentals
Module 13: Introduction to Microsoft Fabric
- Microsoft Fabric architecture
- Fabric workloads
- Fabric workspaces
- Capacity concepts
- Data Engineering
- Data Factory
- Data Warehouse
- Power BI integration
- Real-time and analytics concepts
- Fabric end-to-end analytics architecture
Module 14: Microsoft OneLake
- Understanding OneLake
- OneLake architecture
- Centralized enterprise data
- Files and tables
- OneLake shortcuts
- Connecting data across domains
- Data accessibility and reuse
- OneLake security considerations
Module 15: Fabric Lakehouse
- Understanding Lakehouse architecture
- Creating a Lakehouse
- Managing files and folders
- Tables
- Delta Lake fundamentals
- Managed and external data concepts
- Loading data into a Lakehouse
- Querying Lakehouse data
- SQL analytics endpoint concepts
Day 6 – Microsoft Fabric Data Factory and Data Ingestion
Module 16: Introduction to Data Factory in Microsoft Fabric
- Data Factory architecture
- Data integration concepts
- Connections and data sources
- Destinations
- Copy activities
- Data ingestion patterns
- Batch data ingestion
Module 17: Building Data Pipelines
- Creating pipelines
- Pipeline activities
- Copying data between systems
- Parameters and variables
- Expressions
- Control flow concepts
- Dependencies
- Pipeline execution
- Scheduling and orchestration
Module 18: Dataflows Gen2 and Data Transformation
- Introduction to Dataflows Gen2
- Connecting to data sources
- Data transformation concepts
- Power Query transformations
- Combining datasets
- Data cleansing
- Data type management
- Loading transformed data
- Dataflows vs. pipelines
Day 7 – Fabric Data Engineering with Notebooks and Spark
Module 19: Fabric Notebooks
- Introduction to notebooks
- Notebook environment
- Using Python in Fabric
- Working with Lakehouse data
- Reading files and tables
- Writing transformed data
- Notebook execution
Module 20: Apache Spark Concepts in Microsoft Fabric
- Introduction to distributed data processing
- Spark architecture concepts
- Spark DataFrames
- PySpark fundamentals
- Reading data with Spark
- Filtering and transforming data
- Aggregating data
- Joining datasets
- Writing processed data
Module 21: Delta Lake and Medallion Architecture
- Delta Lake fundamentals
- Delta tables
- Data reliability concepts
- Bronze data layer
- Silver data layer
- Gold data layer
- Transforming data between layers
- Designing reusable data pipelines
- Preparing curated analytical datasets
Day 8 – Fabric Data Warehouse and Integrated Data Solutions
Module 22: Microsoft Fabric Data Warehouse
- Lakehouse vs. Warehouse
- Creating a Fabric Warehouse
- Tables and schemas
- Loading warehouse data
- Querying using T-SQL
- Views and analytical structures
- Data modeling considerations
- Warehouse use cases
Module 23: Integrating SQL, Python, Lakehouse, and Warehouse
- Selecting the appropriate processing technology
- SQL-based transformations
- Python-based transformations
- Spark-based processing
- Moving data between layers
- Pipeline orchestration
- Building reusable transformation processes
Module 24: Designing an End-to-End Fabric Data Pipeline
- Source-system identification
- Data ingestion
- Raw data storage
- Data cleansing
- Data transformation
- Business-rule implementation
- Curated data layer
- Warehouse integration
- Analytics consumption layer
Day 9 – Semantic Modeling and DAX (Data Analysis Expressions)
Module 25: Data Modeling for Analytics
- Analytical data modeling concepts
- Fact and dimension tables
- Star schema
- Relationships
- Cardinality
- Filter direction
- Date dimensions
- Designing models for reporting
- Semantic model concepts
Module 26: DAX Fundamentals
- Introduction to DAX
- DAX syntax
- Calculated columns
- Measures
- Calculated tables concepts
- Operators
- Aggregation functions
- Logical functions
- Mathematical functions
- Date and time functions
Module 27: Intermediate and Advanced DAX
- Row context
- Filter context
- Context transition
- CALCULATE
- FILTER
- ALL and related filter functions
- Iterator functions
- SUMX and related X functions
- Variables
- DIVIDE
- Time intelligence concepts
- Year-to-date calculations
- Previous-period comparisons
- Percentage and variance calculations
- DAX performance considerations
Day 10 – Production Data Engineering, Governance and End-to-End Integration
Module 28: Data Quality and Reliability
- Data quality dimensions
- Data validation
- Completeness and consistency
- Duplicate management
- Schema considerations
- Error handling
- Pipeline failure handling
- Logging concepts
- Data reconciliation
Module 29: Security, Governance, and Data Management
- Data governance fundamentals
- Access management
- Workspace security
- Data permissions
- Sensitive data considerations
- Data lineage
- Metadata concepts
- Governance within Microsoft Fabric
- Enterprise data management practices
Module 30: Monitoring and Optimization
- Pipeline monitoring
- Identifying pipeline failures
- Troubleshooting data loads
- SQL performance considerations
- Spark performance concepts
- Data partitioning concepts
- Optimizing transformation workflows
- Managing scalable data solutions
Module 31: Integrated End-to-End Data Engineering Solution
- Connecting source systems
- Ingesting data through Data Factory
- Landing data in OneLake
- Building Bronze, Silver, and Gold layers
- Transforming data using SQL and Python
- Using Fabric notebooks and Spark
- Loading curated data into Lakehouse or Warehouse
- Building analytical data models
- Creating business calculations using DAX
- Preparing datasets for Power BI and downstream analytics
- Reviewing the complete enterprise data engineering workflow

