Data Engineering on AWS with AWS DMS and Modern Data Services

Inquire now

The Data Engineering on AWS with AWS DMS and Modern Data Services Training Course provides participants with practical knowledge of designing, building, migrating, processing, and managing modern data solutions using Amazon Web Services (AWS).

The course covers the end-to-end data engineering lifecycle, including data ingestion, database migration, data lakes, ETL/ELT processing, data warehousing, analytics, orchestration, monitoring, security, and performance optimization. A major focus is placed on AWS Database Migration Service (AWS DMS) for migrating and continuously replicating data between database platforms with minimal downtime.

Participants will explore commonly used AWS data engineering services including Amazon S3, AWS Glue, AWS DMS, Amazon RDS, Amazon Redshift, Amazon Athena, AWS Lambda, Amazon Kinesis, AWS Step Functions, Amazon CloudWatch, AWS IAM, and related services. The course also introduces architectural patterns for batch and streaming data pipelines and modern AWS data lake and analytics environments.

 

Duration 5 Days – 35 hrs.

 

Objectives

  • Explain the role and responsibilities of a data engineer in an AWS environment.
  • Understand AWS data engineering architecture and common data pipeline patterns.
  • Select appropriate AWS storage, database, processing, and analytics services.
  • Design scalable data ingestion and integration pipelines.
  • Build and manage data lakes using Amazon S3.
  • Understand ETL and ELT approaches for cloud-based data engineering.
  • Create data catalogs and ETL pipelines using AWS Glue.
  • Use AWS DMS to migrate and replicate databases.
  • Understand homogeneous and heterogeneous database migration scenarios.
  • Configure AWS DMS replication instances, endpoints, and migration tasks.
  • Implement full-load and Change Data Capture (CDC) migration strategies.
  • Work with relational databases through Amazon RDS and Amazon Aurora.
  • Design analytical data warehouse solutions using Amazon Redshift.
  • Query data directly from Amazon S3 using Amazon Athena.
  • Understand batch and real-time/streaming data processing architectures.
  • Use Amazon Kinesis for streaming data workloads.
  • Incorporate AWS Lambda and Step Functions into data workflows.
  • Apply security and access controls using AWS IAM and encryption.
  • Monitor data pipelines and AWS resources using Amazon CloudWatch.
  • Apply data quality, reliability, scalability, and performance best practices.
  • Design an end-to-end AWS data engineering solution.

 

Target Audience

  • Data Engineers
  • Database Administrators
  • Database Engineers
  • Data Architects
  • Cloud Engineers
  • Cloud Architects
  • ETL Developers
  • Data Warehouse Developers
  • Data Analysts with technical responsibilities
  • Business Intelligence Developers
  • Software Developers working with data platforms
  • DevOps Engineers supporting data workloads
  • Solutions Architects
  • System Engineers
  • IT Professionals responsible for database migration
  • Technical professionals transitioning into cloud data engineering

 

Prerequisites

  • Basic understanding of databases and relational database concepts.
  • Basic SQL knowledge, including SELECT, JOIN, GROUP BY, and data manipulation.
  • General understanding of ETL/ELT and data integration concepts.
  • Basic familiarity with cloud computing concepts.
  • Basic knowledge of data warehousing is helpful but not mandatory.
  • Familiarity with Python or another scripting language is beneficial.
  • Previous AWS experience is helpful but not required.

Course Outline

Day 1 – AWS Data Engineering Fundamentals and Data Storage

Module 1: Introduction to Modern Data Engineering

  • Role of a data engineer
  • Data engineering lifecycle
  • Structured, semi-structured, and unstructured data
  • Operational versus analytical workloads
  • Batch versus streaming processing
  • ETL versus ELT
  • Data pipelines and data integration
  • Data lakes, data warehouses, and lakehouse concepts

Module 2: AWS Architecture for Data Engineering

  • Overview of AWS global infrastructure
  • Regions and Availability Zones
  • AWS data engineering ecosystem
  • Compute, storage, database, integration, and analytics services
  • Designing highly available data architectures
  • Scalability and fault tolerance
  • Common AWS data engineering reference architectures

Module 3: Amazon S3 for Data Engineering

  • Amazon S3 architecture
  • Buckets and objects
  • Storage classes
  • Organizing data lake structures
  • Data partitioning strategies
  • File formats for data engineering
  • CSV, JSON, Parquet, and ORC
  • Compression considerations
  • Versioning and lifecycle management
  • S3 security and access control

Module 4: Building Data Lakes on AWS

  • Data lake architecture
  • Raw, processed, curated, and consumption layers
  • Data ingestion patterns
  • Data organization and partitioning
  • Metadata management
  • Data lake governance concepts
  • Designing scalable S3-based data platforms

 

Day 2 – AWS Database Migration Service and Database Integration

Module 5: AWS Database Services for Data Engineers

  • Amazon RDS overview
  • Amazon Aurora
  • Relational database engines
  • Amazon DynamoDB overview
  • Operational versus analytical databases
  • Selecting appropriate AWS database services
  • Database connectivity considerations

Module 6: Introduction to AWS Database Migration Service (AWS DMS)

  • AWS DMS architecture
  • Database migration concepts
  • Homogeneous migrations
  • Heterogeneous migrations
  • Source and target endpoints
  • Replication instances
  • Migration tasks
  • Supported migration patterns
  • Network and connectivity considerations

Module 7: Implementing Database Migration with AWS DMS

  • Preparing source and target databases
  • Creating replication instances
  • Configuring source endpoints
  • Configuring target endpoints
  • Testing endpoint connectivity
  • Creating migration tasks
  • Full-load migration
  • Change Data Capture (CDC)
  • Full load plus CDC
  • Table mappings and selection rules
  • Transformation rules

Module 8: Advanced AWS DMS Concepts

  • Continuous data replication
  • Migration monitoring
  • Validation and reconciliation
  • Logging and troubleshooting
  • Performance considerations
  • Handling large tables
  • Managing migration failures
  • Minimizing migration downtime
  • Migration cutover planning
  • AWS Schema Conversion Tool concepts
  • Schema conversion considerations for heterogeneous migrations

 

Day 3 – AWS Glue, ETL/ELT and Data Transformation

Module 9: Introduction to AWS Glue

  • AWS Glue architecture
  • AWS Glue Data Catalog
  • Databases and tables
  • Crawlers
  • Classifiers
  • Metadata discovery
  • Integrating Glue with Amazon S3
  • Schema management

Module 10: Building ETL Pipelines with AWS Glue

  • Creating ETL jobs
  • Data extraction
  • Data transformation
  • Data loading
  • Working with Glue Studio
  • Working with structured and semi-structured datasets
  • Filtering and cleansing data
  • Joining datasets
  • Data type conversion
  • Writing transformed data to Amazon S3

Module 11: Data Quality and Pipeline Reliability

  • Data validation concepts
  • Handling missing and invalid data
  • Duplicate detection
  • Schema consistency
  • Data quality rules
  • Error handling
  • Logging
  • Retry strategies
  • Idempotent pipeline concepts
  • Designing resilient data pipelines

Module 12: Serverless Data Processing

  • AWS Lambda fundamentals
  • Event-driven data processing
  • S3 event triggers
  • Integrating Lambda with data pipelines
  • AWS Step Functions overview
  • Workflow orchestration
  • Coordinating data engineering processes
  • Error handling and workflow control

 

Day 4 – Data Warehousing, Analytics and Streaming Data

Module 13: Amazon Redshift for Data Warehousing

  • Data warehouse architecture
  • Amazon Redshift concepts
  • Provisioned and serverless approaches
  • Loading data into Redshift
  • Data distribution concepts
  • Sort keys and data organization
  • Query performance considerations
  • Redshift Spectrum concepts
  • Integrating Redshift with Amazon S3

Module 14: Serverless Analytics with Amazon Athena

  • Amazon Athena architecture
  • Querying data stored in Amazon S3
  • Integrating Athena with AWS Glue Data Catalog
  • Working with partitions
  • Query optimization
  • Columnar data formats
  • Cost optimization
  • Common analytics use cases

Module 15: Real-Time Data Engineering with Amazon Kinesis

  • Streaming data fundamentals
  • Batch versus streaming architectures
  • Amazon Kinesis ecosystem
  • Kinesis Data Streams
  • Producers and consumers
  • Stream processing concepts
  • Data ingestion patterns
  • Real-time analytics use cases
  • Integrating streaming data with AWS data services

Module 16: Designing Batch and Streaming Pipelines

  • Batch ingestion architecture
  • Streaming ingestion architecture
  • Lambda architecture concepts
  • Event-driven architectures
  • Handling high-volume datasets
  • Selecting between batch and streaming
  • Combining historical and real-time data
  • Designing scalable ingestion pipelines

 

Day 5 – Security, Monitoring, Optimization and End-to-End Architecture

Module 17: Security for AWS Data Engineering

  • AWS Shared Responsibility Model
  • AWS IAM fundamentals
  • Users, roles, and policies
  • Least-privilege access
  • Service roles
  • Encryption at rest
  • Encryption in transit
  • AWS Key Management Service concepts
  • Securing S3 data
  • Securing databases and data pipelines
  • Secrets management concepts
  • Network security considerations

Module 18: Monitoring and Troubleshooting Data Pipelines

  • Amazon CloudWatch overview
  • Metrics and logs
  • Monitoring AWS DMS
  • Monitoring AWS Glue
  • Monitoring database workloads
  • Pipeline observability
  • Identifying failed jobs
  • Troubleshooting connectivity
  • Troubleshooting permissions
  • Troubleshooting data processing failures

Module 19: Performance and Cost Optimization

  • Storage optimization
  • Data partitioning
  • File format selection
  • Compression strategies
  • Query optimization
  • ETL performance considerations
  • AWS DMS performance considerations
  • Redshift optimization concepts
  • Serverless cost considerations
  • AWS data engineering cost-management practices

Module 20: End-to-End AWS Data Engineering Architecture

  • Identifying source systems
  • Data ingestion layer
  • Database migration and CDC layer
  • Data lake storage layer
  • Data catalog and governance layer
  • ETL/ELT transformation layer
  • Data warehouse and analytics layer
  • Batch and streaming integration
  • Security architecture
  • Monitoring and operational considerations
  • Designing an end-to-end AWS data pipeline
  • Reviewing production-ready data engineering patterns

 

Inquire now

Best selling courses

CLOUD COMPUTING

Terraform

Terraform is a configuration orchestration tool for building and managing infrastructure on cloud & data centers. The course is instructor-led, live training (onsite or remote), and is designed for Engineers with little or no previous experience managing infrastructure. The course talks about in-depth Terraform syntax and techniques used to automate the setup and deployment of infrastructure.

Duration  3 days – 21 hrs    Overview    The ITIL Leadership – Digital and IT Strategy training course is designed for senior IT professionals, managers, and leaders who seek to navigate the complex landscape of digital transformation and IT strategy. This course focuses on providing strategic insights, leadership skills, and practical approaches for aligning...

PROGRAMMING / CODING

Spring Architecture and Design

Spring Cloud is a platform for building Java-based distributed systems and microservices. Building complex enterprise applications is challenging. Any change made to a part of the systems could trigger the need for changing the design of the entire system. By the end of this training, participants will have a solid understanding of Service-Oriented Architecture (SOA) and Microservice Architecture as well practical experience using Spring Cloud and related Spring technologies for rapidly developing their own cloud-scale, cloud-ready microservices.

BUSINESS INTELLIGENCE

Dax

Duration 5 days – 35 hrs   Overview The DAX (Data Analysis Expressions) Training Course is designed to provide participants with a comprehensive understanding of DAX, the powerful formula language used in Power BI, Excel, and SQL Server Analysis Services. This course covers the essential concepts, functions, and techniques required to create advanced calculations and...

OPERATING SYSTEMS

Linux Fundamentals

Linux Fundamental provides students a thorough introduction to Linux™ for those who are new to the Linux environment. Delegates will learn how to manage files and directories, utilize the vi editor, work with Linux security mechanisms to protect files and programs, work with the Linux shell to control the flow and processing of data through pipelines, design and write shell programs of moderate complexity, and manage multiple concurrent processes in order to achieve higher utilization of Linux. They will learn how to perform basic operations on the system and how quickly to solve problem.

PROGRAMMING / CODING

Google Apps Script

The Google Apps Script training course give you a detailed knowledge on coding like Automating data calculation, Fetching and sending data from third party software like Trello & Salesforce, connecting different sheets, Documents and other tools, Setting a trigger based on an event. This course is ideal for someone who use google sheets and have no coding background.

This workshop teaches the participants how to design and develop server side applications using the event-driven, non-blocking model framework Node.js. This program inducts the participant in some of the advanced concepts of the JavaScript language so that the participant is well equipped to build end-to-end application using JavaScript.

Duration: 3 days – 21 hrs   Overview This training course is designed to provide participants with a comprehensive understanding of Portfolio Management and Contract Management, focusing on best practices, tools, and techniques. The course covers the strategic alignment of projects within a portfolio, effective management of contracts, risk management, and optimization of resources to...

// BG EARTH WHEN NOT PLAYING

We use cookies on our website to personalize your experience by storing your preferences and recognizing repeat visits. By clicking “Accept”, you agree to the use of all cookies. You can also select “Cookie Settings” to adjust your preferences and provide more specific consent. Cookie Policy