Data Engineering Training Behavioural and case interviews
Data Engineering Training
0% Completed
M1 — Data Engineering Foundations
The data engineering lifecycle
Reading
Batch vs streaming trade-offs
Reading
Source-of-truth modelling concepts
Reading
Lab environment setup
Reading
M2 — SQL and Python for Data Engineers
Window functions and CTEs
Reading
Pandas and Polars essentials
Reading
PySpark in 90 minutes
Reading
Idempotent pipelines
Reading
M3 — Relational Modelling and Warehousing
Star schema and snowflake
Reading
Slowly Changing Dimensions (SCD2)
Reading
Surrogate keys and audit columns
Reading
dbt core models and snapshots
Reading
M4 — ETL and ELT Pipelines
Extract patterns: CDC, APIs, files
Reading
Transform with SQL, Spark, dbt
Reading
Load: warehouse, lake, lakehouse
Reading
Error handling and retries
Reading
M5 — Orchestration with Apache Airflow
DAGs, operators and sensors
Reading
Backfills and SLAs
Reading
Idempotency and retries
Reading
Production-grade Airflow on Kubernetes
Reading
M6 — Streaming with Apache Kafka
Topics, partitions, consumer groups
Reading
Exactly-once semantics
Reading
Schema Registry and Avro
Reading
Stream-table duality with Kafka Connect
Reading
M7 — Distributed Compute with Apache Spark
Spark architecture and DAGs
Reading
Structured Streaming APIs
Reading
Partitioning and shuffle tuning
Reading
Adaptive Query Execution
Reading
M8 — Lakehouse and Delta Lake
Medallion (bronze/silver/gold) architecture
Reading
Delta time travel and OPTIMIZE
Reading
Schema enforcement and MERGE
Reading
Unity Catalog basics
Reading
M9 — Data Quality and Observability
Great Expectations / Soda checks
Reading
Data contracts and lineage
Reading
Monitoring freshness, volume, quality
Reading
Alerting with PagerDuty / Slack
Reading
M10 — CI/CD and Infra-as-Code for Data
Git, CI pipelines and pull-request reviews
Reading
IaC with Terraform for data stacks
Reading
Schema migrations
Reading
Promoting dbt models across envs
Reading
M11 — Performance, Cost and Security
Query optimisation and Z-ordering
Reading
Cost tuning across Snowflake / Databricks
Reading
Row/column-level security
Reading
PII handling and masking
Reading
M12 — Capstone and Interview Prep
End-to-end lakehouse build
Reading
Failure-mode simulation
Reading
Behavioural and case interviews
Reading
Announcements Course Info
PreviousNext