Data engineering course

A data engineering course built on exercises, not slides

14 sequenced tracks, 300+ hands-on exercises, 20–30 minutes a day. You write SQL, break pipelines, fix them, and finish with real projects on Docker, Airflow, dbt, Spark and a cloud warehouse. Start with 7 days free — no credit card.

What you'll build

  • An incremental ingestion job that is safe to re-run (idempotency, watermarks, retries).
  • A dimensional model with a star schema and SCD Type 2 history.
  • A dbt project with tests, snapshots and generated lineage on a medallion layout.
  • An Airflow DAG with backfills, sensors and alerting that survives a failed upstream.
  • A PySpark job you tuned yourself after reading the Spark UI.
  • A Kafka consumer with explicit delivery semantics.

The 14-track curriculum, in order

1. SQL foundationsSELECT to CTEs and window functions — the language you will use every single day.
2. Python for dataIdiomatic Python, typing, pandas, requests, and writing code that survives a code review.
3. Linux & DockerShell fluency, images, volumes, docker-compose — how your pipeline actually ships.
4. Data modelingNormalization, star schema, slowly changing dimensions, grain and keys.
5. WarehousingSnowflake/BigQuery: virtual warehouses, clustering, cost control, query tuning.
6. Ingestion & APIsBatch and incremental extraction, pagination, retries, idempotency.
7. dbtModels, sources, tests, snapshots, incremental strategies, documentation and lineage.
8. Orchestration with AirflowDAGs, sensors, backfills, retries, SLAs and the mistakes that page you at 3am.
9. Spark & PySparkPartitions, shuffles, joins, skew, broadcast, and reading the Spark UI.
10. Streaming with KafkaTopics, partitions, consumer groups, offsets, delivery semantics.
11. Lakehouse & IcebergOpen table formats, ACID on object storage, time travel, schema evolution.
12. Data qualityContracts, expectations, freshness and volume tests, alerting that people trust.
13. Cloud & IaCS3/GCS, IAM, cost awareness, Terraform basics for data infrastructure.
14. Interview prepSQL rounds, system design for pipelines, and defending your architecture choices.

How the lessons work

Every topic follows the same loop: a short explanation with a concrete code example, then exercises where you produce the answer. No 4-hour videos you never finish.

  • Explain — the concept, why it exists, and where it breaks in production.
  • Show — a real snippet (SQL, Python, YAML) you can copy and run.
  • Do — graded exercises, XP, streaks and badges to keep the habit.
  • Break — the daily Bug Hunt: find the defect in real pipeline code.
  • Zoom outarchitecture deep dives into how Netflix, Uber, Airbnb and Spotify run data at scale.

Who it's for

  • Analysts moving into engineering who know SQL but not orchestration or Spark.
  • Software engineers pivoting to data who need modeling and warehouse fundamentals.
  • Career switchers starting from zero who want a path instead of 60 browser tabs.
  • Working data engineers filling gaps before an interview loop.

Free vs member

 Free trialMember
Duration7 days, no cardMonthly, cancel anytime
TracksFull accessFull access
Exercises300+300+
Bug Hunt & badgesYesYes
Architecture deep divesYesYes
New content each monthIncluded

Start today

Open the first SQL lesson, finish it in 20 minutes, and keep the streak. If you want the map first, read the data engineer roadmap — then come back and become a member to unlock the whole path.

FAQ

What does a data engineering course need to cover?
SQL, Python, Linux and Docker, one cloud, a warehouse (Snowflake or BigQuery), orchestration (Airflow), transformations (dbt), distributed processing (Spark), streaming (Kafka), open table formats (Iceberg or Delta), plus data modeling and quality. DataForge covers all of it in 14 sequenced tracks.
Is this data engineering course good for beginners?
Yes. It starts at SQL basics and Python fundamentals and assumes no prior data experience. Each lesson is 20–30 minutes with hands-on exercises, so you can complete it around a full-time job.
How long does the course take?
At 20–30 minutes a day, expect 4–6 months to reach a junior-employable level and 12–18 months to cover the full path including Spark, Kafka and Iceberg.
How much does it cost?
There's a 7-day free trial with no credit card, then a single monthly membership that unlocks every track, all 300+ exercises, the Bug Hunt and the architecture deep dives. Cancel anytime.
Do I get a certificate?
You earn badges and a completion record per track. In hiring, the 2–3 end-to-end pipelines you build here matter far more than a PDF — which is why the course is exercise-first.
Is it better than a bootcamp?
It costs a fraction of a bootcamp, is self-paced, and covers the same stack with more repetition. What a bootcamp adds is cohort pressure and career services — if you can hold a daily habit, this is the better trade.

Ready to become a member?

7 days free. Then less than a coffee per month — cancel anytime.

Become a member — 7 days free