Data Engineering & Analytics

Operational databases are built to run the product. Analytics needs something different: years of history, joins across every system the company owns, and queries that scan billions of rows. Data engineering is the discipline that bridges the two — extracting data from sources, transforming it into trustworthy models, and serving it to dashboards, analysts, and machine learning.

This hub covers the modern data stack end to end. Individual database engines live in Databases; streaming transport lives in Messaging & Event Streaming; modeling and ML live in Data Science & ML.

TL;DR

The Modern Data Stack

Featured Topics

Foundations

Storage & Query

Processing & Orchestration

Warehouse or Lakehouse?

Common Mistakes

🚫 Transforming in the ingestion tool — Business logic hidden in connectors is untested and unreviewable. Load raw, transform in dbt.

🚫 No ownership of datasets — When a dashboard breaks, nobody knows who fixes the source. Assign owners.

🚫 Full reloads forever — Fine at 10k rows, ruinous at 10 billion. Learn incremental models early.

🚫 Spark for small data — A single DuckDB or warehouse query often beats a cluster. Distribute only when you must.

🚫 Skipping data tests — Nulls, duplicates, and broken joins silently corrupt every downstream number.

Learning Path

Beginner

Learn SQL deeply, load a public dataset into a warehouse, and build a star schema following Data Warehousing.

Intermediate

Model it with dbt including tests, then schedule it with Airflow and add incremental loads.

Advanced

Design a lakehouse on Iceberg, query it from Trino and Spark, and run a data quality program with SLAs.

Related Topics