Data Science & Machine Learning

Data science is the practice of asking questions of data and answering them credibly — cleaning messy tables, exploring distributions, testing hypotheses, and communicating what the numbers actually say. Machine learning goes one step further: instead of writing rules, you let a model learn them from examples, then use it to predict, classify, rank, or generate.

This hub covers the classical and deep-learning toolkit and the engineering needed to ship models. Large language models and generative AI have their own hub at AI & Generative AI; pipelines that feed models live in Data Engineering & Analytics.

TL;DR

The ML Workflow

Featured Topics

Foundations

Modeling

Domains

Production

Which Tool for Which Job

Common Mistakes

🚫 Data leakage — Features that encode the answer (or future information) produce great validation scores and useless models.

🚫 Accuracy on imbalanced data — 99% accuracy means nothing when 99% of rows are negative. Use precision, recall, and PR-AUC. See Model Evaluation.

🚫 Skipping the baseline — Without a simple benchmark you can't tell whether a deep model is worth its cost.

🚫 Training/serving skew — Features computed differently in production than in training silently degrade predictions.

🚫 Deploy and forget — Data drifts. Monitor inputs and outcomes and retrain on a schedule or trigger.

Learning Path

Beginner

Learn Python for data science and pandas. Do exploratory analysis on a public dataset and train a first scikit-learn model.

Intermediate

Practice feature engineering and rigorous evaluation. Train a neural network in PyTorch and try an NLP or vision task with pretrained models.

Advanced

Build an MLOps pipeline with tracking, a registry, automated retraining, and drift monitoring. Scale feature pipelines with Spark.

Related Topics