Data Science & Machine Learning
Data science is the practice of asking questions of data and answering them credibly — cleaning messy tables, exploring distributions, testing hypotheses, and communicating what the numbers actually say. Machine learning goes one step further: instead of writing rules, you let a model learn them from examples, then use it to predict, classify, rank, or generate.
This hub covers the classical and deep-learning toolkit and the engineering needed to ship models. Large language models and generative AI have their own hub at AI & Generative AI; pipelines that feed models live in Data Engineering & Analytics.
TL;DR
- Python is the lingua franca — pandas for tables, scikit-learn for classical ML, PyTorch for deep learning.
- Start with a baseline. A simple model you understand beats a complex one you can't evaluate.
- Data and features matter more than algorithms. Feature engineering is where most gains come from.
- Evaluate honestly. Hold out data, pick metrics that match the business cost of errors, and watch for leakage.
- A model in a notebook isn't a product. MLOps covers versioning, deployment, and monitoring.
The ML Workflow
Featured Topics
Foundations
- Data Science — The discipline: questions, statistics, and communicating results
- Python for Data Science — NumPy, Jupyter, and the scientific Python stack
- pandas — DataFrames, cleaning, grouping, and joins
- Feature Engineering — Turning raw columns into signals models can use
Modeling
- scikit-learn — Classical ML: pipelines, models, and cross-validation
- PyTorch — The research-to-production deep learning framework
- TensorFlow — Google's deep learning platform and Keras
- Model Evaluation — Metrics, validation strategies, and avoiding leakage
Domains
- Natural Language Processing — Text classification, embeddings, and transformers
- Computer Vision — CNNs, detection, segmentation, and vision transformers
Production
- MLOps — Experiment tracking, model registries, serving, and drift monitoring
Which Tool for Which Job
Common Mistakes
🚫 Data leakage — Features that encode the answer (or future information) produce great validation scores and useless models.
🚫 Accuracy on imbalanced data — 99% accuracy means nothing when 99% of rows are negative. Use precision, recall, and PR-AUC. See Model Evaluation.
🚫 Skipping the baseline — Without a simple benchmark you can't tell whether a deep model is worth its cost.
🚫 Training/serving skew — Features computed differently in production than in training silently degrade predictions.
🚫 Deploy and forget — Data drifts. Monitor inputs and outcomes and retrain on a schedule or trigger.
Learning Path
Beginner
Learn Python for data science and pandas. Do exploratory analysis on a public dataset and train a first scikit-learn model.
Intermediate
Practice feature engineering and rigorous evaluation. Train a neural network in PyTorch and try an NLP or vision task with pretrained models.
Advanced
Build an MLOps pipeline with tracking, a registry, automated retraining, and drift monitoring. Scale feature pipelines with Spark.
Related Topics
- AI & Generative AI — LLMs, agents, and generative models
- Data Engineering & Analytics — Pipelines and warehouses that feed models
- Python — The primary language of the field
- Databases — Where training data usually starts