100 days of learning the things that actually matter in Data Science.
This is NOT another repository where I implement Linear Regression, Decision Trees, K-Means, CNNs, etc. from scratch.
Instead, this repository is about the small but important things that every Data Scientist should know β concepts, insights, mistakes, interview questions, practical tricks, statistical intuition, and lessons from working with real-world data.
Machine Learning is much more than knowing algorithms.
A good Data Scientist needs to know:
- How to think about a problem
- How to question the data
- How to detect data leakage
- How to choose the right metric
- How to design experiments
- How to interpret model results
- How to avoid misleading conclusions
- How models behave in the real world
- How to communicate findings
- And, of course, how to answer those "simple" interview questions that aren't always simple.
That's what this 100-day journey is about.
- Sampling mistakes
- Correlation vs causation
- Simpson's Paradox
- Selection bias
- Survivorship bias
- P-values and what they actually mean
- Confidence intervals
- Statistical significance vs practical significance
- Outliers
- Missing data
- Data distributions
- Feature leakage
- How A/B tests actually work
- Sample size
- Statistical power
- False positives & false negatives
- Multiple testing
- Peeking at experiments
- Novelty effects
- Guardrail metrics
- Experiment design mistakes
- Data leakage
- Overfitting vs underfitting
- Bias-variance tradeoff
- Cross-validation
- Feature engineering
- Feature selection
- Class imbalance
- Threshold tuning
- Calibration
- Model interpretability
- When a simpler model is better
- Why offline metrics can lie
- Why normalization helps
- Vanishing/exploding gradients
- Batch size trade-offs
- Learning rate intuition
- Dropout
- Batch normalization
- Embeddings
- Transfer learning
- Why training loss can decrease while performance gets worse
- What makes a good ML problem
- Data quality problems
- Defining the target variable
- Choosing the right metric
- Train/validation/test strategy
- Model monitoring
- Concept drift
- Data drift
- Retraining
- Model failure in production
Short explanations of questions that every Data Scientist should be able to answer.
Examples:
- Why do we split data into train and test?
- Why can't we train and test on the same data?
- Why is accuracy bad for imbalanced data?
- Why does scaling matter for some algorithms but not others?
- What is data leakage?
- What is the difference between correlation and causation?
- Why does cross-validation work?
- What happens if you increase model complexity?
- When would you choose precision over recall?
Small techniques that make working with data easier, faster, and more reliable.
| Day | Topic | Category |
|---|---|---|
| 01 | ... | Statistics |
| 02 | ... | Data |
| 03 | ... | ML |
| ... | ... | ... |
| 100 | ... | ... |
Don't just learn algorithms. Learn how to think like a Data Scientist.
Most ML tutorials teach you:
Algorithm β Code β Accuracy
Real-world Data Science is closer to:
Problem β Data β Assumptions β Experiment β Model β Evaluation β Decision
This repository focuses on everything in between.
The goal isn't to become an ML expert in 100 days.
The goal is to build a habit of asking:
"What is something about Data Science that I don't understand yet?"
One question. One concept. One experiment. Every day.
Tarun
B.S. Data Science β IIT Madras
Learning, experimenting and documenting the journey through Machine Learning, Data Science and AI.