01 — Notebook Information & Scope
“All models are wrong, but some are useful.” — George E. P. Box. The objective of predictive modeling is not to achieve zero in-sample training error, but to minimize out-of-sample generalization error on unseen data.
- Domain: Analytics & Quantitative Methods
- Subject: Predictive Modeling
- Core Reference Model: Bias-Variance Decomposition & Elastic Net Regularization
02 — The Bias-Variance Trade-off
The expected test Mean Squared Error (MSE) of a supervised learning model decomposes into three irreducible quantities:
Bias-Variance Decomposition
FORMULAAs model complexity increases, bias declines while variance increases. Regularization techniques penalize extreme model coefficients to strike the optimal bias-variance balance.
$\mathrm{Bias}[\hat{f}(x)]^2$Error introduced by approximating a complicated real-world problem with a simpler model$\mathrm{Var}[\hat{f}(x)]$Amount by which the model estimate would change if estimated using a different training set$\sigma^2$Irreducible random error inherent in the true data generating process03 — Regularized Regression: Ridge, Lasso & Elastic Net
L1 vs L2 Regularization Dynamics
COMPARISONShrinks coefficients smoothly toward zero but never forces them to exact zero. Retains all predictors in the final model; well-suited when many predictors have small, non-zero effects.
Enforces sparsity by driving coefficients to exact zero when penalty lambda is sufficiently large. Performs automated feature selection; well-suited for high-dimensional sparse data.