Theory and mathematical foundations of statistical learning and artificial intelligence.
The data science workflow, the distinction between supervised and unsupervised learning, and the prediction vs. inference framing that runs through the entire course.
The formal Y = f(X) + epsilon framework, mean squared error, the bias-variance decomposition, polynomial curve fitting, classification error rate, and the probability foundations the rest of the course depends on.
Simple and multiple regression, ordinary least squares, qualitative predictors, interaction terms, polynomial extensions, and model diagnostics. Follows ISLR Ch. 3.
Logistic regression, maximum likelihood estimation, multiclass logistic regression, linear and quadratic discriminant analysis, Naive Bayes, and ROC curves. Follows ISLR Ch. 4.
Validation set approach, LOOCV, k-fold cross-validation, and the bootstrap. Best subset selection, forward and backward stepwise selection, with Cp, AIC, BIC, and adjusted R² as criteria. Follows ISLR Ch. 5 and Ch. 6.
Ridge regression, the Lasso, adaptive Lasso, group Lasso, elastic net, adaptive elastic net, bridge regression, and total variation regularization, with tuning parameter selection strategies.
Principal component analysis, principal components regression, partial least squares, K-means clustering, Fuzzy C-means, and hierarchical clustering with dendrograms. Follows ISLR Ch. 6 and Ch. 10.
Regression trees, classification trees, recursive binary splitting, cost-complexity pruning, the Gini index, and entropy as impurity measures. Follows ISLR Ch. 8.
Bagging, random forests with variable importance and out-of-bag error, AdaBoost, and gradient boosting — how combining weak learners produces strong predictors. Follows ISLR Ch. 8.
Definition and history of AI, major application domains, societal impact, the principal subfields of AI, and the components that make up a modern AI system.
Formalising problems as state spaces, uninformed strategies (BFS, DFS), and informed strategies including best-first search and the A* algorithm with heuristics.
Definition and sources of big data, the 5 Vs, the challenges of scale, and the Hadoop ecosystem including HDFS, MapReduce, YARN, and Hive.
End-to-end analyses on real datasets using methods from Modules 2 through 9, with complete Python code. Covers regression, classification, regularisation, clustering, and tree methods on applied problems.