Recall:
"Use" a model
"Learn" a model
Regression
Algorithm
\(\mathcal{D}_\text{train}\)
🧠⚙️
hypothesis class
loss function
hyperparameters
regressor
"Use" a model
"Learn" a model
train, optimize, tune, adapt ...
adjust/update/find \(\theta\)
gradient-based
Regression
Algorithm
\(\mathcal{D}_\text{train}\)
🧠⚙️
hypothesis class
loss function
hyperparameters
regressor
predict, test, evaluate, infer ...
plug in the \(\theta\) found
no gradients involved
Recall:
Today:
{"good", "better", "best", ...}
\(\{0,1\}\)
\(\{😍, 🥺\}\)
{"fish", "grizzly", "chameleon", ...}
Classification
Algorithm
🧠⚙️
hypothesis class
loss function
hyperparameters
classifier
\(\mathcal{D}_\text{train}\)
linear regressor
linear binary classifier
features
parameters
linear combo
predict
\(x \in \mathbb{R}^d\)
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
\(\theta^T x +\theta_0\)
\(g = z\)
\(=z\)
if \(z > 0\)
otherwise
\(1\)
\(0\)
Today, \(z\) stands for \(\theta^T x +\theta_0\).
\(g=\)
label
\(y\in \mathbb{R}\)
\(y\in \{0,1\}\)
e.g., \(\theta = 1,\ \theta_0 = -1\)
\({J}_{01}(\theta)\) is very hard to optimize (NP-hard) 🥺
linear binary classifier
features
parameters
linear combo
predict
\(x \in \mathbb{R}^d\)
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
\(\theta^T x +\theta_0\)
\(=z\)
loss
linear regressor
optimize
method
\(y \in \mathbb{R}\)
\(y \in \{0,1\}\)
training error almost "flat" w.r.t. \(\theta,\) so the gradient gives very little info
\(\mathcal{L}_{01}(g, y)\) is "flat" and discrete in \(g\)
\(g\) is "flat" and discrete in \(\theta\)
label
linear binary classifier
features
parameters
linear combo
predict
\(x \in \mathbb{R}^d\)
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
\(\theta^T x +\theta_0\)
\(=z\)
linear logistic binary classifier
if \(z > 0\)
otherwise
\(1\)
\(0\)
if \(\sigma(z) > 0.5\)
otherwise
\(1\)
\(0\)
\(\sigma(z)\): confidence, or estimated probability that \(x\) belongs to the positive class
label
\(y \in \{0,1\}\)
linear logistic binary classifier
e.g., one feature
e.g., two features
features: \(x \in \mathbb{R}^d\)
parameters: \(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
the logit \(z\):
apply sigmoid:
Predict 1 if \(\sigma(z) > 0.5\), else 0.
the separator is linear in \(x\)!
One training point, label \(y = 1\)
😍
🥺
😍
🥺
A smooth loss \(\mathcal{L}(g, y)\) that shrinks as \(g\) nears \(y\)?
negative
log
likelihood
😍
🥺
😍
🥺
One training point, label \(y = 0\)
\(1-g\): the model's estimated probability that \(x\) belongs to the negative class.
Because the true label \(y\) is \(0\) or \(1\),
Read as: \(\sum\) (true label for class \(k\)) \(\cdot\) \(-\log\)(estimated probability of class \(k\)).
Since \(y \in \{0,1\}\), only the true class's term survives.
linear binary classifier
linear logistic binary classifier
features
label
parameters
linear combo
predict
loss
optimize via
\(x \in \mathbb{R}^d\)
\(y \in \{0,1\}\)
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
\(\theta^T x +\theta_0 = z\)
no efficient method (NP-hard)
gradient descent
One training point, \(x = 1\), label \(y = 1\)
Video edited from: HBO, Silicon Valley, 2015
ad as seen on a MBTA train, 2026
🌭
\(x\)
\(\theta^T x +\theta_0\)
\(z \in \mathbb{R}\)
for two classes, {hotdog, not_hotdog}, one scalar logit \(z\) suffices
scalar logit,
raw score for hotdog
\(1-\sigma(z):\) estimated probability that \(x\) belongs to not_hotdog
\(\sigma(z):\) estimated probability that \(x\) belongs to hotdog
sigmoid
normalizing (squashing)
implicitly determines
for \(K > 2\) classes, one scalar \(z\) no longer suffices, so we use \(K\) logits, one per class
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
🌭
\(x\)
\(\theta^T x +\theta_0\)
\(z \in \mathbb{R^3}\)
estimated probability that \(x\) belongs to hotdog
normalizing (squashing)
for \(K\) classes, use \(K\) logit scores.
e.g. \(K = 3\): \(\{\)hotdog, pizza, veggie\(\}\)
… to pizza
… to veggie
in general \(K\) logits
one raw score per class
\(\theta \in \mathbb{R}^{d \times K},\)
\(\theta_0 \in \mathbb{R}^{K}\)
outputs lie in \((0, 1)\) and sum to \(1\)
max among the \(K\) logits
"soft" max: the largest logit gets the largest probability
e.g.,
sigmoid
predict the class with the largest softmax score
softmax:
predict positive if \(\sigma(z)>0.5 = \sigma(0)\)
unifying rule: predict the class with the largest logit (= largest softmax score)
implicit logit for the negative class
features
parameters
linear combo
predict
\(x \in \mathbb{R}^d\)
\(\theta \in \mathbb{R}^d, \theta_0 \in \mathbb{R}\)
\(\theta^T x +\theta_0\)
\(=z \in \mathbb{R}\)
linear logistic
binary classifier
one-out-of-\(K\) classifier
\(\theta \in \mathbb{R}^{d \times K},\)
\(=z \in \mathbb{R}^{K}\)
\(\theta^T x +\theta_0\)
predict positive if \(\sigma(z)>\sigma(0)\)
predict the class with the largest softmax score
\(\theta_0 \in \mathbb{R}^{K}\)
One-hot encoding:
Training data
| \(x\) | \(y\) | |||
| ( | 🌭 | , | "hotdog" | ) |
| ( | 🍕 | , | "pizza" | ) |
| ( | 🥗 | , | "veggie" | ) |
| ( | 🥦 | , | "veggie" | ) |
| \(\vdots\) | ||||
Training data
| \(x\) | \(y\) | |||
| ( | 🌭 | , | \(\begin{bmatrix}1\\0\\0\end{bmatrix}\) | ) |
| ( | 🍕 | , | \(\begin{bmatrix}0\\1\\0\end{bmatrix}\) | ) |
| ( | 🥗 | , | \(\begin{bmatrix}0\\0\\1\end{bmatrix}\) | ) |
| ( | 🥦 | , | \(\begin{bmatrix}0\\0\\1\end{bmatrix}\) | ) |
| \(\vdots\) | ||||
in general, for \(K\) classes:
Only the true class's term survives the \(K\)-term sum, since every other \(y_k=0\)
Negative log-likelihood multi-class loss (also called cross-entropy)
\(y:\) one-hot encoding label
\(y_{{k}}:\) \(k\)th entry in \(y\), either 0 or 1
\(g:\) softmax output
\(g_{{k}}:\) estimated probability that \(x\) belongs to class \(k\)
🌭
current prediction \(g=\text{softmax}(z)\)
end-to-end pipeline
contrived predictions
To reduce the loss, \(g_{\text{hotdog}}\) needs to go up.
That signal flows smoothly back to \(\theta\) through \(-\!\log\) and softmax, so gradient descent can train \(\theta\).
🌭
current prediction \(g=\text{softmax}(z)\)
end-to-end pipeline
contrived predictions
Linear (binary) classifiers
separatornormal vector0-1 loss
Linear logistic (binary) classifiers
sigmoidlogitnegative log-likelihoodlinearly separable
Linear multi-class classifiers
softmaxone-hot encodingcross-entropy
\(\mathcal{L}_{\text{nll}}(g, y) = -\left[y \log g + (1-y) \log (1-g)\right], \quad g = \sigma(\theta^T x + \theta_0)\)
Linear (binary) classifiers
to use: predict 1 when \(z = \theta^T x + \theta_0 > 0\); the separator \(z = 0\) is a hyperplane with normal \(\theta\).
to learn: the 0-1 loss is flat in \(\theta\), so there is no gradient to follow.
Linear logistic (binary) classifiers
to use: \(g = \sigma(z)\) is a probability, and \(g > 0.5\) exactly when \(z > 0\), so the separator stays linear.
to learn: the negative log-likelihood is smooth, so gradient descent works; on separable data, a ridge penalty keeps \(\theta\) finite.
Linear multi-class classifiers
to use: softmax turns one score per class into a distribution; sigmoid is its two-class case.
to learn: one-hot labels and the cross-entropy loss, \(-\log\) of the true class's probability.