Machine Learning Diagrams

Last updated:

Rules Versus Examples

Flow

A hand-written spam filter beside one trained from labeled emails.

A hand-written program compared with a trained model In a hand-written program, a programmer writes rules, and the program applies them to each new email to answer spam or not spam. Every new spam tactic needs a new rule. In machine learning, a training algorithm reads ten thousand emails that people have already labeled and produces a model of learned values. The model then takes the place of the program: a new email goes in and an answer comes out. The programmer chooses the examples and how success is measured instead of writing the rules. HAND-WRITTEN PROGRAM Rules written by a programmer New email subject, sender, links Program applies the rules the programmer wrote Answer spam or not spam every new spam tactic needs a new rule MACHINE LEARNING Labeled examples 10,000 emails marked by people Training algorithm finds the separating patterns New email subject, sender, links Model learned values Answer spam or not spam the programmer chooses the examples and how success is measured

Loss on a Line Through House Prices

Chart

Two candidate lines through the same eight houses, scored by loss.

Two lines fitted to house prices and the loss of each Eight houses are plotted by floor area and sale price. On the left, a poor guess, price = 0.5 times area plus 170, leaves long vertical gaps between each house and the line, and its average squared error is 1,676. On the right, the line after training, price = 1.5 times area plus 56, runs through the middle of the points with short gaps, and its average squared error is 97. A poor guess price = 0.5 × area + 170 50 100 150 200 100 200 300 price ($k) floor area (m²) average squared error: 1,676 After training price = 1.5 × area + 56 50 100 150 200 100 200 300 price ($k) floor area (m²) average squared error: 97 one house's error the model's line

Gradient Descent on a Loss Curve

Chart

Steps down the same loss curve at three learning rates.

Gradient descent at three learning rates Each panel plots loss against one weight as a U-shaped curve with its lowest point in the middle, and marks the steps gradient descent takes from the same starting weight on the left. With a learning rate that is too small, six steps creep partway down the left side and are still far from the bottom. With a learning rate about right, the steps reach the bottom in four moves. With a learning rate that is too large, each step jumps past the bottom to a higher point on the other side, and the steps climb out of the curve. Learning rate too small loss weight w lowest loss six steps and still far from the bottom Learning rate about right loss weight w lowest loss settles at the bottom in four steps Learning rate too large loss weight w lowest loss overshoots further on every step and diverges

The Training Loop

Flow

Predict, measure the loss, adjust the parameters, then repeat.

The training loop A batch of training examples supplies features to the model, which uses its current parameters to predict a price for each house. The predictions and the batch's labels, the correct prices, both feed the loss, a single number measuring how far off the predictions are. The optimizer takes the loss, computes the gradient, and steps each parameter. The updated model then receives the next batch, and the loop repeats. Batch of examples features and labels for, say, 32 houses Model current parameters Predictions a price for each house Loss one number: how far the predictions are from the labels Optimizer computes the gradient and steps each parameter features loss updated parameters, then the next batch labels, the correct prices

A Small Decision Tree for House Prices

Structure

Three learned questions sorting houses into four price leaves.

A decision tree that predicts house prices The root asks whether floor area is under 120 square meters, splitting 200 training houses into 96 and 104. The smaller houses are split by whether they were built over 30 years ago, giving leaves that predict 165 thousand dollars for the older ones and 205 thousand for the newer ones. The larger houses are split by whether they have four or more bedrooms, giving leaves that predict 330 thousand and 265 thousand. Each leaf predicts the average price of the training houses that ended there, and a new house is sorted down the tree to the leaf it lands in. Floor area under 120 m²? all 200 houses Built over 30 years ago? 96 houses Four or more bedrooms? 104 houses yes no $165k average of 41 houses $205k average of 55 houses $330k average of 38 houses $265k average of 66 houses yes no yes no A house that hasn't sold is sorted down the tree and gets the price of the leaf it lands in.

A Unit and a Small Network

Structure

One unit's weighted sum and bend, then units stacked in layers.

A single neural network unit and a small network built from them On the left, one unit takes floor area, bedrooms, and age, multiplies each by its own weight, adds them together with a bias, and passes the total through a bend called ReLU, which turns a negative total into zero, to produce one output number. On the right, a small network has three input units for area, bedrooms, and age, two hidden layers of four units each, and one output unit for price. Every unit in one layer connects to every unit in the next, and every connecting line carries one learned weight. ONE UNIT A SMALL NETWORK floor area × w₁ bedrooms × w₂ age × w₃ sum + b Bend ReLU Output one number Multiply each input by its weight, add them up with a bias b, then bend: ReLU turns a negative total into zero. area beds age input hidden 1 hidden 2 price Every line carries one learned weight. The darker lines are the four weights applied to floor area.

Classification, Regression, and Clustering

Chart

What each task learns, drawn on the same kind of scatter plot.

What classification, regression, and clustering each learn Three scatter plots. In classification, emails are plotted by number of links and by words like free, colored as spam or not spam, and the model learns a dashed boundary between the two colors. In regression, sold houses are plotted by floor area and price, and the model learns a line through them. In clustering, customers are plotted by orders per month and average order value with no labels at all, and the algorithm finds three groups on its own, circled A, B, and C. Classification links in the email words like 'free' spam not spam boundary labels are categories, so it learns a boundary Regression floor area price sold house learned line labels are numbers, so it learns a line or curve Clustering orders per month average order value A B C customer group found no labels at all, so it finds the groups itself

The Reinforcement Learning Loop

Flow

An agent acts, and the environment returns a state and reward.

The reinforcement learning loop An agent, here a game-playing program, chooses an action according to its policy and sends it to the environment, the game itself. The environment responds with a new state, the board after the move, and a reward, the points won or lost, which is often zero. The agent chooses its next action from the new state, and the loop repeats. Agent the game-playing program chooses moves by its policy Environment the game itself responds to every move action: a move new state: the board after the move reward: points won or lost, often zero

State Values in a Small Grid

Chart

Each cell's value when reward is discounted by 0.9 per step.

Discounted state values in a grid with a wall A grid of five columns and three rows has the goal in the top-right corner, worth plus ten, and a wall filling the top two cells of the middle column. Each other cell shows its value with a discount of 0.9 per step. Cells next to the goal are worth 10.0, cells two steps away 9.0, then 8.1, 7.3, 6.6, 5.9, 5.3, and 4.8 for the top-left corner. The cell just left of the wall is worth only 5.3 although it is close to the goal in a straight line, because the path around the wall is seven steps long. An arrow in each cell points toward the neighbor with the highest value. 4.8 5.3 5.9 5.3 5.9 6.6 wall wall 7.3 10.0 9.0 8.1 GOAL +10 10.0 9.0 Setup move one cell per step +10 for reaching the goal 0 for every other move discount γ = 0.9 Reading a cell value = reward for the move + 0.9 × value of the cell it leads to best move from here

Holding Data Out for Evaluation

Structure

A train, validation, and test split, and five-fold cross-validation.

A holdout split and five-fold cross-validation At the top, one dataset is divided into a training set of 70 percent that fits the parameters, a validation set of 15 percent that tunes and compares models, and a test set of 15 percent used for one final check. Below, five-fold cross-validation divides the non-test data into five parts. In each of five rounds a different part is used for validation and the other four for training, producing scores of 0.81, 0.84, 0.79, 0.83, and 0.83, which average to 0.82. The test set stays untouched until the very end. ONE HOLDOUT SPLIT Training set 70% Validation 15% Test 15% fits the parameters tunes and compares one final check FIVE-FOLD CROSS-VALIDATION ON THE NON-TEST DATA round 1 validate train train train train 0.81 round 2 train validate train train train 0.84 round 3 train train validate train train 0.79 round 4 train train train validate train 0.83 round 5 train train train train validate 0.83 0.82 average score Test set untouched until the very end

Underfitting, a Good Fit, and Overfitting

Chart

Three models of rising flexibility fitted to the same points.

Underfit, good fit, and overfit models on the same data The same curved data appears in three panels, with filled training points used for fitting and hollow validation points held out. A straight line underfits: it misses the curve, so both training and validation error are high. A gentle curve fits well: it passes near both kinds of point, so both errors are low. A curve that passes through every training point overfits: training error is zero, but the curve swings far from the held-out validation points between them, so validation error is high. Underfit a straight line training error: high validation error: high Good fit a gentle curve training error: low validation error: low Overfit a curve through every point training error: zero validation error: high training point, used to fit validation point, held out the model

The Confusion Matrix for a Fraud Model

Structure

A thousand transactions sorted by prediction and truth.

A confusion matrix for a fraud model with precision and recall Of 1,000 transactions, 10 are fraud. The matrix sorts them by what the model predicted and what was true. It caught 8 fraudulent transactions, missed 2, raised 4 false alarms on legitimate ones, and correctly left 986 legitimate ones alone. Precision reads down the predicted-fraud column: 8 of the 12 flagged were fraud, 67 percent. Recall reads across the actually-fraud row: 8 of the 10 frauds were caught, 80 percent. Accuracy is 99.4 percent, but a model that flags nothing scores 99.0 percent accuracy with 0 percent recall. 1,000 TRANSACTIONS, 10 OF THEM FRAUD predicted fraud predicted legitimate actually fraud actually legitimate 8 fraud caught 2 fraud missed 4 false alarms 986 legitimate, left alone precision = 8 / (8 + 4) = 67% of everything flagged, how much was fraud recall = 8 / (8 + 2) = 80% of all fraud, how much was caught accuracy = (8 + 986) / 1,000 = 99.4% a model that flags nothing: 99.0% accuracy, 0% recall

Model Deployment Patterns

Flow

How five deployment patterns split live traffic between two models.

How shadow, canary, blue-green, A/B, and bandit deployments route traffic Five rows, each routing live traffic to a current model and a new model. Shadow: all users are served by the current model while the new model receives a copy of the traffic and its predictions are logged, not used. Canary: most traffic goes to the current model and a small share goes to the new one, rising as metrics hold. Blue-green: traffic switches from the current environment to the new one at once, and rollback is instant. A/B test: traffic splits between the two models long enough to compare outcomes. Multi-armed bandit: traffic shifts toward whichever model performs better while the test runs. PATTERN WHERE LIVE TRAFFIC GOES WHAT USERS SEE Shadow copy of live traffic Live traffic Current model New model current model's predictions new: logged, not used Canary small share, then more Live traffic Current model New model most users a small share, rising as metrics hold Blue-green switch over at once Live traffic Current model New model before the switch everyone, after one cutover, with instant rollback A/B test split to compare outcomes Live traffic Current model New model one group another group, long enough to compare outcomes Multi-armed bandit shift toward the better one Live traffic Current model New model less, if it performs worse more, while it performs better Line width shows the share of live traffic. Dashed: a copy that reaches a model but not users. Pale: traffic before the switch.

Data Drift and Concept Drift

Chart

Inputs shifting versus the right answer shifting.

Data drift compared with concept drift Left: data drift. A loan model's training data centers on applicants aged 30 to 50, while live inputs center on applicants in their early twenties. The learned relationship may still hold, but the model predicts where it saw little data, and the shift is detectable without labels. Right: concept drift. For the same transaction pattern, the chance it is fraud was low at training and is high now. The inputs look familiar and the model is wrong anyway, and detecting it needs the correct answers. Data drift the inputs move applicant age 20 30 40 50 60 70 training data live inputs The learned relationship may still hold, but the model now predicts where it saw little data. Detectable without labels. Concept drift the right answer moves transaction pattern chance it is fraud at training: legitimate now: fraud same pattern The inputs look familiar, and the model is wrong anyway. Detecting it needs the correct answers.

Found this useful? Share it:

Share on LinkedIn