Most of what gets called AI today is machine learning: programs that learn a pattern from examples instead of following rules someone wrote by hand. A handful of ideas explain how nearly all of it works, and once they click, the jargon in any ML tutorial stops being a wall.
What AI actually means
AI stands for artificial intelligence: getting computers to do tasks that normally need human judgement, such as recognising a face, translating a sentence or spotting a fraudulent payment. It is a broad umbrella. Older AI systems were mostly hand-written rules ('if the email contains these words, mark it as spam'), and they broke as soon as the world did something the rules didn't expect.
Machine learning flips that around. Instead of writing the rules, you give the program lots of examples and let it work out the rules itself. The result of that process is a model: a function with adjustable numbers inside (its parameters) that takes an input and produces a prediction.
Automation is a different thing. A script that renames files every night is automated, but it doesn't learn anything and never gets better at the job. Keeping those two apart saves a lot of confusion when people call every clever program 'AI'.
Supervised and unsupervised learning
The biggest split in machine learning is whether your examples come with answers.
Supervised learning
In supervised learning, every training example has a label: the right answer for that input. A dataset of emails, each marked 'spam' or 'not spam', is labelled data. So is a table of houses with their sale prices. The model sees the inputs, makes a guess, compares it with the label and adjusts itself to do better next time. It is 'supervised' because the labels act like a teacher marking its work.
Most practical ML is supervised, because most business questions have a known answer you can collect examples of. The catch is that labels cost time and money: someone has to mark all those emails.
Unsupervised learning
In unsupervised learning, the data has no labels at all. The model's job is to find structure on its own. The classic example is clustering: give it a year of customer purchases and it groups customers who behave alike, without anyone saying what the groups should be. Other uses include spotting unusual transactions and squeezing many measurements down into a few that capture most of the variation.
The trade-off is that nobody can say whether an unsupervised result is 'right'. The groups it finds may be useful, or they may be patterns a human has to interpret carefully.
Features and feature vectors
A model can't look at a house or a flower. It reads numbers, so every example has to be described as a set of measurable inputs called features. For a house, the features might be floor area, number of bedrooms and distance to the station. The price you want to predict is not a feature: it is the label, or target.
Put all of one example's features in a fixed order and you get a feature vector. The well-known iris flower dataset describes each flower with four measurements, so each flower becomes a feature vector of four numbers:
# sepal length, sepal width,
# petal length, petal width (cm)
flower = [5.1, 3.5, 1.4, 0.2]A whole dataset is then a table of these vectors, one row per example, with a matching column of labels for supervised learning. The order matters: if petal length is in the third slot for one flower, it must be in the third slot for every flower, or the model learns nonsense.
Choosing good features is often worth more than choosing a clever algorithm. A model predicting house prices without floor area will struggle however sophisticated it is.
Classification and regression
Supervised problems come in two main shapes, depending on what the label looks like.
- Classification predicts a category. Is this email spam or not? Which of three iris species is this flower? Is this X-ray normal or abnormal? The answer comes from a fixed set of options.
- Regression predicts a number on a continuous scale: a house price, tomorrow's temperature, how many minutes a delivery will take.
The difference changes how you measure success. For classification you ask how often the model picked the right category. For regression you ask how far off its numbers were. Classification is not the same as sorting data into order: it assigns each example a category, and the examples stay where they are.
How a model learns: loss and gradient descent
Training a supervised model is a loop:
- The model makes predictions for a batch of examples.
- A loss function measures how wrong those predictions are, as a single number.
- The training algorithm nudges the parameters to make that number smaller.
- Repeat, many times.
The loss function
The loss turns 'how bad is the model?' into something a computer can minimise. For regression, a common choice is mean squared error: take each prediction's distance from the right answer, square it and average the lot. Squaring makes every error positive and punishes big misses much more than small ones. Classification models usually use a different loss, but the idea is the same: lower means better.
Gradient descent
Gradient descent is the most common way to lower the loss. Picture the loss as a hilly landscape, where your position is the model's current parameters and your height is the loss. The gradient tells you which direction is uphill, so you take a small step the opposite way, downhill. Then you check the slope again and step again, until the ground flattens out.
The size of each step is the learning rate. Too small and training crawls. Too large and each step overshoots the bottom, so the loss bounces around or even grows.
Here is gradient descent learning a single parameter, w, for the rule y = w * x. The data follows y = 2x, and the model starts from a bad guess of 0:
xs = [1, 2, 3, 4]
ys = [2, 4, 6, 8] # true rule: y = 2x
w = 0.0 # starting guess
lr = 0.01 # learning rate
for step in range(200):
grad = 0
for x, y in zip(xs, ys):
# slope of squared error for w
grad += 2 * (w * x - y) * x
grad /= len(xs)
w -= lr * grad # step downhill
print(round(w, 3)) # 2.0Each pass works out the slope of the mean squared error with respect to w, then moves w a little against it. Real models have millions of parameters instead of one, but every one of them is adjusted the same way.
Test sets and overfitting
A model that does well on the examples it trained on has proved very little. The real question is how it does on data it has never seen, because that is what it will face once it is in use.
The test set
So before training, you split the data. Most of it becomes the training set, which the model learns from. A slice, often around a fifth, is held back as the test set. The model never sees it during training, and you only use it at the end to measure performance. It is the ML version of an exam with questions the student hasn't memorised.
scikit-learn does the split in one call:
from sklearn.datasets import load_iris
from sklearn.model_selection import (
train_test_split,
)
from sklearn.tree import DecisionTreeClassifier
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = (
train_test_split(
X, y, test_size=0.2, random_state=0
)
)
model = DecisionTreeClassifier()
model.fit(X_train, y_train)
print(model.score(X_train, y_train))
print(model.score(X_test, y_test))The two scores are the point. The first says how well the model fits what it has seen; the second is the honest estimate of how it will do in the real world.
Overfitting
Overfitting is when a model memorises its training data, noise and quirks included, instead of learning the general pattern. It scores brilliantly on the training set and badly on anything new. A decision tree with no depth limit is prone to it: it can keep splitting until it has a branch for almost every training example.
The tell-tale sign is a big gap between training and test scores. Common fixes are more training data, a simpler model (such as a depth limit on that tree), fewer or better features, and stopping training before the model starts chasing noise. Its opposite, underfitting, is a model too simple to capture the pattern at all, and it scores poorly on both sets.
Common mistakes
- Letting test data leak into training. If the model has seen the test set, even indirectly, its test score is meaningless. Split first, then do anything else.
- Treating the label as a feature. Feeding the answer in as an input gives a perfect score and a useless model.
- Judging a model on training accuracy. A near-perfect training score is more often a warning than a win.
- Picking a learning rate by luck. If the loss jumps around or grows, the steps are too big; if it barely moves, they are too small.
Key takeaways
- Machine learning is the part of AI that learns rules from examples instead of having them written by hand.
- Supervised learning uses labelled examples; unsupervised learning finds patterns in data with no labels.
- Each example is described by features, collected into a feature vector the model can read.
- A loss function measures how wrong the model is, and gradient descent lowers it one small step at a time.
- Always judge a model on a held-back test set: a big gap between training and test scores means it is overfitting.