One-hot encoding turns a column of labels, such as a cat's favourite fish, into numbers a machine learning model can use without inventing an order that isn't there. It is one of the first steps in preparing most datasets, and getting it wrong makes a model worse without any error message.
Why models need numbers
A machine learning model is maths: it multiplies inputs by weights and measures distances between rows, and that only works on numbers. A column like 'favourite fish', with values 'tuna', 'salmon' and 'sardine', is categorical data: each value is a label for a group, not an amount of anything.
Before training, every categorical column has to become numbers somehow. The question is how to do that without changing what the data means.
The problem with numbering the categories
The obvious fix is to give each label a number: tuna is 1, salmon is 2, sardine is 3. This is called label encoding (or integer encoding), and it takes one line of code. It also tells the model things that aren't true.
To a model, 3 is bigger than 1, and 2 sits halfway between them. So it now 'knows' that sardine is more than tuna, and that salmon is the average of the two. None of that is in the data. The numbers are just labels, not rankings.
You can see the damage in two common kinds of model:
- Linear models learn one weight per input column. With a single fish column, the fish's effect on the prediction is that weight times 1, 2 or 3. Salmon's effect is forced to sit exactly between tuna's and sardine's, whatever the real data says.
- Distance-based models, such as k-nearest neighbours, compare rows by how far apart their values are. Tuna (1) is one step from salmon (2) but two steps from sardine (3), so the model treats tuna and sardine as the least alike, with nothing in the data to support it.
Tree-based models, such as random forests, split on thresholds rather than multiplying, so arbitrary numbers hurt them less. They still split the codes in order, so one split can group tuna with salmon, but never tuna with sardine.
How one-hot encoding works
One-hot encoding gives each category its own column. For every row, the column matching its value gets a 1 and every other column gets a 0. Only one column is 'hot'; the rest stay 'cold'.
With the columns in the order tuna, salmon, sardine:
tuna -> 1, 0, 0
salmon -> 0, 1, 0
sardine -> 0, 0, 1Each fish is now a separate yes-or-no flag. No value is bigger than another, and every pair of fish is exactly the same distance apart. A linear model gets one weight per fish, so it can learn each fish's effect on its own.
In pandas
For exploring data, pandas.get_dummies does it in one call:
import pandas as pd
df = pd.DataFrame(
{"fish": ["tuna", "salmon", "sardine", "tuna"]}
)
encoded = pd.get_dummies(
df, columns=["fish"], dtype=int
)
print(encoded) fish_salmon fish_sardine fish_tuna
0 0 0 1
1 1 0 0
2 0 1 0
3 0 0 1The new columns come out in alphabetical order, and the original fish column is replaced. Without dtype=int, recent versions of pandas fill them with True and False instead of 1 and 0.
In scikit-learn
For a model you will train and then use on new data, scikit-learn's OneHotEncoder is the safer choice, because it remembers the categories it learned:
from sklearn.preprocessing import OneHotEncoder
train = [["tuna"], ["salmon"], ["sardine"]]
enc = OneHotEncoder(
handle_unknown="ignore", sparse_output=False
)
enc.fit(train)
print(enc.categories_)
print(enc.transform([["salmon"], ["cod"]]))[array(['salmon', 'sardine', 'tuna'], dtype=object)]
[[1. 0. 0.]
[0. 0. 0.]]It learns the categories from the training data with fit, then applies the same columns to anything you transform later. Salmon becomes 1, 0, 0 in its sorted column order. Cod was never seen in training, so with handle_unknown="ignore" it becomes a row of zeros. The default, handle_unknown="error", raises an error on unseen values instead.
The dummy variable trap
The one-hot columns for a single feature always add up to exactly 1. That means any one column can be worked out from the others: if a fish isn't tuna or salmon, it must be sardine.
For most models this redundancy is harmless. For a plain linear regression with an intercept, it causes perfect multicollinearity: the columns plus the intercept describe the same thing twice, so there is no single best set of weights. This is known as the dummy variable trap.
The fix is to drop one column (drop_first=True in pandas, drop="first" in scikit-learn). The dropped category becomes the baseline, and each remaining weight says how much that fish differs from it. Tree models and regularised models usually don't need this, and keeping every column makes the result easier to read.
When there are too many categories
One-hot encoding adds one column per category. Three fish make three columns. A thousand fish make a thousand columns, and every row has 999 zeros in it.
That is called high cardinality. The dataset grows wide and memory-hungry, rare categories have too few rows to learn a useful weight, and a model with that many inputs overfits more easily. Scikit-learn's encoder returns a sparse matrix by default, which stores only the 1s, so memory is less of a problem than it looks. The other problems remain.
When a column has many categories, the usual options are:
- Group the rare ones into a single 'other' column. Scikit-learn's encoder can do this with
min_frequencyormax_categories. - Target encoding replaces each category with the average outcome for that category in the training data. It is compact, but it can leak the answer into the input unless it is done carefully.
- Feature hashing squeezes any number of categories into a fixed number of columns, at the cost of some categories sharing a column.
- Learned embeddings give each category a short list of numbers that a neural network learns during training, so similar categories end up close together.
Embeddings grow straight out of one-hot encoding. A language model's vocabulary is tens of thousands of tokens, and looking up a token's embedding gives the same result as multiplying its one-hot vector by a table of weights, without ever building the huge vector. It is how a large language model turns each token of your prompt into numbers.
When the order is real
One-hot encoding is for categories with no natural order. Some categories do have one, such as T-shirt sizes (small, medium, large) or star ratings. There, numbering them in their real order keeps information the model can use, and one-hot encoding would throw it away.
This is ordinal encoding. In scikit-learn, OrdinalEncoder takes the order explicitly, so you decide that small is 0, medium is 1 and large is 2, rather than leaving it to alphabetical order (large, medium, small).
Common mistakes
- Encoding the training and test data separately: if a fish appears only in the training data,
get_dummieson the test data produces fewer columns, and the model breaks or silently reads the wrong ones. Fit one encoder on the training data and reuse it everywhere. - Forgetting about new categories: production data will eventually contain a value the model has never seen. Decide up front whether that should fail loudly or become a row of zeros.
- Treating numeric codes as numbers: postcodes and store IDs are made of digits, but they are labels. A model that sees store 400 as 'more' than store 12 has learned nothing true.
- One-hot encoding free text or IDs: a column where nearly every row is unique turns into thousands of columns that each describe one row.
Key takeaways
- Models need numbers, but numbering categories 1, 2, 3 invents an order and distances that aren't real.
- One-hot encoding gives each category its own 0-or-1 column, so exactly one is hot in every row.
- Fit the encoder on training data and reuse it, and decide how unseen categories are handled.
- It suits columns with a small number of unordered categories; for many categories, group, hash, target-encode or use embeddings.
- When the order is real, use ordinal encoding instead.