Fine-tuning means taking an AI model that has already been trained and training it a little more, on a smaller dataset built for one specific job. It is how a general model that 'knows language' becomes one that answers your vet clinic's support chats in your clinic's style.
Starting from a foundation model
Training a large language model from nothing takes huge amounts of text, time and hardware. The result is a pre-trained model, also called a foundation model. It already knows language, code, patterns and a great deal of general knowledge, but it isn't specialised in anything.
Say Kitty runs a vet clinic. A foundation model can chat politely, but it doesn't know her booking rules, it answers in a generic tone, and it gets confused by her business's own phrases, like 'premium salmon experience'. Fine-tuning keeps everything the model already knows and adjusts it towards her job, instead of starting from scratch.
What fine-tuning changes
A model's behaviour lives in its parameters, also known as weights: billions of numbers learned during training. Fine-tuning shows the model many examples of the job done well and nudges those numbers so that its answers look more like the examples.
Because it changes the weights, fine-tuning is best at teaching behaviour:
- Style and tone: always friendly, short and in British English.
- Structure: always reply in a set format, such as a summary line then numbered steps.
- Narrow tasks: sorting messages into your own categories, or pulling the same fields out of every form.
It is much weaker at teaching facts, especially facts that change. That difference decides when to fine-tune at all.
Full fine-tuning
The most direct method is full fine-tuning: every parameter in the model is updated during training.
It is very powerful, because the whole model can adapt. It is also very expensive. Training needs memory not just for the weights but for extra numbers that track each one as it changes, so a large model may need several high-end GPUs and a lot of time. And each fine-tuned version is a full copy of the model, often many gigabytes.
PEFT: parameter-efficient fine-tuning
So most people use PEFT, parameter-efficient fine-tuning: methods that train only a small number of parameters and leave the rest alone.
LoRA: freeze the model, train a tiny add-on
The most popular PEFT method is LoRA (low-rank adaptation). It freezes the base model, so its original weights never change, and trains only a tiny add-on: two thin matrices of numbers next to each of the model's chosen weight matrices. When the model runs, the add-on's small adjustment is added to the frozen weights.
The numbers show why this is cheap. One weight matrix in a model might be 4,096 by 4,096 numbers:
Full matrix: 4,096 × 4,096 = 16,777,216
LoRA add-on, rank 8:
4,096 × 8 + 8 × 4,096 = 65,536The add-on is about 0.4% of the size of the matrix it adjusts. Train only those, across the model, and you need far less VRAM (the memory on a graphics card), far less time, and a far smaller cloud bill. The result is a small adapter file you can load on top of the original model, so one base model can serve several fine-tuned jobs.
In Python, Hugging Face's PEFT library sets this up in a few lines:
from peft import LoraConfig, get_peft_model
config = LoraConfig(
r=8, # rank: the add-on's size
lora_alpha=16, # how strongly it applies
target_modules=["q_proj", "v_proj"],
task_type="CAUSAL_LM",
)
model = get_peft_model(base_model, config)
model.print_trainable_parameters()The last line prints how few of the model's parameters will actually be trained, usually well under 1%.
QLoRA: shrink the base model too
QLoRA goes further. It shrinks the frozen base model to lower precision, storing each weight in 4 bits instead of 16, to save memory, then trains a LoRA add-on on top. That cuts the memory needed for the base model to about a quarter, which can bring fine-tuning within reach of a single graphics card. The cost is some extra computation during training.
The real secret: data quality
The method matters less than the examples. A few hundred clean, consistent examples can beat thousands of messy ones, because the model learns whatever patterns the data contains, including the mistakes.
Training data for a chat model is usually a file of example conversations, each one a small JSON record. Two records from Kitty's dataset might look like this (in the real file, each record sits on a single line):
{"messages": [
{"role": "user", "content": "Can I book a check-up for Mochi?"},
{"role": "assistant", "content": "Of course. Which day suits you?"}
]}
{"messages": [
{"role": "user", "content": "What's the premium salmon experience?"},
{"role": "assistant", "content": "It's our gourmet meal plan for recovering cats."}
]}Good datasets share a few habits:
- Keep the format consistent. Same structure, same tone, same length of answer in every example.
- Show what you want, not just facts. Fine-tune for style or structure.
- Remove the bad examples. Typos, wrong answers and contradictions get learned too.
- Hold some back. Keep examples the model never trains on, and test it on those.
When not to fine-tune
If Kitty needs exact facts from private documents, such as today's opening hours, prices or a patient's history, fine-tuning is the wrong tool. Facts baked into weights go out of date and can come back slightly wrong. RAG (retrieval-augmented generation) is often the better way: the app looks up the relevant documents when a question comes in and passes them to the model with the question. Change a document, and the next answer changes with it.
In practice, try the cheaper options in order:
- A better prompt, with clear instructions and examples. Often enough on its own; see context engineering.
- RAG, when the model needs your facts.
- Fine-tuning, when you need consistent behaviour that prompting can't get reliably.
So Kitty doesn't fine-tune unless she has no other choice. Fine-tuning is part of the wider toolkit of AI engineering, not the first step.
Common mistakes
- Fine-tuning to add knowledge. Use RAG for facts. Fine-tune for style or structure.
- Training on messy data. Inconsistent examples produce an inconsistent model.
- Training too long on too little. The model memorises the examples instead of learning the pattern. Check it on the held-back set.
- Skipping the comparison. Test the fine-tuned model against the original with a good prompt. Sometimes the prompt alone was enough.
Key takeaways
- Fine-tuning retrains an already trained model on a smaller, task-specific dataset.
- Full fine-tuning updates every parameter: very powerful, and very expensive.
- PEFT methods like LoRA freeze the base model and train a tiny add-on, saving memory, time and money.
- QLoRA also shrinks the base model to lower precision, saving more memory.
- Data quality beats quantity, and for exact facts from private documents, RAG is often the better way.