A large language model (LLM) is a program that writes by guessing the next small chunk of text, then the next, then the next. That one trick is behind chat assistants that draft emails, explain Python and summarise long documents, and knowing how it works tells you when to trust the answer and when to check it.
What the name means
LLM stands for large language model, and each word says something about how it works.
- Large: it is trained on huge amounts of text, and the model itself is huge too. What it learns is stored as billions of numbers called parameters, which are adjusted during training until the model gets good at guessing text.
- Language: it works with words. Text goes in, text comes out. Some models also handle images or audio, but text is the core.
- Model: it is a mathematical model that predicts what comes next. It doesn't look answers up; it calculates which continuation is most likely.
Tokens: the chunks a model reads
An LLM doesn't read whole words or letters. It splits text into tokens: small chunks that might be a whole word, a piece of a word, or a punctuation mark.
A common word such as 'cat' is usually a single token. A long or unusual word may be split into several pieces, and a full stop or a comma is often a token of its own. Exactly where the splits fall depends on the model's tokeniser, the part that turns text into tokens and back.
Tokens matter in practice for two reasons:
- Limits. A model can only consider a certain amount of text at once, called its context window, and that limit is counted in tokens.
- Cost. Paid AI APIs usually charge per token, for the text you send and the text you get back.
How a reply is built
When you send a prompt, the model doesn't plan the whole answer and then write it out. It works one token at a time:
- Read every token so far: your prompt plus anything it has already written.
- Work out a score for every token it knows, saying how likely each one is to come next.
- Pick one of the likely tokens and add it to the end.
- Go back to step 1 with the longer text, until the reply is finished.
So if you type 'The cat sat on the', the model scores possible next tokens. 'mat' might score highly, 'sofa' a little lower, and 'spreadsheet' very low. It picks one, adds it, and predicts again from the longer sentence.
Here is that loop, with each new token fed back in as input for the next guess:
Building a reply one token at a time
- Comparing
- Writing
- Done
Step 1 of 5: The prompt is five tokens. The model reads all of them before guessing.
Why the same prompt can give different answers
The model doesn't always take the single top-scoring token. Most chat tools pick from the likely few with a bit of randomness, controlled by a setting usually called temperature. Low temperature keeps it close to the top choice, so answers are more predictable. Higher temperature lets it take less likely tokens, so answers vary more. That is why asking the same question twice can give two different replies, an idea covered in deterministic vs non-deterministic.
A tiny next-word predictor
You can see the core idea in a few lines of Python. This toy counts which word follows which in a small piece of text, then always picks the most common follower:
from collections import Counter, defaultdict
text = ("the cat sat on the mat . "
"the cat sat on the sofa . "
"the dog sat on the mat .")
words = text.split()
follows = defaultdict(Counter)
for a, b in zip(words, words[1:]):
follows[a][b] += 1
def next_word(word):
return follows[word].most_common(1)[0][0]
out = ["the", "cat"]
for _ in range(4):
out.append(next_word(out[-1]))
print(" ".join(out))
# the cat sat on the catIt gets 'sat on the' right, then says 'the cat sat on the cat'. The toy only looks at the last word, and after 'the' it has seen 'cat' as often as anything else, so it repeats it.
A real LLM is the same idea scaled up enormously. Instead of counting pairs of words, it has learned billions of patterns from huge amounts of text, and it looks at every token in the context, not just the last one. That is why it can tell that 'sat on the' calls for somewhere to sit, keep track of who 'she' refers to three sentences back, and follow an instruction you gave at the start of the prompt.
What LLMs are good at
Because they have seen so much text, LLMs are strong wherever a plausible, well-shaped piece of writing is useful as a starting point:
- Drafting: emails, documentation, outlines and first versions of almost anything.
- Explaining: walking through a concept, a regular expression or an error message in plain words.
- Brainstorming: names, test cases, edge cases and alternative approaches.
- Coding help: boilerplate, small functions, refactoring ideas and translating code between languages.
- Summarising and translating: condensing a long text, or moving it between languages.
In each case you can check the result: you read the draft, run the code, or compare the summary with the original.
Where they go wrong
An LLM predicts what text sounds right. Usually that is also what is right, but not always. When it isn't, the model can hallucinate: produce something false, such as a function that doesn't exist or a reference that was never written, in the same confident tone as everything else. There is more on why in AI hallucinations.
A few other limits follow from how it works:
- No live knowledge. What it learned is fixed at training time. Anything newer, or anything it never saw, it can only guess at, unless the tool around it fetches fresh information.
- No built-in memory. The model only sees what is in its context window. Chat apps fake memory by sending earlier messages back in with each new one.
- Wording matters. A vague prompt gets a vague answer. Giving the model the right context, as described in context engineering, makes a large difference.
The rule of thumb is simple: use an LLM as a fast, capable assistant, and keep your own judgement switched on.
Common mistakes
- Treating it as a search engine. It generates likely text; it doesn't look facts up. Check names, numbers, links and dates.
- Running code you haven't read. Generated code can look right and still be wrong. Read it and test it.
- Thinking it 'knows' what it says. It doesn't think like a human. Confidence in its tone tells you nothing about accuracy.
- Pasting in more than it can see. A huge document can overflow the context window, and the parts that don't fit are simply not considered.
Key takeaways
- LLM stands for large language model: trained on huge amounts of text, working with words, predicting what comes next.
- It builds a reply one token at a time, feeding each new token back in before guessing the next.
- Tokens are small chunks of text: words, pieces of words and punctuation.
- LLMs are great for drafting, explaining, brainstorming and coding help.
- They can hallucinate, so check anything that matters.